Coding and decoding method, encoder, decoder, code stream and storage medium
Patent Information
- Application Number
- CN202280100853.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2025-05-23
AI Technical Summary
In video coding and decoding systems, different coding parameters and error distributions of adjacent coding units lead to block effects, which affect image quality and coding and decoding performance. The existing low-complexity preset network model has a single processing method during filtering, resulting in coding and decoding effects. Not good.
By parsing the code stream, we determine the syntax element identification information of the component to be filtered, determine whether to use the preset network model for filtering, select quantization parameters according to different rate distortion costs, and flexibly select filtering methods to improve encoding and decoding efficiency.
Without increasing the complexity of the model, the compression performance and encoding and decoding efficiency of different components to be filtered are improved, and the quantization parameter information is flexibly selected to adapt to different scenarios.
Smart Images

Figure CN120035991A_ABST
Abstract
Description
Coding and decoding method, encoder, decoder, code stream and storage medium Technical Field
[0001] The embodiments of the present application relate to the field of image processing technology, and in particular to a coding and decoding method, an encoder, a decoder, a bit stream, and a storage medium. Background Art
[0002] In video codec systems, most video codecs use a block-based hybrid coding framework. Each video frame is divided into several Coding Tree Units (CTUs), which can be further divided into several rectangular Coding Units (CUs), which can be rectangular or square blocks. Because adjacent CUs use different coding parameters, such as different transform processes, different quantization parameters (QPs), different prediction methods, and different reference image frames, and because the error size and distribution characteristics introduced by each CU are independent of each other, the discontinuity of adjacent CU boundaries causes blocking artifacts, which affects the subjective and objective quality of the reconstructed image and even the prediction accuracy of subsequent codecs.
[0003] In this way, during the encoding and decoding process, loop filters are used to improve the subjective and objective quality of the reconstructed image. The neural network-based loop filtering method has the most outstanding coding performance. In related technologies, for different test conditions and quantization parameters, loop filtering can be performed using only a simplified low-complexity preset network model. When using the low-complexity preset network model for filtering, quantization parameter information is added as an additional input, that is, the quantization parameter information is used as the input of the network to improve the generalization ability of the neural preset network model, so as to achieve good coding performance without switching the preset network model.
[0004] However, when a low-complexity preset network model is used for filtering, different color components are processed in a single way, and the selection during filtering is not flexible enough, resulting in poor encoding and decoding effects during encoding and decoding.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide a coding and decoding method, an encoder, a decoder, a code stream, and a storage medium, which can make the selection of input information for filtering more flexible without increasing the complexity, thereby improving the coding and decoding efficiency.
[0007] The technical solution of the embodiment of the present application can be implemented as follows:
[0008] In a first aspect, an embodiment of the present application provides a filtering method, applied to a decoder, the method comprising:
[0009] Parsing the bitstream to determine first syntax element identification information of a component to be filtered of a current frame or a current slice; the first syntax element identification information is used to determine whether the component to be filtered of each block in the current frame or the current slice is all filtered based on a preset network model;
[0010] When the first syntax element identification information indicates that there is a to-be-filtered component of a partitioned block in the current frame or the current slice that allows filtering using a preset network model, determining quantization parameter information of the to-be-filtered component and determining second syntax element identification information;
[0011] Based on the second syntax element identification information, the quantization parameter information and the preset network model, the current block of the current frame or the current slice is filtered to obtain a filtered reconstructed value of the to-be-filtered component of the current block.
[0012] In a second aspect, an embodiment of the present application provides a filtering method, applied to an encoder, the method comprising:
[0013] Determining a first rate-distortion cost of a current frame or a current slice; the first rate-distortion cost is obtained by filtering all to-be-filtered components of all divided blocks included in the current frame or the current slice without using a preset network model;
[0014] Determining a second rate-distortion cost of a current frame or a current slice; the second rate-distortion cost is obtained by filtering all to-be-filtered components of all divided blocks included in the current frame or the current slice using the preset network model;
[0015] Determining a third rate-distortion cost for a current frame or a current slice; the third rate-distortion cost is obtained by allowing filtering of a to-be-filtered component of at least one partition block in the current frame or the current slice using the preset network model; the at least one partition block is a portion of the partition blocks in the current frame or the current slice;
[0016] First syntax element identification information of the to-be-filtered component of the current frame or current slice is determined according to the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost.
[0017] In a third aspect, an embodiment of the present application provides a decoder, including:
[0018] A parsing part is configured to parse the bitstream and determine first syntax element identification information of a component to be filtered of a current block; the first syntax element identification information is used to determine whether each block in the current frame or current slice is filtered based on a preset network model;
[0019] a first determining part configured to, when the first syntax element identification information indicates that a to-be-filtered component of a partitioned block in the current frame or the current slice allows filtering using a preset network model, determine quantization parameter information of the to-be-filtered component and determine second syntax element identification information;
[0020] The first filtering part is configured to filter the current block of the current frame or the current slice based on the second syntax element identification information, the quantization parameter information and the preset network model to obtain a filtered reconstructed value of the to-be-filtered component of the current block.
[0021] In a fourth aspect, an embodiment of the present application provides a decoder, including:
[0022] a first memory configured to store a computer program executable on the first processor;
[0023] The first processor is configured to execute the method described in the first aspect when running the computer program.
[0024] In a fifth aspect, an embodiment of the present application provides an encoder, including:
[0025] The second determining part is configured to determine a first rate-distortion cost of the current frame or the current slice; the first rate-distortion cost is obtained by filtering all to-be-filtered components of all divided blocks included in the current frame or the current slice without using a preset network model;
[0026] The second filtering part is configured to determine a second rate-distortion cost for the current frame or the current slice; the second rate-distortion cost is obtained by filtering the to-be-filtered components of all the partitions included in the current frame or the current slice using the preset network model; determine a third rate-distortion cost for the current frame or the current slice; the third rate-distortion cost is obtained by allowing the to-be-filtered components of at least one partition in the current frame to be filtered using the preset network model; the at least one partition is a portion of the partitions in the current frame or the current slice;
[0027] The second determining part is further configured to determine first syntax element identification information of the to-be-filtered component of the current frame or the current slice according to the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost.
[0028] In a sixth aspect, an embodiment of the present application provides an encoder, including:
[0029] a second memory configured to store a computer program executable on the second processor;
[0030] The second processor is configured to execute the method described in the second aspect when running the computer program.
[0031] In a seventh aspect, an embodiment of the present application provides a computer storage medium storing a computer program, which implements the method described in the first aspect when executed by a first processor, or implements the method described in the second aspect when executed by a second processor.
[0032] In an eighth aspect, an embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following: at least one of quantization parameter information or quantization parameter index, first syntax element identification information of the to-be-filtered component of the current frame or current slice, second syntax element identification information of the to-be-filtered component of the current block, third syntax element identification information of the current block contained in the current frame or the current slice, fourth syntax element identification information of the current video sequence, and filtered reconstructed values of each partitioned block included in the current frame or the current slice; wherein the current block is any one of the partitioned blocks.
[0033] Embodiments of the present application provide a filtering method, an encoder, a decoder, a bitstream, and a storage medium. In the decoder, first syntax element identification information of a component to be filtered of a current frame or a current slice is determined by parsing the bitstream; the first syntax element identification information is used to determine whether the components to be filtered of each block in the current frame or the current slice are all filtered based on a preset network model; when the first syntax element identification information indicates that there are components to be filtered of divided blocks in the current frame or the current slice that allow filtering using the preset network model, quantization parameter information of the components to be filtered is determined, and second syntax element identification information is determined; based on the second syntax element identification information, the quantization parameter information, and the preset network model, the current block of the current frame or the current slice is filtered to obtain a filtered reconstructed value of the component to be filtered of the current block. In the encoder, a first rate-distortion cost of the current frame or current slice is determined; the first rate-distortion cost is obtained by not filtering the to-be-filtered components of all the partitioned blocks included in the current frame or current slice using a preset network model; a second rate-distortion cost of the current frame or current slice is determined; the second rate-distortion cost is obtained by filtering the to-be-filtered components of all the partitioned blocks included in the current frame or current slice using a preset network model; a third rate-distortion cost of the current frame or current slice is determined; the third rate-distortion cost is obtained by allowing the to-be-filtered components of at least one partitioned block in the current frame or current slice to be filtered using a preset network model; at least one partitioned block is a partial partitioned block in the current frame or current slice; based on the first rate-distortion cost, the second rate-distortion cost and the third rate-distortion cost, the first syntax element identification information of the to-be-filtered components of the current frame or current slice is determined.
[0034] The present application determines the encoding method with the lowest frame-level rate-distortion cost or slice-level rate-distortion cost through multiple model reasoning, from the cases where each partition block is not filtered, each partition block is filtered, and each partition block is partially filtered. The quantization parameter information corresponding to each component to be filtered can be different types of quantization parameter information under different circumstances. In other words, the quantization parameter information corresponding to different components to be filtered of the current frame or current slice can be different (the quantization parameter information of each component to be filtered is the optimal quantization parameter information obtained by the encoder through multiple model reasoning during encoding), and each component to be filtered can be filtered and reconstructed to obtain the filtered value of each component to be filtered. Therefore, each component to be filtered of the current block can use its own quantization parameter information to implement filtering of each component to be filtered in the same preset network model, thereby improving the compression performance of each component to be filtered. On the basis of ensuring that the complexity of the model is not increased, the selection of input information (quantization parameter information) for filtering different components to be filtered is more flexible, thereby improving the coding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] 1A-1C are exemplary component distribution diagrams in different color formats provided by an embodiment of the present application;
[0036] FIG2 is a schematic diagram of an exemplary division of coding units provided in an embodiment of the present application;
[0037] FIG3A is a first schematic diagram of a network architecture of an exemplary neural network model provided in an embodiment of the present application;
[0038] FIG3B is a first exemplary structure of a residual block according to an embodiment of the present application;
[0039] FIG4 is a second schematic diagram of a network architecture of an exemplary neural network model provided in an embodiment of the present application;
[0040] FIG5A is a third schematic diagram of a network architecture of an exemplary neural network model provided in an embodiment of the present application;
[0041] FIG5B is a second exemplary structure of a residual block provided in an embodiment of the present application;
[0042] FIG6 is a structural diagram of an exemplary video encoding system provided in an embodiment of the present application;
[0043] FIG7 is a structural diagram of an exemplary video decoding system provided in an embodiment of the present application;
[0044] FIG8 is a schematic diagram of a flowchart of a decoding method provided in an embodiment of the present application;
[0045] FIG9 is a schematic diagram of a flow chart of an encoding method provided in an embodiment of the present application;
[0046] FIG10A is a first schematic diagram of selecting quantization parameter information of a block under different components to be filtered according to an exemplary embodiment of the present application;
[0047] FIG10B is a second schematic diagram of selecting quantization parameter information of a block under different components to be filtered according to an exemplary embodiment of the present application;
[0048] FIG10C is a third exemplary diagram of selecting quantization parameter information of a block under different components to be filtered provided by an embodiment of the present application;
[0049] FIG11 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application;
[0050] FIG12 is a schematic diagram of the hardware structure of a decoder provided in an embodiment of the present application;
[0051] FIG13 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;
[0052] FIG14 is a schematic diagram of the hardware structure of an encoder provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In the embodiments of the present application, digital video compression technology primarily compresses large amounts of digital video data for easier transmission and storage. With the surge in Internet video usage and increasing demand for higher video clarity, while existing digital video compression standards can save significant amounts of video data, there is still a need for better digital video compression technologies to reduce bandwidth and traffic pressures associated with digital video transmission.
[0054] During digital video encoding, the encoder reads unequal pixels from raw video sequences in different color formats, including luminance and chrominance components. This means the encoder reads a black-and-white or color image. The image is then divided into blocks and encoded by the encoder, which typically uses a hybrid frame coding scheme, typically involving intra-frame and inter-frame prediction, transform and quantization, inverse transform and inverse quantization, loop filtering, and entropy coding. Intra-frame prediction refers only to information from the same frame, predicting the pixels within the current block to eliminate spatial redundancy. Inter-frame prediction can reference information from different frames, using motion estimation to search for the motion vector that best matches the current block, eliminating temporal redundancy. Transform and quantization convert the predicted image blocks to the frequency domain, redistributing the energy. Combined with quantization, this process removes information that is insensitive to the human eye, eliminating visual redundancy. Entropy coding eliminates character redundancy based on the current context model and the probabilistic information of the binary bitstream. Loop filtering primarily processes the inverse-transformed and inverse-quantized pixels to compensate for distortion and provide a better reference for subsequent pixel encoding.
[0055] Currently, the scenario in which filtering processing can be performed can be the reference software test platform HPM based on AVS or the VVC reference software test platform (VVC TEST MODEL, VTM) based on Versatile Video Coding (VVC), which is not limited in the embodiments of the present application.
[0056] In a video image, a first video component, a second video component, and a third video component are generally used to represent a current block (Coding Block, CB); wherein, the three image components are a luminance component, a blue chrominance component, and a red chrominance component, respectively. The luminance component is usually represented by the symbol Y, the blue chrominance component is usually represented by the symbol Cb or U, and the red chrominance component is usually represented by the symbol Cr or V; in this way, the video image can be represented in YCbCr format, or in YUV format, or even in RGB format or YCgCo format, and the embodiments of the present application do not impose any restrictions thereon.
[0057] Video compression technology primarily compresses large amounts of digital video data for easier transmission and storage. With the surge in internet video usage and increasing demand for higher-quality video, while existing digital video compression standards can save significant amounts of video data, there is a continued need for better digital video compression technologies to reduce the bandwidth and traffic pressure of digital video transmission. During video encoding, the encoder reads unequal pixels from the original video sequence in different color formats, including the luminance and chrominance color components. In other words, the encoder reads a black-and-white or color image. It then divides the image into blocks, and the block data is passed to the encoder for encoding.
[0058] Typically, digital video compression technology operates on image data encoded in the YCbCr (YUV) color encoding format. The YUV ratio is typically 4:2:0, 4:2:2, or 4:4:4. Y represents luminance (Luma), Cb (U) represents blue chrominance, and Cr (V) represents red chrominance. U and V represent chroma, used to describe color and saturation. Figures 1A through 1C illustrate the distribution of components in different color formats, with white representing the Y component and dark gray representing the UV components. As shown in Figure 1A, a 4:2:0 color format indicates four luminance components and two chrominance components for every four pixels (YYYYCbCr). As shown in Figure 1B, a 4:2:2 format indicates four luminance components and four chrominance components for every four pixels (YYYYCbCrCbCr). As shown in Figure 1C, a 4:4:4 format indicates a full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0059] Currently, common video codec standards all use a block-based hybrid coding framework. Each video frame is divided into square Largest Coding Units (LCUs) of the same size (e.g., 128×128, 64×64, etc.). Each LCU can be further divided into rectangular Coding Units (CUs) according to a rule. Coding Units may also be divided into smaller Prediction Units (PUs). Specifically, the hybrid coding framework may include modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module may include intra-frame prediction and inter-frame prediction, and inter-frame prediction may include motion estimation and motion compensation. Since there is a strong correlation between adjacent pixels within a video frame, the use of intra-frame prediction in video codec technology can eliminate spatial redundancy between adjacent pixels. Inter-frame prediction can refer to image information from different frames and use motion estimation to search for the motion vector information that best matches the current partition block to eliminate temporal redundancy; the transformation converts the predicted image block to the frequency domain, redistributes the energy, and combines it with quantization to remove information that the human eye is not sensitive to, thereby eliminating visual redundancy; entropy coding can eliminate character redundancy based on the current context model and the probability information of the binary code stream.
[0060] It should be noted that during the video encoding process, the encoder first reads the image information and divides the image into several coding tree units (CTUs). A coding tree unit can be further divided into several coding units (CUs). These coding units can be rectangular blocks or square blocks. The specific relationship can be shown in Figure 2.
[0061] During intra-frame prediction, the current coding unit cannot reference information from different frames and can only use adjacent coding units from the same frame as reference information for prediction. This means that, based on the current left-to-right, top-to-bottom coding order, the current coding unit can reference the upper-left coding unit, the upper coding unit, and the left coding unit as reference information to predict the current coding unit. The current coding unit then serves as reference information for the next coding unit, thus predicting the entire image. If the input digital video is in color format, the current mainstream digital video encoder input source is in YUV 4:2:0 format, meaning that every four pixels in the image are composed of four Y components and two UV components. The encoder encodes the Y and UV components separately, using slightly different encoding tools and techniques. The decoder also decodes the video according to the different formats.
[0062] Intra-frame prediction in digital video encoding and decoding primarily uses information from adjacent blocks in the current frame to predict the current block. The residual information is then calculated between the predicted block and the original image block. This residual information is then transformed and quantized before being transmitted to the decoder. After receiving and parsing the bitstream, the decoder performs an inverse transform and quantization to obtain the residual information. This residual information is then superimposed on the predicted image block obtained by the decoder to create the reconstructed image block.
[0063] Currently common video codec standards (such as H.266 / VVC) all adopt a block-based hybrid coding framework. Each video frame is divided into square largest coding units (LCUs) of the same size (e.g., 128x128, 64x64, etc.). Each LCU can be divided into rectangular coding units (CUs) according to a rule. Coding units may also be divided into prediction units (PUs) and transform units (TUs). The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module includes intra-frame prediction and inter-frame prediction. Inter-frame prediction includes motion estimation and motion compensation. Because adjacent pixels in a video frame are strongly correlated, intra-frame prediction is used in video codecs to eliminate spatial redundancy between adjacent pixels. Because adjacent frames in a video have strong similarities, inter-frame prediction is used in video codecs to eliminate temporal redundancy between adjacent frames, thereby improving coding efficiency.
[0064] The basic process of a video codec is as follows. On the encoder side, a frame is divided into blocks. Intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block for the current block. The prediction block is subtracted from the original image block to obtain a residual block. This residual block is transformed and quantized to obtain a quantization coefficient matrix. This quantization coefficient matrix is entropy-encoded and output to the bitstream. On the decoder side, intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block for the current block. The bitstream is then parsed to obtain a quantization coefficient matrix. This quantization coefficient matrix is inversely quantized and inversely transformed to obtain a residual block. The prediction block and residual block are added together to obtain a reconstructed block. The reconstructed blocks form a reconstructed image, which is then subjected to image-based or block-based loop filtering to obtain a decoded image. The encoder side also performs similar operations to the decoder side to obtain a decoded image. The decoded image can serve as a reference frame for inter-frame prediction in subsequent frames. The block division information determined by the encoder, as well as information about the prediction, transform, quantization, entropy coding, loop filtering, and other modes or parameters, are output to the bitstream if necessary. The decoding end determines the same block division information as the encoding end by parsing and analyzing the existing information, as well as the mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, thereby ensuring that the decoded image obtained by the encoding end is the same as the decoded image obtained by the decoding end. The decoded image obtained by the encoding end is also usually called a reconstructed image. The current block can be divided into prediction units during prediction, and can be divided into transformation units during transformation. The division of prediction units and transformation units can be different. The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. The current block (current block) can be the current coding unit (CU) or the current prediction unit (PU) or the current transformation unit (TU), etc., and is not limited in the embodiments of the present application.
[0065] JVET, the international video coding standards development organization, has established two exploratory experimental groups, one based on neural network coding and the other beyond VVC, and has also set up several corresponding expert discussion groups.
[0066] The above-mentioned exploratory experimental group that surpasses VVC aims to explore higher coding efficiency based on the latest codec standard H.266 / VVC with strict performance and complexity requirements. The coding method studied by this group is closer to VVC and can be called a traditional coding method. At present, the performance of the algorithm reference model of this exploratory experiment has surpassed the coding performance of the latest VVC reference model VTM by about 15%.
[0067] The method studied by the first exploratory group is an intelligent coding method based on neural networks. Deep learning and neural networks are currently hot topics across various industries, especially in computer vision, where deep learning-based methods often have overwhelming advantages. Experts from the JVET standards organization have introduced neural networks to the field of video codecs. Leveraging the powerful learning capabilities of neural networks, neural network-based coding tools often achieve very high coding efficiency. In the early stages of VVC standard development, many companies focused on coding tools based on deep learning, proposing methods including neural network-based intra-frame prediction, neural network-based inter-frame prediction, and neural network-based loop filtering. Among them, the neural network-based loop filtering method has the most outstanding coding performance. After multiple research and exploration meetings, the coding performance has reached over 8%. The neural network-based loop filtering scheme studied by the first exploratory group has currently achieved coding performance as high as 12%, almost contributing to half a generation of coding performance.
[0068] This embodiment of the present application improves upon the exploratory experiments conducted at the JVET conference and proposes a neural network (NN)-based loop filtering enhancement solution. The following will first briefly introduce the neural network-based loop filtering solution currently used at the JVET conference, followed by a detailed description of the improved methods employed in this embodiment of the present application.
[0069] In the related art, exploration of neural network-based loop filtering solutions has primarily focused on two approaches: the first employing a multi-model intra-frame switchable scheme, and the second employing a non-intra-frame switchable model. Regardless of the approach, the neural network architecture remains largely unchanged, and the tool is incorporated into the in-loop filtering of traditional hybrid coding frameworks. Therefore, the fundamental processing unit for both approaches is the coding tree unit, which is the maximum coding unit size.
[0070] The biggest difference between the first scheme in which multiple models can be switched within a frame and the second scheme in which models cannot be switched within a frame is that the first scheme can switch the neural network model at will when encoding and decoding the current frame, while the second scheme cannot switch the neural network model. Taking the first scheme as an example, when encoding a frame of an image, each coding tree unit has multiple candidate neural network models to choose from. The encoder selects which neural network model to use for the current coding tree unit with the best filtering effect, and then writes the neural network model index into the bitstream. That is, in this scheme, if the coding tree unit needs to be filtered, it is necessary to first transmit a coding tree unit-level usage flag, and then transmit the neural network model index. If filtering is not required, it is only necessary to transmit a coding tree unit-level usage flag; after parsing the index value, the decoder loads the neural network model corresponding to the index into the current coding tree unit to filter the current coding tree unit.
[0071] Taking the second scheme as an example, when encoding a frame of image, the available neural network model for each coding tree unit in the current frame is fixed, and each coding tree unit uses the same neural network model, that is, the second scheme does not have a model selection process at the encoding end; the decoding end parses and obtains the usage flag of whether the current coding tree unit uses the neural network-based loop filter. If the usage flag is true, the pre-set model (the same as the encoding end) is used to filter the coding tree unit. If the usage flag is false, no additional operation is performed.
[0072] For the first multi-model intra-frame switchable solution, it has strong flexibility at the coding tree unit level and can adjust the model according to local details, that is, local optimization to achieve a better global effect. Usually, this solution has more neural network models, and different neural network models are trained under different quantization parameters for the JVET general test conditions. At the same time, different coding frame types may also require different neural network models to achieve better results. Taking a filter in the related technology as an example, the filter uses up to 22 neural network models to cover different coding frame types and different quantization parameters, and the model switching is performed at the coding tree unit level. This filter can provide up to 10% more coding performance based on VVC.
[0073] For the second solution in which the model cannot be switched within a frame, although the solution has two neural network models overall, the model is not switched within the frame. The solution makes a judgment at the encoding end. If the current encoding frame type is an I frame, the neural network model corresponding to the I frame is imported, and only the neural network model corresponding to the I frame is used in the current frame; if the current encoding frame type is a B frame, the neural network model corresponding to the B frame is imported, and similarly only the neural network model corresponding to the B frame is used in the current frame. This solution can provide 8.65% encoding performance based on VVC. Although it is slightly lower than the first solution, the overall performance is a coding efficiency that is almost impossible to achieve compared to traditional encoding tools.
[0074] While the first approach offers greater flexibility and higher encoding performance, it suffers from a significant hardware implementation drawback. Hardware experts are concerned about the code for intra-frame model switching. Switching models at the coding tree unit level means, in the worst case, that the decoder must reload the neural network model for each coding tree unit processed. This, in addition to the hardware implementation complexity, creates an additional burden on current high-performance graphics processing units (GPUs). Furthermore, the presence of multiple models requires a large number of parameters to be stored, which is a significant hardware implementation overhead. However, the second approach, a neural network loop filter, further explores the powerful generalization capabilities of deep learning. It uses a variety of information as input, rather than simply reconstructed samples. This increased information provides more information for neural network learning, enhancing model generalization and eliminating many unnecessary redundant parameters. Continuously updated approaches have now emerged that can adapt to different test conditions and quantization parameters using a single, simplified, low-complexity neural network model. Compared to the first approach, this eliminates the overhead of constantly reloading the model and the need for larger storage space to accommodate a large number of parameters.
[0075] The following will introduce the neural network architecture of these two solutions.
[0076] Refer to Figure 3A, which shows a schematic diagram of the network architecture of a neural network model. As shown in Figure 3A, the main structure of the network architecture can be composed of multiple residual blocks (ResBlocks). The composition structure of the residual block is shown in Figure 3B. In Figure 3B, a single residual block is composed of multiple convolutional layers (Conv) connected to a convolutional attention mechanism module (Convolutional Blocks Attention Module, CBAM) layer. As an attention mechanism module, CBAM is mainly responsible for further extraction of detail features. In addition, the residual block also has a direct skip connection (Skip Connection) structure between the input and output. Here, the multiple convolutional layers in Figure 3B include the first convolutional layer, the second convolutional layer and the third convolutional layer, and the first convolutional layer is also connected to an activation layer. For example, the size of the first convolutional layer is 1×1×k×n, the size of the second convolutional layer is 1×1×n×k, and the size of the third convolutional layer is 3×3×k×k, where k and n are positive integers. The activation layer can include a rectified linear unit (ReLU) function, also known as a linear rectifier function, which is a commonly used activation function in current neural network models. ReLU is actually a ramp function that is simple and converges quickly.
[0077] As for Figure 3A, there is also a skip connection structure in the network architecture, which connects the input reconstructed YUV information to the output of the pixel reorganization (Pixel Shuffle) module. The main function of Pixel Shuffle is to obtain a high-resolution feature map through convolution and multi-channel reorganization of low-resolution feature maps; as an upsampling method, it can effectively amplify the reduced feature map. In addition, the input of the network architecture mainly includes reconstructed YUV information (rec_yuv), predicted YUV information (pred_yuv), and YUV information with partitioning information (par_yuv). All inputs undergo simple convolution and activation operations and are then spliced (Cat) and sent to the main structure. It is worth noting that the processing of YUV information with partitioning information may be different in I frames and B frames. I frames require input of YUV information with partitioning information, while B frames do not.
[0078] In summary, for every JVET-required parameter point in each I-frame and B-frame, the first solution has a corresponding neural network model. Furthermore, because the three YUV color components are primarily composed of two channels, luminance and chrominance, the color components differ.
[0079] Refer to Figure 4, which shows a schematic diagram of the network architecture of another neural network model. As shown in Figure 4, the network architecture of the first scheme is basically the same as the second scheme in the main structure. The difference is that the input of the second scheme adds quantization parameter information as an additional input compared to the first scheme. The first scheme mentioned above loads different neural network models according to different quantization parameter information to achieve more flexible processing and more efficient encoding effects, while the second scheme uses quantization parameter information as the input of the network to improve the generalization ability of the neural network, so that the model can adapt to different quantization parameter conditions and provide good filtering performance.
[0080] As can be seen from Figure 4, there are two quantization parameters that enter the neural network model as input, one is BaseQP and the other is SliceQP. BaseQP here indicates the sequence-level quantization parameter set by the encoder when encoding the video sequence, that is, the quantization parameter point required by the JVET pass test, and is also the parameter used to determine the neural network model in the first solution. SliceQP is the quantization parameter of the current frame. The quantization parameter of the current frame can be different from the sequence-level quantization parameter. This is because in the video encoding process, the quantization conditions of the B frame are different from those of the I frame, and the quantization parameters are different at different time domain levels. Therefore, SliceQP is generally different from BaseQP in the B frame. Therefore, in the relevant technology, the input of the neural network model of the I frame only requires SliceQP, while the neural network model of the B frame requires both BaseQP and SliceQP as input.
[0081] In addition, the second solution differs from the first solution in one respect. The output of the model in the first solution generally does not require additional processing. Specifically, if the model output is residual information, it is superimposed with the reconstructed samples of the current coding tree unit and used as the output of the neural network-based loop filter tool. If the model output is a complete reconstructed sample, the model output is the output of the neural network-based loop filter tool. The output of the second solution generally requires scaling. For example, the model outputs residual information of the current coding tree unit. This residual information is scaled and then superimposed with the reconstructed sample information of the current coding tree unit. This scaling factor is obtained by the encoder and written into the code stream for transmission to the decoder.
[0082] In summary, it is precisely because the quantization parameters serve as additional information that the reduction in the number of models is possible, making it the most popular solution at the current JVET conference. Furthermore, general neural network-based loop filtering solutions may not be identical to the two solutions described above. While the specific details may differ, the underlying concepts remain largely the same. For example, the differences in solution 2 can be reflected in the design of the neural network architecture, such as the convolution size of ResBlocks, the number of convolution layers, and whether an attention module is included. They can also be reflected in the inputs to the neural network, which can even include more additional information, such as the boundary strength value of the deblocking filter.
[0083] In addition, neural network-based loop filtering solutions can be further optimized. For example, a solution can be implemented where a single model processes images of different frame types and different quantization parameter configurations, and the same model simultaneously outputs filtering results for different color components. The main model framework is shown in Figure 5A. The structure of the residual block is detailed in Figure 5B. The input to the neural network model consists of the reconstructed sample YUV, the predicted sample YUV, BaseQP, SliceQP, and the frame type. Using a single neural network model, a single inference cycle outputs filtered image results for all three color components, significantly reducing model storage and computational complexity.
[0084] The first and second solutions significantly reduce the implementation complexity of neural network loop filtering technology while maintaining relatively good performance. However, regardless of whether the neural network loop filter is a single model or a dual model, processing both the luma and chroma components is performed by a single model. Through parameter training, good luma performance can be maintained. However, if the loop filter is improved by using four neural network models, giving the chroma component an independent neural network model, the chroma performance can be improved by an average of 2-5%. In other words, if the luma performance is not transferred, there is still room for improvement in the chroma performance. Therefore, by adding an independent neural network model for the chroma component or by adding an additional channel BaseQP input to control the chroma quantization parameter, the compression performance of the chroma component can be improved. However, adding models and an additional channel BaseQP input increases the computational complexity of the model.
[0085] In an embodiment of the present application, a solution of multiple inferences of the same model is proposed to adapt to the requirements for quantization parameters between different color components. For the neural preset network model architecture shown in Figure 5A. For a quantization parameter BaseQP or sliceQP, the preset network model outputs different filtering results of three color components in the inference stage. For another different quantization parameter, the preset network model infers the current image again to obtain another set of filtering results of three different color components, and compares the filtering results of different quantization parameters of the three color components by rate-distortion cost. The best result among each color component is taken as the output of the preset network model, so that the compression performance of each color component can be well supported without increasing the complexity of the model and channel.
[0086] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0087] The embodiment of the present application provides a video coding system. FIG6 is a schematic diagram of the composition structure of the video coding system of the embodiment of the present application. The video coding system 10 includes: a transform and quantization unit 101, an intra-frame estimation unit 102, an intra-frame prediction unit 103, a motion compensation unit 104, a motion estimation unit 105, an inverse transform and inverse quantization unit 106, a filter control analysis unit 107, a filtering unit 108, a coding unit 109 and a decoded image cache unit 110, etc., wherein the filtering unit 108 can implement DBF filtering / SAO filtering / ALF filtering, and the coding unit 109 can implement header information encoding and context-based adaptive binary arithmetic coding (Context-based Adaptive Binary Arithmetic Coding, CABAC). For the input original video signal, the coding tree unit (Coding Tree A video coding block can be obtained by dividing the video coding block into a plurality of frames (CTUs) and then the residual pixel information obtained after intra-frame or inter-frame prediction is transformed by the transform and quantization unit 101, including transforming the residual information from the pixel domain to the transform domain and quantizing the obtained transform coefficients to further reduce the bit rate; the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to perform intra-frame prediction on the video coding block; specifically, the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to determine the intra-frame prediction mode to be used to encode the video coding block; the motion compensation unit 104 and the motion estimation unit 105 are used to perform inter-frame prediction coding of the received video coding block relative to one or more blocks in one or more reference frames to provide temporal prediction information; the motion estimation performed by the motion estimation unit 105 is a process of generating a motion vector, which can estimate the motion of the video coding block, and then the motion compensation unit 104 calculates the motion vector based on the motion vector determined by the motion estimation unit 105. After determining the intra-frame prediction mode, the intra-frame prediction unit 103 is further configured to provide the selected intra-frame prediction data to the encoding unit 109, and the motion estimation unit 105 also sends the calculated motion vector data to the encoding unit 109. In addition, the inverse transform and inverse quantization unit 106 is configured to reconstruct the video coding block and reconstruct a residual block in the pixel domain. The reconstructed residual block is subjected to the filter control analysis unit 107 and the filtering unit 108 to remove the block effect artifacts. The reconstructed residual block is then added to a predictive block in the frame of the decoded image buffer unit 110 to generate a reconstructed video coding block. The encoding unit 109 is configured to encode various coding parameters and quantized transform coefficients. In the CABAC-based coding algorithm, the context content can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-frame prediction mode, and output the code stream of the video signal. The decoded image buffer unit 110 is configured to store the reconstructed video coding block for prediction reference.As the video image encoding proceeds, new reconstructed video encoding blocks are continuously generated, and these reconstructed video encoding blocks are stored in the decoded image buffer unit 110 .
[0088] An embodiment of the present application provides a video decoding system. FIG7 is a schematic diagram of the composition structure of the video decoding system of the embodiment of the present application. The video decoding system 20 includes: a decoding unit 201, an inverse transform and inverse quantization unit 202, an intra-frame prediction unit 203, a motion compensation unit 204, a filtering unit 205 and a decoded image cache unit 206, etc., wherein the decoding unit 201 can implement header information decoding and CABAC decoding, and the filtering unit 205 can implement DBF filtering / SAO filtering / ALF filtering. After the input video signal is encoded as shown in FIG3A , a code stream of the video signal is output; the code stream is input to the video decoding system 20 and first passes through the decoding unit 201 to obtain decoded transform coefficients; the transform coefficients are processed by the inverse transform and inverse quantization unit 202 to generate residual blocks in the pixel domain; the intra-frame prediction unit 203 can be used to generate prediction data for the current video decoding block based on the determined intra-frame prediction mode and data from the previously decoded blocks of the current frame or picture; the motion compensation unit 204 determines the prediction information for the video decoding block by analyzing the motion vector and other associated syntax elements, and uses The prediction information is used to generate a predictive block for the video decoding block being decoded; a decoded video block is formed by summing the residual block from the inverse transform and inverse quantization unit 202 with the corresponding predictive block generated by the intra-frame prediction unit 203 or the motion compensation unit 204; the decoded video signal passes through the filtering unit 205 to remove blocking artifacts, thereby improving video quality; the decoded video block is then stored in the decoded image buffer unit 206, which stores reference images used for subsequent intra-frame prediction or motion compensation, and is also used for outputting the video signal, thereby obtaining the restored original video signal.
[0089] It should be noted that the filtering method provided in the embodiment of the present application can be applied to the filtering unit 108 shown in FIG6 (indicated by a black bold box), and can also be applied to the filtering unit 205 shown in FIG7 (indicated by a black bold box). In other words, the filtering method in the embodiment of the present application can be applied to both a video encoding system (referred to as "encoder") and a video decoding system (referred to as "decoder"), and can even be applied to both a video encoding system and a video decoding system simultaneously, but this is not limited here.
[0090] The embodiments of the present application can use different input parameters to adjust multiple model inference processes based on a solution in which one model processes images of different frame types, has different quantization parameter configurations, and simultaneously outputs different color component filtering results by the same model, thereby providing the encoder with more selectivity and possibilities and improving encoding performance without increasing the complexity of the model.
[0091] An embodiment of the present application provides a filtering method applied to a decoder. As shown in FIG8 , the method may include:
[0092] S101. Parse a bitstream to determine first syntax element identification information of a component to be filtered of a current block; the first syntax element identification information is used to determine whether each block in a current frame or current slice is filtered based on a preset network model.
[0093] In an embodiment of the present application, a loop filtering method based on a neural network model can be applied. Specifically, a loop filtering method can be obtained based on a neural network model with multiple input parameters.
[0094] On the decoding end, the decoder uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. Simultaneously, the decoder parses the bitstream to obtain quantization parameter information, dequantizes and inversely transforms the quantization parameter information to obtain a residual block. The prediction block and residual block are added together to obtain a reconstructed block, which is then used to form the reconstructed image. The decoder then performs loop filtering on the reconstructed image, either on an image-based or block-based basis, to produce the decoded image.
[0095] It should be noted that, since the original image can be divided into CTUs (coding tree units), or CTUs can be divided into CUs, and CUs can be divided into TUs, etc.; therefore, the filtering method of the embodiment of the present application can be applied not only to CU-level loop filtering (the block partitioning information in this case is CU partitioning information), but also to CTU-level loop filtering (the block partitioning information in this case is CTU partitioning information) or TU-level loop filtering, and the embodiment of the present application is not limited. That is, in the embodiment of the present application, a block can refer to a CTU, or a CU or a TU, and the embodiment of the present application is not limited.
[0096] In an embodiment of the present application, when the decoder performs loop filtering on the reconstructed image of the current frame or current slice, the decoder can first parse the bitstream to obtain the sequence-level enable flag (sps_nnlf_enable_flag), i.e., the fourth syntax element identification information. The sequence-level enable flag is a switch that determines whether the filtering function is enabled for the entire video sequence to be processed. Here, the fourth syntax element identification information can determine whether the filtering function is enabled for the entire video sequence to be processed.
[0097] It should be noted that for the fourth syntax element identification information (sps_nnlf_enable_flag), the fourth syntax element identification information is determined by the value of the sequence-level enable flag obtained by decoding. In some embodiments, the decoder can parse the bitstream to determine the fourth syntax element identification information of the video sequence to be processed. Specifically, the decoder can obtain the value of the fourth syntax element identification information.
[0098] In some embodiments of the present application, if the value of the fourth grammatical element identification information is the first value, it is determined that the fourth grammatical element identification information indicates that the filtering function is turned on for the video sequence to be processed, that is, the sequence level allows the use of the identification bit to represent permission; if the value of the fourth grammatical element identification information is the second value, it is determined that the fourth grammatical element identification information indicates that the filtering function is not turned on for the video sequence to be processed, that is, the sequence level allows the use of the identification bit to represent disallowance.
[0099] In the embodiment of the present application, the first value and the second value are different, and the first value and the second value can be in parameter form or in digital form. Specifically, the fourth syntax element identification information can be a parameter written in the profile or a flag value, which is not specifically limited here.
[0100] For example, taking flag as an example, there are two ways to set the flag: an enable flag (enable_flag) and a disable flag (disable_flag). Assuming that the value of the enable flag is a first value and the value of the disable flag is a second value; then for the first value and the second value, the first value can be set to 1 and the second value can be set to 0; or, the first value can also be set to true (true) and the second value can also be set to false (false); however, this embodiment of the application is not specifically limited to this.
[0101] In this embodiment of the present application, when the fourth syntax element identification information indicates that the sequence including the current frame or current slice allows filtering using the preset network model, that is, when the value of the fourth syntax element identification information is the first value, the bitstream is parsed to determine the first syntax element identification information of the to-be-filtered component of the current block. The fourth syntax element identification information is used to determine whether each block in the current frame or current slice is to be filtered based on the preset network model.
[0102] It should be noted that in the embodiments of the present application, the video sequence to be processed can be first divided into slices to obtain multiple slices, and each slice can be further divided. When the decoder performs decoding at the frame level or the slice level, the decoding method provided in the embodiments of the present application can be used.
[0103] In some embodiments of the present application, when the sequence level allows the use of identification bits to represent permission, the decoder parses the first syntax element identification information of the to-be-filtered component of the current frame or current slice where the current block is located, and obtains the frame-level switch identification bit or slice-level switch identification bit (i.e., the first syntax element identification information) based on a preset network model.
[0104] The frame-level switch flag or the slice-level switch flag indicates whether each block of the current frame or the current slice is filtered based on a preset network model.
[0105] It should be noted that in the embodiment of the present application, the component to be filtered may refer to a color component. The frame-level switch flag or the slice-level switch flag may correspond to each video component. When the video component is a color component, that is, the component to be filtered may be a color component, the frame-level switch flag or the slice-level switch flag may also indicate whether all blocks under the current color component (component to be filtered) are filtered using a neural network-based loop filtering technology.
[0106] The color component may include at least one of the following: a first color component, a second color component, and a third color component. The first color component may be a luminance color component, and the second color component and the third color component may be chrominance color components (for example, the second color component is a blue chrominance color component, and the third color component is a red chrominance color component; or the second color component is a red chrominance color component, and the third color component is a blue chrominance color component).
[0107] Exemplarily, taking the frame-level switch flag as an example, if the component to be filtered is a luminance color component, the first syntax element identification information may be ph_nnlf_luma_ctrl_flag; if the component to be filtered is a chrominance color component, the first syntax element identification information may be ph_nnlf_chroma_ctrl_flag.
[0108] That is, different first syntax element identification information is set for different color components in the current frame or current slice. After parsing the bitstream, the decoder can first determine the first syntax element identification information of the component to be filtered. Only when the first syntax element identification information indicates that the component to be filtered in the current frame or current slice is allowed to be filtered using the preset network model, the decoder needs to decode to obtain the value of the first syntax element identification information.
[0109] In some embodiments of the present application, the decoder parses the bitstream to determine the first syntax element identification information of the component to be filtered of the current frame or current slice, including: parsing the bitstream to obtain the value of the first syntax element identification information of the component to be filtered of the current frame or current slice.
[0110] If the first syntax element identification information (the value of the first syntax element identification information) is the third value, it is determined that the first syntax element identification information indicates that the current frame or current slice is not allowed to be filtered using the preset network model.
[0111] If the first syntax element identification information (the value of the first syntax element identification information) is the fourth value, it is determined that the first syntax element identification information indicates that all to-be-filtered components of all partitioned blocks in the current frame or current slice are allowed to be filtered using the preset network model.
[0112] If the first syntax element identification information (the value of the first syntax element identification information) is the fifth value, it is determined that the first syntax element identification information indicates that at least one partition block in the current frame or current slice allows filtering using the preset network model. The at least one partition block is a portion of the partition blocks in the current frame or current slice.
[0113] In the embodiment of the present application, the value of the first syntax element identification information can be divided into three cases: the first case is that the current frame or current slice does not allow the use of the preset network model for filtering; the second case is that all partitioned blocks in the current frame or current slice allow the use of the preset network model for filtering; and the third case is that some partitioned blocks in the current frame or current slice allow the use of the preset network model for filtering. For the above three different cases, the decoder can use three different values to represent them, each value representing a case.
[0114] The value of the first syntax element identification information in the first case is represented by the third value, the value of the first syntax element identification information in the second case is represented by the fourth value, and the value of the first syntax element identification information in the third case is represented by the fifth value. For example, the third value may be 0, the fourth value may be 1, and the fifth value may be 2. The manner of setting the third to fifth values is not limited in this embodiment of the application.
[0115] It should be noted that when the value of the first syntax element identification information is not the third value, it is determined that the first syntax element identification information indicates that there are partitioned blocks in the current frame or current slice for filtering; that is, for the second and third cases, the filtering situation of the current block can be determined by the second syntax element identification information. In the second case, the second syntax element identification information indicates that all blocks in the current frame or current slice are filtered. In the third case, there are some blocks that use the loop filtering technology based on the neural network, and there are some blocks that do not use the loop filtering technology based on the neural network. It is necessary to further analyze the block-level usage identification bits of all blocks in the current frame to determine (the second syntax element identification information). Among them, the current block is any partitioned block in the current frame or current slice, and the current block or current slice is divided into multiple partitioned blocks.
[0116] It should be noted that, in the embodiment of the present application, the filtering method provided in the embodiment of the present application can be used for different components to be filtered, and the first syntax element identification information corresponding to each filtering component can be determined for different filtering components.
[0117] S102. When the first syntax element identification information indicates that there is a to-be-filtered component of a partitioned block in the current frame or current slice that allows filtering using a preset network model, determine quantization parameter information of the to-be-filtered component and determine second syntax element identification information.
[0118] In an embodiment of the present application, if the first syntax element identification information is not the third value, the first syntax element identification information indicates that there are components to be filtered in the partitioned blocks in the current frame or current slice, which are allowed to be filtered using a preset network model, that is, the first element identification information indicates that all components to be filtered in the partitioned blocks in the current frame or current slice are allowed to be filtered using a preset network model, or indicates that there are at least one partitioned block (at least part of the partitioned blocks) in the current frame or current slice, which are allowed to be filtered using a preset network model.
[0119] At this time, it is necessary to continue parsing the code stream to determine the second syntax element identification information of the component to be filtered of the current block. In addition, the current block here refers to the partition block currently to be loop filtered, which can be any of the partition blocks included in the current frame or current slice.
[0120] In an embodiment of the present application, for any component to be filtered, i.e., a color component, the decoder can decode the first syntax element identification information of each component to be filtered. For a component to be filtered, if the value of the corresponding first syntax element identification information is the third value, it indicates that the current block can be filtered using the preset network model for the component to be filtered. Therefore, the frame-level or slice-level quantization parameter information for the component to be filtered is definitely present in the bitstream. Therefore, when the first syntax element identification information is not the third value, the quantization parameter information of the component to be filtered can be determined from the bitstream.
[0121] In some embodiments of the present application, the process of determining the quantization parameter information of the component to be filtered may include the following two methods:
[0122] (1) Analyze the bitstream and determine the quantization parameter index of the component to be filtered of the current block;
[0123] According to the quantization parameter index, the quantization parameter information of the to-be-filtered component corresponding to the current block is determined from the quantization parameter candidate set.
[0124] It should be noted that in the embodiment of the present application, the first quantization parameter index and the second quantization parameter index are written into the bitstream. In this case, the decoder can obtain the first quantization parameter index by parsing the bitstream; then, based on the first quantization parameter index, it can determine the block quantization parameter value of the first color component (corresponding to the luma color component) from the first quantization parameter candidate set. The decoder can also obtain the second quantization parameter index by parsing the bitstream; then, based on the second quantization parameter index, it can determine the block quantization parameter value of the second color component (corresponding to the chroma color component) from the second quantization parameter candidate set.
[0125] (2) Analyze the bitstream and determine the quantization parameter information of the component to be filtered corresponding to the current block.
[0126] It should be noted that in the embodiment of the present application, the block quantization parameter value of the first color component and the block quantization parameter information of the second color component are written into the bitstream. In this way, the decoder can directly determine the block quantization parameter information of the first color component and the block quantization parameter information of the second color component by parsing the bitstream.
[0127] In an embodiment of the present application, the decoder can parse the quantization parameter index corresponding to each of the different components to be filtered from the bit stream, and then determine the quantization parameter information of the components to be filtered from the quantization parameter candidate set based on the quantization parameter index, that is, obtain the quantization parameter information of the current block under the components to be filtered, or the decoder can directly parse the quantization parameter information of the components to be filtered of the current block or current slice from the bit stream.
[0128] It should be noted that for the current frame or current slice, the quantization parameter information corresponding to different components to be filtered can be different. This is because during encoding, the encoder traverses the quantization parameters for different components to be filtered and performs multiple model inferences to find the quantization parameter information with the lowest rate-distortion cost for each component to be filtered. Therefore, the quantization parameter information corresponding to different components to be filtered can be different.
[0129] Exemplarily, in the embodiment of the present application, the quantization parameter index of the luminance color component can be represented by nnlf_luma_qp_index, and the two chrominance color components can be represented by nnlf_chroma1_qp_index and nnlf_chroma2_qp_index.
[0130] Among them, when the first syntax element identification information indicates that there is a luminance color component of the partitioned block in the current frame or current slice and allows filtering using a preset network model, the quantization parameter information of the luminance color component is determined, which can be expressed as:
[0131] if (ph_nnlf_luma_ctrl_flag!=0) / / Not the first case
[0132] nnlf_luma_qp_index
[0133] When the first syntax element identification information indicates that the chrominance color component of the partitioned block in the current frame or current slice allows filtering using a preset network model, the quantization parameter information of the chrominance color component is determined, which can be expressed as:
[0134] if (ph_nnlf_luma_ctrl_flag!=0) / / Not the first case
[0135] nnlf_chroma1_qp_index
[0136] nnlf_chroma2_qp_index
[0137] It should be noted that, in the embodiment of the present application, the decoder may further determine the value of the second syntax element identification information by determining whether the value of the first syntax element identification information is the fourth value or the fifth value.
[0138] In some embodiments of the present application, when the value of the first syntax element identification information is the fourth value, the first syntax element identification information indicates that the components to be filtered of all divided blocks in the current frame or current slice are allowed to be filtered using a preset network model, the quantization parameter information of the components to be filtered is determined from the bitstream, and the first value is assigned to the second syntax element identification information.
[0139] It should be noted that when the value of the first syntax element identification information is the fourth value, all blocks belonging to the current frame or current slice are filtered using the preset network model. Therefore, the decoder assigns the first value to the second syntax element identification information of each divided block belonging to the current frame or current slice, indicating that all blocks of the current frame or current slice are filtered using the preset network model.
[0140] It should be noted that the decoder parses the first syntax element identification information. If it is the second case mentioned above, then under the component to be filtered, the decoder traverses the various partitioned blocks of the current frame or current slice, and determines the second syntax element identification information corresponding to each partitioned block under the component to be filtered as the first value, that is, it is determined to use the preset network model for filtering at the block level.
[0141] In the embodiment of the present application, the second syntax element identification information is a block-level usage identification bit. Under different components to be filtered, each partition block corresponds to a block-level usage identification bit of a different component to be filtered.
[0142] Exemplarily, in the embodiment of the present application, the block-level usage flag corresponding to the luminance color component is represented as: ctb_nnlf_luma_flag; the block-level usage flag corresponding to the chrominance color component is represented as ctb_nnlf_chroma1_flag and ctb_nnlf_chroma2_flag.
[0143] For each component to be filtered, when the first syntax element identification information indicates the second situation, in addition to determining the quantization parameter information of the component to be filtered based on the bit stream, the decoder can also determine the second syntax element identification information of each partition block under each component to be filtered as the first value.
[0144] Exemplarily, the value of the first syntax element identification information is the fourth value, which is represented as: if (ph_nnlf_luma_ctrl_flag==1).
[0145] The first syntax element identification information indicates that the luminance and color components of all partitioned blocks in the current frame or current slice are allowed to be filtered using a preset network model, and the second syntax element identification information that assigns the first value to the luminance and color components is represented as follows:
[0146]
[0147] The first syntax element identification information indicates that the chrominance color components of all partitioned blocks in the current frame or current slice are allowed to be filtered using a preset network model, and the second syntax element identification information assigning the first value to the chrominance color component is represented as:
[0148]
[0149] In some embodiments of the present application, the value of the first syntax element identification information is the fifth value, the first syntax element identification information indicates that there is at least one partition block in the current frame or the current slice whose to-be-filtered component allows filtering using a preset network model, quantization parameter information of the to-be-filtered component is determined from the bitstream, and the second syntax element identification information is determined from the bitstream; the at least one partition block is a partial partition block in the current frame or the current slice.
[0150] It should be noted that when the value of the first syntax element identification information is the fifth value, there are some partitioned blocks among all the blocks belonging to the current frame or the current slice that are filtered using the preset network model. Therefore, the decoder needs to parse the second syntax element identification information of each partitioned block under the component to be filtered from the bitstream to determine whether the current block or each block is filtered using the preset network model.
[0151] In this embodiment of the present application, the decoder parses the first syntax element identification information. If the third case is described above, the decoder then parses the bitstream to determine the second syntax element identification information corresponding to each partition block of the decoder for the component to be filtered. Based on the value of the second syntax element identification information corresponding to each partition block for each component to be filtered, it is determined whether each partition block is filtered using the preset network model.
[0152] In the embodiment of the present application, the second syntax element identification information is a block-level usage identification bit. Under different components to be filtered, each partition block corresponds to a block-level usage identification bit of a different component to be filtered.
[0153] Exemplarily, in the embodiment of the present application, the block-level usage flag corresponding to the luminance color component is represented as: ctb_nnlf_luma_flag; the block-level usage flag corresponding to the chrominance color component is represented as ctb_nnlf_chroma1_flag and ctb_nnlf_chroma2_flag.
[0154] For each component to be filtered, when the first syntax element identification information indicates the third situation, the decoder can not only determine the quantization parameter information of the component to be filtered based on the bitstream, but also obtain the second syntax element identification information under each component to be filtered by parsing the bitstream.
[0155] In an embodiment of the present application, if the value of the second syntax element identification information of the component to be filtered is a first value, it is determined that the second syntax element identification information indicates that the partition block is filtered using a preset network model under the component to be filtered, that is, the block-level usage identification bit of the component to be filtered is used; if the value of the second syntax element identification information of the component to be filtered is a second value, it is determined that the second syntax element identification information indicates that the partition block is not filtered using the preset network model under the component to be filtered, that is, the block-level usage identification bit of the component to be filtered is not used.
[0156] In the embodiment of the present application, the first value and the second value are different, and the first value and the second value can be in parameter form or in digital form. Specifically, the second syntax element identification information can be a parameter written in the profile or a flag value, which is not specifically limited here.
[0157] For example, taking flag as an example, there are two ways to set the flag: an enable flag (enable_flag) and a disable flag (disable_flag). Assuming that the value of the enable flag is a first value and the value of the disable flag is a second value; then for the first value and the second value, the first value can be set to 1 and the second value can be set to 0; or, the first value can also be set to true (true) and the second value can also be set to false (false); however, this embodiment of the application is not specifically limited to this.
[0158] Exemplarily, the value of the first syntax element identification information is the fifth value expressed as: if (ph_nnlf_luma_ctrl_flag == 2), or expressed as else except the first and second cases.
[0159] Exemplarily, the first syntax element identification information (fifth value) indicates that there is at least one luma and color component of a partition in the current frame or current slice, which allows filtering using a preset network model, and the luma and color components of each partition are determined from the bitstream. The second syntax element identification information is expressed as:
[0160]
[0161] or,
[0162]
[0163] The first syntax element identification information (fifth value) indicates that there is at least one chrominance color component of a partitioned block in the current frame or current slice that allows filtering using a preset network model, and the chrominance color component of each partitioned block is determined from the bitstream. The second syntax element identification information is expressed as:
[0164]
[0165] or,
[0166]
[0167] In some embodiments of the present application, if the first syntax element identification information is a third value, indicating that the current block is not filtered using the preset network model for the component to be filtered, the second syntax element identification information is assigned a second value. The second value of the second syntax element identification information indicates that the current block is not filtered based on the preset network model.
[0168] Exemplarily, the value of the first syntax element identification information is the third value, which is expressed as: if (ph_nnlf_luma_ctrl_flag == 0), or is expressed as else except the third and second cases.
[0169] The first syntax element identification information is the third value, indicating that the current block is not filtered using the preset network model under the brightness and color components. The second syntax element identification information of the brightness and color components is assigned the second value as follows:
[0170]
[0171] or,
[0172]
[0173] The first syntax element identification information is the third value, indicating that the current block is not filtered using the preset network model under the chrominance color component, and the second syntax element identification information of the chrominance color component is assigned the second value as follows:
[0174]
[0175] or,
[0176]
[0177] S103 : Filter the current block of the current frame or current slice based on the second syntax element identification information, the quantization parameter information, and the preset network model to obtain a filtered reconstructed value of the to-be-filtered component of the current block.
[0178] In the embodiment of the present application, the decoder can obtain the reconstructed value of the to-be-filtered component of the current block by parsing the bitstream.
[0179] It should be noted that in the decoding method of the embodiment of the present application, each component to be filtered is decoded separately to obtain the filtered reconstructed value of each component to be filtered. Based on the second syntax element identification information, quantization parameter information, the reconstructed value of the component to be filtered of the current block, and the preset network model for each component to be filtered, the current block of the current frame or current slice is filtered to obtain the filtered reconstructed value of the component to be filtered of the current block.
[0180] In an embodiment of the present application, the preset network model may be a neural network model, and the neural network model includes at least: a convolutional layer, an activation layer, a splicing layer, and a skip connection layer. The convolutional layer may also be replaced with a depthwise separable convolution, which is not limited in the embodiment of the present application.
[0181] In an embodiment of the present application, when the second syntax element identification information indicates that the component to be filtered of the current block is filtered using a preset network model, the reconstructed value of the component to be filtered and the quantization parameter corresponding to the quantization parameter index are input into the preset network model, and the current block of the current frame or current slice is filtered to obtain the filtered reconstructed value of the component to be filtered of the current block.
[0182] It should be noted that when the decoder determines that the current block is filtered using the preset network model under the component to be filtered, the current block can be filtered to obtain a filtered reconstructed value of the component to be filtered of the current block.
[0183] In some embodiments of the present application, filtering is performed on a current block of a current frame based on the second syntax element identification information, the quantization parameter information, and a preset network model to obtain first residual information of a to-be-filtered component of the current block;
[0184] Based on the first residual information and the reconstructed value of the to-be-filtered component of the current block, a filtered reconstructed value of the to-be-filtered component of the current block is determined.
[0185] It should be noted that the first residual information of the component to be filtered is the residual information corresponding to each color component. The decoder determines the reconstructed value of the color component of the current block based on the block-level usage identifier corresponding to each color component. If the block-level usage identifier corresponding to the color component is used, the filtered reconstructed value corresponding to the color component is the sum of the reconstructed value of the color component of the current block and the first residual information of the filtered output under the color component. If the block-level usage identifier corresponding to the color component is not used, the reconstructed value corresponding to the color component is the filtered reconstructed value of the color component of the current block.
[0186] In an embodiment of the present application, the decoder filters the current block of the current frame or current slice based on the quantization parameter information and the preset network model to obtain the reconstructed value of the component to be filtered of the current block before obtaining the first residual information of the component to be filtered of the current block.
[0187] Among them, when the second syntax element identification information corresponding to the color component indicates that the current block is filtered using a preset network model, the decoder uses the preset network model and combines the quantization parameter information to filter the reconstruction value corresponding to the color component of the current block to obtain the first residual information of the current block under the component to be filtered, and then determines the filtered reconstruction value of the color component of the current block based on the first residual information and the reconstruction value of the color component of the current block.
[0188] When the second syntax element identification information corresponding to the color component indicates that the current block is not filtered using the preset network model, the reconstructed value corresponding to the color component is the filtered reconstructed value of the color component of the current block.
[0189] In some embodiments of the present application, if the first syntax element identification information is a third value, indicating that the current block is not filtered using the preset network model for the component to be filtered, the second syntax element identification information is assigned a second value. The second value of the second syntax element identification information indicates that the current block is not filtered based on the preset network model. In this way, after determining the reconstructed value of the component to be filtered of the current block, the decoder directly determines the reconstructed value of the component to be filtered of the current block as the filtered reconstructed value of the component to be filtered of the current block.
[0190] It is understandable that during the decoding process, the decoder can determine the post-filtering reconstruction value corresponding to each component to be filtered, based on the quantization parameter information corresponding to each component to be filtered. In other words, the quantization parameter information corresponding to different components to be filtered in the current frame or current slice can be different (and is the optimal quantization parameter information of each component to be filtered obtained by the encoder through multiple model reasoning during encoding), and each component to be filtered can be subjected to respective filtering processes to obtain the post-filtering reconstruction value of each component to be filtered. Therefore, each component to be filtered in the current block can use its own quantization parameter information to implement filtering of each component to be filtered in the same preset network model, thereby improving the compression performance of each component to be filtered, and making the selection of input information (quantization parameter information) for filtering different components to be filtered more flexible while ensuring that the complexity of the model is not increased, thereby improving decoding efficiency.
[0191] In some embodiments of the present application, new syntax elements are introduced, such as first syntax element identification information and second syntax element identification information of the component to be filtered. In some embodiments, the component to be filtered includes at least a luminance color component and a chrominance color component. It should be noted that in embodiments of the present application, third syntax element identification information of the component to be filtered may also be introduced. The third syntax element identification information is an allowable use flag of the component to be filtered of the current frame or current slice. The third syntax element identification information of the component to be filtered of the current block is a block-level allowable use flag of the component to be filtered of the current block.
[0192] In an embodiment of the present application, if the value of the third syntax element identification information is a first value, it is determined that the third syntax element identification information indicates that the to-be-filtered component of the current block included in the current frame or current slice is allowed to be filtered using a preset network model; if the value of the third syntax element identification information is a second value, it is determined that the third syntax element identification information indicates that the to-be-filtered component of the current block included in the current frame or current slice may not be filtered using a preset network model.
[0193] In this embodiment of the present application, the first value and the second value are different, and the first value and the second value can be in parameter form or in numerical form. The third syntax element identification information can be a parameter written in the profile or a flag value, which is not specifically limited here.
[0194] In an embodiment of the present application, when the first syntax element identification information indicates that there are to-be-filtered components of the partitioned blocks in the current frame or the current slice, which allow filtering using a preset network model, the decoder may first parse the third syntax element identification information and then parse the second syntax element identification information; or it may first parse the second syntax element identification information and then parse the third syntax element identification information, which is not limited in the embodiment of the present application.
[0195] In an embodiment of the present application, the process of the decoder determining the filtered reconstructed value of the to-be-filtered component of the current block may also include: when the first syntax element identification information indicates that the to-be-filtered component of the divided block in the current frame or current slice is allowed to be filtered using a preset network model, and the third syntax element identification information indicates that the to-be-filtered component of the current block included in the current frame or current slice is allowed to be filtered using a preset network model, the current block of the current frame or current slice is filtered based on the quantization parameter information and the preset network model to obtain the filtered reconstructed value of the to-be-filtered component of the current block.
[0196] In an embodiment of the present application, the process of the decoder determining the filtered reconstructed value of the component to be filtered of the current block may also include: when the third syntax element identification information indicates that the component to be filtered of the current block included in the current frame or current slice is allowed to be filtered using a preset network model, and the second syntax element identification information indicates that the component to be filtered of the current block is filtered using a preset network model, filtering the current block of the current frame or current slice based on the quantization parameter information and the preset network model to obtain the filtered reconstructed value of the component to be filtered of the current block.
[0197] It should be noted that, when the second syntax element identification information indicates that the to-be-filtered component of the current block is not filtered using the preset network model, the third syntax element identification information may not be parsed.
[0198] It should be noted that, when the third syntax element identification information is the first value, it can be directly determined to filter the current block without parsing the second syntax element identification information.
[0199] In an embodiment of the present application, the process of the decoder determining the filtered reconstructed value of the to-be-filtered component of the current block may also include: when the first syntax element identification information indicates that the to-be-filtered component of the divided block in the current frame or current slice is allowed to be filtered using a preset network model, and the third syntax element identification information indicates that the to-be-filtered component of the current block included in the current frame or current slice is allowed to be filtered using a preset network model, the current block of the current frame or current slice is filtered based on the quantization parameter information and the preset network model to obtain the filtered reconstructed value of the to-be-filtered component of the current block.
[0200] It should be noted that the inference time of neural network-based loop filter models is relatively long, several times or even dozens of times longer than that of traditional codecs. Given the limitations of current hardware, limiting the inference time on the decoder side can significantly reduce decoding time by limiting the scope of use of the current frame or slice, such as using coding tree units or partition blocks that offer significant performance improvements. This may require the introduction of additional block-level syntax element identifiers, such as the third syntax element identifier.
[0201] In some embodiments of the present application, the decoder filters the current block of the current frame or current slice based on the second syntax element identification information, the quantization parameter index and the preset network model, and obtains the predicted value of the component to be filtered of the current block, at least one of the block partitioning information and the deblocking filter boundary strength, and the reconstructed value of the component to be filtered of the current block before obtaining the filtered reconstructed value of the component to be filtered of the current block.
[0202] In some embodiments of the present application, the decoder uses a preset network model, combined with the predicted value of the component to be filtered of the current block, block partition information and at least one of the deblocking filter boundary strength, and quantization parameter information, to filter the reconstructed value of the component to be filtered of the current block to obtain the filtered reconstructed value of the component to be filtered of the current block.
[0203] That is to say, the decoder obtains at least one of the predicted value of the component to be filtered of the current block, the block division information of the component to be filtered, and the deblocking filter boundary strength, as well as the reconstructed value of the component to be filtered of the current block. It should be noted that in the filtering process, the input parameters input to the network filtering model may include: under each component to be filtered, the predicted value of the current block, the block division information, the deblocking filter boundary strength, the reconstructed value of the current block, and the quantization parameter information in this application. This application does not limit the type of information of the input parameters when filtering each component to be filtered. However, the predicted value of the current block, the block division information, and the deblocking filter boundary strength of each component to be filtered are not necessarily required every time, and need to be determined according to the actual situation.
[0204] It is understandable that the input parameters of the decoder when filtering can be different, thereby increasing the diversity of operations during filtering. For the preset network model, its input may include: the reconstructed value of the component to be filtered (represented by rec_yuv), the quantization parameter information of the luminance color component or the quantization parameter information of the chrominance color component; its output may be: the reconstructed value of the component to be filtered (represented by output_yuv). Since the embodiment of the present application removes non-important input elements such as predicted YUV information and YUV information with partitioning information, the computational complexity of network model reasoning can be reduced, which is beneficial to the implementation of the decoding end and reduces the decoding time. In addition, in the embodiment of the present application, the input of the preset network model may also include the quantization parameter information of the luminance color component, which may include SliceQP or BaseQP; the quantization parameter information of the chrominance color component may include SliceQP or BaseQP, which is determined by the parsed code stream.
[0205] In the embodiment of the present application, through the above implementation, the filtered reconstructed values of each to-be-filtered component of the current block can be determined based on the quantization parameter information corresponding to each to-be-filtered component. Specifically, the steps of parsing the bitstream and determining the filtered reconstructed values of each to-be-filtered component of the current block are repeated for each to-be-filtered component, thereby obtaining the filtered reconstructed values of each to-be-filtered component. Each to-be-filtered component corresponds to its own quantization parameter information. In this way, the filtered reconstructed value of the current block can be determined based on the filtered reconstructed values of each to-be-filtered component of the current block.
[0206] It should be noted that the filtered reconstructed value of the current block may be determined by merging different filtered components of the filtered reconstructed values of the respective filtered components of the current block.
[0207] In some embodiments of the present application, when the decoder determines the post-filtering reconstruction values of each component to be filtered of the current block, the decoder traverses each partitioned block in the current frame or current slice, takes each partitioned block as the current block in turn, and repeatedly performs the steps of parsing the code stream and determining the post-filtering reconstruction values of the components to be filtered of the current block to obtain the post-filtering reconstruction values of each component to be filtered corresponding to each partitioned block; according to the post-filtering reconstruction values of each component to be filtered corresponding to each partitioned block, the post-filtering reconstruction value of each partitioned block is determined, and thus the reconstructed image of the current frame or current slice is determined based on the post-filtering reconstruction values of each partitioned block.
[0208] It should be noted that for the current frame or current slice, the current frame or current slice may include many partitioned blocks. These partitioned blocks are then traversed, and each partitioned block is used as the current block in turn. The decoding method process of the embodiment of the present application is repeatedly executed to obtain the filtered reconstruction value corresponding to each partitioned block; based on these obtained filtered reconstruction values, the reconstructed image of the current frame can be determined. In addition, it should be noted that the decoder can also continue to traverse other loop filtering tools and output a complete reconstructed image after completion. The specific process is not closely related to the embodiment of the present application and is therefore not described in detail here.
[0209] In some embodiments of the present application, when the first syntax element identification information indicates that there is a component to be filtered in a divided block in the current frame or current slice that allows filtering using a preset network model, and the current frame or current slice is a first type frame, the quantization parameter information of the component to be filtered and the second syntax element identification information are determined.
[0210] It should be noted that the first type of frame may include an I frame or a bottom layer B frame.
[0211] In some embodiments, due to the relatively long inference time of the neural network-based loop filter model, which can be several times or even dozens of times longer than the runtime of traditional codecs, this technique is not used for decoding inference time due to the limited support of current hardware. This not only reduces decoding time but also saves bits of quantization parameter transmission for I and B frames, further improving compression efficiency.
[0212] An embodiment of the present application provides a filtering method, which is applied to an encoder. As shown in FIG9 , the method may include:
[0213] S201. Determine a first rate-distortion cost of a current frame or a current slice. The first rate-distortion cost is obtained by filtering all to-be-filtered components of all divided blocks included in the current frame or the current slice without using a preset network model.
[0214] In an embodiment of the present application, the encoder traverses intra-frame or inter-frame prediction to obtain a prediction block for each block. The residual of each block can be obtained by subtracting the original image block from the prediction block. The residual obtains a frequency domain residual coefficient by various transformation modes, which is then quantized and inversely quantized. After inverse transformation, the distortion residual information is obtained. The distortion residual information is superimposed on the prediction block to obtain a reconstructed block of the current block. After the image is encoded, the loop filter module filters the image at the block level (for example, the coding tree unit level) as the basic unit. The embodiment of the present application can describe the coding tree unit as a block, but the block is not limited to CTU, but can also be CU or TU, which is not limited in the embodiment of the present application. The encoder obtains the fourth syntax element identification information based on the preset network model, that is, the sequence level allows the use of the flag bit, that is, sps_nnlf_enable_flag. If the sequence level allows the use of the flag bit to be allowed, the use of the neural network-based loop filtering technology is allowed; if the sequence level allows the use of the flag bit to be not allowed, the use of the neural network-based loop filtering technology is not allowed. The sequence level allows the use of flags that need to be written into the bitstream when encoding a video sequence.
[0215] In an embodiment of the present application, the fourth syntax element identification information is determined, and the sequence level allows the use of the identification bit; when the sequence level allows the use of the identification bit to represent permission, the encoder attempts the loop filtering technology based on the preset network model, and the encoder obtains the original value of the current block in the current frame or current slice, the reconstructed value and quantization parameter information of the current block under each component to be filtered; if the sequence level based on the preset network model allows the use of the identification bit as not allowed, the encoder does not try the network-based loop filtering technology; continue to try other loop filtering tools, such as LF filtering, and output a complete reconstructed image after completion.
[0216] In an embodiment of the present application, when the obtained fourth syntax element identification information represents that the sequence containing the current frame or current slice allows filtering using a preset network model, the original values of each partition block of the component to be filtered in the current frame or current slice, the reconstructed values of each partition block and the original quantization parameter information are obtained.
[0217] In some embodiments of the present application, the components to be filtered of all the partitioned blocks included in the current frame or current slice are not filtered using a preset network model, and a first model inference is performed (i.e., the first case). In this way, the encoder performs rate-distortion cost calculation based on the original value of each partitioned block and the reconstructed value of each partitioned block, determines the rate-distortion cost of each component to be filtered in the current frame or current slice, and determines the first rate-distortion cost (frame-level cost or slice-level cost) of the current frame or current slice based on the sum of the rate-distortion costs of each component to be filtered.
[0218] It should be noted that in the embodiments of the present application, the component to be filtered may refer to a color component. When the video component is a color component, that is, the component to be filtered may be a color component, and the frame-level switch flag or the slice-level switch flag may also indicate whether all blocks under the current color component (component to be filtered) are filtered using the neural network-based loop filtering technology.
[0219] The color component may include at least one of the following: a first color component, a second color component, and a third color component. The first color component may be a luminance color component, and the second color component and the third color component may be chrominance color components (for example, the second color component is a blue chrominance color component, and the third color component is a red chrominance color component; or the second color component is a red chrominance color component, and the third color component is a blue chrominance color component).
[0220] In an embodiment of the present application, the encoder performs rate-distortion cost calculation based on the original value of each partition block and the reconstructed value of each partition block under the luminance color component to determine the frame-level rate-distortion cost or slice-level rate-distortion cost of the luminance color component of the current frame or current slice, and performs rate-distortion cost calculation based on the original value of each partition block and the reconstructed value of each partition block under the two chrominance color components to determine the frame-level rate-distortion cost or slice-level rate-distortion cost of the two chrominance color components of the current frame or current slice. Based on the frame-level rate-distortion cost or slice-level rate-distortion cost of the luminance color component and the frame-level rate-distortion cost or slice-level rate-distortion cost of the two chrominance color components, a first rate-distortion cost of the current frame or current slice is determined.
[0221] It should be noted that for the fourth syntax element identification information (sps_nnlf_enable_flag), the fourth syntax element identification information is determined by the value of the sequence-level enable flag obtained by decoding. In some embodiments, the encoder can obtain the value of the fourth syntax element identification information and write the fourth syntax element identification information into the bitstream.
[0222] In some embodiments of the present application, if the value of the fourth grammatical element identification information is the first value, it is determined that the fourth grammatical element identification information indicates that the filtering function is turned on for the video sequence to be processed, that is, the sequence level allows the use of the identification bit to represent permission; if the value of the fourth grammatical element identification information is the second value, it is determined that the fourth grammatical element identification information indicates that the filtering function is not turned on for the video sequence to be processed, that is, the sequence level allows the use of the identification bit to represent disallowance.
[0223] In the embodiment of the present application, the first value and the second value are different, and the first value and the second value can be in parameter form or in digital form. Specifically, the fourth syntax element identification information can be a parameter written in the profile or a flag value, which is not specifically limited here.
[0224] For example, taking flag as an example, there are two ways to set the flag: an enable flag (enable_flag) and a disable flag (disable_flag). Assuming that the value of the enable flag is a first value and the value of the disable flag is a second value; then for the first value and the second value, the first value can be set to 1 and the second value can be set to 0; or, the first value can also be set to true (true) and the second value can also be set to false (false); however, this embodiment of the application is not specifically limited to this.
[0225] In this embodiment of the present application, when the fourth syntax element identification information indicates that the sequence including the current frame or current slice allows filtering using the preset network model, that is, when the value of the fourth syntax element identification information is the first value, the bitstream is parsed to determine the first syntax element identification information of the to-be-filtered component of the current block. The fourth syntax element identification information is used to determine whether each block in the current frame or current slice is to be filtered based on the preset network model.
[0226] It should be noted that in the embodiments of the present application, the video sequence to be processed can be first divided into slices to obtain multiple slices, and each slice can be further divided. When the encoder encodes at the frame level or the slice level, the encoding method provided in the embodiments of the present application can be used.
[0227] Exemplarily, costOrg is used to represent the first rate-distortion cost obtained by filtering all the to-be-filtered components of all the partitioned blocks included in the current frame or current slice without using the preset network model. It should be noted that costOrg can be obtained by summing the rate-distortion costs of different to-be-filtered components.
[0228] In some embodiments of the present application, when the fourth syntax element identification information represents that a sequence containing the current frame or the current slice is allowed to be filtered using a preset network model, and the current frame or the current slice is a first type frame, the first rate-distortion cost, the second rate-distortion cost and the third rate-distortion cost in the above-mentioned encoding method are determined.
[0229] It should be noted that the first type of frame may include an I frame or a bottom layer B frame.
[0230] In some embodiments, due to the relatively long inference time of the neural network-based loop filter model, which can be several times or even dozens of times longer than the runtime of traditional codecs, this technique is not used for encoding inference time due to the limited support of current hardware. This not only reduces encoding time but also saves bits in transmitting quantization parameters for I and B frames, further improving compression efficiency.
[0231] S202 , determining a second rate-distortion cost of the current frame or current slice; the second rate-distortion cost is obtained by filtering all to-be-filtered components of all divided blocks included in the current frame or current slice using a preset network model.
[0232] In an embodiment of the present application, all components to be filtered of all divided blocks included in the current frame or current slice are filtered using a preset network model, and a second model inference (i.e., the second case) is performed to determine the second rate-distortion cost of the current frame or current slice.
[0233] In some embodiments of the present application, the encoder determines at least two types of original quantization parameter information in the original quantization parameter information; under each type of original quantization parameter information, filters the reconstructed values of the components to be filtered of each partitioned block included in the current frame or current slice and each type of original quantization parameter information based on a preset network model to obtain the original filtered reconstructed values of the components to be filtered of each partitioned block included in the current frame or current slice; performs rate-distortion cost calculation based on the original values of the components to be filtered of each partitioned block included in the current frame or current slice and the original filtered reconstructed values of the components to be filtered of each partitioned block to obtain a fourth rate-distortion cost of the components to be filtered under each type of original quantization parameter information; determines a third minimum rate-distortion cost having the minimum rate-distortion cost from the at least two obtained fourth rate-distortion costs; and determines a second rate-distortion cost based on the third minimum rate-distortion cost of each component to be filtered.
[0234] In an embodiment of the present application, the original quantization parameter information may include at least two types, such as SliceQP or BaseQP, or other forms of quantization parameters. During the second model inference process, the encoder may perform at least two model inference passes for each component to be filtered, i.e., each color component, using each original quantization parameter, and obtain the fourth rate-distortion cost corresponding to each original quantization parameter. Thus, for each component to be filtered (color component), at least two fourth rate-distortion costs corresponding to at least two types of original quantization parameter information may be obtained during the second model inference process. For each component to be filtered (color component), a third minimum rate-distortion cost with the minimum rate-distortion cost is determined from the at least two fourth rate-distortion costs, thereby obtaining the third minimum rate-distortion cost for each component to be filtered. The third minimum rate-distortion costs for different components to be filtered are summed to obtain the second rate-distortion cost. For different components to be filtered, the original quantization parameters corresponding to their second rate-distortion costs are different. That is, during the second inference process, the optimal original quantization parameters for filtering the preset network model may be different for different components to be filtered.
[0235] It should be noted that the second rate-distortion cost is a frame-level cost or a slice-level cost, which is the frame-level cost or the slice-level cost obtained by summing the third minimum rate-distortion costs of the components to be filtered of the current frame or current slice determined by the encoder.
[0236] For example, for a luminance color component, such as the Y component, the rate-distortion costs corresponding to the two original quantization parameters are represented as costYcand1 and costYcand2, respectively. The smallest of costYcand1 and costYcand2 is determined as the third minimum rate-distortion cost for the Y component. The second rate-distortion cost can be represented as costFrameBest.
[0237] S203. Determine a third rate-distortion cost of the current frame or current slice; the third rate-distortion cost is obtained by allowing a to-be-filtered component of at least one partition block in the current frame or current slice to be filtered using a preset network model; the at least one partition block is a partial partition block in the current frame or current slice.
[0238] In an embodiment of the present application, the encoder may determine the first sub-rate-distortion cost of the component to be filtered corresponding to each divided block when the first rate-distortion cost is obtained; determine the second sub-rate-distortion cost of the component to be filtered corresponding to each divided block when the second rate-distortion cost is obtained; determine the first minimum rate-distortion cost of the component to be filtered of the current frame or current slice based on the first sub-rate-distortion cost and the second sub-rate-distortion cost; and determine the third rate-distortion cost (i.e., the third rate-distortion cost in the third case) based on the first minimum rate-distortion cost of each component to be filtered.
[0239] In an embodiment of the present application, the encoder obtains the block-level rate-distortion cost of each block under different components to be filtered during the first model inference, that is, the first sub-rate-distortion cost; and the encoder obtains the block-level rate-distortion cost of each block under different components to be filtered at the optimal original quantization parameter information during the second model inference, that is, the second sub-rate-distortion cost. For the third case, for each block under different components to be filtered, the corresponding first sub-rate-distortion cost and the second sub-rate-distortion cost are compared to determine the first minimum rate-distortion cost with the minimum rate-distortion cost under each component to be filtered, the block-level rate-distortion cost. In this way, the first minimum rate-distortion cost with the minimum rate-distortion cost in the first model inference and the second model inference for each block under the component to be filtered is obtained. Then, the first minimum rate-distortion cost of each block under different components to be filtered is added or summed to obtain the third rate-distortion cost of the current frame or current slice.
[0240] In the embodiment of the present application, the encoder traverses each component to be filtered, and determines the sum of the first minimum rate-distortion cost of each divided block as the third rate-distortion cost of the current frame or current slice.
[0241] For example, during the first model inference process, the block-level rate-distortion cost for the component to be filtered, i.e., the first sub-rate-distortion cost, can be expressed as costCTUrec. For example, costCTUrec for the luma color component and costCTUrec for the chroma color component. During the second model inference process, the block-level rate-distortion cost for the component to be filtered, i.e., the second sub-rate-distortion cost, can be expressed as costCTUcnn. For example, costCTUcnn for the luma color component and costCTUcnn for the chroma color component.
[0242] In an embodiment of the present application, for each component to be filtered, if the first sub-rate distortion cost corresponding to any partition block is less than the second sub-rate distortion cost, the second syntax element identification information of the component to be filtered of any partition block is determined to be a second value; the second value indicates that the partition block is not allowed to be filtered using a preset network model; if the first sub-rate distortion cost corresponding to any partition block is greater than or equal to the second sub-rate distortion cost, the second syntax element identification information of the component to be filtered of any partition block is determined to be a first value; the first value indicates that the partition block is allowed to be filtered using a preset network model.
[0243] In an embodiment of the present application, for each component to be filtered, if the first sub-rate distortion cost corresponding to any partitioned block is less than the second sub-rate distortion cost, it indicates that the coding performance of the block is better when not filtered than when filtered. Therefore, it is determined that the block does not use the preset network model for filtering under the component to be filtered, and the block-level usage flag is set to the second value, that is, false or no. If the first sub-rate distortion cost corresponding to any partitioned block is greater than or equal to the second sub-rate distortion cost, it indicates that the coding performance of the block is better when filtered than when not filtered. Therefore, it is determined that the block uses the preset network model for filtering under the component to be filtered, and the block-level usage flag is set to the first value, that is, true. Each partitioned block determines the block-level usage flag of each block according to the above principle.
[0244] Exemplarily, after the encoder determines the first minimum rate-distortion cost of each block of each component to be filtered, it can traverse the first minimum rate-distortion cost of each component to be filtered and each block to obtain a third rate-distortion cost (frame-level rate-distortion cost or slice-level rate-distortion cost). The third rate-distortion cost can be expressed as costCTUBest.
[0245] It should be noted that the third rate-distortion cost determined in the third case is obtained when some blocks are filtered and some blocks are not filtered. That is, the third case is when some divided blocks in the current frame or current slice are filtered.
[0246] For example, as shown in Figures 10A and 10B, quantization parameters QP1 and QP2 corresponding to QP index 1 and QP index 2 are used to traverse the Y, U, and V components of the current block, obtaining first sub-rate-distortion costs CostY1, CostU1, and CostV1 corresponding to Y1, U1, and V1, and second sub-rate-distortion costs CostY2, CostU2, and CostV2 corresponding to Y2, U2, and V2. By comparing the first sub-rate-distortion costs with the second sub-rate-distortion costs, it is determined that CostY1 is less than CostY2, CostU1 is greater than CostU2, and CostV1 is less than CostV2. The resulting quantization parameter information corresponding to each color component of the current block is the quantization parameter information corresponding to CostY1, CostU2, and CostV1 obtained in the case shown in Figure 10C.
[0247] S204: Determine first syntax element identification information of the to-be-filtered component of the current frame or current slice according to the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost.
[0248] In an embodiment of the present application, the encoder determines a second minimum rate-distortion cost with the minimum rate distortion based on the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost. Based on the second minimum rate-distortion cost, the first syntax element identification information of the current frame or current slice can be determined, and then the first syntax element identification information of each component to be filtered is consistent with the value of the first syntax element identification information of the current frame or current slice, thereby obtaining the first syntax element identification information of the component to be filtered of the current frame or current slice.
[0249] In the embodiment of the present application, after the encoder determines the first syntax element identification information of the component to be filtered, it will eventually write the first syntax element identification information of the component to be filtered into the bitstream and transmit it to the decoder for decoding.
[0250] In this embodiment of the present application, the process of the encoder determining the first syntax element identification information of the component to be filtered of the current frame or current slice according to the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost includes:
[0251] Determining a second minimum rate-distortion cost having the minimum rate-distortion cost from the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost;
[0252] If the second minimum rate-distortion cost is the first rate-distortion cost, determining that the first syntax element identification information of the to-be-filtered component of the current frame or current slice is a third value;
[0253] If the second minimum rate-distortion cost is the second rate-distortion cost, determining that the first syntax element identification information of the to-be-filtered component of the current frame or current slice is a fourth value;
[0254] If the second minimum rate-distortion cost is the third rate-distortion cost, the first syntax element identification information of the to-be-filtered component of the current frame or current slice is determined to be a fifth value.
[0255] It should be noted that, in the first, second, and third cases mentioned above, the encoder determines which case has the smallest frame-level rate-distortion cost or slice-level rate-distortion cost, and then determines the value of the first syntax element identification information of the component to be filtered of the current frame or current slice.
[0256] In the embodiment of the present application, the value of the first syntax element identification information can be divided into three cases: the first case is that the current frame or current slice does not allow the use of the preset network model for filtering; the second case is that all partitioned blocks in the current frame or current slice allow the use of the preset network model for filtering; and the third case is that some partitioned blocks in the current frame or current slice allow the use of the preset network model for filtering. For the above three different cases, the encoder can use three different values to represent them, each value representing a different case.
[0257] In the embodiment of the present application, the first case is represented by the third value, the second case is represented by the fourth value, and the third case is represented by the fifth value. For example, the third value can be 0, the fourth value can be 1, and the fifth value can be 2. The configuration of the third to fifth values is not limited in the embodiment of the present application.
[0258] In some embodiments of the present application, when the encoder determines that the first syntax element identification information of the to-be-filtered component of the current frame or current slice is the fifth value, the encoder writes the second syntax element identification information of each partition block into the bitstream.
[0259] It should be noted that when the encoder determines that the first syntax element identification information of the to-be-filtered component of the current frame or current slice is the fifth value, that is, the encoder determines that some of the partitioned blocks in the current frame or current slice are allowed to be filtered using a preset network model, then it is necessary to write the second syntax element identification information corresponding to each partitioned block in the current frame or current slice, that is, the block-level usage identification bit into the bitstream and transmit it to the decoder for use in decoding.
[0260] In some embodiments of the present application, if the second minimum rate-distortion cost is the second rate-distortion cost or the third rate-distortion cost, an original quantization parameter information corresponding to the second rate-distortion cost or the third rate-distortion cost is determined as the quantization parameter information of the component to be filtered, and the quantization parameter information is written into the bitstream; alternatively, a quantization parameter index corresponding to the quantization parameter information is written into the bitstream.
[0261] It should be noted that in the second and third cases, the encoder needs to use the original quantization parameter information corresponding to each to-be-filtered component when the second rate-distortion cost or the third rate-distortion cost is obtained as the input for filtering, that is, to determine the quantization parameter information of the to-be-filtered component, and needs to write the quantization parameter information into the bitstream and transmit it to the decoder for use during decoding. Alternatively, the encoder writes the index corresponding to the quantization parameter information of the to-be-filtered component into the bitstream and transmits it to the decoder for use during decoding, which is not limited in this embodiment of the present application.
[0262] At the same time, in an embodiment of the present application, if the second minimum rate-distortion cost is the second rate-distortion cost or the third rate-distortion cost, the original filtered reconstruction values of each partitioned block corresponding to the second rate-distortion cost or the third rate-distortion cost are determined as the filtered reconstruction values of the to-be-filtered components of each partitioned block, and the filtered reconstruction values are written into the bitstream.
[0263] It is understood that during the encoding process, the encoder determines the encoding method that minimizes the frame-level rate-distortion cost or the slice-level rate-distortion cost by performing multiple model inferences based on the following scenarios: no filtering of each partitioned block, filtering of each partitioned block, and partial filtering of each partitioned block. The quantization parameter information corresponding to each component to be filtered can be different types of quantization parameter information under different circumstances. In other words, the quantization parameter information corresponding to different components to be filtered in the current frame or current slice can be different (the optimal quantization parameter information for each component to be filtered obtained by the encoder through multiple model inferences during encoding), and each component to be filtered can be filtered to obtain a filtered reconstructed value. Therefore, each component to be filtered in the current block can use its own quantization parameter information to implement filtering of each component to be filtered in the same preset network model, thereby improving the compression performance of each component to be filtered. While ensuring that the complexity of the model is not increased, the selection of input information (quantization parameter information) for filtering different components to be filtered is more flexible, thereby improving encoding efficiency.
[0264] In some embodiments of the present application, when the acquired fourth syntax element identification information indicates that a sequence including the current frame or the current slice allows filtering using a preset network model, obtaining predicted values of each partitioned block of a component to be filtered in the current frame or the current slice, at least one of block partitioning information and a deblocking filter boundary strength, and reconstructed values and original quantization parameter information of each partitioned block of the component to be filtered in the current frame or the current slice; and determining at least two types of original quantization parameter information in the original quantization parameter information;
[0265] Under each type of original quantization parameter information, input each partition prediction value, block partition information, and at least one of a deblocking filter boundary strength, each partition reconstruction value, and each type of original quantization parameter information into a preset network model for filtering, thereby determining an original filtered reconstruction value of each partition block of the component to be filtered;
[0266] Performing rate-distortion cost calculation based on the original values of the components to be filtered of each divided block included in the current frame and the original filtered reconstructed values of the components to be filtered of each divided block to obtain a fourth rate-distortion cost of the components to be filtered under each original quantization parameter information;
[0267] determining a third minimum rate-distortion cost having the minimum rate-distortion cost from the obtained at least two fourth rate-distortion costs;
[0268] A second rate-distortion cost is determined based on the third minimum rate-distortion cost of each to-be-filtered component.
[0269] It should be noted that this implementation process is consistent with the implementation principle in S202. The difference is that the input parameters during filtering are diversified. It combines the predicted values of each partitioned block of the component to be filtered, block partitioning information and at least one of the deblocking filter boundary strengths, and the reconstructed values and original quantization parameter information of each partitioned block of the component to be filtered in the current frame or current slice.
[0270] It is understandable that the input parameters of the encoder when filtering can be different, thereby increasing the diversity of operations during filtering. For the preset network model, its input may include: the reconstructed value of the component to be filtered (represented by rec_yuv), the quantization parameter information of the luminance color component or the quantization parameter information of the chrominance color component; its output may be: the reconstructed value of the component to be filtered after filtering (represented by output_yuv). Since the embodiment of the present application removes non-important input elements such as predicted YUV information and YUV information with partitioning information, the amount of computation of network model reasoning can be reduced, which is beneficial to the implementation of the encoding end and reduces the encoding time.
[0271] In some embodiments of the present application, the encoder may further determine third syntax element identification information; the third syntax element identification information is used to indicate that the to-be-filtered component of the current block is allowed to be filtered using a preset network model; and the third syntax element identification information is written into the bitstream.
[0272] The third syntax element identification information of the to-be-filtered component of the previous block is the block-level usage permission flag of the to-be-filtered component of the current block.
[0273] In an embodiment of the present application, if the value of the third syntax element identification information is a first value, it is determined that the third syntax element identification information indicates that the to-be-filtered component of the current block included in the current frame or current slice is allowed to be filtered using a preset network model; if the value of the third syntax element identification information is a second value, it is determined that the third syntax element identification information indicates that the to-be-filtered component of the current block included in the current frame or current slice may not be filtered using a preset network model.
[0274] In this embodiment of the present application, the first value and the second value are different, and the first value and the second value can be in parameter form or in numerical form. The third syntax element identification information can be a parameter written in the profile or a flag value, which is not specifically limited here.
[0275] It should be noted that when the second syntax element identification information is the first value, the third syntax element identification information is the first value. When the second syntax element identification information is the second value, the third syntax element identification information may be the first value, and may be the second value. When the third syntax element identification information is the second value, the second syntax element identification information is the second value.
[0276] It should be noted that the inference time of neural network-based loop filter models is relatively long, several times or even dozens of times longer than the runtime of traditional codecs. Given the limitations of current hardware, limiting the inference time on the encoder side can significantly reduce encoding time by limiting the scope of use of the current frame or slice, such as using coding tree units or partition blocks that offer significant performance improvements. This may require the introduction of additional block-level syntax element identifiers, such as the third syntax element identifier.
[0277] In another embodiment of the present application, based on the decoding method and encoding method described in the aforementioned embodiment, taking the three YUV color components as an example, the embodiment of the present application proposes a scheme of multiple inferences with the same model to adapt to the requirements of different color components for quantization parameters (quantization parameter information). Taking the preset network model as shown in Figure 5A as an example, for a quantization parameter BaseQP or sliceQP, the neural network filter model outputs the filtering results of the three YUV color components respectively during the inference stage. For another different quantization parameter (the other of BaseQP or sliceQP), the preset network model again infers the current image to obtain another set of filtering results of the three YUV color components, and compares the filtering results of the three YUV components with different quantization parameters through rate-distortion cost, and takes the best result between each color component as the output of the neural network filter model.
[0278] The encoder can select multiple quantization parameters. The coding tree unit or coding unit (the basic unit of the neural network loop filter, i.e., the partition block) infers and calculates the filtered reconstructed sample block based on each quantization parameter. After all quantization parameters are inferred, the optimal quantization parameter for the current color component is selected based on the rate-distortion cost at the coding tree unit (CTB) level, frame level, or slice level. The quantization parameter is written into the bitstream for transmission to the decoder through quantization parameter indexing or direct binarization.
[0279] The decoding end parses the quantization parameter information of the current coding tree unit level (CTB level), frame level (frame level) or slice level (slice level), and uses the neural network-based loop filtering technology to filter according to all the quantization parameters obtained by the analysis to obtain corresponding filtering results. According to the quantization parameter information identified by each color component, the corresponding result is used to replace the reconstructed sample as the current neural network-based loop filtering technology output sample.
[0280] For example, in a specific embodiment, for the encoding end, the specific process is as follows:
[0281] The encoder performs intra-frame or inter-frame prediction to obtain a prediction block for each coding unit. The residual for each coding unit is obtained by subtracting the original image block from the prediction block. This residual is transformed using various transform modes to obtain frequency-domain residual coefficients. This residual is then quantized and inverse-quantized, and inversely transformed to obtain distortion residual information. This distortion residual information is then superimposed with the prediction block to produce the reconstructed block. After the image is encoded, the loop filter module filters the encoded image at the coding tree unit level. This is where the technology proposed in this article is applied. The sequence-level enable flag (sps_nnlf_enable_flag) is used to enable the use of neural network-based loop filtering. If this flag is true, neural network-based loop filtering is enabled; if it is false, neural network-based loop filtering is disabled. The sequence-level enable flag must be written into the bitstream when encoding a video sequence.
[0282] Step 1: If the flag indicating the use of loop filtering based on the neural network model is true, the encoder attempts the loop filtering technology based on the neural network model, i.e., executes step 2. If the flag indicating the use of loop filtering based on the neural network model is false, the encoder does not attempt the loop filtering technology based on the neural network model, i.e., skips step 2 and directly executes step 3.
[0283] Step 2: Initialize the neural network-based loop filtering technology and load the neural network model applicable to the current frame.
[0284] First round:
[0285] The encoder calculates the cost information of not using the loop filtering technology based on the neural network model. That is, it uses the reconstructed samples of the coding tree unit to be used as the input of the neural network model and the original image samples of the coding tree unit to calculate the rate-distortion cost (first rate-distortion cost), which is recorded as costOrg;
[0286] Second round:
[0287] The encoder attempts a neural network-based loop filtering technique, traversing two quantization parameter candidates separately. The reconstructed YUV sample of the current coding tree unit, the predicted YUV sample, the current frame type, and the quantization parameter are input into a pre-loaded network model for inference. The neural network loop filtering model outputs the reconstructed sample block of the current coding tree unit, completing the filtering of the current frame or slice. The three YUV color components are traversed. For example, for the Y component, the rate-distortion costs corresponding to the two quantization parameters are costYcand1 and costYcand2, respectively, resulting in two rate-distortion costs for each color component. The reconstructed sample (filtered result) corresponding to the minimum cost for each color component is selected as the filtered reconstructed sample for the Y component of the current frame, and the quantization parameter index (nnlf_luma_qp_index) corresponding to the Y component is recorded. After traversing the three color components and determining the corresponding quantization parameters, the total rate-distortion cost of the three color components is calculated and labeled as costFrameBest.
[0288] Round 3:
[0289] The encoder attempts to optimize the selection at the coding tree unit level. In the second round of the encoder's attempt at neural network loop filtering, it directly defaults to using this technology for all coding tree units in the current frame. Luma and chroma are each controlled by a frame-level switch flag, while the coding tree unit level does not need to transmit the use flag. This round attempts a combination of switches at the coding tree unit level, and each component can be controlled individually. The encoder traverses the coding tree units and calculates the rate-distortion cost of the reconstructed samples without the use of neural network loop filtering and the original samples of the current coding tree unit, which is recorded as costCTUrec. The rate-distortion cost of the reconstructed samples with the use of neural network loop filtering and the original samples of the current coding tree unit is calculated, which is recorded as costCTUcnn.
[0290] For the luminance component, if the costCTUrec of the current luminance component is smaller than the costCTUcnn of the luminance component, the usage flag (ctb_nnlf_luma_flag) of the neural network loop filter at the coding tree unit level of the luminance component is set to no; otherwise, the ctb_nnlf_luma_flag is set to true.
[0291] For the chroma component (two chroma components, here denoted as chroma1 and chroma2), if the costCTUrec of the current chroma component is smaller than the costCTUcnn of the chroma component, the use flag (ctb_nnlf_chroma1_flag or ctb_nnlf_chroma2_flag) of the neural network loop filter at the coding tree unit level of the chroma component is set to no; otherwise, the ctb_nnlf_chroma1_flag or ctb_nnlf_chroma2_flag is set to true.
[0292] If all coding tree units in the current frame have been traversed, the rate-distortion cost of the reconstructed sample of the current frame and the original image sample in this case is calculated and recorded as costCTUBest.
[0293] Traverse each color component. If the value of costOrg is the smallest, set the frame-level switch flag (first syntax element identification information) of the neural network loop filter corresponding to the color component to the first case (third value, 0) and write it into the bitstream. If the value of costFrameBest is the smallest, set the frame-level switch flag (ph_nnlf_luma_ctrl_flag / ph_nnlf_chroma_ctrl_flag) of the neural network loop filter corresponding to the color component to the second case (fourth value, 1) and write it into the bitstream. At the same time, write the quantization parameter index (nnlf_luma_qp_index / nnlf_chroma1_qp_index / nnlf_chroma2_qp_index) of different color components into the bitstream. If the value of costCTUBest is the smallest, the frame-level switch flag based on the neural network loop filter corresponding to the color component is set to the third case (fifth fourth value, 2) and written into the bitstream. At the same time, the use flag based on the neural network loop filter at the coding unit level is written into the bitstream, and the quantization parameter index of different color components (nnlf_luma_qp_index / nnlf_chroma1_qp_index / nnlf_chroma2_qp_index) is written into the bitstream.
[0294] Step 3: The encoder continues to try other loop filtering tools and outputs a complete reconstructed image after completion.
[0295] For example, in a specific embodiment, for the decoding end, the specific process is as follows:
[0296] The decoding end parses the sequence-level flag. If sps_nnlf_enable_flag is true, it indicates that the current bitstream allows the use of loop filtering technology based on the neural network model, and the subsequent decoding process needs to parse the relevant syntax elements. Otherwise, it indicates that the current bitstream does not allow the use of loop filtering technology based on the neural network model, and the subsequent decoding process does not need to parse the relevant syntax elements. The relevant syntax elements are defaulted to the initial value or to the false state.
[0297] Step 1. The decoder parses the syntax elements of the current frame and obtains the frame-level switch flag (first syntax element identification information) based on the neural network model. If the frame-level flag is not all the first case, execute step 2; otherwise, skip step 2 and execute step 3.
[0298] Step 2: If the frame-level switch flag is the first case, it means that the current color component does not use the neural network-based loop filtering technology, then all the coding tree units in the current frame use the flag bit (ctb_nnlf_luma_flag / ctb_nnlf_chroma1_flag / ctb_nnlf_chroma2_flag) to be negated, and continue to traverse other color components;
[0299] If the frame-level switch flag is the second case, it means that all coding tree units under the current color component are filtered using the neural network-based loop filtering technology, that is, the coding tree unit level usage flag (ctb_nnlf_luma_flag / ctb_nnlf_chroma1_flag / ctb_nnlf_chroma2_flag) of all coding tree units in the current frame under the color component is automatically set to true;
[0300] If the frame-level switch flag is the second case, it means that some coding tree units in the current color component use the neural network-based loop filtering technology, while some coding tree units do not use the neural network-based loop filtering technology. Therefore, if the frame-level switch flag is the third case, it is necessary to further analyze the coding tree unit-level usage flags (ctb_nnlf_luma_flag / ctb_nnlf_chroma1_flag / ctb_nnlf_chroma2_flag) of all coding tree units in the current frame for this color component.
[0301] If the frame-level switch flag ph_nnlf_luma_ctrl_flag or ph_nnlf_chroma_ctrl_flag is in the second or third case, the decoder can parse the quantization parameter index (nnlf_luma_qp_index / nnlf_chroma1_qp_index / nnlf_chroma2_qp_index) of the corresponding color component from the bitstream. At the same time, based on the quantization parameter index, the decoder uses a neural network-based loop filtering technique to infer the filtered reconstructed samples under different quantization parameters for the current frame image or current slice. The coding tree unit level flag is used to determine whether to filter the current coding tree unit. If filtering is required, the corresponding filtered reconstructed sample is obtained according to the quantization parameter index of the current color component as the output sample of the current color component of the coding tree unit; if filtering is not required, the original reconstructed sample is used as the output sample of the current color component of the coding tree unit.
[0302] After traversing all coding tree units of the current frame or current slice, the neural network-based loop filtering module ends.
[0303] Step 3: The decoder continues to traverse other loop filtering tools and outputs a complete reconstructed image after completion.
[0304] In another specific embodiment, a brief description of the parsing process at the decoding end is shown in Table 1, wherein bold fonts indicate grammatical elements that need to be parsed.
[0305] Table 1
[0306]
[0307]
[0308]
[0309] In yet another embodiment of the present application, the embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded may include at least one of the following: at least one of quantization parameter information or quantization parameter index, first syntax element identification information of the to-be-filtered component of the current frame or current slice, second syntax element identification information of the to-be-filtered component of the current block, third syntax element identification information of the current block contained in the current frame or current slice, fourth syntax element identification information of the current video sequence, and filtered reconstructed values of each partitioned block included in the current frame or current slice; wherein the current block is any one of the partitioned blocks.
[0310] In the embodiment of the present application, the specific implementation of the aforementioned embodiment is described in detail through the above embodiment. According to the technical solution of the aforementioned embodiment, it can be seen that the embodiment of the present application uses different quantization parameters corresponding to different filter components as input to improve encoding performance, and introduces new syntax elements. In this way, while maintaining only one model, the selection of different quantization parameters is increased, so that the luminance color component and the chrominance color component have more choices and adaptations. Through the rate-distortion optimization calculation on the encoding end, the decoding end does not need to store multiple neural network models to achieve a more flexible configuration, which is conducive to improving encoding performance; at the same time, the present technical solution can also remove non-important input elements such as prediction information YUV, division information YUV, etc., reducing the amount of calculation for inference of the preset network model, which is beneficial to the implementation of the encoding and decoding end and reducing the encoding and decoding time.
[0311] In another embodiment of the present application, based on the same inventive concept as the above embodiment, see Figure 11, which shows a schematic diagram of the composition structure of a decoder provided by an embodiment of the present application. As shown in Figure 11, the decoder 1 may include:
[0312] The parsing section 10 is configured to parse the bitstream and determine first syntax element identification information of a component to be filtered of a current block; the first syntax element identification information is used to determine whether each block in the current frame or current slice is filtered based on a preset network model;
[0313] The first determining part 11 is configured to determine quantization parameter information of the to-be-filtered component and determine second syntax element identification information when the first syntax element identification information indicates that the to-be-filtered component of the partitioned block in the current frame or the current slice allows filtering using a preset network model;
[0314] The first filtering part 12 is configured to filter the current block of the current frame or the current slice based on the second syntax element identification information, the quantization parameter information and the preset network model to obtain a filtered reconstructed value of the to-be-filtered component of the current block.
[0315] In some embodiments of the present application, the parsing part 10 is further configured to determine the quantization parameter information of the component to be filtered from the bitstream when the first syntax element identification information indicates that the component to be filtered of all divided blocks in the current frame or the current slice is allowed to be filtered using the preset network model, and the first determination part 11 is further configured to assign a first value to the second syntax element identification information.
[0316] In some embodiments of the present application, the parsing part 10 is further configured to determine the quantization parameter information of the component to be filtered from the bitstream and determine the second syntax element identification information from the bitstream when the first syntax element identification information indicates that there is at least one partition block in the current frame or the current slice that allows filtering using the preset network model; the at least one partition block is a partial partition block in the current frame or the current slice.
[0317] In some embodiments of the present application, the first determining part 11 is further configured to assign a second value to the second syntax element identification information when the first syntax element identification information indicates that the current frame or the current slice does not allow filtering using a preset network model;
[0318] After determining the reconstructed value of the component to be filtered of the current block, the reconstructed value of the component to be filtered of the current block is directly determined as the filtered reconstructed value of the component to be filtered of the current block.
[0319] In some embodiments of the present application, the parsing part 10 is further configured to parse the code stream to determine the reconstructed value of the to-be-filtered component of the current block.
[0320] In some embodiments of the present application, the first filtering part 12 is further configured to input the reconstructed value of the component to be filtered and the quantization parameter corresponding to the quantization parameter index into the preset network model when the second syntax element identification information indicates that the component to be filtered of the current block is filtered using a preset network model, filter the current block of the current frame or the current slice, and obtain the filtered reconstructed value of the component to be filtered of the current block.
[0321] In some embodiments of the present application, the first determining part 11 is further configured to, if the first syntax element identification information is a third value, determine that the first syntax element identification information indicates that the current frame or the current slice is not allowed to be filtered using the preset network model;
[0322] If the first syntax element identification information is a fourth value, determining that the first syntax element identification information indicates that all to-be-filtered components of all partitioned blocks in the current frame or the current slice are allowed to be filtered using the preset network model;
[0323] If the first syntax element identification information is the fifth value, it is determined that the first syntax element identification information indicates that there is at least one partition block in the current frame or the current slice that allows filtering using the preset network model.
[0324] In some embodiments of the present application, the parsing part 10 is further configured to parse the code stream to determine the quantization parameter index of the component to be filtered of the current block; and determine the quantization parameter information of the component to be filtered corresponding to the current block from the quantization parameter candidate set based on the quantization parameter index.
[0325] In some embodiments of the present application, the parsing part 10 is further configured to parse the code stream to determine the quantization parameter information of the component to be filtered corresponding to the current block.
[0326] In some embodiments of the present application, the first determination part 11 is further configured to determine the quantization parameter information of the component to be filtered and the second syntax element identification information when the first syntax element identification information indicates that there is a component to be filtered in a divided block in the current frame or the current slice, which allows filtering using a preset network model, and the current frame or the current slice is a first type frame.
[0327] In some embodiments of the present application, the parsing part 10 is further configured to determine third syntax element identification information of the to-be-filtered component of the current block from the bitstream;
[0328] The first filtering part 12 is further configured to, when the first syntax element identification information indicates that there is a to-be-filtered component of the divided block in the current frame or the current slice, which allows filtering using a preset network model, and the third syntax element identification information indicates that the to-be-filtered component of the current block included in the current frame or the current slice allows filtering using the preset network model, filter the current block of the current frame or the current slice based on the quantization parameter information and the preset network model to obtain the filtered reconstructed value of the to-be-filtered component of the current block; or,
[0329] When the third syntax element identification information indicates that the component to be filtered of the current block included in the current frame or the current slice is allowed to be filtered using the preset network model, and the second syntax element identification information indicates that the component to be filtered of the current block is filtered using the preset network model, the current block of the current frame or the current slice is filtered based on the quantization parameter information and the preset network model to obtain the filtered reconstructed value of the component to be filtered of the current block.
[0330] In some embodiments of the present application, the parsing part 10 is further configured to not parse the third syntax element identification information when the second syntax element identification information indicates that the to-be-filtered component of the current block is not filtered using a preset network model.
[0331] In some embodiments of the present application, the first determining part 11 is further configured to obtain at least one of a predicted value of a component to be filtered of the current block, block partition information, and a deblocking filter boundary strength, and a reconstructed value of the component to be filtered of the current block;
[0332] The first filtering part 12 is also configured to use the preset network model, combined with the predicted value of the component to be filtered of the current block, the block division information and at least one of the deblocking filter boundary strength, and the quantization parameter information, to filter the reconstructed value of the component to be filtered of the current block to obtain the filtered reconstructed value of the component to be filtered of the current block.
[0333] In some embodiments of the present application, the first filtering part 12 is further configured to filter the current block of the current frame based on the second syntax element identification information, the quantization parameter information and the preset network model to obtain first residual information of the to-be-filtered component of the current block;
[0334] The filtered reconstructed value of the component to be filtered of the current block is determined based on the first residual information and the reconstructed value of the component to be filtered of the current block.
[0335] In some embodiments of the present application, the parsing part 10 is further configured to parse out fourth syntax element identification information before determining the first syntax element identification information of the to-be-filtered component of the current block;
[0336] When the fourth syntax element identification information indicates that a sequence including the current frame or the current slice allows filtering using the preset network model, the first syntax element identification information is parsed.
[0337] In some embodiments of the present application, the first determining part 11 is further configured to traverse each divided block in the current frame or the current slice, take each divided block as the current block in turn, and repeatedly perform the steps of parsing the bitstream and determining the filtered reconstructed value of the to-be-filtered component of the current block, so as to obtain the filtered reconstructed value of each to-be-filtered component corresponding to each divided block;
[0338] The reconstructed image of the current frame or the current slice is determined according to the filtered reconstructed values of the components to be filtered corresponding to the divided blocks.
[0339] In some embodiments of the present application, the first determining part 11 is further configured to traverse each to-be-filtered component of the current block, and repeat the steps of parsing the bitstream and determining the filtered reconstructed value of each to-be-filtered component of the current block for each to-be-filtered component in turn, so as to obtain the filtered reconstructed value of each to-be-filtered component; wherein each to-be-filtered component corresponds to its own quantization parameter information;
[0340] The filtered reconstructed value of the current block is determined according to the filtered reconstructed values of the components to be filtered of the current block.
[0341] The embodiment of the present application further provides a decoder 1, as shown in FIG12 , the decoder 1 may include:
[0342] A first memory 13 configured to store a computer program that can be executed on the first processor 14;
[0343] The first processor 14 is configured to execute the decoding method of the decoder when running the computer program.
[0344] It is understandable that during the decoding process, the decoder can determine the post-filtering reconstruction value corresponding to each component to be filtered, based on the quantization parameter information corresponding to each component to be filtered. In other words, the quantization parameter information corresponding to different components to be filtered in the current frame or current slice can be different (and is the optimal quantization parameter information of each component to be filtered obtained by the encoder through multiple model reasoning during encoding), and each component to be filtered can be subjected to respective filtering processes to obtain the post-filtering reconstruction value of each component to be filtered. Therefore, each component to be filtered in the current block can use its own quantization parameter information to implement filtering of each component to be filtered in the same preset network model, thereby improving the compression performance of each component to be filtered, and making the selection of input information (quantization parameter information) for filtering different components to be filtered more flexible while ensuring that the complexity of the model is not increased, thereby improving decoding efficiency.
[0345] In another embodiment of the present application, based on the same inventive concept as the above embodiment, see Figure 13, which shows a schematic diagram of the composition structure of an encoder provided by an embodiment of the present application. As shown in Figure 13, the encoder 2 may include:
[0346] The second determining part 20 is configured to determine a first rate-distortion cost of the current frame or the current slice; the first rate-distortion cost is obtained by filtering all the to-be-filtered components of the divided blocks included in the current frame or the current slice without using the preset network model;
[0347] The second filtering part 21 is configured to determine a second rate-distortion cost for the current frame or the current slice; the second rate-distortion cost is obtained by filtering the to-be-filtered components of all the partitions included in the current frame or the current slice using the preset network model; determine a third rate-distortion cost for the current frame or the current slice; the third rate-distortion cost is obtained by allowing the to-be-filtered components of at least one partition in the current frame to be filtered using the preset network model; the at least one partition is a portion of the partitions in the current frame or the current slice;
[0348] The second determining part 20 is further configured to determine first syntax element identification information of the to-be-filtered component of the current frame or the current slice according to the first rate-distortion cost, the second rate-distortion cost and the third rate-distortion cost.
[0349] In some embodiments of the present application, the second determination part 20 is further configured to obtain the original values of each partitioned block of the component to be filtered in the current frame or current slice, the reconstructed values of each partitioned block and the original quantization parameter information when the obtained fourth syntax element identification information represents that the sequence containing the current frame or current slice allows filtering using the preset network model.
[0350] In some embodiments of the present application, the second determining part 20 is further configured to perform rate-distortion cost calculation based on the original values of the respective partitioned blocks and the reconstructed values of the respective partitioned blocks to determine a first rate-distortion cost of the current frame or current slice.
[0351] In some embodiments of the present application, the second determining part 20 is further configured to determine at least two types of original quantization parameter information in the original quantization parameter information;
[0352] The second filtering part 21 is further configured to filter, based on the preset network model, the reconstructed values of the components to be filtered of each divided block included in the current frame or the current slice and each original quantization parameter information under each original quantization parameter information, to obtain the original filtered reconstructed values of the components to be filtered of each divided block included in the current frame or the current slice;
[0353] The second determination part 20 is further configured to perform rate-distortion cost calculation based on the original values of the components to be filtered of each divided block included in the current frame or the current slice and the original filtered reconstructed values of the components to be filtered of each divided block, to obtain a fourth rate-distortion cost of the components to be filtered under each original quantization parameter information; determine a third minimum rate-distortion cost with the minimum rate-distortion cost from the at least two fourth rate-distortion costs obtained; and determine the second rate-distortion cost based on the third minimum rate-distortion cost of each component to be filtered.
[0354] In some embodiments of the present application, the second determining part 20 is further configured to determine the first sub-rate-distortion cost of the to-be-filtered component corresponding to each partitioned block when the first rate-distortion cost is obtained; determine the second sub-rate-distortion cost of the to-be-filtered component corresponding to each partitioned block when the second rate-distortion cost is obtained; determine the first minimum rate-distortion cost of the to-be-filtered component of the current frame or the current slice based on the first sub-rate-distortion cost and the second sub-rate-distortion cost; and determine the third rate-distortion cost based on the first minimum rate-distortion cost of each component to be filtered.
[0355] In some embodiments of the present application, the second determining part 20 is further configured to traverse each component to be filtered and determine the sum of the first minimum rate-distortion cost of each divided block as the third rate-distortion cost of the current frame or the current slice.
[0356] In some embodiments of the present application, the second determining part 20 is further configured to:
[0357] If the first sub-rate distortion cost corresponding to any partition block is less than the second sub-rate distortion cost, determining that the second syntax element identification information of the to-be-filtered component of the any partition block is a second value; the second value indicates that the partition block is not allowed to be filtered using the preset network model;
[0358] If the first sub-rate distortion cost corresponding to any partition block is greater than or equal to the second sub-rate distortion cost, then determine that the second syntax element identification information of the to-be-filtered component of any partition block is a first value; the first value indicates that the partition block allows filtering using the preset network model.
[0359] In some embodiments of the present application, the second determining part 20 is further configured to:
[0360] determining a second minimum rate-distortion cost having the minimum rate-distortion cost from among the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost;
[0361] If the second minimum rate-distortion cost is the first rate-distortion cost, determining that the first syntax element identification information of the to-be-filtered component of the current frame or the current slice is a third value;
[0362] If the second minimum rate-distortion cost is the second rate-distortion cost, determining that the first syntax element identification information of the to-be-filtered component of the current frame or the current slice is a fourth value;
[0363] If the second minimum rate-distortion cost is the third rate-distortion cost, the first syntax element identification information of the to-be-filtered component of the current frame or the current slice is determined to be a fifth value.
[0364] In some embodiments of the present application, the encoder 2 further includes a writing portion 22;
[0365] The writing part 22 is configured to write the second syntax element identification information of each partition block into the bitstream when determining that the first syntax element identification information of the to-be-filtered component of the current frame or the current slice is the fifth value.
[0366] In some embodiments of the present application, the encoder 2 further includes a writing part 22 for writing the first syntax element identification information into a bitstream.
[0367] In some embodiments of the present application, the encoder 2 further includes a writing portion 22;
[0368] The second determining portion 20 is further configured to, if the second minimum rate-distortion cost is the second rate-distortion cost or the third rate-distortion cost, determine an original quantization parameter information corresponding to the second rate-distortion cost or the third rate-distortion cost as the quantization parameter information of the component to be filtered;
[0369] The writing part 22 is configured to write the quantization parameter information into the bitstream; or write the quantization parameter index corresponding to the quantization parameter information into the bitstream.
[0370] In some embodiments of the present application, the second determining part 20 is further configured to, when the acquired fourth syntax element identification information indicates that the sequence including the current frame or the current slice allows filtering using the preset network model, acquire at least one of the predicted value, block partition information, and deblocking filter boundary strength of each partitioned block of the component to be filtered in the current frame or the current slice, the reconstructed value, and original quantization parameter information of each partitioned block of the component to be filtered in the current frame or the current slice; and determine at least two types of original quantization parameter information among the original quantization parameter information;
[0371] The second filtering part 21 is further configured to input the respective partition prediction values, the block partition information, and at least one of the deblocking filter boundary strength, the respective partition reconstruction values, and each of the original quantization parameter information into the preset network model for filtering under each type of original quantization parameter information, to determine the original filtered reconstruction value of each partition block of the component to be filtered;
[0372] The second determination part 20 is further configured to perform rate-distortion cost calculation based on the original values of the components to be filtered of each divided block included in the current frame and the original filtered reconstructed values of the components to be filtered of each divided block, to obtain a fourth rate-distortion cost of the components to be filtered under each original quantization parameter information; determine a third minimum rate-distortion cost with the minimum rate-distortion cost from the at least two fourth rate-distortion costs obtained; and determine the second rate-distortion cost based on the third minimum rate-distortion costs of each component to be filtered.
[0373] In some embodiments of the present application, the encoder 2 further includes a writing portion 22;
[0374] The second determining portion 20 is further configured to, if the second minimum rate-distortion cost is the second rate-distortion cost or the third rate-distortion cost, determine the original filtered reconstructed value of each partition block corresponding to the second rate-distortion cost or the third rate-distortion cost as the filtered reconstructed value of the to-be-filtered component of each partition block;
[0375] The writing part 22 is configured to write the filtered reconstructed value into the bitstream.
[0376] In some embodiments of the present application, the encoder 2 further includes a writing portion 22;
[0377] The second determining part 20 is further configured to determine fourth syntax element identification information;
[0378] The fourth syntax element identification information represents that when the sequence containing the current frame or the current slice allows filtering using the preset network model, and the current frame or the current slice is a first type frame, the first rate-distortion cost, the second rate-distortion cost and the third rate-distortion cost in the encoding method are determined.
[0379] The writing part 22 is configured to write the fourth syntax element identification information into the bitstream.
[0380] In some embodiments of the present application, the encoder 2 further includes a writing portion 22;
[0381] The second determining part 20 is further configured to determine third syntax element identification information; the third syntax element identification information is used to indicate that the to-be-filtered component of the current block allows filtering using the preset network model;
[0382] The writing part 22 is configured to write the third syntax element identification information into the bitstream.
[0383] The embodiment of the present application provides an encoder 2, as shown in FIG14 , the encoder 2 may include:
[0384] a second memory 23 configured to store a computer program capable of running on the second processor 24;
[0385] The second processor 24 is configured to execute the encoding method described by the encoder when running the computer program.
[0386] It is understood that during the encoding process, the encoder determines the encoding method that minimizes the frame-level rate-distortion cost or the slice-level rate-distortion cost by performing multiple model inferences based on the following scenarios: no filtering of each partitioned block, filtering of each partitioned block, and partial filtering of each partitioned block. The quantization parameter information corresponding to each component to be filtered can be different types of quantization parameter information under different circumstances. In other words, the quantization parameter information corresponding to different components to be filtered in the current frame or current slice can be different (the optimal quantization parameter information for each component to be filtered obtained by the encoder through multiple model inferences during encoding), and each component to be filtered can be filtered to obtain a filtered reconstructed value. Therefore, each component to be filtered in the current block can use its own quantization parameter information to implement filtering of each component to be filtered in the same preset network model, thereby improving the compression performance of each component to be filtered. While ensuring that the complexity of the model is not increased, the selection of input information (quantization parameter information) for filtering different components to be filtered is more flexible, thereby improving encoding efficiency.
[0387] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which implements the decoding method described in the decoder when executed by a first processor, or implements the encoding method described in the encoder when executed by a second processor.
[0388] The various components in the embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of software functional modules.
[0389] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned computer-readable storage medium includes: ferromagnetic random access memory (FRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface storage, optical disk, or compact disc read-only memory (CD-ROM), etc. Various media that can store program codes are not limited in the embodiments of the present disclosure.
[0390] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability
[0391] An embodiment of the present application provides a coding and decoding method, an encoder, a decoder, a bitstream, and a storage medium. In the decoder, by parsing the bitstream, first syntax element identification information of a component to be filtered of a current frame or a current slice is determined; the first syntax element identification information is used to determine whether the components to be filtered of each block in the current frame or the current slice are all filtered based on a preset network model; when the first syntax element identification information indicates that there are components to be filtered of divided blocks in the current frame or the current slice that allow filtering using the preset network model, quantization parameter information of the components to be filtered is determined, and second syntax element identification information is determined; based on the second syntax element identification information, the quantization parameter information, and the preset network model, the current block of the current frame or the current slice is filtered to obtain a filtered reconstructed value of the component to be filtered of the current block. In the encoder, a first rate-distortion cost of the current frame or current slice is determined; the first rate-distortion cost is obtained by not filtering the to-be-filtered components of all the partitioned blocks included in the current frame or current slice using a preset network model; a second rate-distortion cost of the current frame or current slice is determined; the second rate-distortion cost is obtained by filtering the to-be-filtered components of all the partitioned blocks included in the current frame or current slice using a preset network model; a third rate-distortion cost of the current frame or current slice is determined; the third rate-distortion cost is obtained by allowing the to-be-filtered components of at least one partitioned block in the current frame or current slice to be filtered using a preset network model; at least one partitioned block is a partial partitioned block in the current frame or current slice; based on the first rate-distortion cost, the second rate-distortion cost and the third rate-distortion cost, the first syntax element identification information of the to-be-filtered components of the current frame or current slice is determined.
[0392] The present application determines the encoding method with the lowest frame-level rate-distortion cost or slice-level rate-distortion cost through multiple model reasoning, from the cases where each partition block is not filtered, each partition block is filtered, and each partition block is partially filtered. The quantization parameter information corresponding to each component to be filtered can be different types of quantization parameter information under different circumstances. In other words, the quantization parameter information corresponding to different components to be filtered of the current frame or current slice can be different (the quantization parameter information of each component to be filtered is the optimal quantization parameter information obtained by the encoder through multiple model reasoning during encoding), and each component to be filtered can be filtered and reconstructed to obtain the filtered value of each component to be filtered. Therefore, each component to be filtered of the current block can use its own quantization parameter information to implement filtering of each component to be filtered in the same preset network model, thereby improving the compression performance of each component to be filtered. On the basis of ensuring that the complexity of the model is not increased, the selection of input information (quantization parameter information) for filtering different components to be filtered is more flexible, thereby improving the coding and decoding efficiency.
Claims
1. A decoding method, applied to a decoder, comprising: Parse the code stream to determine the first syntax element identification information of the component to be filtered of the current frame or current slice; The first syntax element identification information is used to determine whether the to-be-filtered components of each block in the current frame or current slice are all filtered based on the preset network model; When the first syntax element identification information indicates that there is a to-be-filtered component of a partitioned block in the current frame or the current slice that allows filtering using a preset network model, determining quantization parameter information of the to-be-filtered component and determining second syntax element identification information; Based on the second syntax element identification information, the quantization parameter information and the preset network model, the current block of the current frame or the current slice is filtered to obtain a filtered reconstructed value of the to-be-filtered component of the current block.
2. The method according to claim 1, wherein When the first syntax element identification information indicates that a to-be-filtered component of a partitioned block in the current frame or the current slice allows filtering using a preset network model, determining quantization parameter information of the to-be-filtered component and determining the second syntax element identification information include: When the first syntax element identification information indicates that the components to be filtered of all partitioned blocks in the current frame or the current slice are allowed to be filtered using the preset network model, the quantization parameter information of the components to be filtered is determined from the bitstream, and a first value is assigned to the second syntax element identification information.
3. The method according to claim 1, wherein When the first syntax element identification information indicates that a to-be-filtered component of a partitioned block in the current frame or the current slice allows filtering using a preset network model, determining quantization parameter information of the to-be-filtered component and determining the second syntax element identification information include: When the first syntax element identification information indicates that there is at least one partition block in the current frame or the current slice whose to-be-filtered component allows filtering using the preset network model, the quantization parameter information of the to-be-filtered component is determined from the bitstream, and the second syntax element identification information is determined from the bitstream; the at least one partition block is a partial partition block in the current frame or the current slice.
4. The method according to claim 1, wherein The method further comprises: When the first syntax element identification information indicates that the current frame or the current slice does not allow filtering using a preset network model, assigning a second value to the second syntax element identification information; After determining the reconstructed value of the component to be filtered of the current block, the reconstructed value of the component to be filtered of the current block is directly determined as the filtered reconstructed value of the component to be filtered of the current block.
5. The method according to claim 1, wherein The method further comprises: Parse the code stream to determine the reconstructed value of the to-be-filtered component of the current block.
6. The method according to claim 5, wherein: The filtering, based on the second syntax element identification information, the quantization parameter index, and the preset network model, on the current block of the current frame or the current slice to obtain a filtered reconstructed value of a to-be-filtered component of the current block, includes: When the second syntax element identification information indicates that the component to be filtered of the current block is filtered using a preset network model, the reconstructed value of the component to be filtered and the quantization parameter corresponding to the quantization parameter index are input into the preset network model, and the current block of the current frame or the current slice is filtered to obtain the filtered reconstructed value of the component to be filtered of the current block.
7. The method according to claim 1, wherein The determining of the first syntax element identification information of the component to be filtered of the current frame or current slice includes: If the first syntax element identification information is a third value, determining that the first syntax element identification information indicates that the current frame or the current slice is not allowed to be filtered using the preset network model; If the first syntax element identification information is a fourth value, determining that the first syntax element identification information indicates that all to-be-filtered components of all partitioned blocks in the current frame or the current slice are allowed to be filtered using the preset network model; If the first syntax element identification information is the fifth value, it is determined that the first syntax element identification information indicates that there is at least one partition block in the current frame or the current slice that allows filtering using the preset network model.
8. The method according to claim 1, wherein The determining of the quantization parameter information of the component to be filtered includes: Parse the code stream and determine the quantization parameter index of the component to be filtered of the current block; According to the quantization parameter index, the quantization parameter information of the to-be-filtered component corresponding to the current block is determined from a quantization parameter candidate set.
9. The method according to claim 1, wherein: The determining of the quantization parameter information of the component to be filtered includes: Parse the code stream to determine the quantization parameter information of the to-be-filtered component corresponding to the current block.
10. The method according to claim 1, wherein The method further comprises: When the first syntax element identification information indicates that there is a component to be filtered in the divided block in the current frame or the current slice, which allows filtering using a preset network model, and the current frame or the current slice is a first type frame, determine the quantization parameter information of the component to be filtered, as well as the second syntax element identification information.
11. The method according to claim 1, wherein The method further comprises: Determine third syntax element identification information of the to-be-filtered component of the current block from a bitstream; When the first syntax element identification information indicates that there is a to-be-filtered component of a divided block in the current frame or the current slice, which allows filtering using a preset network model, and the third syntax element identification information indicates that the to-be-filtered component of a current block included in the current frame or the current slice allows filtering using the preset network model, filtering the current block of the current frame or the current slice based on the quantization parameter information and the preset network model to obtain the filtered reconstructed value of the to-be-filtered component of the current block; or When the third syntax element identification information indicates that the component to be filtered of the current block included in the current frame or the current slice is allowed to be filtered using the preset network model, and the second syntax element identification information indicates that the component to be filtered of the current block is filtered using the preset network model, the current block of the current frame or the current slice is filtered based on the quantization parameter information and the preset network model to obtain the filtered reconstructed value of the component to be filtered of the current block.
12. The method according to claim 1, wherein The method further comprises: When the second syntax element identification information indicates that the to-be-filtered component of the current block is not filtered using a preset network model, the third syntax element identification information is not parsed.
13. The method according to any one of claims 1 to 3, 7 to 10, wherein: Before filtering the current block of the current frame or the current slice based on the second syntax element identification information, the quantization parameter index, and the preset network model to obtain a filtered reconstructed value of a to-be-filtered component of the current block, the method further includes: Obtaining at least one of a predicted value of a component to be filtered of a current block, block partition information, and a deblocking filter boundary strength, and a reconstructed value of the component to be filtered of the current block; The filtering, based on the second syntax element identification information, the quantization parameter index, and the preset network model, on the current block of the current frame or the current slice to obtain a filtered reconstructed value of a to-be-filtered component of the current block, includes: Using the preset network model, combined with the predicted value of the component to be filtered of the current block, the block division information and at least one of the deblocking filter boundary strength, and the quantization parameter information, the reconstructed value of the component to be filtered of the current block is filtered to obtain the filtered reconstructed value of the component to be filtered of the current block.
14. The method according to any one of claims 1 to 3, 7 to 10, wherein: The filtering of the current block of the current frame based on the second syntax element identification information, the quantization parameter index, and the preset network model to obtain a filtered reconstructed value of a to-be-filtered component of the current block includes: Filtering the current block of the current frame based on the second syntax element identification information, the quantization parameter information, and the preset network model to obtain first residual information of a to-be-filtered component of the current block; The filtered reconstructed value of the component to be filtered of the current block is determined based on the first residual information and the reconstructed value of the component to be filtered of the current block.
15. The method according to claim 1, wherein Before determining the first syntax element identification information of the to-be-filtered component of the current block, the method further includes: Parsing the fourth syntax element identification information; When the fourth syntax element identification information indicates that a sequence including the current frame or the current slice allows filtering using the preset network model, the first syntax element identification information is parsed.
16. The method according to any one of claims 1 to 12 and 15, wherein: The method further comprises: Traversing each divided block in the current frame or the current slice, taking each divided block as the current block in turn, and repeatedly performing the steps of parsing the bitstream and determining the filtered reconstructed value of the to-be-filtered component of the current block, so as to obtain the filtered reconstructed value of each to-be-filtered component corresponding to each divided block; The reconstructed image of the current frame or the current slice is determined according to the filtered reconstructed values of the components to be filtered corresponding to the divided blocks.
17. The method according to claim 1, wherein The method further comprises: Traversing each to-be-filtered component of the current block, and repeating the steps of parsing the bitstream for each to-be-filtered component in turn to determine a post-filtering reconstructed value of each to-be-filtered component of the current block, so as to obtain a post-filtering reconstructed value of each to-be-filtered component; wherein each to-be-filtered component corresponds to its own quantization parameter information; The filtered reconstructed value of the current block is determined according to the filtered reconstructed values of the components to be filtered of the current block.
18. A coding method, applied to an encoder, comprising: determining a first rate-distortion cost for a current frame or a current slice; The first rate-distortion cost is obtained by filtering all to-be-filtered components of all divided blocks included in the current frame or the current slice without using a preset network model; Determining a second rate-distortion cost of a current frame or a current slice; the second rate-distortion cost is obtained by filtering all to-be-filtered components of all divided blocks included in the current frame or the current slice using the preset network model; Determining a third rate-distortion cost of a current frame or a current slice; the third rate-distortion cost is obtained when a to-be-filtered component of at least one divided block in the current frame or the current slice is allowed to be filtered using the preset network model; The at least one partition block is a partial partition block in the current frame or the current slice; First syntax element identification information of the to-be-filtered component of the current frame or current slice is determined according to the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost.
19. The method according to claim 18, wherein The method further comprises: When the fourth syntax element identification information obtained represents that the sequence containing the current frame or current slice allows filtering using the preset network model, the original values of each partition block of the to-be-filtered component in the current frame or current slice, the reconstructed values of each partition block and the original quantization parameter information are obtained.
20. The method according to claim 19, wherein The determining of the first rate-distortion cost of the current frame or the current slice includes: A rate-distortion cost is calculated based on the original value of each partition block and the reconstructed value of each partition block to determine a first rate-distortion cost of the current frame or the current slice.
21. The method according to claim 19, wherein The determining a second rate-distortion cost of the current frame or the current slice includes: determining at least two types of original quantization parameter information among the original quantization parameter information; Under each type of original quantization parameter information, filtering, based on the preset network model, the reconstructed value of the to-be-filtered component of each divided block included in the current frame or the current slice and each type of original quantization parameter information to obtain the original filtered reconstructed value of the to-be-filtered component of each divided block included in the current frame or the current slice; performing rate-distortion cost calculation based on original values of the components to be filtered of each divided block included in the current frame or the current slice and original filtered reconstructed values of the components to be filtered of each divided block, to obtain a fourth rate-distortion cost of the components to be filtered under each original quantization parameter information; Determining a third minimum rate-distortion cost with the minimum rate-distortion cost from the obtained at least two fourth rate-distortion costs; The second rate-distortion cost is determined based on the third minimum rate-distortion cost of each to-be-filtered component.
22. The method according to claim 18, wherein The determining a third rate-distortion cost of the current frame or the current slice includes: determining, when obtaining the first rate-distortion cost, first sub-rate-distortion costs of the components to be filtered corresponding to the respective divided blocks; When the second rate-distortion cost is determined, the second sub-rate-distortion cost of the to-be-filtered component corresponding to each divided block is obtained; Determining a first minimum rate-distortion cost of a to-be-filtered component of the current frame or the current slice based on the first sub-rate-distortion cost and the second sub-rate-distortion cost; The third rate-distortion cost is determined based on the first minimum rate-distortion cost of each to-be-filtered component.
23. The method according to claim 22, wherein The determining, based on the first minimum rate-distortion cost of each to-be-filtered component, a third rate-distortion cost of the current frame or the current slice includes: Each to-be-filtered component is traversed, and the sum of the first minimum rate-distortion costs of each divided block is determined as the third rate-distortion cost of the current frame or the current slice.
24. The method according to claim 22 or 23, wherein The method further comprises: If the first sub-rate distortion cost corresponding to any partition block is less than the second sub-rate distortion cost, determining that the second syntax element identification information of the to-be-filtered component of the any partition block is a second value; the second value indicates that the partition block is not allowed to be filtered using the preset network model; If the first sub-rate distortion cost corresponding to any partition block is greater than or equal to the second sub-rate distortion cost, then determine that the second syntax element identification information of the to-be-filtered component of any partition block is a first value; the first value indicates that the partition block allows filtering using the preset network model.
25. The method according to claim 18, wherein The determining, according to the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost, first syntax element identification information of the to-be-filtered component of the current frame or the current slice includes: determining a second minimum rate-distortion cost having the minimum rate-distortion cost from among the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost; If the second minimum rate-distortion cost is the first rate-distortion cost, determining that the first syntax element identification information of the to-be-filtered component of the current frame or the current slice is a third value; If the second minimum rate-distortion cost is the second rate-distortion cost, determining that the first syntax element identification information of the to-be-filtered component of the current frame or the current slice is a fourth value; If the second minimum rate-distortion cost is the third rate-distortion cost, the first syntax element identification information of the to-be-filtered component of the current frame or the current slice is determined to be a fifth value.
26. The method according to claim 18, wherein The method further comprises: When it is determined that the first syntax element identification information of the to-be-filtered component of the current frame or the current slice is the fifth value, the second syntax element identification information of each partition block is written into the bitstream.
27. The method according to claim 18, wherein The method further comprises: The first syntax element identification information is written into the bitstream.
28. The method according to claim 18 or 25, wherein The method further comprises: If the second minimum rate-distortion cost is the second rate-distortion cost or the third rate-distortion cost, an original quantization parameter information corresponding to the second rate-distortion cost or the third rate-distortion cost is determined as the quantization parameter information of the component to be filtered, and the quantization parameter information is written into the bitstream; alternatively, a quantization parameter index corresponding to the quantization parameter information is written into the bitstream.
29. The method according to claim 18, wherein The determining of a second rate-distortion cost of a component to be filtered of the current frame or the current slice includes: When the obtained fourth syntax element identification information indicates that a sequence including the current frame or the current slice allows filtering using the preset network model, obtaining at least one of a predicted value, block partitioning information, and a deblocking filter boundary strength for each partitioned block of a component to be filtered in the current frame or the current slice, and reconstructed values and original quantization parameter information for each partitioned block of the component to be filtered in the current frame or the current slice; determining at least two types of original quantization parameter information among the original quantization parameter information; Under each type of original quantization parameter information, inputting the respective partition prediction values, at least one of the block partition information and the deblocking filter boundary strength, the respective partition reconstruction values and each type of original quantization parameter information into the preset network model for filtering, thereby determining an original filtered reconstruction value of each partition block of the component to be filtered; performing rate-distortion cost calculation based on original values of the components to be filtered of each divided block included in the current frame and original filtered reconstructed values of the components to be filtered of each divided block, to obtain a fourth rate-distortion cost of the components to be filtered under each original quantization parameter information; Determining a third minimum rate-distortion cost with the minimum rate-distortion cost from the obtained at least two fourth rate-distortion costs; The second rate-distortion cost is determined based on the third minimum rate-distortion cost of each to-be-filtered component.
30. The method of claim 18, wherein The method further comprises: If the second minimum rate-distortion cost is the second rate-distortion cost or the third rate-distortion cost, the original filtered reconstruction values of each partitioned block corresponding to the second rate-distortion cost or the third rate-distortion cost are determined as the filtered reconstruction values of the to-be-filtered components of each partitioned block, and the filtered reconstruction values are written into the bitstream.
31. The method according to claim 18, wherein The method further comprises: Determine fourth syntax element identification information; and write the fourth syntax element identification information into a bitstream. The fourth syntax element identification information represents that when the sequence containing the current frame or the current slice allows filtering using the preset network model, and the current frame or the current slice is a first type frame, the first rate-distortion cost, the second rate-distortion cost and the third rate-distortion cost in the encoding method are determined.
32. The method of claim 18, wherein: Determining third syntax element identification information; the third syntax element identification information is used to indicate that the to-be-filtered component of the current block allows filtering using the preset network model; The third syntax element identification information is written into the bitstream.
33. A decoder, comprising: A parsing part is configured to parse the code stream and determine the first syntax element identification information of the to-be-filtered component of the current block; The first syntax element identification information is used to determine whether each block in the current frame or current slice is filtered based on a preset network model; a first determining part configured to, when the first syntax element identification information indicates that a to-be-filtered component of a partitioned block in the current frame or the current slice allows filtering using a preset network model, determine quantization parameter information of the to-be-filtered component and determine second syntax element identification information; The first filtering part is configured to filter the current block of the current frame or the current slice based on the second syntax element identification information, the quantization parameter information and the preset network model to obtain a filtered reconstructed value of the to-be-filtered component of the current block.
34. An encoder, comprising: A second determining part is configured to determine a first rate-distortion cost of a current frame or a current slice; The first rate-distortion cost is obtained by filtering all to-be-filtered components of all divided blocks included in the current frame or the current slice without using a preset network model; The second filtering part is configured to determine a second rate-distortion cost for the current frame or the current slice; the second rate-distortion cost is obtained by filtering the to-be-filtered components of all the partitioned blocks included in the current frame or the current slice using the preset network model; determine a third rate-distortion cost for the current frame or the current slice; the third rate-distortion cost is obtained by allowing the to-be-filtered components of at least one partitioned block in the current frame to be filtered using the preset network model; The at least one partition block is a partial partition block in the current frame or the current slice; The second determining part is further configured to determine first syntax element identification information of the to-be-filtered component of the current frame or the current slice according to the first rate-distortion cost, the second rate-distortion cost, and the third rate-distortion cost.
35. A decoder, comprising: a first memory configured to store a computer program executable on the first processor; The first processor is configured to execute the method according to any one of claims 1 to 17 when running the computer program.
36. An encoder, comprising: a second memory configured to store a computer program executable on the second processor; The second processor is configured to perform the method according to any one of claims 18 to 32 when running the computer program.
37. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, which implements the method according to any one of claims 1 to 17 when executed by a first processor, or implements the method according to any one of claims 18 to 32 when executed by a second processor.
38. A code stream, the code stream being generated by bit encoding based on information to be encoded; wherein, The information to be encoded includes at least one of the following: at least one of quantization parameter information or quantization parameter index, first syntax element identification information of the component to be filtered of the current frame or current slice, second syntax element identification information of the component to be filtered of the current block, third syntax element identification information of the current block contained in the current frame or the current slice, fourth syntax element identification information of the current video sequence, and filtered reconstructed values of each partitioned block included in the current frame or the current slice; wherein the current block is any one of the partitioned blocks.