Encoding method, decoding method, bitstream, encoder, decoder, and storage medium

By applying an enhanced model based on transformer network in video processing technology, the reconstruction value of image components is enhanced in quality, and the problems of insufficient encoding and decoding efficiency and video compression performance in the prior art are solved, and a more efficient encoding and decoding process and better video compression effect are achieved.

WO2025129410A1PCT designated stage expired Publication Date: 2025-06-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Patent Information

Application Number
PCT/CN2023/139615
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the existing video processing technology, the loop filtering module and the post-processing module perform filtering operations on the encoding and decoding ends respectively, which cannot effectively improve the encoding and decoding efficiency and video compression performance. Intra-frame post-processing technology does not consider global dependencies, and the encoding performance improvement is limited.

Method used

A codec is proposed. Using the enhancement model based on the transformer network structure, the reconstruction value of image components is enhanced in quality, and combined with prediction mode and preset quantization parameters, to improve the codec efficiency and video compression performance.

Benefits of technology

By applying the enhanced model during the encoding and decoding process, the reconstruction quality of image components is improved, more accurate reference frames are provided, and the encoding and decoding efficiency and video compression performance are effectively improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023139615_26062025_PF_FP_ABST
    Figure CN2023139615_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present application are an encoding method, a decoding method, a bitstream, an encoder, a decoder, and a storage medium. The decoding method comprises: decoding a bitstream at a decoding end, and determining decoding parameters of a current block, wherein the decoding parameters comprise a predicted residual, a prediction mode for the current block, an identification parameter, and a preset quantization parameter; on the basis of the prediction mode and the predicted residual, determining a first reconstruction value for an image component of the current block; and when the identification parameter indicates using an enhancement model to decode the current block, determining a reconstruction value for the image component of the current block on the basis of the enhancement model, wherein input parameters for the enhancement model comprise the first reconstruction value and the preset quantization parameter, and the enhancement model comprises a transformer network structure.
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method, code stream, encoder, decoder and storage medium Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium. Background Art

[0002] In the field of video processing technology, the loop filter module, at the encoder end, filters a frame of video after it has been reconstructed, improving its quality and providing a more accurate reference frame for inter-frame prediction of subsequent frames. The post-processing module, at the decoder end, filters a frame of reconstructed video, improving its quality at the decoder end. While both modules contribute to improved coding efficiency, the former filters a frame after it has been fully encoded, while the latter filters a frame after it has been fully decoded. Neither module considers enhancing the quality of a frame during encoding.

[0003] However, intra-frame post-processing technology only uses traditional convolutional neural networks to enhance the quality of reconstructed values ​​without considering the acquisition of global dependencies, resulting in limited improvement in encoding performance.

[0004] In other words, common quality enhancement methods in the encoding and decoding process cannot effectively improve encoding and decoding efficiency and video compression performance.

[0005] Summary of the Invention

[0006] The embodiments of the present application provide a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium, which can effectively improve coding and decoding efficiency and video compression performance.

[0007] The technical solution of the embodiment of the present application can be implemented as follows:

[0008] In a first aspect, an embodiment of the present application provides a decoding method, applied to a decoder, the method comprising:

[0009] Decoding the bitstream to determine decoding parameters of the current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter;

[0010] determining a first reconstructed value of an image component of the current block according to the prediction mode and the prediction residual;

[0011] When the identification parameter indicates that the current block is decoded using an enhancement model, a reconstruction value of the image component of the current block is determined based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure.

[0012] In a second aspect, an embodiment of the present application provides an encoding method, applied to an encoder, the method comprising:

[0013] Determining a first reconstructed value of an image component of the current block according to a prediction mode and a prediction residual of the current block;

[0014] Determining a reconstruction value of an image component corresponding to the current block based on the enhancement model; wherein input parameters of the enhancement model include the first reconstruction value and a preset quantization parameter, and the enhancement model includes a transformer network structure;

[0015] determining an identification parameter according to the first reconstruction value and the reconstruction value of the image component;

[0016] Writing the decoding parameters of the current block into a bitstream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter.

[0017] In a third aspect, an embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least: a prediction residual, a prediction mode of a current block, the identification parameter, and the preset quantization parameter.

[0018] In a fourth aspect, an embodiment of the present application provides an encoder, the encoder including a first determining unit; wherein,

[0019] The first determination part is configured to determine a first reconstruction value of an image component of the current block according to a prediction mode and a prediction residual of the current block; determine a reconstruction value of an image component corresponding to the current block based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and a preset quantization parameter, and the enhancement model includes a transformer network structure; determine an identification parameter according to the first reconstruction value and the reconstruction value of the image component; and write decoding parameters of the current block into a bitstream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter.

[0020] In a fifth aspect, an embodiment of the present application provides an encoder, the encoder including a first memory and a first processor; wherein,

[0021] The first memory is used to store a computer program that can be run on the first processor;

[0022] The first processor is configured to execute the encoding method described above when running the computer program.

[0023] In a sixth aspect, an embodiment of the present application provides a decoder, the decoder including a second determining unit; wherein,

[0024] The second determination unit is configured to decode the code stream and determine the decoding parameters of the current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; determine a first reconstruction value of the image component of the current block according to the prediction mode and the prediction residual; and determine the reconstruction value of the image component of the current block based on the enhancement model when the identification parameter indicates that the current block is to be decoded using an enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure.

[0025] In a seventh aspect, an embodiment of the present application provides a decoder, the decoder including a second memory and a second processor; wherein,

[0026] The second memory is used to store a computer program that can be run on the second processor;

[0027] The second processor is configured to execute the above-mentioned decoding method when running the computer program.

[0028] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements the decoding method as described in the first aspect, or implements the encoding method as described in the second aspect.

[0029] The embodiments of the present application provide a coding and decoding method, a code stream, an encoder, a decoder, and a storage medium. At the decoding end, the code stream is decoded to determine the decoding parameters of the current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; a first reconstruction value of an image component of the current block is determined based on the prediction mode and the prediction residual; when the identification parameter indicates that an enhancement model is used to decode the current block, the reconstruction value of the image component of the current block is determined based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure. At the encoding end, the first reconstruction value of the image component of the current block is determined based on the prediction mode and the prediction residual of the current block; the reconstruction value of the image component corresponding to the current block is determined based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure; the identification parameter is determined based on the first reconstruction value and the reconstruction value of the image component; and the decoding parameters of the current block are written into the code stream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter. It can be seen that in the embodiment of the present application, after completing the prediction of the image component of the current block based on the prediction mode and determining the corresponding first reconstruction value, it is possible to further use an enhancement model to perform quality enhancement processing on the first reconstruction value, wherein the enhancement model may include a transformer based on a multi-head self-attention mechanism. In other words, the encoding and decoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding process to enhance the quality of the reconstruction value of the image component, which is beneficial to improving the reconstruction quality of the image component and providing a more accurate reference for the intra-frame prediction of the subsequent coding blocks of the frame, thereby effectively improving the encoding and decoding efficiency and video compression performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] FIG1 is a block diagram of a video encoding system according to an embodiment of the present invention;

[0031] FIG2 is a block diagram of a video decoding system according to an embodiment of the present application;

[0032] Figure 3 is a common SmallNN network model structure diagram;

[0033] Figure 4 is a common algorithm flow chart;

[0034] Figure 5 is a neural network model based on the neuron attention mechanism;

[0035] Figure 6 is a schematic diagram of the network structure of Restormer;

[0036] FIG7 is a first schematic diagram of an implementation flow of a decoding method proposed in an embodiment of the present application;

[0037] FIG8 is a schematic diagram of the connection processing proposed in an embodiment of the present application;

[0038] FIG9 is a first structural diagram of a shallow layer extraction module proposed in an embodiment of the present application;

[0039] FIG10 is a second structural diagram of the shallow layer extraction module proposed in an embodiment of the present application;

[0040] FIG11 is a schematic diagram of shallow feature extraction proposed in an embodiment of the present application;

[0041] FIG12 is a first structural diagram of a deep extraction module proposed in an embodiment of the present application;

[0042] FIG13 is a second structural diagram of the deep extraction module proposed in an embodiment of the present application;

[0043] FIG14 is a schematic diagram of deep feature extraction proposed in an embodiment of the present application;

[0044] FIG15 is a schematic diagram of the structure of the self-attention mechanism module proposed in an embodiment of the present application;

[0045] FIG16 is a schematic diagram of the structure of the feedforward network module proposed in an embodiment of the present application;

[0046] FIG17 is a schematic diagram of the structure of the enhanced model proposed in an embodiment of the present application;

[0047] FIG18 is a second schematic diagram of the implementation flow of the decoding method proposed in an embodiment of the present application;

[0048] FIG19 is a schematic diagram of the encoding and decoding method proposed in an embodiment of the present application;

[0049] FIG20 is a schematic diagram of an implementation flow of the encoding method proposed in an embodiment of the present application;

[0050] FIG21 is a schematic diagram of the structure of an encoder proposed in an embodiment of the present application;

[0051] FIG22 is a schematic diagram of the specific hardware structure of the encoder proposed in an embodiment of the present application;

[0052] FIG23 is a schematic diagram of the structure of a decoder according to an embodiment of the present application;

[0053] FIG24 is a schematic diagram of the specific hardware structure of the decoder proposed in an embodiment of the present application;

[0054] FIG25 is a schematic diagram of the composition structure of the encoding and decoding system proposed in an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0057] In the following description, reference is made to "some embodiments," which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. It should also be noted that the terms "first, second, and third" in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that "first, second, and third" may be interchanged in a specific order or sequential order where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0058] In a video image, a first image component, a second image component, and a third image component are generally used to represent a coding block (CB); wherein the three image components are a luminance component, a blue chrominance component, and a red chrominance component, respectively. Specifically, the luminance component is usually represented by the symbol Y, the blue chrominance component is usually represented by the symbol Cb or U, and the red chrominance component is usually represented by the symbol Cr or V; thus, the video image can be represented in either the YCbCr format or the YUV format.

[0059] It's understandable that the goal of video coding is to minimize distortion in the reconstructed video while maintaining a constant bitrate, effectively improving its quality. Therefore, improving reconstructed video quality is a key issue addressed by the next-generation video coding standard, H.266 / Versatile Video Coding (VVC).

[0060] Referring to Figure 1, which shows a schematic block diagram of the composition of the video coding system proposed in an embodiment of the present application. As shown in Figure 1, the video coding system 10 includes a transform and quantization unit 101, an intra-frame estimation unit 102, an intra-frame prediction unit 103, a motion compensation unit 104, a motion estimation unit 105, an inverse transform and inverse quantization unit 106, a filter control analysis unit 107, a filtering unit 108, an encoding unit 109, and a decoded image cache unit 110, etc., wherein the filtering unit 108 can implement deblocking filtering and sample adaptive offset (SAO) filtering, and the encoding unit 109 can implement header information encoding and context-based adaptive binary arithmetic coding (CABAC).For the input original video signal, a video coding block can be obtained by dividing it into coding tree units (CTUs). Then, the residual pixel information obtained after intra-frame or inter-frame prediction is transformed by the transform and quantization unit 101, including transforming the residual information from the pixel domain to the transform domain and quantizing the obtained transform coefficients to further reduce the bit rate; the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to perform intra-frame prediction on the video coding block; specifically, the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to determine the intra-frame prediction mode to be used to encode the video coding block; the motion compensation unit 104 and the motion estimation unit 105 are used to perform inter-frame prediction coding on the received video coding block relative to one or more blocks in one or more reference frames to provide temporal prediction information; the motion estimation performed by the motion estimation unit 105 is the process of generating a motion vector, which can estimate the motion of the video coding block, and the motion compensation unit 104 then calculates the motion vector based on the motion vector determined by the motion estimation unit 105. After determining the intra-frame prediction mode, the intra-frame prediction unit 103 is further configured to provide the selected intra-frame prediction data to the encoding unit 109, and the motion estimation unit 105 also sends the calculated motion vector data to the encoding unit 109. In addition, the inverse transform and inverse quantization unit 106 is configured to reconstruct the video coding block and reconstruct a residual block in the pixel domain. The reconstructed residual block is subjected to the filter control analysis unit 107 and the filtering unit 108 to remove the block effect artifacts. The reconstructed residual block is then added to a predictive block in the frame of the decoded image buffer unit 110 to generate a reconstructed video coding block. The encoding unit 109 is configured to encode various coding parameters and quantized transform coefficients. In the CABAC-based coding algorithm, the context content can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-frame prediction mode, and output the code stream of the video signal. The decoded image buffer unit 110 is configured to store the reconstructed video coding block for prediction reference. As the video image encoding proceeds, new reconstructed video encoding blocks are continuously generated, and these reconstructed video encoding blocks are stored in the decoded image buffer unit 110 .

[0061] Refer to Figure 2, which shows a block diagram of the composition of the video decoding system proposed in an embodiment of the present application. As shown in Figure 2, the video decoding system 20 includes a decoding unit 201, an inverse transform and inverse quantization unit 202, an intra-frame prediction unit 203, a motion compensation unit 204, a filtering unit 205 and a decoded image cache unit 206, etc., wherein the decoding unit 201 can implement header information decoding and CABAC decoding, and the filtering unit 205 can implement deblocking filtering and SAO filtering. After the input video signal is encoded and processed in Figure 1, the code stream of the video signal is output; the code stream is input into the video decoding system 20, and first passes through the decoding unit 201 to obtain the decoded transform coefficients; the transform coefficients are processed by the inverse transform and inverse quantization unit 202 to generate residual blocks in the pixel domain; the intra-frame prediction unit 203 can be used to generate prediction data for the current video decoding block based on the determined intra-frame prediction mode and the data of the previously decoded block from the current frame or picture; the motion compensation unit 204 is to determine the prediction information for the video decoding block by analyzing the motion vector and other associated syntax elements, and use The prediction information is used to generate a predictive block for the video decoding block being decoded; a decoded video block is formed by summing the residual block from the inverse transform and inverse quantization unit 202 with the corresponding predictive block generated by the intra-frame prediction unit 203 or the motion compensation unit 204; the decoded video signal passes through the filtering unit 205 to remove blocking artifacts, thereby improving video quality; the decoded video block is then stored in the decoded image buffer unit 206, which stores reference images used for subsequent intra-frame prediction or motion compensation, and is also used for outputting the video signal, thereby obtaining the restored original video signal.

[0062] Currently, the VVC coding framework involves two modules involved in enhancing the quality of reconstructed video: the loop filter module and the post-processing module. These modules enhance the quality of the reconstructed video after the video frame sequence is encoded, but cannot provide a more accurate reference for intra-frame prediction. Intra-frame post-processing technology has emerged to address this issue. It proposes enhancing the quality of reconstructed coding tree units (CTUs) during the encoding of the encoded frame sequence, providing an effective reference for intra-frame prediction. The following briefly introduces these three technologies.

[0063] (1) Loop filter module

[0064] The in-loop filter module is used to filter the reconstructed frames at the encoding end. This operation not only helps to improve the quality of the current reconstructed video frame, but also helps to provide a more accurate reference frame for inter-frame prediction of subsequent video frames.

[0065] The position of the loop filtering module in the VVC coding framework is the filtering unit 108 in Figure 1.

[0066] (2) Post-processing module

[0067] The function of the post-processing module is to filter the reconstructed video frames at the decoding end. This operation helps to improve the quality of the reconstructed video at the decoding end.

[0068] For example, some technologies have proposed content-adaptive post-processing filters based on neural networks, wherein a network model SmallNN and a network model BigNN are proposed for post-processing of Class A and Classes B, C, D, and F, respectively.

[0069] Figure 3 is a common SmallNN network model structure diagram. As shown in Figure 3, the input of SmallNN consists of the luminance component Y and the chrominance components Cb, Cr and the normalized QP. The first module includes a convolution layer (excluding bias), a bias layer and a nonlinear activation function LeakyReLU. The structures of the following four modules are the same as that of the first module. The blocks are connected by skip connections. Each of the first five modules includes 64 convolution kernels. The last module is similar to the first module and includes 3 convolution kernels.

[0070] BigNN has a similar structure to SmallNN, the only difference is that each of the first five modules includes 512 convolution kernels.

[0071] Figure 4 is a common algorithm flow chart. As shown in Figure 4, the algorithm flow is roughly divided into three steps: the first step is the pre-training of the network model; the second step is to fine-tune the pre-trained network model at the encoding end; and the third step is the specific encoding and decoding operations.

[0072] The first step is to pre-train the network model offline using the BVI-DVC dataset.

[0073] The second step involves online fine-tuning the pre-trained network model from the first step using the input video sequence. After fine-tuning, the weight update is calculated as the difference between the bias of the pre-trained network parameters and the bias of the fine-tuned network parameters. Finally, the weight update is compressed and encoded using the LZMA2 algorithm.

[0074] The content of the third step is as follows: At the decoding end, the weight update encoded in the second step is first decoded and added to the bias term of the original neural network model to obtain a neural network-based post-processing filter.

[0075] (3) Intra-frame post-processing technology

[0076] In previous work, a neural network model based on the neuron attention mechanism was proposed to enhance the quality of the luminance component of the reconstructed CTU during the encoding process, thereby improving the encoding efficiency.

[0077] Figure 5 shows a neural network model based on the neuron attention mechanism. As shown in Figure 5, the NACNN (Neuron Attention-based CNN) neural network model based on the neuron attention mechanism enhances the quality of the luminance component of the CTU. The network model inputs the CTU luminance component I1 and the normalized QP value, and outputs the quality-enhanced luminance component O1. The network model consists of a shallow feature extraction module, a deep feature extraction module, and a reconstruction module.

[0078] First, in the shallow feature extraction module, two consecutive convolution operations are used to perform simple feature extraction on the input luminance component, preparing for the subsequent extraction of deep features. Second, in the deep feature extraction module, multiple cascaded multi-scale and neuron attention (MSNA) modules are used to extract and weight the output features of the shallow feature extraction module to better explore the relationship between the reconstructed and true luminance components of the input. Finally, in the reconstruction module, the features output by the deep feature extraction module are further nonlinearly mapped to obtain the final quality-enhanced luminance component.

[0079] The implementation process of this algorithm can be roughly divided into three steps:

[0080] Step 1: Online neural network training. Encode and decode the original video sequence to obtain a compressed video sequence and construct a training dataset. Using the PyTorch framework, the neural network is trained using the training dataset.

[0081] Step 2: Embed the neural network into the test platform. Using the libtorch library, embed the trained neural network into the encoder and decoder of the VVC Test Model (VTM) software test platform.

[0082] Step 3: Offline neural network testing: Use the codec embedded with the neural network to perform encoding and decoding operations on the standard test sequences in the general test conditions specified by the video coding standard to analyze the encoding performance.

[0083] With the rapid development of computers and storage devices in recent years, deep learning-based image processing technology has seen significant development, with image quality enhancement techniques based on convolutional neural networks and transformers gaining significant attention. Image restoration, denoising, and enhancement refer to the process of recovering the original image from a damaged or distorted image. The goal of image restoration, denoising, and enhancement is to minimize or remove these impairments, remove noise and interfering signals, and restore the true information in the image, making the image closer to the original in terms of quality and information content, thereby improving image quality and usability. Transformers have recently been widely used in image restoration, denoising, and enhancement tasks, with Restormer being the most effective.

[0084] Figure 6 shows the Restormer network architecture. As shown in Figure 6, the Transformer (Restormer) network model for image restoration enhances the quality of degraded, high-resolution images and restores their original information. In the network model, the image input consists of three color channels: RGB, and the output is also in three color spaces. The network model consists of several stacked transformer blocks and a series of upsampling and downsampling operations.

[0085] First, the network input undergoes a simple feature extraction operation using a 3×3 convolution. Next, the extracted features are fed into a stack of multi-layered transformer blocks. Connections are introduced between the layers of the transformer blocks to facilitate the aggregation of deep features from different levels. The resulting deep features are then fused using a 3×3 convolution. Finally, the fused features are added to the network input to produce the final, quality-enhanced network output.

[0086] As can be seen, the transformer block is the core module of the algorithm, consisting of an attention mechanism module and a feedforward network. Depth-wise convolution is introduced within the attention mechanism module, enabling convolution operations on a single channel, effectively reducing the number of parameters. Furthermore, by appropriately adjusting the size of the query (Q), key (K), and value (V), the calculation of the transposed attention mechanism weighting matrix is ​​implemented, further reducing the number of parameters while maintaining the global receptive field.

[0087] It's important to note that both the loop filtering module and the post-processing module improve the quality of reconstructed video frames. The former, on the encoding side, filters a frame after it's been reconstructed, improving its quality and providing a more accurate reference frame for inter-frame prediction of subsequent frames. The latter, on the decoding side, filters a frame of reconstructed video, improving its quality at the decoding end. While both modules contribute to improved coding efficiency, the former filters a frame after it's fully encoded, while the latter filters it after it's fully decoded. Neither module considers enhancing the quality of a frame during encoding.

[0088] Intra-frame post-processing technology only uses traditional convolutional neural networks to enhance the quality of reconstructed CTUs, without considering the acquisition of global dependencies, resulting in limited improvement in coding performance. In addition, the Peak Signal to Noise Ratio (PSNR) judgment condition is used without considering the impact of the coding flag bit on the bit rate, which is even more detrimental to further improvement of coding performance.

[0089] However, the common transformer used for image denoising and restoration does not take into account the actual quantization distortion in the encoding and decoding process, and cannot effectively achieve quality enhancement in the encoding and decoding process.

[0090] To solve the above problems, embodiments of the present application provide a coding and decoding method, a bitstream, an encoder, a decoder, and a storage medium. At the decoding end, the bitstream is decoded to determine decoding parameters of the current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; a first reconstruction value of an image component of the current block is determined based on the prediction mode and the prediction residual; when the identification parameter indicates that an enhancement model is used to decode the current block, the reconstruction value of the image component of the current block is determined based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure. At the encoding end, the first reconstruction value of the image component of the current block is determined based on the prediction mode and the prediction residual of the current block; the reconstruction value of the image component corresponding to the current block is determined based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure; the identification parameter is determined based on the first reconstruction value and the reconstruction value of the image component; and the decoding parameters of the current block are written into the bitstream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter. It can be seen that in the embodiment of the present application, after completing the prediction of the image component of the current block based on the prediction mode and determining the corresponding first reconstruction value, it is possible to further use an enhancement model to perform quality enhancement processing on the first reconstruction value, wherein the enhancement model may include a transformer based on a multi-head self-attention mechanism. In other words, the encoding and decoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding process to enhance the quality of the reconstruction value of the image component, which is beneficial to improving the reconstruction quality of the image component and providing a more accurate reference for the intra-frame prediction of the subsequent coding blocks of the frame, thereby effectively improving the encoding and decoding efficiency and video compression performance.

[0091] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0092] It should be noted that the embodiments of the present application can be applied to both encoders and decoders, and can even be applied to both encoders and decoders at the same time, but this is not specifically limited here.

[0093] It should also be noted that when the method of the embodiment of the present application is applied to the encoder, the "current block" specifically refers to the encoding block in the image to be encoded that is currently to be intra-frame predicted; when the method of the embodiment of the present application is applied to the decoder, the "current block" specifically refers to the decoding block in the image to be decoded that is currently to be intra-frame predicted.

[0094] FIG7 is a schematic diagram of a first implementation flow of a decoding method proposed in an embodiment of the present application. As shown in FIG7 , the decoding method performed by the decoder may include the following steps:

[0095] Step 101: Decode the code stream to determine decoding parameters of the current block, wherein the decoding parameters include prediction residual, prediction mode of the current block, identification parameters, and preset quantization parameters.

[0096] In an embodiment of the present application, the decoder may first decode the code stream to determine decoding parameters of the current block, wherein the decoding parameters may be used to perform decoding processing on the current block.

[0097] It should be noted that a video image can be divided into multiple image blocks. Each image block to be decoded can be referred to as a decoding block, and the current block here specifically refers to the decoding block currently to be predicted. The current block can be a CTU, or even a coding unit (CU), prediction unit (PU), etc., and this embodiment of the application does not impose any limitation.

[0098] It should be noted that, in the embodiment of the present application, the decoding parameters of the current block may at least include the following parameters: prediction residual, prediction mode of the current block, identification parameter, and preset quantization parameter.

[0099] Furthermore, in an embodiment of the present application, the decoding method may be applied to an intra-frame decoding process. Specifically, the decoding method may include a post-processing method for intra-frame reconstruction values.

[0100] It can be understood that, in the embodiments of the present application, the prediction mode of the current block can be any one of the intra-frame prediction modes.

[0101] It can be understood that, in the embodiment of the present application, the identification parameter can be used to determine whether to use the enhancement model to further decode the current block.

[0102] It is understood that in the embodiment of the present application, the preset quantization parameter may be a normalized quantization parameter (QP), wherein the value of the preset quantization parameter may be any value and is not specifically limited in the present application.

[0103] For example, in some embodiments, the preset quantization parameter value, ie, QP, can be set to 22, 27, 32, 37, or 42.

[0104] Step 102: Determine a first reconstructed value of an image component of the current block according to the prediction mode and the prediction residual.

[0105] In an embodiment of the present application, after decoding the code stream and determining the decoding parameters of the current block including the prediction mode of the current block, the decoder can further determine the first reconstructed value of the image component of the current block based on the prediction mode and prediction residual of the current block.

[0106] It can be understood that, in the embodiment of the present application, the image component of the current block may be the chrominance component of the current block or the luminance component of the current block, and the present application does not make any specific limitation.

[0107] Exemplarily, in some embodiments, in a video image, a first image component, a second image component, and a third image component are generally used to represent a coding block (Coding Block, CB); wherein the three image components are a luminance component, a blue chrominance component, and a red chrominance component, respectively. Specifically, the luminance component is usually represented by the symbol Y, the blue chrominance component is usually represented by the symbol Cb or U, and the red chrominance component is usually represented by the symbol Cr or V; in this way, the video image can be represented in YCbCr format or in YUV format.

[0108] Furthermore, in an embodiment of the present application, the decoder may also decode the code stream to determine the prediction residual of the image component of the current block.

[0109] Furthermore, in an embodiment of the present application, when determining the first reconstruction value of the image component of the current block based on the prediction mode and the prediction residual, the prediction value of the image component of the current block can be first determined based on the prediction mode; and then the first reconstruction value can be determined based on the prediction value and the prediction residual.

[0110] Exemplarily, in some embodiments, the sum of the prediction value of the image component of the current block and the prediction residual of the corresponding image component may be determined as the first reconstructed value of the image component.

[0111] It can be understood that in the embodiments of the present application, the first reconstruction value can be understood as the initial reconstruction value of the image component of the current block, that is, the first reconstruction value is the reconstruction value of the image component directly determined based on the prediction mode and prediction residual of the current block.

[0112] Step 103: When the identification parameter indicates that the current block is decoded using an enhancement model, a reconstruction value of the image component of the current block is determined based on the enhancement model; wherein the input parameters of the enhancement model include a first reconstruction value and a preset quantization parameter, and the enhancement model includes a transformer network structure.

[0113] In an embodiment of the present application, after decoding the code stream and determining the decoding parameters of the current block including the identification parameters and the preset quantization parameters, if it is determined to use the enhancement model based on the identification parameters, that is, when the identification parameters indicate to use the enhancement model to decode the current block, the reconstructed value of the image component of the current block can be further determined based on the enhancement model.

[0114] It should be noted that, in the embodiment of the present application, the enhancement model can be used to perform quality enhancement on the first reconstruction value of the image component of the current block, thereby obtaining the enhanced reconstruction value of the image component.

[0115] Furthermore, in an embodiment of the present application, the input parameters of the enhancement model may include a first reconstruction value and a preset quantization parameter. After the first reconstruction value of the image component of the current block is quality enhanced by the enhancement model, an enhanced reconstruction value may be obtained, that is, the output result of the enhancement model may be the reconstruction value of the image component of the current block.

[0116] It should be noted that, in the embodiments of the present application, the enhanced model may include a transformer network structure.

[0117] Furthermore, in the embodiments of the present application, the transformer is a deep learning model that is different from convolutional neural networks and is widely used in fields such as natural language processing (NLP). Specifically, the transformer uses a self-attention mechanism to simultaneously focus on all positions in the input sequence, thereby better understanding the context of the sequence data.

[0118] It can be understood that in the embodiments of the present application, the enhancement model can be understood as a transformer-based reconstructed CTU quality enhancement network, which is used for intra-frame post-processing of the reconstructed CTU to improve the quality of the reconstructed CTU, thereby improving the prediction accuracy of subsequent coding blocks and ultimately improving the coding performance.

[0119] That is, in the embodiments of the present application, an enhancement model including a transformer network structure can be used for intra-frame post-processing of reconstructed CTUs. The enhancement model leverages the transformer's feature extraction and global attention mechanism to enhance the quality of the current CTU, thereby improving the intra-frame prediction accuracy of subsequent CTUs.

[0120] Exemplarily, in some embodiments, the enhancement model may be composed of a cascade of a shallow extraction module and a deep extraction module; wherein both the shallow extraction module and the deep extraction module include a transformer network structure.

[0121] It should be noted that, in the embodiment of the present application, the shallow layer extraction module in the enhancement model can be used to extract shallow features. Specifically, the shallow layer extraction module is used to perform multi-level and multi-scale feature extraction processing.

[0122] It should be noted that, in the embodiment of the present application, the deep layer extraction module in the enhanced model can be used to extract deep features. Specifically, the deep layer extraction module is used to perform multi-level and multi-scale feature extraction and feature fusion processing.

[0123] It can be understood that, in the embodiment of the present application, the enhanced model may further include a first connection layer.

[0124] Furthermore, in an embodiment of the present application, when determining the reconstruction value of the image component of the current block based on the enhancement model, the first reconstruction value and the preset quantization parameter can be first spliced ​​through the first connection layer to determine the first spliced ​​component; then the first spliced ​​component is input into the shallow extraction module to output the first feature information; then the first feature information is input into the deep extraction module to output the second feature information; finally, the reconstruction value of the image component can be determined based on the first reconstruction value and the second feature information.

[0125] That is to say, in an embodiment of the present application, for the first reconstruction value of the input image component, shallow feature extraction can be first performed through a shallow extraction module to obtain first feature information of the image component; then, based on the first feature information, deep feature extraction can be further performed through a deep extraction module to obtain second feature information of the image component; finally, the first reconstruction value and the second feature information can be combined to complete the quality enhancement of the image component of the current block and obtain the reconstruction value of the image component.

[0126] It should be noted that, in an embodiment of the present application, for the first connection layer in the enhancement model, its input may be the first reconstruction value and preset quantization parameter of the image component of the current block, and its output may be the corresponding first spliced ​​component.

[0127] It should be noted that in the embodiments of the present application, in order to fully reduce the impact of the quantization operation on the quality of encoding the current block, the quantization parameter QP can be considered as prior information for quality enhancement, so the input of the first connection layer is the normalized reconstruction value (first reconstruction value) and the normalized quantization parameter (preset quantization parameter).

[0128] Exemplarily, in some embodiments, assuming that the image component is a Y component, the corresponding first reconstruction value is the normalized reconstructed Y component I′1∈R 1×128×128 , the preset quantization parameter is QP′∈R composed of normalized QP 1×128×128 .

[0129] For example, in some embodiments, FIG8 is a schematic diagram of the connection processing proposed in an embodiment of the present application. As shown in FIG8 , assuming that the image component is a Y component, the first reconstruction value I′1 and the preset quantization parameter QP′ are input to the first connection layer, the two normalized components are spliced ​​on the channel, and the corresponding first spliced ​​component is output, wherein the first spliced ​​component can be the spliced ​​normalized component I1∈R 2×128×128 .

[0130] Furthermore, in an embodiment of the present application, the shallow extraction module includes a first convolutional layer, a first feature extraction module and a second feature extraction module, wherein the first feature extraction module and the second feature extraction module both include a transformer network structure.

[0131] Exemplarily, in some embodiments, Figure 9 is a structural schematic diagram of the shallow extraction module proposed in an embodiment of the present application. As shown in Figure 9, the shallow extraction module may include a first convolutional layer, a first feature extraction module and a second feature extraction module, wherein the first convolutional layer may be a 3×3 convolution, and the first feature extraction module and the second feature extraction module respectively include two transformer network structures, that is, the first feature extraction module and the second feature extraction module can constitute 4 cascaded transformer blocks.

[0132] It should be noted that, in an embodiment of the present application, the shallow extraction module may further include a ReLU activation function and a downsampling operation.

[0133] Exemplarily, in some embodiments, Figure 10 is a second structural schematic diagram of the shallow extraction module proposed in an embodiment of the present application. As shown in Figure 10, the shallow extraction module may include a first convolutional layer, a ReLU activation function, a first feature extraction module and a second feature extraction module, wherein the first feature extraction module and the second feature extraction module respectively include a downsampling operation.

[0134] Furthermore, in an embodiment of the present application, when the first spliced ​​component is input into the shallow extraction module and the first feature information is output, the first feature can be first determined through the first convolutional layer based on the first spliced ​​component; then the first feature can be input into the first transformer network in the first feature extraction module to output the second feature; then based on the second feature, the third feature can be determined through the second transformer network in the first feature extraction module; then the third feature can be input into the third transformer network in the second feature extraction module to output the fourth feature; finally, based on the fourth feature, the first feature information can be determined through the fourth transformer network in the second feature extraction module.

[0135] For example, in some embodiments, FIG11 is a schematic diagram of shallow feature extraction proposed in an embodiment of the present application. As shown in FIG11 , after the first reconstruction value and the preset quantization parameter are concatenated through the first connection layer and the first concatenated component I1 is output, the first feature M0∈R can be determined by sequentially passing through the first convolution layer and the ReLU activation function. 32×128×128 , that is, the spliced ​​normalized component I1 is input into a single 3×3 convolution layer. After 3×3 convolution and nonlinear mapping operations, the shallow feature extraction and the increase of the number of channels are realized to obtain the shallow feature M0. Then, with the help of downsampling operation, the first feature M0 is input into the four cascaded transformer networks for feature extraction, and the second features M1∈R of different levels and scales are obtained respectively. 32×128×128 , the third feature M2∈R 64×64×64 , the fourth feature M3∈R 64×64×64 , first feature information M4∈R 128×32×32 Among them, the four transformer networks are each processed in pairs, with a total of two stages (i.e., the first feature extraction module and the second feature extraction module) to extract features of different levels and scales from the input features.

[0136] It should be noted that in the embodiment of the present application, when determining the first feature through the first convolution layer based on the first spliced ​​component, the first spliced ​​component, that is, the spliced ​​normalized component I1, can be subjected to a 3×3 convolution and a nonlinear mapping operation to obtain the first feature, that is, the shallow feature M0, as follows: M0(I1)=σ(W1(I1)) (1)

[0137] Among them, σ represents the ReLU activation function and W1 represents the 3×3 convolution kernel.

[0138] It should be noted that in an embodiment of the present application, when feature extraction of different levels and scales is performed based on the first feature extraction module and the second feature extraction module with the help of a downsampling operation, the first feature M0 can be first input into the first transformer network in the first feature extraction module to output the second feature M1. Then, based on the second feature M1, the third feature M2 is determined through the downsampling operation and the second transformer network in the first feature extraction module in sequence, and then the third feature M2 is input into the third transformer network in the second feature extraction module to output the fourth feature M3. Then, based on the fourth feature M3, the first feature information M4 is determined through the downsampling operation and the fourth transformer network in the second feature extraction module in sequence.

[0139] For example, in some embodiments, the first feature M0 can be input into the multi-level feature extraction module of the first stage, and passed through the first transformer block, i.e., the first transformer network in the first feature extraction module, to obtain feature M1; feature M1 is downsampled, and the width and height of the feature are reduced to half of the original size, and its channel dimension is doubled. The downsampled feature M1 is then input into the second transformer block, i.e., the second transformer network in the first feature extraction module, to obtain feature M2. Where M1 and M2 are features of different levels and scales, and the acquisition process is as follows:

[0140] Here, Transformer(·) represents the Transformer operation, W2 represents the 3×3 convolution kernel, Pixelunshuffle(·) represents the inverse pixel shuffle operation, and downs□mple(·) represents the downsampling operation. The downsampling operation consists of a 3×3 convolution operation with half the number of channels, followed by an inverse pixel shuffle operation with a sampling factor of 2 (pixelunshuffle).

[0141] For example, in some embodiments, feature M2 is input into the second-stage multi-level feature extraction module to continue feature extraction at different levels. After passing through the first transformer block, i.e., the third transformer network in the second feature extraction module, feature M3 is obtained; feature M3 is still downsampled, with its width and height dimensions reduced to half of the original size, and its channel dimension doubled. The downsampled feature M3 is then input into the second transformer block, i.e., the fourth transformer network in the second feature extraction module, to obtain feature M4. The specific acquisition process is as follows:

[0142] Furthermore, in an embodiment of the present application, the deep extraction module includes a second connection layer, a third connection layer, a fourth connection layer, a third feature extraction module, a fourth feature extraction module, a second convolutional layer, and a third convolutional layer, wherein the third feature extraction module and the fourth feature extraction module both include a transformer network structure.

[0143] Exemplarily, in some embodiments, Figure 12 is a structural schematic diagram of the deep extraction module proposed in an embodiment of the present application. As shown in Figure 12, the second convolution layer in the deep extraction module can be a 1×1 convolution, the third convolution layer can be a 3×3 convolution, and the third feature extraction module and the fourth feature extraction module respectively include two transformer network structures, that is, the third feature extraction module and the fourth feature extraction module can constitute 4 cascaded transformer blocks, and the second connection layer, the third connection layer, and the fourth connection layer can be respectively placed before the three transformer networks.

[0144] It should be noted that, in an embodiment of the present application, the deep extraction module may further include a ReLU activation function and an upsampling operation.

[0145] Exemplarily, in some embodiments, Figure 13 is a second structural diagram of the deep extraction module proposed in an embodiment of the present application. As shown in Figure 13, the third feature extraction module and the fourth feature extraction module in the deep extraction module can respectively include an upsampling operation, and the deep extraction module can include a ReLU activation function.

[0146] Furthermore, in an embodiment of the present application, when the first feature information is input into the deep extraction module and the second feature information is output, the first feature information and the fourth feature can be first input into the second connection layer to output the second spliced ​​component; then, based on the second spliced ​​component, the fifth feature can be determined through the second convolutional layer and the fifth transformer network in the third feature extraction module; then, the fifth feature and the third feature can be input into the third connection layer to output the third spliced ​​component; then, based on the third spliced ​​component, the sixth feature can be determined through the second convolutional layer and the sixth transformer network in the third feature extraction module; then, the sixth feature and the second feature can be input into the fourth connection layer to output the fourth spliced ​​component; then, based on the fourth spliced ​​component, the seventh feature can be determined through the second convolutional layer and the seventh transformer network in the fourth feature extraction module; then, the seventh feature can be input into the eighth transformer network in the fourth feature extraction module to output the eighth feature; finally, based on the eighth feature, the second feature information can be determined through the third convolutional layer.

[0147] Exemplarily, in some embodiments, FIG14 is a schematic diagram of the deep feature extraction proposed in the embodiment of the present application. As shown in FIG14, after completing the shallow feature extraction and obtaining the second feature M1, the third feature M2, the fourth feature M3, and the first feature information M4 of different levels and scales, the first feature information M4 can be upsampled first, and then the fourth feature M3 and the upsampled first feature information M4∈ are channel-spliced ​​using the second connection layer to obtain a second spliced ​​component, and then the second spliced ​​component is input into the fifth transformer network in the third feature extraction module, and the fifth feature M5∈R is output. 64×64×64 ; Then, the third connection layer can be used to perform channel splicing on the third feature M2 and the fifth feature M5 to obtain the third spliced ​​component, and then the third spliced ​​component is input into the sixth transformer network in the third feature extraction module to output the sixth feature M6∈R 64×64×64 ; Then, the sixth feature M6 can be upsampled first, and the fourth connection layer can be used to perform channel splicing on the second feature M1 and the upsampled sixth feature M6 to obtain the fourth spliced ​​component, and then the fourth spliced ​​component is input into the seventh transformer network in the fourth feature extraction module to obtain the seventh feature M7∈R 32×128×128 ; Then the seventh feature M7 can be input into the eighth transformer network in the fourth feature extraction module, and the eighth feature M8∈R 32×128×128 Finally, based on the eighth feature, the second feature information, namely the quality-enhanced brightness component C1∈R 1×128×128 .

[0148] It can be understood that in the embodiment of the present application, the role of the deep extraction module is to use the same number of cascaded transformer blocks to perform channel splicing on features M1, M2, M3, and M4 of different levels and scales, and then continue to perform multi-scale feature extraction and feature weighting to complete the deep mining of connection features and the fusion of multi-level features to better restore the lost detail information and thus reconstruct the image components.

[0149] It should be noted that in the embodiments of this application, the operations in the deep extraction module can be understood as the inverse operations in the shallow extraction module. The four transformer networks in the deep extraction module still process two by two at a time, extracting features at different levels and scales from the input features in two stages, resulting in features of different scales.

[0150] It can be understood that in the embodiment of the present application, the deep extraction module is composed of a two-stage multi-level feature fusion module and a single convolutional layer, wherein the two-stage multi-level feature fusion module realizes the fusion of features of different levels and sizes, and the single 3×3 convolutional layer realizes the deep fusion of features and the reduction of the number of channels to form the final residual CTU.

[0151] For example, in some embodiments, the first feature information M4 is further input into four cascaded transformer networks for feature extraction. Upsampling is performed to achieve scale matching between pairs of features, and the pairwise scale-matched features are then spliced ​​on the channel. The spliced ​​components are then input into four cascaded transformer networks for feature extraction at different levels and scales, thereby obtaining features M5, M6, M7, and M8 at different levels and scales.

[0152] Exemplarily, in some embodiments, feature M4 is first upsampled, the width and height of the feature are doubled, and its channel dimension is halved; and the upsampled feature M4 is spliced ​​with feature M3 in the channel dimension; the spliced ​​features are fused and the number of channels is halved by means of 1×1 convolution, and the spliced ​​features with halved channels are input into the multi-level feature extraction module of the third stage, and pass through the first transformer block, i.e., the fifth transformer network in the third feature extraction module, to obtain feature M5; after channel splicing of feature M5 and feature M2, the spliced ​​features are fused and the number of channels is halved by means of 1×1 convolution, and input into the second transformer block, i.e., the sixth transformer network in the third feature extraction module, to obtain feature M6. The specific process is as follows:

[0153] Here, Transformer (·) represents the Transformer operation, W3 represents the 3×3 convolution kernel, Pixelunshuffle (·) represents the inverse pixel shuffle operation, and upsample (·) represents the upsampling operation. The upsampling operation consists of a 3×3 convolution operation that halves the number of channels, followed by a pixel shuffle operation with a sampling factor of 2.

[0154] For example, in some embodiments, feature M6 is upsampled, with its width and height doubled and its channel dimension halved. The upsampled feature M6 is then concatenated with feature M1 in terms of the channel dimension. 1×1 convolution is used to fuse the concatenated features and halve the number of channels. The concatenated features with halved channels are then input into the fourth-stage multi-level feature extraction module, passing through the first transformer block, i.e., the seventh transformer network in the fourth feature extraction module, to obtain feature M7. Feature M7 is then input into the second transformer block, i.e., the eighth transformer network in the fourth feature extraction module, to obtain feature M8. The specific process is as follows:

[0155] Among them, W4 represents a 3×3 convolution kernel.

[0156] For example, in some embodiments, after obtaining the eighth feature, that is, the deep feature M8, the deep feature M8 is input into a single 3×3 convolution layer, and after 3×3 convolution and nonlinear mapping operations, the final fusion of features and reduction of the number of channels are achieved to obtain the final second feature information, that is, the brightness component after quality enhancement. The formula is as follows: C1=σ(W2(M8)) (6)

[0157] Among them, σ represents the ReLU activation function and W2 represents the 3×3 convolution kernel.

[0158] Furthermore, in an embodiment of the present application, when determining the reconstruction value of an image component based on the first reconstruction value and the second feature information, the second feature information extracted by the shallow extraction module and the deep extraction module in the enhancement model can be used to perform quality enhancement on the initial first reconstruction value, thereby obtaining an enhanced image component, that is, the reconstruction value of the image component.

[0159] For example, in some embodiments, the second feature information C1 may be added to the input brightness component I′1 to obtain the final output reconstruction value O1∈R 1×128×128 The process is obtained from the following formula: O1=C1+I′1 (7)

[0160] Furthermore, in an embodiment of the present application, the transformer network structure may include a self-attention mechanism module and a feedforward network.

[0161] That is to say, in the embodiments of the present application, for any transformer network in the shallow extraction module and the deep extraction module, it is composed of a cascade of a self-attention module and a feedforward neural network (FFN). Among them, the output of the attention mechanism module in the current transformer block is the input of the feedforward network, and the output of the feedforward network is the final output of the current transformer block. The attention mechanism module is used to realize feature extraction and weighted fusion, and the feedforward network is used to further perform feature fusion and nonlinear mapping on the weighted fusion features obtained in the attention mechanism module.

[0162] Exemplarily, in some embodiments, FIG15 is a structural diagram of the self-attention mechanism module proposed in the embodiment of the present application. As shown in FIG15 , the self-attention mechanism module is used to realize feature extraction and weighted fusion. The module consists of three parts: layer normalization operation, two-head attention mechanism and single 1×1 convolution operation. The layer normalization operation completes the normalization of the input features in the channel dimension, ensuring that the feature point values ​​of the input features are always within a reasonable range, creating a prerequisite for obtaining the subsequent residual features; the two-head attention mechanism module divides the features into two on the channel, and calculates the weighted features of the two parts in the upper and lower branches respectively, and then splices the two weighted features on the channel to obtain the summary residual features. In addition, the attention matrix transpose calculation operation is introduced in the attention mechanism module, which not only realizes multi-level feature extraction, but also effectively reduces the number of parameters; the single 1×1 convolution operation completes the change in the number of channels of the summary residual features. The specific processing process is as follows:

[0163] The first step is to input features i∈[1,8] is normalized in the channel dimension, and the normalized features are divided into two equal parts in the channel dimension to obtain Where C is the number of channels, H is the height, and W is the width.

[0164] The second step is to and Using three 1×1 and 3×3 convolution operations, we get In each convolution operation, unless otherwise specified, the feature size and number of channels do not change before and after processing.

[0165] The third step is to merge and transpose the above features to achieve dimension change and obtain Will After matrix multiplication, the Softmax function is used for normalization to obtain the attention matrix Similarly, the attention matrix can be obtained Then use and Perform matrix calculations to realize matrix The weighted operation of Similarly, we can get The weighted features are then spliced ​​in the channel dimension and the dimension is changed to obtain

[0166] Among them, Softmax(·) represents the softmax activation function, · represents the matrix multiplication operation, α is a learnable scaling parameter used to adjust the size of the matrix product, and Reshape(·) represents the dimension change operation on the matrix.

[0167] The fourth step is to input feature M i-1 and features Add them together to get the output of the self-attention mechanism module

[0168] Furthermore, in an embodiment of the present application, the output of the attention mechanism module can be input into the feedforward network module for further processing, and finally the output of the transformer block can be obtained.

[0169] Exemplarily, in some embodiments, FIG16 is a schematic diagram of the structure of the feedforward network module proposed in the embodiment of the present application. As shown in FIG16 , the feedforward network is used to further perform feature fusion and nonlinear mapping on the weighted fusion features obtained in the attention mechanism module. This module draws on the design framework of the channel attention mechanism and implements further feature extraction and weighted fusion of the input features of the module in two branches. The input of this module is the output of the attention mechanism module. The output of the attention mechanism module is input into the feedforward network module for further processing to obtain the output of the feedforward network, which is the final output M of the current transformer block. i '. The specific processing process is as follows:

[0170] The first step is to input features i∈[1,8] is normalized in the channel dimension, and the normalized features are divided into two equal parts in the channel dimension to obtain Where C is the number of channels, H is the height, and W is the width.

[0171] The second step is to and Using three 1×1 and 3×3 convolution operations, we get In each convolution operation, unless otherwise specified, the feature size and number of channels do not change before and after processing.

[0172] The third step is to merge and transpose the above features to achieve dimension change and obtain Will After matrix multiplication, the Softmax function is used for normalization to obtain the attention matrix Similarly, the attention matrix can be obtained Then use and Perform matrix calculations to realize matrix The weighted operation of Similarly, we can get The weighted features are then spliced ​​in the channel dimension and the dimension is changed to obtain As shown in formula (8), Softmax(·) represents the softmax activation function, · represents the matrix multiplication operation, α is a learnable scaling parameter used to adjust the size of the matrix product, and Reshape(·) represents the dimension change operation of the matrix.

[0173] The fourth step is to input feature M i-1 and features Add them together to get the output F of the self-attention mechanism module i-1 ∈R C×H×W .

[0174] Furthermore, in an embodiment of the present application, the output of the attention mechanism module can be input into the feedforward network module for further processing, and finally the output of the transformer block can be obtained.

[0175] Exemplarily, in some embodiments, as shown in FIG16 , the feedforward network is used to further perform feature fusion and nonlinear mapping on the weighted fusion features obtained in the attention mechanism module. This module draws on the design framework of the channel attention mechanism and implements further feature extraction and weighted fusion of the input features of the module in two branches. The input of this module is the output of the attention mechanism module, and the output of the attention mechanism module is input into the feedforward network module for further processing to obtain the output of the feedforward network, which is the final output M of the current transformer block. i '. The specific processing process is as follows:

[0176] The first step is to focus on the output F of the attention mechanism module i-1 The input is fed into the feedforward network module and multi-layer convolution operation is performed in two ways to achieve further feature extraction of the input features of the module. right Use Sigmoid nonlinear monotone excitation function to process and get the weight. Weighted to get features

[0177] In the second step, the input feature F i-1 and feature G i Add them together to get the output of the feedforward network, which is the final output of the current transformer block

[0178] It is worth noting that in the embodiment of the present application, since the network structure is intended to achieve feature extraction of different levels and sizes, multiple transformer blocks are used in the network, and the size of the input of each transformer block is also different, that is, the input feature M i-1 Size i∈[1,8] is different, and the sizes of various intermediate features in the corresponding transformer blocks are also different. The final output M of each transformer block is i The sizes of ' are also different.

[0179] That is, in the embodiment of the present application, the input sizes in each transformer block are different, namely Correspondingly, the sizes of various intermediate features in each transformer block can be obtained.

[0180] Furthermore, in the embodiments of the present application, the enhancement model can be understood as a transformer---Enhanceformer based on a multi-head attention mechanism (Multi-head Attention-based transformer, Enhanceformer), and the role of the enhancement model is to enhance the quality of the reconstructed values ​​of the image components.

[0181] Exemplarily, in some embodiments, Figure 17 is a structural diagram of the enhancement model proposed in an embodiment of the present application. As shown in Figure 17, the input of the enhancement model is the first reconstruction value of the image component of the current block, such as the initial reconstruction value I′1 of the brightness component of the CTU, and the preset quantization parameter QP′, and the output is the quality-enhanced reconstruction value of the image component of the current block, such as the enhanced reconstruction value of the brightness component of the CTU, for example, O1.

[0182] It should be noted that in the embodiments of the present application, the enhancement model is mainly composed of a shallow extraction module and a deep extraction module. The overall network presents a symmetrical structure, with the shallow extraction module on the left and the deep extraction module on the right. Among them, the shallow extraction module is used to extract multi-level features, and the deep extraction module is used to fuse multi-level features.

[0183] A convolutional layer feature extraction / fusion module and a two-stage multi-level feature extraction / fusion module are introduced into the shallow extraction module and the deep extraction module, respectively. The feature extraction / fusion module of each stage is composed of two transformer blocks (transformer networks).

[0184] It is understood that in the embodiments of the present application, the role of the shallow extraction module is to extract features at different levels and scales from the input image components, preparing for the fusion of layer-by-layer deep features. Specifically, the shallow extraction module consists of a single 3×3 convolutional layer and a two-stage multi-level feature extraction module. The single 3×3 convolutional layer performs simple feature extraction of the input image components to obtain shallow features; the two-stage multi-level feature extraction module further extracts shallow features in stages, ultimately achieving feature extraction at different levels and scales.

[0185] It can be understood that in the embodiment of the present application, the role of the deep extraction module is to use the same number of cascaded transformer blocks to perform channel splicing on features M1, M2, M3, and M4 of different levels and scales, and then continue to perform multi-scale feature extraction and feature weighting to complete the deep mining of connection features and the fusion of multi-level features to better restore the lost detail information and thus reconstruct the image components.

[0186] It should be noted that in the embodiments of the present application, in order to fully reduce the impact of the quantization operation on the quality of encoding the current block, the quantization parameter QP can be considered as prior information for quality enhancement, so the input is a normalized reconstruction value (first reconstruction value) and a normalized quantization parameter (preset quantization parameter).

[0187] It should be noted that in the embodiments of the present application, the enhanced model can be trained using a training set, and the enhanced model can also be tested using a test set. The enhanced model can be used to process videos in the YUV420p101e color sampling format. Therefore, before constructing the training and test sets, all types of data used need to be uniformly processed into this format.

[0188] It should be noted that in the embodiments of the present application, the data set for training and testing the enhanced model can be obtained under the configuration conditions of full frame (Al11lntra, AI), and the offline encoding test can be completed under the AI ​​configuration conditions.

[0189] Exemplarily, in some embodiments, the training set of the network model can be composed of 800 images used for training in the DIV2K dataset. Using VTM-11.0_NNVC-6.0, these 800 images are encoded with QP set to 22, 27, 32, 37, 42, AI mode and closed loop filter LMCS, DB, SAO and ALF configurations, resulting in a total of 4,000 encoded images. These 4,000 encoded images are then cropped into small blocks of 128×128, and after removing small blocks whose content remains unchanged before and after encoding and decoding, a total of 607,652 block images are obtained. These 607,652 encoded images are used as input to the network model, and their corresponding unencoded original images are used as true values ​​to constitute the final training set.

[0190] The network model's test set consists of 26 images, including the first frame of all sequences from Class A1, Class A2, Class B, Class C, Class D, Class E, and Class F, under the general test conditions specified by the video coding standard. These 26 images were encoded using VTM-11.0_NNVC-6.0 with QP settings of 22, 27, 32, 37, and 42, AI mode, and with LMCS, DB, SAO, and ALF disabled. This yielded a total of 130 encoded images. These 130 encoded images were then cropped into 128×128 tiles. After removing tiles whose content remained unchanged before and after encoding, a total of 19,714 tiled images were obtained. These 19,714 encoded images served as the input to the network model, and the corresponding unencoded original images served as the ground truth, forming the final test set.

[0191] For example, in some embodiments, the loss function of the network training of the enhanced model is MSELoss. The initial learning rate is set to 1×10 -5 , a total of 30 epochs, every 10 epochs the learning rate drops to the original Adam was used as the optimizer and training was performed in batch mode with a batch size of 16. Training was completed on a Pytorch 1.11.0 platform using an NVIDIA RTX 4090 GPU.

[0192] For example, in some embodiments, after the training of the enhancement model is completed, the trained network model parameters can be embedded into the VVC coding framework to achieve improved coding efficiency.

[0193] For example, in some embodiments, the offline trained network model can be embedded into both ends of the encoder and decoder respectively with the help of the libtorch1.11.0 library.

[0194] Furthermore, in an embodiment of the present application, FIG18 is a second schematic diagram of an implementation flow of the decoding method proposed in an embodiment of the present application. As shown in FIG18 , after determining the first reconstructed value of the image component of the current block according to the prediction mode and the prediction residual, that is, after step 102, the method for the decoder to perform decoding processing may include the following steps:

[0195] Step 104: When the identification parameter indicates that the enhancement model is not used to decode the current block, determine the first reconstruction value as the reconstruction value of the image component of the current block.

[0196] In an embodiment of the present application, after decoding the code stream and determining the decoding parameters of the current block including the identification parameters, if it is determined based on the identification parameters not to use the enhancement model, that is, when the identification parameters indicate not to use the enhancement model to decode the current block, the first reconstruction value of the image component of the current block determined using the prediction mode and the prediction residual can be directly determined as the reconstruction value of the image component of the current block.

[0197] It is understandable that, in the embodiment of the present application, after decoding the code stream and determining the identification parameter, it is possible to further determine whether to use the enhancement model to decode the current block based on the value of the identification parameter.

[0198] Exemplarily, in some embodiments, when the value of the identification parameter is the first value, it can be determined that the identification parameter indicates that the enhancement model is used to decode the current block;

[0199] Exemplarily, in some embodiments, when the value of the identification parameter is the second value, it can be determined that the identification parameter indicates not to use the enhancement model to decode the current block.

[0200] It should be noted that, in the embodiments of the present application, an identification parameter can be used to indicate whether the current block uses an enhancement model. Furthermore, the first value and the second value are different, and the first value and the second value can be in parameter form or in numerical form. Typically, the identification parameter can be a parameter written in the profile, but the identification parameter can also be a flag, which is not limited here.

[0201] It should also be noted that if the identification parameter is a flag, then in a specific example, the first value can be set to 1 and the second value can be set to 0; in another specific example, the first value can also be set to true and the second value can also be set to false; even in another specific example, the first value can also be set to 0 and the second value can also be set to 1; or, the first value can also be set to false and the second value can also be set to true. The first value and the second value in the embodiment of the present application are not limited in any way.

[0202] Taking the first value as 1 and the second value as 0 as an example, in the embodiment of the present application, if the value of the identification parameter is 1, it can be determined that the current block uses the enhancement model to perform image component enhancement processing. Otherwise, if the value of the identification parameter is 0, it can be determined that the current block does not use the enhancement model to perform image component enhancement processing.

[0203] In summary, the decoding method proposed through steps 101 to 104 above primarily includes a transformer-based quality enhancement technique that can enhance the quality of reconstructed values ​​during intra-frame encoding and decoding. First, a transformer-based Enhancer model based on a multi-head self-attention mechanism is proposed. This network model enhances the quality of the image components of the reconstructed values. Second, the trained network model parameters are embedded into the VVC coding framework to improve coding efficiency.

[0204] For example, in some embodiments, FIG19 is a schematic diagram of the encoding and decoding method proposed in an embodiment of the present application. As shown in FIG19 , at the encoding end, a rate-distortion cost can be used to determine whether to use an enhancement model (Enhanceformer) to enhance the quality of the reconstructed value of the image component of the current block. After obtaining a first reconstructed value of the image component of the current block based on the prediction mode and enhancing the quality of the first reconstructed value using the enhancement model and a preset quantization parameter to obtain the reconstructed value of the image component, a first generation value of the first reconstructed value and a second generation value of the reconstructed value of the image component can be determined respectively. If the first generation value is not less than the second generation value, the enhancement model can be selected to be used, and the identification parameter can be set to 1. Otherwise, the enhancement model can be selected not to be used, and the identification parameter can be set to 0. Finally, the prediction mode, preset quantization parameter, identification parameter, and prediction residual can be written into the bitstream and transmitted to the decoding end.

[0205] Correspondingly, at the decoding end, after decoding the identification parameter from the input bitstream, if the value of the identification parameter is 1, after determining the first reconstruction value of the image component of the current block based on the prediction mode and the prediction residual, it is necessary to use the enhancement model to perform quality enhancement on the first reconstruction value; if the value of the identification parameter is 0, there is no need to use the enhancement model for quality enhancement processing, and the first reconstruction value can be directly used as the reconstruction value of the image component.

[0206] It can be understood that the decoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding and decoding process, which is beneficial to improving the reconstruction quality of the image component of the current block, and is beneficial to providing a more accurate reference for the intra-frame prediction of subsequent coding blocks of the frame, thereby achieving an improvement in coding efficiency.

[0207] Exemplarily, the decoding method proposed in the embodiment of the present application can use the current reconstructed CTU and its γ-shaped area to enhance the quality of the current CTU to eliminate the blocking effect at the edge of the CU.

[0208] Exemplarily, the decoding method proposed in the embodiment of the present application may use other encoding prior information, such as the predicted value, residual value, and partition information of the current CTU, to assist in enhancing the quality of the reconstructed CTU.

[0209] Exemplarily, the decoding method proposed in the embodiment of the present application can be run in RA, LP and LB configuration modes.

[0210] For example, the decoding method proposed in the embodiments of this application is implemented in the latest VVC test software platform VTM-11.0_NNVC-6.0. The test sequences used are the Class A1, Class A2, Class B, Class C, Class D, Class E, and Class F sequences given in the general test conditions. The results in Table 1 were obtained with QP settings of 22, 27, 32, 37, and 42, and encoding in AI mode.

[0211] Class represents the video category, Sequence represents the specific test sequence, and Y, Cb, and Cr represent the performance of the three video components, luma and chroma. The values ​​in the table represent BD-rate, a measure of algorithm performance that indicates the change in bitrate and Peak Signal to Noise Ratio (PSNR) (or SSIM) compared to the original encoding algorithm. A negative value indicates improved performance, and a larger absolute value indicates a greater improvement.

[0212] It can be seen from Table 1 that on the latest version of NNVC code platform, this encoding algorithm significantly reduces the encoding bit rate of the Y component in each test sequence under the same quality (PSNR or SSIM) conditions.

[0213] Table 1 BD-rate test results of the proposed method in VTM-11.0_NNVC-6.0

[0214] It is understandable that the decoding method proposed in the embodiment of the present application is used for post-processing operations of intra-frame reconstruction values ​​(such as reconstructed CTUs). Therefore, this decoding method does not conflict with the Low Complexity Operation Point (LOP) or High Performance Operating Point (HOP) methods specified in the standard, and the two methods can be deployed and used simultaneously. Therefore, the processing performance of this decoding method is compared when LOP and HOP are enabled.

[0215] For example, Table 2 shows the performance comparison of the decoding method with or without HOP enabled when QP is 22, 27, 32, 37, and 42. Table 3 shows the comparison of the encoding results when the decoding method is enabled and LOP is enabled.

[0216] Table 2 BD-rate test results of the proposed method under HOP-on condition

[0217] Table 3 BD-rate test results of only enabling this solution versus only enabling the LOP method

[0218] Among them, as can be seen from Table 2, when HOP is turned on, turning on this decoding method will result in a negative coding gain. As can be seen from Table 3, turning on only the LOP method has a greater coding gain than turning on only this decoding method. On the one hand, this is because the input during network training is a data set obtained by encoding and decoding without turning on the HOP / LOP method, so turning on this decoding method will result in a negative coding gain. On the other hand, both the HOP and LOP methods are loop filtering methods based on neural networks, which are used to enhance the quality of the entire frame of video after the encoding of a frame is completed. However, this decoding method only enhances the quality of the complete CTU (size is 128×128) in the video frame. The incomplete CTU (size is less than 128×128) on the right and bottom of the video frame is not processed and cannot provide forward coding gain. In addition, this intra-frame post-processing method is used to restore block-level information. The receptive field is the current reconstructed CTU, and the entire frame image information cannot be obtained. Therefore, when the CTU processed by this decoding method is used for subsequent HOP processing of the entire frame image, the intra-frame post-processing results of this paper, as shown in Table 2, may even have a negative gain effect on the subsequent HOP processing results.

[0219] The embodiment of the present application provides a decoding method, at the decoding end, decoding the code stream, determining the decoding parameters of the current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; determining a first reconstruction value of the image component of the current block based on the prediction mode and the prediction residual; when the identification parameter indicates that the current block is to be decoded using an enhancement model, determining the reconstruction value of the image component of the current block based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure. It can be seen that in the embodiment of the present application, after completing the prediction of the image component of the current block based on the prediction mode and determining the corresponding first reconstruction value, the enhancement model can be further used to perform quality enhancement processing on the first reconstruction value, wherein the enhancement model can include a transformer based on a multi-head self-attention mechanism. In other words, the encoding and decoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding process to enhance the quality of the reconstruction value of the image component, which is beneficial to improving the reconstruction quality of the image component and providing a more accurate reference for the intra-frame prediction of subsequent coding blocks of the frame, thereby effectively improving the encoding and decoding efficiency and video compression performance.

[0220] Based on the above embodiment, another embodiment of the present application proposes an encoding method, wherein FIG20 is a schematic diagram of an implementation flow of the encoding method proposed in the embodiment of the present application. As shown in FIG20 , the encoding method of the encoder may include the following steps:

[0221] Step 201: Determine a first reconstructed value of an image component of a current block according to a prediction mode and a prediction residual of the current block.

[0222] In an embodiment of the present application, the encoder may first determine a first reconstructed value of an image component of the current block according to a prediction mode and a prediction residual of the current block.

[0223] It should be noted that for a video image, the video image can be divided into multiple image blocks, each of which can be encoded as a coding block, and the current block here specifically refers to the coding block currently to be predicted. The current block can be a CTU, or even a CU, PU, ​​etc., and this embodiment of the application does not impose any limitation.

[0224] Furthermore, in an embodiment of the present application, the encoding method may be applied to an intra-frame encoding process. Specifically, the encoding method may include a post-processing method for intra-frame reconstruction values.

[0225] It can be understood that, in the embodiments of the present application, the prediction mode of the current block can be any one of the intra-frame prediction modes.

[0226] It can be understood that, in the embodiment of the present application, the image component of the current block may be the chrominance component of the current block or the luminance component of the current block, and the present application does not make any specific limitation.

[0227] Exemplarily, in some embodiments, in a video image, a first image component, a second image component, and a third image component are generally used to represent a coding block (Coding Block, CB); wherein the three image components are a luminance component, a blue chrominance component, and a red chrominance component, respectively. Specifically, the luminance component is usually represented by the symbol Y, the blue chrominance component is usually represented by the symbol Cb or U, and the red chrominance component is usually represented by the symbol Cr or V; in this way, the video image can be represented in YCbCr format or in YUV format.

[0228] Furthermore, in an embodiment of the present application, when determining the first reconstruction value of the image component of the current block based on the prediction mode and the prediction residual, the prediction value of the image component of the current block can be first determined based on the prediction mode; and then the first reconstruction value can be determined based on the prediction value.

[0229] Exemplarily, in some embodiments, the sum of the prediction value of the image component of the current block and the prediction residual of the corresponding image component may be determined as the first reconstructed value of the image component.

[0230] It can be understood that in the embodiment of the present application, the first reconstruction value can be understood as the initial reconstruction value of the image component of the current block, that is, the first reconstruction value is the reconstruction value of the image component directly determined based on the prediction mode of the current block.

[0231] Step 202: Determine a reconstruction value of an image component corresponding to the current block based on an enhancement model; wherein the input parameters of the enhancement model include a first reconstruction value and a preset quantization parameter, and the enhancement model includes a transformer network structure.

[0232] In an embodiment of the present application, after determining a first reconstruction value of an image component of the current block according to a prediction mode and a prediction residual of the current block, a reconstruction value of an image component corresponding to the current block may be further determined based on an enhancement model.

[0233] It is understood that, in the embodiment of the present application, the preset quantization parameter may be a normalized quantization parameter QP, wherein the value of the preset quantization parameter may be any value and is not specifically limited in the present application.

[0234] For example, in some embodiments, the preset quantization parameter value, ie, QP, can be set to 22, 27, 32, 37, or 42.

[0235] It should be noted that, in the embodiment of the present application, the enhancement model can be used to perform quality enhancement on the first reconstruction value of the image component of the current block, thereby obtaining the enhanced reconstruction value of the image component.

[0236] Furthermore, in an embodiment of the present application, the input parameters of the enhancement model may include a first reconstruction value and a preset quantization parameter. After the first reconstruction value of the image component of the current block is quality enhanced by the enhancement model, an enhanced reconstruction value may be obtained, that is, the output result of the enhancement model may be the reconstruction value of the image component of the current block.

[0237] It should be noted that, in the embodiments of the present application, the enhanced model may include a transformer network structure.

[0238] Furthermore, in the embodiments of the present application, the transformer is a deep learning model that is different from convolutional neural networks and is widely used in fields such as natural language processing (NLP). Among them, the transformer can simultaneously focus on all positions in the input sequence through the application of the self-attention mechanism, thereby better understanding the context of the sequence data.

[0239] It can be understood that in the embodiments of the present application, the enhancement model can be understood as a transformer-based reconstructed CTU quality enhancement network, which is used for intra-frame post-processing of the reconstructed CTU to improve the quality of the reconstructed CTU, thereby improving the prediction accuracy of subsequent coding blocks and ultimately improving the coding performance.

[0240] That is, in the embodiments of the present application, an enhancement model including a transformer network structure can be used for intra-frame post-processing of reconstructed CTUs. The enhancement model leverages the transformer's feature extraction and global attention mechanism to enhance the quality of the current CTU, thereby improving the intra-frame prediction accuracy of subsequent CTUs.

[0241] Exemplarily, in some embodiments, the enhancement model may be composed of a cascade of a shallow extraction module and a deep extraction module; wherein both the shallow extraction module and the deep extraction module include a transformer network structure.

[0242] It should be noted that, in the embodiment of the present application, the shallow layer extraction module in the enhancement model can be used to extract shallow features. Specifically, the shallow layer extraction module is used to perform multi-level and multi-scale feature extraction processing.

[0243] It should be noted that, in the embodiment of the present application, the deep layer extraction module in the enhanced model can be used to extract deep features. Specifically, the deep layer extraction module is used to perform multi-level and multi-scale feature extraction and feature fusion processing.

[0244] It can be understood that, in the embodiment of the present application, the enhanced model may further include a first connection layer.

[0245] Furthermore, in an embodiment of the present application, when determining the reconstruction value of the image component of the current block based on the enhancement model, the first reconstruction value and the preset quantization parameter can be first spliced ​​through the first connection layer to determine the first spliced ​​component; then the first spliced ​​component is input into the shallow extraction module to output the first feature information; then the first feature information is input into the deep extraction module to output the second feature information; finally, the reconstruction value of the image component can be determined based on the first reconstruction value and the second feature information.

[0246] That is to say, in an embodiment of the present application, for the first reconstruction value of the input image component, shallow feature extraction can be first performed through a shallow extraction module to obtain first feature information of the image component; then, based on the first feature information, deep feature extraction can be further performed through a deep extraction module to obtain second feature information of the image component; finally, the first reconstruction value and the second feature information can be combined to complete the quality enhancement of the image component of the current block and obtain the reconstruction value of the image component.

[0247] It should be noted that, in an embodiment of the present application, for the first connection layer in the enhancement model, its input may be the first reconstruction value and preset quantization parameter of the image component of the current block, and its output may be the corresponding first spliced ​​component.

[0248] It should be noted that in the embodiments of the present application, in order to fully reduce the impact of the quantization operation on the quality of encoding the current block, the quantization parameter QP can be considered as prior information for quality enhancement, so the input of the first connection layer is the normalized reconstruction value (first reconstruction value) and the normalized quantization parameter (preset quantization parameter).

[0249] Exemplarily, in some embodiments, assuming that the image component is a Y component, the corresponding first reconstruction value is the normalized reconstructed Y component I′1∈R 1×128×128 , the preset quantization parameter is QP′∈R composed of normalized QP 1×128×128 .

[0250] For example, in some embodiments, as shown in FIG8 , assuming that the image component is a Y component, the first reconstruction value I′1 and the preset quantization parameter QP′ are input to the first connection layer, the two normalized components are spliced ​​on the channel, and the corresponding first spliced ​​component is output, wherein the first spliced ​​component can be the spliced ​​normalized component I1∈R 2×128×128 .

[0251] Furthermore, in an embodiment of the present application, the shallow extraction module includes a first convolutional layer, a first feature extraction module and a second feature extraction module, wherein the first feature extraction module and the second feature extraction module both include a transformer network structure.

[0252] Exemplarily, in some embodiments, as shown in Figure 9, the shallow extraction module may include a first convolutional layer, a first feature extraction module and a second feature extraction module, wherein the first convolutional layer may be a 3×3 convolution, and the first feature extraction module and the second feature extraction module respectively include two transformer network structures, that is, the first feature extraction module and the second feature extraction module can constitute 4 cascaded transformer blocks.

[0253] It should be noted that, in an embodiment of the present application, the shallow extraction module may further include a ReLU activation function and a downsampling operation.

[0254] Exemplarily, in some embodiments, as shown in FIG10 , the shallow extraction module may include a first convolutional layer, a ReLU activation function, a first feature extraction module, and a second feature extraction module, wherein the first feature extraction module and the second feature extraction module each include a downsampling operation.

[0255] Furthermore, in an embodiment of the present application, when the first spliced ​​component is input into the shallow extraction module and the first feature information is output, the first feature can be first determined through the first convolutional layer based on the first spliced ​​component; then the first feature can be input into the first transformer network in the first feature extraction module to output the second feature; then based on the second feature, the third feature can be determined through the second transformer network in the first feature extraction module; then the third feature can be input into the third transformer network in the second feature extraction module to output the fourth feature; finally, based on the fourth feature, the first feature information can be determined through the fourth transformer network in the second feature extraction module.

[0256] For example, in some embodiments, as shown in FIG11 , after the first reconstruction value and the preset quantization parameter are concatenated through the first connection layer and the first concatenated component I1 is output, the first feature M0∈R can be determined by sequentially passing through the first convolution layer and the ReLU activation function. 32×128×128 , that is, the spliced ​​normalized component I1 is input into a single 3×3 convolution layer. After 3×3 convolution and nonlinear mapping operations, the shallow feature extraction and the increase of the number of channels are realized to obtain the shallow feature M0. Then, with the help of downsampling operation, the first feature M0 is input into the four cascaded transformer networks for feature extraction, and the second features M1∈R of different levels and scales are obtained respectively. 32×128×128 , the third feature M2∈R 64×64×64 , the fourth feature M3∈R 64×64×64 , first feature information M4∈R 128×32×32 Among them, the four transformer networks are each processed in pairs, with a total of two stages (i.e., the first feature extraction module and the second feature extraction module) to extract features of different levels and scales from the input features.

[0257] It should be noted that in the embodiment of the present application, when determining the first feature through the first convolution layer based on the first concatenated component, the first concatenated component, i.e., the normalized component I1 after concatenation, can be subjected to a 3×3 convolution and nonlinear mapping operation to obtain the first feature, i.e., the shallow feature M0, as shown in formula (1). Wherein, σ represents the ReLU activation function, and W1 represents the 3×3 convolution kernel.

[0258] It should be noted that in an embodiment of the present application, when feature extraction of different levels and scales is performed based on the first feature extraction module and the second feature extraction module with the help of a downsampling operation, the first feature M0 can be first input into the first transformer network in the first feature extraction module to output the second feature M1. Then, based on the second feature M1, the third feature M2 is determined through the downsampling operation and the second transformer network in the first feature extraction module in sequence, and then the third feature M2 is input into the third transformer network in the second feature extraction module to output the fourth feature M3. Then, based on the fourth feature M3, the first feature information M4 is determined through the downsampling operation and the fourth transformer network in the second feature extraction module in sequence.

[0259] For example, in some embodiments, the first feature M0 can be input into the multi-level feature extraction module of the first stage, and passed through the first transformer block, i.e., the first transformer network in the first feature extraction module, to obtain feature M1; the feature M1 is downsampled, and the width and height of the feature are reduced to half of the original size, and its channel dimension is doubled. Then, the downsampled feature M1 is input into the second transformer block, i.e., the second transformer network in the first feature extraction module, to obtain feature M2. M1 and M2 are features of different levels and scales, and the acquisition process is as shown in formula (2). Transformer (·) represents the Transformer operation, W2 represents the 3×3 convolution kernel, pixelunshuffle (·) represents the inverse pixel reshuffling operation, and downsample (·) represents the downsampling operation. The downsampling operation is specifically composed of a 3×3 convolution operation with half the number of channels and a pixelunshuffle inverse operation (pixelunshuffle) with a sampling factor of 2.

[0260] For example, in some embodiments, feature M2 is input into the second-stage multi-level feature extraction module to continue feature extraction at different levels. After passing through the first transformer block, i.e., the third transformer network in the second feature extraction module, feature M3 is obtained; feature M3 is still downsampled, with its width and height dimensions reduced to half of the original size, and its channel dimension doubled. The downsampled feature M3 is then input into the second transformer block, i.e., the fourth transformer network in the second feature extraction module, to obtain feature M4. The specific acquisition process is shown in formula (3).

[0261] Furthermore, in an embodiment of the present application, the deep extraction module includes a second connection layer, a third connection layer, a fourth connection layer, a third feature extraction module, a fourth feature extraction module, a second convolutional layer, and a third convolutional layer, wherein the third feature extraction module and the fourth feature extraction module both include a transformer network structure.

[0262] Exemplarily, in some embodiments, as shown in Figure 12, the second convolution layer in the deep extraction module can be a 1×1 convolution, the third convolution layer can be a 3×3 convolution, the third feature extraction module and the fourth feature extraction module respectively include two transformer network structures, that is, the third feature extraction module and the fourth feature extraction module can constitute 4 cascaded transformer blocks, and the second connection layer, the third connection layer, and the fourth connection layer can be respectively placed before the three transformer networks.

[0263] It should be noted that, in an embodiment of the present application, the deep extraction module may further include a ReLU activation function and an upsampling operation.

[0264] Exemplarily, in some embodiments, as shown in FIG13 , the third feature extraction module and the fourth feature extraction module in the deep extraction module may each include an upsampling operation, and the deep extraction module may include a ReLU activation function.

[0265] Furthermore, in an embodiment of the present application, when the first feature information is input into the deep extraction module and the second feature information is output, the first feature information and the fourth feature can be first input into the second connection layer to output the second spliced ​​component; then, based on the second spliced ​​component, the fifth feature can be determined through the second convolutional layer and the fifth transformer network in the third feature extraction module; then, the fifth feature and the third feature can be input into the third connection layer to output the third spliced ​​component; then, based on the third spliced ​​component, the sixth feature can be determined through the second convolutional layer and the sixth transformer network in the third feature extraction module; then, the sixth feature and the second feature can be input into the fourth connection layer to output the fourth spliced ​​component; then, based on the fourth spliced ​​component, the seventh feature can be determined through the second convolutional layer and the seventh transformer network in the fourth feature extraction module; then, the seventh feature can be input into the eighth transformer network in the fourth feature extraction module to output the eighth feature; finally, based on the eighth feature, the second feature information can be determined through the third convolutional layer.

[0266] Exemplarily, in some embodiments, as shown in FIG14 , after completing shallow feature extraction and obtaining the second feature M1, the third feature M2, the fourth feature M3, and the first feature information M4 of different levels and scales, the first feature information M4 can be upsampled first, and then the fourth feature M3 and the upsampled first feature information M4∈ are channel-spliced ​​using the second connection layer to obtain a second spliced ​​component, and then the second spliced ​​component is input into the fifth transformer network in the third feature extraction module, and the fifth feature M5∈R is output. 64×64×64 Then, the third connection layer can be used to perform channel splicing on the third feature M2 and the fifth feature M5 to obtain the third spliced ​​component, and then the third spliced ​​component is input into the sixth transformer network in the third feature extraction module to output the sixth feature. Next, the sixth feature M6 can be upsampled first, and the fourth connection layer can be used to perform channel splicing on the second feature M1 and the upsampled sixth feature M6 to obtain the fourth spliced ​​component, and then the fourth spliced ​​component is input into the seventh transformer network in the fourth feature extraction module to obtain the seventh feature M7∈R 32×128×128 ; Then the seventh feature M7 can be input into the eighth transformer network in the fourth feature extraction module, and the eighth feature M8∈R 32×128×128 Finally, based on the eighth feature, the second feature information, namely the quality-enhanced brightness component C1∈R 1×128×128 .

[0267] It can be understood that in the embodiment of the present application, the role of the deep extraction module is to use the same number of cascaded transformer blocks to perform channel splicing on features M1, M2, M3, and M4 of different levels and scales, and then continue to perform multi-scale feature extraction and feature weighting to complete the deep mining of connection features and the fusion of multi-level features to better restore the lost detail information and thus reconstruct the image components.

[0268] It should be noted that in the embodiments of this application, the operations in the deep extraction module can be understood as the inverse operations in the shallow extraction module. The four transformer networks in the deep extraction module still process two by two at a time, extracting features at different levels and scales from the input features in two stages, resulting in features of different scales.

[0269] It can be understood that in the embodiment of the present application, the deep extraction module is composed of a two-stage multi-level feature fusion module and a single convolutional layer, wherein the two-stage multi-level feature fusion module realizes the fusion of features of different levels and sizes, and the single 3×3 convolutional layer realizes the deep fusion of features and the reduction of the number of channels to form the final residual CTU.

[0270] For example, in some embodiments, the first feature information M4 is further input into four cascaded transformer networks for feature extraction. Upsampling is performed to achieve scale matching between pairs of features, and the pairwise scale-matched features are then spliced ​​on the channel. The spliced ​​components are then input into four cascaded transformer networks for feature extraction at different levels and scales, thereby obtaining features M5, M6, M7, and M8 at different levels and scales.

[0271] For example, in some embodiments, feature M4 is first upsampled, with the width and height of the feature doubled and its channel dimension halved; the upsampled feature M4 is concatenated with feature M3 in terms of channel dimension; the concatenated features are fused and the number of channels is halved using 1×1 convolution, and the concatenated features with the number of channels halved are then input into the multi-level feature extraction module of the third stage, and after passing through the first transformer block, i.e., the fifth transformer network in the third feature extraction module, feature M5 is obtained; after channel concatenation of feature M5 and feature M2, the concatenated features are fused and the number of channels is halved using 1×1 convolution, and then input into the second transformer block, i.e., the sixth transformer network in the third feature extraction module, to obtain feature M6. The specific process is shown in formula (4). Wherein, Transformer(·) represents the Transformer operation, W3 represents the 3×3 convolution kernel, Pixelunshuffle(·) represents the inverse pixel shuffling operation, and upsample(·) represents the upsampling operation. The upsampling operation is specifically composed of a 3×3 convolution operation that halves the number of channels and a pixel shuffle operation with a sampling factor of 2.

[0272] For example, in some embodiments, feature M6 is upsampled, with the width and height of the feature doubled and its channel dimension halved; the upsampled feature M6 is concatenated with feature M1 in terms of the channel dimension; 1×1 convolution is used to achieve fusion of the concatenated features and halve the number of channels; the concatenated features with halved channels are then input into the multi-level feature extraction module of the fourth stage, and after passing through the first transformer block, i.e., the seventh transformer network in the fourth feature extraction module, feature M7 is obtained; feature M7 is input into the second transformer block, i.e., the eighth transformer network in the fourth feature extraction module, to obtain feature M8. The specific process is as shown in formula (5). Wherein, W4 represents a 3×3 convolution kernel.

[0273] For example, in some embodiments, after obtaining the eighth feature, that is, the deep feature M8, the deep feature M8 is input into a single 3×3 convolution layer, and after 3×3 convolution and nonlinear mapping operations, the final fusion of features and reduction of the number of channels are achieved to obtain the final second feature information, that is, the quality-enhanced brightness component C1∈R 1×128×128 , as shown in formula (6). Where σ represents the ReLU activation function and W2 represents the 3×3 convolution kernel.

[0274] Furthermore, in an embodiment of the present application, when determining the reconstruction value of an image component based on the first reconstruction value and the second feature information, the second feature information extracted by the shallow extraction module and the deep extraction module in the enhancement model can be used to perform quality enhancement on the initial first reconstruction value, thereby obtaining an enhanced image component, that is, the reconstruction value of the image component.

[0275] For example, in some embodiments, the second feature information C1 can be combined with the input brightness component I ′ 1 is added to obtain the final output reconstruction value O1∈R 1×128×128 The process is as shown in formula (7).

[0276] Furthermore, in an embodiment of the present application, the transformer network structure may include a self-attention mechanism module and a feedforward network.

[0277] That is to say, in the embodiments of the present application, for any transformer network in the shallow extraction module and the deep extraction module, it is composed of a cascade of a self-attention module and a feedforward neural network (FFN). Among them, the output of the attention mechanism module in the current transformer block is the input of the feedforward network, and the output of the feedforward network is the final output of the current transformer block. The attention mechanism module is used to realize feature extraction and weighted fusion, and the feedforward network is used to further perform feature fusion and nonlinear mapping on the weighted fusion features obtained in the attention mechanism module.

[0278] Exemplarily, in some embodiments, as shown in FIG15 , a self-attention mechanism module is used to realize feature extraction and weighted fusion. The module consists of three parts: layer normalization operation, two-head attention mechanism and single 1×1 convolution operation. The layer normalization operation completes the normalization of the input features in the channel dimension, ensuring that the feature point values ​​of the input features are always within a reasonable range, creating a prerequisite for obtaining the subsequent residual features; the two-head attention mechanism module divides the features into two on the channel, and calculates the weighted features of the two parts in the upper and lower branches respectively, and then splices the two weighted features on the channel to obtain the summary residual features. In addition, the attention matrix transpose calculation operation is introduced in the attention mechanism module, which not only realizes multi-level feature extraction, but also effectively reduces the number of parameters; the single 1×1 convolution operation completes the change in the number of channels of the summary residual features. The specific processing process is as follows:

[0279] The first step is to input features i∈[1,8] is normalized in the channel dimension, and the normalized features are divided into two equal parts in the channel dimension to obtain Where C is the number of channels, H is the height, and W is the width.

[0280] The second step is to and Using three 1×1 and 3×3 convolution operations, we get In each convolution operation, unless otherwise specified, the feature size and number of channels do not change before and after processing.

[0281] The third step is to merge and transpose the above features to achieve dimension change and obtain Will After matrix multiplication, the Softmax function is used for normalization to obtain the attention matrix Similarly, the attention matrix can be obtained Then use and Perform matrix calculations to realize matrix The weighted operation of Similarly, we can get The weighted features are then spliced ​​in the channel dimension and the dimension is changed to obtain As shown in formula (8), Softmax(·) represents the softmax activation function, · represents the matrix multiplication operation, α is a learnable scaling parameter used to adjust the size of the matrix product, and Reshape(·) represents the dimension change operation of the matrix.

[0282] The fourth step is to input feature M i-1 and features Add them together to get the output F of the self-attention mechanism module i-1 ∈R C×H×W .

[0283] Furthermore, in an embodiment of the present application, the output of the attention mechanism module can be input into the feedforward network module for further processing, and finally the output of the transformer block can be obtained.

[0284] Exemplarily, in some embodiments, as shown in FIG16 , the feedforward network is used to further perform feature fusion and nonlinear mapping on the weighted fusion features obtained in the attention mechanism module. This module draws on the design framework of the channel attention mechanism and implements further feature extraction and weighted fusion of the input features of the module in two branches. The input of this module is the output of the attention mechanism module, and the output of the attention mechanism module is input into the feedforward network module for further processing to obtain the output of the feedforward network, which is the final output M of the current transformer block. i '. The specific processing process is as follows:

[0285] The first step is to focus on the output F of the attention mechanism module i-1 The input is fed into the feedforward network module and multi-layer convolution operation is performed in two ways to achieve further feature extraction of the input features of the module. right Use Sigmoid nonlinear monotone excitation function to process and get the weight. Weighted to get features

[0286] In the second step, the input feature F i-1 and feature G i Add them together to get the output of the feedforward network, which is the final output of the current transformer block

[0287] It is worth noting that in the embodiment of the present application, since the network structure is intended to achieve feature extraction of different levels and sizes, multiple transformer blocks are used in the network, and the size of the input of each transformer block is also different, that is, the input feature M i-1 Size i∈[1,8] is different, and the sizes of various intermediate features in the corresponding transformer blocks are also different. The final output M of each transformer block is i The sizes of ' are also different.

[0288] That is, in the embodiment of the present application, the input sizes in each transformer block are different, namely {R 32×128×128 , R 64×64×64 , R 64×64×64 , R 128×32×32 , R 64×64×64 , R 64×64×64 , R 32×128×128 , R 32×128×128}, the size of various intermediate features in each transformer block can be obtained accordingly.

[0289] Furthermore, in the embodiments of the present application, the enhancement model can be understood as a transformer---Enhanceformer based on a multi-head self-attention mechanism, and the role of the enhancement model is to enhance the quality of the reconstructed values ​​of the image components.

[0290] Exemplarily, in some embodiments, as shown in FIG17 , the input of the enhancement model is the first reconstruction value of the image component of the current block, such as the initial reconstruction value I′1 of the luminance component of the CTU, and a preset quantization parameter QP′, and the output is the quality-enhanced reconstruction value of the image component of the current block, such as the enhanced reconstruction value of the luminance component of the CTU, for example, O1.

[0291] It should be noted that in the embodiments of the present application, the enhancement model is mainly composed of a shallow extraction module and a deep extraction module. The overall network presents a symmetrical structure, with the shallow extraction module on the left and the deep extraction module on the right. Among them, the shallow extraction module is used to extract multi-level features, and the deep extraction module is used to fuse multi-level features.

[0292] A convolutional layer feature extraction / fusion module and a two-stage multi-level feature extraction / fusion module are introduced into the shallow extraction module and the deep extraction module, respectively. The feature extraction / fusion module of each stage is composed of two transformer blocks (transformer networks).

[0293] It is understood that in the embodiments of the present application, the role of the shallow extraction module is to extract features at different levels and scales from the input image components, preparing for the fusion of layer-by-layer deep features. Specifically, the shallow extraction module consists of a single 3×3 convolutional layer and a two-stage multi-level feature extraction module. The single 3×3 convolutional layer performs simple feature extraction of the input image components to obtain shallow features; the two-stage multi-level feature extraction module further extracts shallow features in stages, ultimately achieving feature extraction at different levels and scales.

[0294] It should be noted that in the embodiments of the present application, in order to fully reduce the impact of the quantization operation on the quality of encoding the current block, the quantization parameter QP can be considered as prior information for quality enhancement, so the input is a normalized reconstruction value (first reconstruction value) and a normalized quantization parameter (preset quantization parameter).

[0295] It should be noted that in the embodiments of the present application, the enhanced model can be trained using a training set, and the enhanced model can also be tested using a test set. The enhanced model can be used to process videos in the YUV420p101e color sampling format. Therefore, before constructing the training and test sets, all types of data used need to be uniformly processed into this format.

[0296] It should be noted that in the embodiments of the present application, the data set used for training and testing of the enhanced model can be obtained under the configuration conditions of AI in the full frame, and the offline encoding test can be completed under the AI ​​configuration conditions.

[0297] Exemplarily, in some embodiments, the training set of the network model can be composed of 800 images used for training in the DIV2K dataset. Using VTM-11.0_NNVC-6.0, these 800 images are encoded with QP set to 22, 27, 32, 37, 42, AI mode and closed loop filter LMCS, DB, SAO and ALF configurations, resulting in a total of 4,000 encoded images. These 4,000 encoded images are then cropped into small blocks of 128×128, and after removing small blocks whose content remains unchanged before and after encoding and decoding, a total of 607,652 block images are obtained. These 607,652 encoded images are used as input to the network model, and their corresponding unencoded original images are used as true values ​​to constitute the final training set.

[0298] The network model's test set consists of 26 images, including the first frame of all sequences from Class A1, Class A2, Class B, Class C, Class D, Class E, and Class F, under the general test conditions specified by the video coding standard. These 26 images were encoded using VTM-11.0_NNVC-6.0 with QP settings of 22, 27, 32, 37, and 42, AI mode, and with LMCS, DB, SAO, and ALF disabled. This yielded a total of 130 encoded images. These 130 encoded images were then cropped into 128×128 tiles. After removing tiles whose content remained unchanged before and after encoding, a total of 19,714 tiled images were obtained. These 19,714 encoded images served as the input to the network model, and the corresponding unencoded original images served as the ground truth, forming the final test set.

[0299] For example, in some embodiments, the loss function of the network training of the enhanced model is MSELoss. The initial learning rate is set to 1×10 -5 , a total of 30 epochs, every 10 epochs the learning rate drops to the original Adam was used as the optimizer and training was performed in batch mode with a batch size of 16. Training was completed on a Pytorch 1.11.0 platform using an NVIDIA RTX 4090 GPU.

[0300] For example, in some embodiments, after the training of the enhancement model is completed, the trained network model parameters can be embedded into the VVC coding framework to achieve improved coding efficiency.

[0301] For example, in some embodiments, the offline trained network model can be embedded into both ends of the encoder and decoder respectively with the help of the libtorch1.11.0 library.

[0302] Step 203: Determine an identification parameter according to the first reconstruction value and the reconstruction value of the image component.

[0303] In an embodiment of the present application, after determining the first reconstruction value of the image component of the current block based on the prediction mode and prediction residual of the current block, and determining the reconstruction value of the image component corresponding to the current block based on the enhancement model, the identification parameter can be further determined based on the first reconstruction value and the reconstruction value of the image component.

[0304] It can be understood that, in the embodiment of the present application, the identification parameter can be used to determine whether to use the enhancement model to further decode the current block.

[0305] Furthermore, in an embodiment of the present application, when determining the identification parameter corresponding to the current block based on the first reconstruction value and the reconstruction value of the image component, the first generation value corresponding to the first reconstruction value and the second generation value corresponding to the reconstruction value of the image component can be determined respectively; then, when the first generation value is greater than the second generation value, the value of the identification parameter can be set to indicate that the current block is decoded using an enhanced model; or, when the first generation value is less than or equal to the second generation value, the value of the identification parameter can be set to indicate that the current block is not decoded using an enhanced model.

[0306] Illustratively, in some embodiments, a rate-distortion cost algorithm may be used to calculate the first generation value and the second cost value.

[0307] That is, in the embodiment of the present application, at the encoding end, assuming that the current block is a CTU, the rate-distortion cost J can be used to determine whether to use the enhancement model to perform quality enhancement processing on the reconstructed value. The calculation method of J is as follows: J = SSE (x, y) + λ CTU R extraflag (9)

[0308] Among them, x is the CTU before quality enhancement, y is the CTU after quality enhancement, λ CTU is the Lagrangian factor at the CTU level, R extraflag The bitrate spent on encoding the extra marker bits.

[0309] It can be understood that in an embodiment of the present application, if the first generation value is greater than the second generation value, then it can be considered that better encoding and decoding performance can be obtained by using the enhancement model to enhance the reconstructed value of the image component. Therefore, the value of the identification parameter can be set to indicate the use of the enhancement model to decode the current block, for example, the value of the identification parameter is set to the first value.

[0310] It can be understood that in an embodiment of the present application, if the first generation value is less than or equal to the second generation value, then it can be considered that using the enhancement model to enhance the reconstructed value of the image component cannot obtain better encoding and decoding performance. Therefore, the value of the identification parameter can be set to indicate that the enhancement model is not used to decode the current block, for example, the value of the identification parameter is set to the second value.

[0311] It should be noted that, in the embodiments of the present application, an identification parameter can be used to indicate whether the current block uses an enhancement model. Furthermore, the first value and the second value are different, and the first value and the second value can be in parameter form or in numerical form. Typically, the identification parameter can be a parameter written in the profile, but the identification parameter can also be a flag, which is not limited here.

[0312] It should also be noted that if the identification parameter is a flag, then in a specific example, the first value can be set to 1 and the second value can be set to 0; in another specific example, the first value can also be set to true and the second value can also be set to false; even in another specific example, the first value can also be set to 0 and the second value can also be set to 1; or, the first value can also be set to false and the second value can also be set to true. The first value and the second value in the embodiment of the present application are not limited in any way.

[0313] Taking the first value as 1 and the second value as 0 as an example, in an embodiment of the present application, if it is determined that the current block uses an enhancement model to perform enhancement processing of image components, then the value of the identification parameter can be set to 1; otherwise, if it is determined that the current block does not use an enhancement model to perform enhancement processing of image components, then the value of the identification parameter can be set to 0.

[0314] Accordingly, in an embodiment of the present application, after determining whether to use the enhancement model for quality enhancement, an additional flag bit, such as the value of an identification parameter, can be written into the code stream and transmitted to the decoder. The decoder then selects whether to use the enhancement model for quality enhancement based on the specific value of the additional flag bit. In other words, at the decoder, after determining the identification parameter in the decoded code stream, the decoder can further determine whether to use the enhancement model to decode the current block based on the value of the identification parameter.

[0315] Exemplarily, in some embodiments, at the encoding end, the proposed enhancement model Enhanceformer is used in the intra-frame encoding process. Whenever a CTU is reconstructed, the Enhanceformer is used to perform quality enhancement on its image components, such as the luminance component. If the rate-distortion cost of the enhanced luminance component is less than the rate-distortion cost of the luminance component before enhancement, the reconstructed luminance value of the CTU is updated to the luminance value after quality enhancement by the Enhanceformer, and a 1-bit flag is added to the CTU to mark whether the CTU uses the Enhanceformer to perform quality enhancement on the luminance component. If used, the CTU flag is 1, otherwise it is 0.

[0316] Step 204: Write the decoding parameters of the current block into the bitstream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter.

[0317] In an embodiment of the present application, after determining the identification parameter based on the first reconstruction value and the reconstruction value of the image component, the decoding parameters of the current block can be written into the bitstream; wherein the decoding parameters may at least include the prediction mode, identification parameter, and preset quantization parameter of the current block.

[0318] It can be understood that in an embodiment of the present application, decoding parameters including prediction mode, identification parameters, and preset quantization parameters can be written into the code stream and transmitted to the decoding end, so that the decoder can decode the image component of the current block according to the decoding parameters.

[0319] In summary, the encoding method proposed through steps 201 to 204 above mainly includes a transformer-based quality enhancement technology that can enhance the quality of reconstructed values ​​during the intra-frame encoding and decoding process. First, a transformer-based Enhancer model based on a multi-head self-attention mechanism is proposed. The function of this network model is to enhance the quality of the image components of the reconstructed value. Secondly, the trained network model parameters are embedded in the VVC coding framework to achieve improved coding efficiency.

[0320] It can be understood that the encoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding and decoding process, which is beneficial to improving the reconstruction quality of the image component of the current block, and is beneficial to providing a more accurate reference for the intra-frame prediction of subsequent encoding blocks of the frame, thereby achieving an improvement in coding efficiency.

[0321] Exemplarily, the encoding method proposed in the embodiment of the present application can use the current reconstructed CTU and its γ-shaped area to enhance the quality of the current CTU to eliminate the blocking effect at the edge of the CU.

[0322] Exemplarily, the encoding method proposed in the embodiment of the present application may use other encoding prior information, such as the predicted value, residual value, and partition information of the current CTU, to assist in enhancing the quality of the reconstructed CTU.

[0323] Exemplarily, the encoding method proposed in the embodiment of the present application can be run in RA, LP and LB configuration modes.

[0324] The embodiment of the present application provides an encoding method, at the encoding end, determining a first reconstruction value of an image component of a current block based on a prediction mode and a prediction residual of the current block; determining a reconstruction value of an image component corresponding to the current block based on an enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and a preset quantization parameter, and the enhancement model includes a transformer network structure; determining an identification parameter based on the first reconstruction value and the reconstruction value of the image component; writing decoding parameters of the current block into a bitstream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter. It can be seen that in the embodiment of the present application, after completing the prediction of the image component of the current block based on the prediction mode and determining the corresponding first reconstruction value, the enhancement model can be further used to perform quality enhancement processing on the first reconstruction value, wherein the enhancement model can include a transformer based on a multi-head self-attention mechanism. That is, the encoding and decoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding process to enhance the quality of the reconstruction value of the image component, which is beneficial to improving the reconstruction quality of the image component and providing a more accurate reference for the intra-frame prediction of subsequent coding blocks of the frame, thereby effectively improving encoding and decoding efficiency and video compression performance.

[0325] In another embodiment of the present application, based on the same inventive concept as the above embodiment, see FIG21, which shows a schematic diagram of the structure of the encoder 210 proposed in the embodiment of the present application. As shown in FIG8, the encoder 210 may include: a first determining unit 2101; wherein,

[0326] The first determination unit 2101 is configured to determine a first reconstruction value of an image component of the current block according to a prediction mode and a prediction residual of the current block; determine a reconstruction value of the image component corresponding to the current block based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and a preset quantization parameter, and the enhancement model includes a transformer network structure; determine an identification parameter according to the first reconstruction value and the reconstruction value of the image component; and write decoding parameters of the current block into a bitstream; wherein the decoding parameters include a prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter.

[0327] It should be noted that, in the embodiment of the present application, the encoder 210 can also be regarded as a data processing mode (or "entropy encoder"), which is used to encode the values ​​of the syntax elements to be encoded.

[0328] It is understood that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular device. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.

[0329] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0330] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 210. The computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, it implements the encoding method described in any one of the aforementioned embodiments.

[0331] Based on the composition of the encoder 210 and the computer-readable storage medium, refer to Figure 22, which shows a specific hardware structure diagram of the encoder 210 provided in an embodiment of the present application. As shown in Figure 22, the encoder 210 may include: a first communication interface 2101, a first memory 2102 and a first processor 2103; each component is coupled together through a first bus system 2104. It can be understood that the first bus system 2104 is used to achieve connection and communication between these components. In addition to the data bus, the first bus system 2104 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the first bus system 2104 in Figure 9. Among them,

[0332] The first communication interface 2101 is used to receive and send signals when sending and receiving information with other external network elements;

[0333] A first memory 2102 is used to store computer programs that can be run on the first processor 2103;

[0334] The first processor 2103 is configured to, when running the computer program, perform the following steps: determining a first reconstruction value of an image component of the current block based on a prediction mode and a prediction residual of the current block; determining a reconstruction value of an image component corresponding to the current block based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and a preset quantization parameter, and the enhancement model includes a transformer network structure; determining an identification parameter based on the first reconstruction value and the reconstruction value of the image component; and writing decoding parameters of the current block into a bitstream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter.

[0335] It is understood that the first memory 2102 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The first memory 2102 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0336] The first processor 2103 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the first processor 2103. The above-mentioned first processor 2103 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. The storage medium is located in the first memory 2102 , and the first processor 2103 reads the information in the first memory 2102 and completes the steps of the above method in combination with its hardware.

[0337] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0338] Optionally, as another embodiment, the first processor 2103 is further configured to execute the encoding method described in any one of the aforementioned embodiments when running the computer program.

[0339] This embodiment provides an encoder that determines a first reconstruction value of an image component of a current block based on a prediction mode and a prediction residual of the current block; determines a reconstruction value of an image component corresponding to the current block based on an enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and a preset quantization parameter, and the enhancement model includes a transformer network structure; determines an identification parameter based on the first reconstruction value and the reconstruction value of the image component; and writes decoding parameters of the current block into a bitstream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter. It can be seen that in the embodiment of the present application, after completing the prediction of the image component of the current block based on the prediction mode and determining the corresponding first reconstruction value, the enhancement model can be further used to perform quality enhancement processing on the first reconstruction value, wherein the enhancement model can include a transformer based on a multi-head self-attention mechanism. In other words, the encoding and decoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding process to enhance the quality of the reconstruction value of the image component, which is beneficial to improving the reconstruction quality of the image component and providing a more accurate reference for the intra-frame prediction of subsequent coding blocks of the frame, thereby effectively improving the encoding and decoding efficiency and video compression performance.

[0340] In another embodiment of the present application, based on the same inventive concept as the above embodiment, refer to FIG23 , which shows a schematic diagram of the structure of the decoder 230 proposed in the embodiment of the present application. As shown in FIG23 , the decoder 230 may include: a second determining unit 2301; wherein,

[0341] The second determination unit 2301 is configured to decode the code stream and determine the decoding parameters of the current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; determine a first reconstruction value of the image component of the current block according to the prediction mode and the prediction residual; when the identification parameter indicates that the current block is to be decoded using an enhancement model, determine the reconstruction value of the image component of the current block based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure.

[0342] It should be noted that, in the embodiment of the present application, the decoder 230 can also be regarded as a data processing mode (or "entropy decoder"), which is used to decode the values ​​of the syntax elements to be decoded.

[0343] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.

[0344] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, this embodiment provides a computer-readable storage medium for use in decoder 230. The computer-readable storage medium stores a computer program that, when executed by the second processor, implements any of the methods described in the aforementioned embodiments.

[0345] Based on the composition of the decoder 230 and the computer-readable storage medium, refer to Figure 24, which shows a specific hardware structure diagram of the decoder 230 provided in an embodiment of the present application. As shown in Figure 24, the decoder 230 may include: a second communication interface 2301, a second memory 2302 and a second processor 2303; each component is coupled together through a second bus system 2304. It can be understood that the second bus system 2304 is used to achieve connection and communication between these components. In addition to the data bus, the second bus system 2304 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the second bus system 2304 in Figure 11. Among them,

[0346] The second communication interface 2301 is used to receive and send signals during the process of sending and receiving information between other external network elements;

[0347] The second memory 2302 is used to store computer programs that can be run on the second processor 2303;

[0348] The second processor 2303 is configured to, when running the computer program, perform the following steps: configuring to decode the code stream, determining decoding parameters of the current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; determining a first reconstruction value of the image component of the current block according to the prediction mode and the prediction residual; and determining a reconstruction value of the image component of the current block based on the enhancement model when the identification parameter indicates that the current block is to be decoded using an enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure.

[0349] Optionally, as another embodiment, the second processor 2303 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.

[0350] It can be understood that the hardware functions of the second memory 2302 are similar to those of the first memory 2102, and the hardware functions of the second processor 2303 are similar to those of the first processor 2103; they will not be described in detail here.

[0351] The present embodiment provides a decoder that decodes a bitstream and determines decoding parameters of a current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; determines a first reconstruction value of an image component of the current block based on the prediction mode and the prediction residual; and when the identification parameter indicates that an enhancement model is used to decode the current block, determines the reconstruction value of the image component of the current block based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure. Thus, in the embodiment of the present application, after completing the prediction of the image component of the current block based on the prediction mode and determining the corresponding first reconstruction value, the enhancement model can be further selected to perform quality enhancement processing on the first reconstruction value, wherein the enhancement model can include a transformer based on a multi-head self-attention mechanism. That is, the encoding and decoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding process to enhance the quality of the reconstruction value of the image component, which is beneficial to improving the reconstruction quality of the image component and providing a more accurate reference for the intra-frame prediction of subsequent coded blocks of the frame, thereby effectively improving the encoding and decoding efficiency and video compression performance.

[0352] In yet another embodiment of the present application, referring to FIG25 , a schematic diagram of the structure of a coding and decoding system proposed in an embodiment of the present application is shown. As shown in FIG25 , a coding and decoding system 250 may include an encoder 210 and a decoder 230 .

[0353] In an embodiment of the present application, the encoder 210 may be the encoder described in any one of the aforementioned embodiments, and the decoder 230 may be the decoder described in any one of the aforementioned embodiments.

[0354] Furthermore, an embodiment of the present application also proposes a code stream, wherein the code stream is generated by bit encoding based on the information to be encoded; wherein the information to be encoded includes at least: a prediction residual, a prediction mode of the current block, the identification parameter, and the preset quantization parameter.

[0355] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0356] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0357] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0358] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0359] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0360] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability

[0361] The embodiments of the present application provide a coding and decoding method, a code stream, an encoder, a decoder, and a storage medium. At the decoding end, the code stream is decoded to determine the decoding parameters of the current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; a first reconstruction value of an image component of the current block is determined based on the prediction mode and the prediction residual; when the identification parameter indicates that an enhancement model is used to decode the current block, the reconstruction value of the image component of the current block is determined based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure. At the encoding end, the first reconstruction value of the image component of the current block is determined based on the prediction mode and the prediction residual of the current block; the reconstruction value of the image component corresponding to the current block is determined based on the enhancement model; wherein the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure; the identification parameter is determined based on the first reconstruction value and the reconstruction value of the image component; and the decoding parameters of the current block are written into the code stream; wherein the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, and the preset quantization parameter. It can be seen that in the embodiment of the present application, after completing the prediction of the image component of the current block based on the prediction mode and determining the corresponding first reconstruction value, it is possible to further use an enhancement model to perform quality enhancement processing on the first reconstruction value, wherein the enhancement model may include a transformer based on a multi-head self-attention mechanism. In other words, the encoding and decoding method proposed in the embodiment of the present application can apply the enhancement model to the intra-frame encoding process to enhance the quality of the reconstruction value of the image component, which is beneficial to improving the reconstruction quality of the image component and providing a more accurate reference for the intra-frame prediction of the subsequent coding blocks of the frame, thereby effectively improving the encoding and decoding efficiency and video compression performance.

Claims

1. A decoding method, applied to a decoder, the method comprising: Decoding a bitstream to determine decoding parameters of a current block, wherein the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; Determining a first reconstruction value of an image component of the current block according to the prediction mode and the prediction residual; When the identification parameter indicates that an enhancement model is used to decode the current block, determining a reconstruction value of the image component of the current block based on the enhancement model; wherein input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure.

2. The method according to claim 1, wherein, The enhancement model is composed of a cascaded shallow extraction module and a deep extraction module; wherein both the shallow extraction module and the deep extraction module include the transformer network structure, The shallow extraction module is configured to perform multi-level and multi-scale feature extraction processing; The deep extraction module is configured to perform multi-level and multi-scale feature extraction and feature fusion processing.

3. The method according to claim 2, wherein, The enhancement model includes a first connection layer.

4. The method according to claim 3, wherein The determining a reconstruction value of the image component of the current block based on the enhancement model includes: Concatenating the first reconstruction value and the preset quantization parameter through the first connection layer to determine a first concatenated component; Inputting the first concatenated component into the shallow extraction module to output first feature information; Inputting the first feature information into the deep extraction module to output second feature information; Determining a reconstruction value of the image component according to the first reconstruction value and the second feature information.

5. The method according to claim 4, wherein, The shallow extraction module includes a first convolutional layer, a first feature extraction module, and a second feature extraction module, wherein both the first feature extraction module and the second feature extraction module include the transformer network structure.

6. The method according to claim 5, wherein The inputting the first concatenated component into the shallow extraction module to output first feature information includes: Determining a first feature based on the first concatenated component through the first convolutional layer; Inputting the first feature into a first transformer network in the first feature extraction module to output a second feature; Determining a third feature based on the second feature through a second transformer network in the first feature extraction module; Inputting the third feature into a third transformer network in the second feature extraction module to output a fourth feature; Determining the first feature information based on the fourth feature through a fourth transformer network in the second feature extraction module.

7. The method according to claim 4, wherein The deep extraction module includes a second connection layer, a third connection layer, a fourth connection layer, a third feature extraction module, a fourth feature extraction module, a second convolutional layer, and a third convolutional layer, wherein both the third feature extraction module and the fourth feature extraction module include the transformer network structure.

8. The method according to claim 7, wherein The inputting the first feature information into the deep extraction module to output second feature information includes: Input the first feature information and the fourth feature into the second connection layer to output a second concatenated component; Based on the second concatenated component, determine a fifth feature through the second convolutional layer and the fifth Transformer network in the third feature extraction module; Input the fifth feature and the third feature into the third connection layer to output a third concatenated component; Based on the third concatenated component, determine a sixth feature through the second convolutional layer and the sixth Transformer network in the third feature extraction module; Input the sixth feature and the second feature into the fourth connection layer to output a fourth concatenated component; Based on the fourth concatenated component, determine a seventh feature through the second convolutional layer and the seventh Transformer network in the fourth feature extraction module; Input the seventh feature into the eighth Transformer network in the fourth feature extraction module to output an eighth feature; Based on the eighth feature, determine the second feature information through the third convolutional layer.

9. The method according to any one of claims 1-8, wherein, The Transformer network structure includes a self-attention mechanism module and a feed-forward network.

10. The method according to claim 1, wherein The method further includes: The prediction residual is the prediction residual of the image component of the current block.

11. The method according to claim 10, wherein, The determining the first reconstruction value of the image component of the current block according to the prediction mode and the prediction residual includes: Determine the predicted value of the image component of the current block according to the prediction mode; Determine the first reconstruction value according to the predicted value and the prediction residual.

12. The method according to any one of claims 1-10, wherein, The method further includes: In the case where the identification parameter indicates not to use the enhancement model to decode the current block, determine the first reconstruction value as the reconstruction value of the image component of the current block.

13. The method according to claim 12, wherein The method further includes: In the case where the value of the identification parameter is a first value, determine that the identification parameter indicates to use the enhancement model to decode the current block; In the case where the value of the identification parameter is a second value, determine that the identification parameter indicates not to use the enhancement model to decode the current block.

14. An encoding method applied to an encoder, the method includes: Determine the first reconstruction value of the image component of the current block according to the prediction mode and the prediction residual of the current block; Determine the reconstruction value of the image component corresponding to the current block based on the enhancement model; wherein, the input parameters of the enhancement model include the first reconstruction value and a preset quantization parameter, and the enhancement model includes a Transformer network structure; Determine an identification parameter according to the first reconstruction value and the reconstruction value of the image component; Write the decoding parameters of the current block into the code stream; wherein, the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, the preset quantization parameter.

15. The method according to claim 14, wherein, The enhancement model is composed of a shallow extraction module and a deep extraction module in cascade; wherein, both the shallow extraction module and the deep extraction module include the Transformer network structure, The shallow extraction module is used to perform multi-level and multi-scale feature extraction processing; The deep extraction module is used to perform multi-level and multi-scale feature extraction and feature fusion processing.

16. The method according to claim 15, wherein The enhancement model includes a first connection layer.

17. The method according to claim 16, wherein Determining the reconstruction value of the image component of the current block based on the enhancement model includes: Concatenating the first reconstruction value and the preset quantization parameter through the first connection layer to determine a first concatenated component; Inputting the first concatenated component into the shallow extraction module to output first feature information; Inputting the first feature information into the deep extraction module to output second feature information; Determining the reconstruction value of the image component according to the first reconstruction value and the second feature information.

18. The method according to claim 17, wherein, The shallow extraction module includes a first convolutional layer, a first feature extraction module, and a second feature extraction module, where both the first feature extraction module and the second feature extraction module include the transformer network structure.

19. The method according to claim 18, wherein, Inputting the first concatenated component into the shallow extraction module to output first feature information includes: Determining a first feature based on the first concatenated component through the first convolutional layer; Inputting the first feature into the first transformer network in the first feature extraction module to output a second feature; Determining a third feature based on the second feature through the second transformer network in the first feature extraction module; Inputting the third feature into the third transformer network in the second feature extraction module to output a fourth feature; Determining the first feature information based on the fourth feature through the fourth transformer network in the second feature extraction module.

20. The method according to claim 17, wherein The deep extraction module includes a second connection layer, a third connection layer, a fourth connection layer, a third feature extraction module, a fourth feature extraction module, a second convolutional layer, and a third convolutional layer, where both the third feature extraction module and the fourth feature extraction module include the transformer network structure.

21. The method according to claim 20, wherein, Inputting the first feature information into the deep extraction module to output second feature information includes: Inputting the first feature information and the fourth feature into the second connection layer to output a second concatenated component; Determining a fifth feature based on the second concatenated component through the second convolutional layer and the fifth transformer network in the third feature extraction module; Inputting the fifth feature and the third feature into the third connection layer to output a third concatenated component; Determining a sixth feature based on the third concatenated component through the second convolutional layer and the sixth transformer network in the third feature extraction module; Inputting the sixth feature and the second feature into the fourth connection layer to output a fourth concatenated component; Determining a seventh feature based on the fourth concatenated component through the second convolutional layer and the seventh transformer network in the fourth feature extraction module; Network Input the seventh feature into the eighth Transformer network in the fourth feature extraction module to output the eighth feature; Based on the eighth feature, determine the second feature information through the third convolutional layer.

22. The method according to any one of claims 14-21, wherein The Transformer network structure includes a self-attention mechanism module and a feed-forward network.

23. The method according to claim 14, wherein, Determining the identification parameter according to the first reconstruction value and the reconstruction value of the image component includes: Determine the first-generation value corresponding to the first reconstruction value and the second-generation value corresponding to the reconstruction value of the image component; When the first-generation value is greater than the second-generation value, set the value of the identification parameter to indicate that the enhancement model is used to decode the current block; When the first-generation value is less than or equal to the second-generation value, set the value of the identification parameter to indicate that the enhancement model is not used to decode the current block.

24. The method according to claim 23, wherein, The method further includes: When the first-generation value is greater than or equal to the second-generation value, determine the prediction residual of the image component of the current block according to the predicted value of the image component and the initial value of the image component of the current block; Write the prediction residual into the bitstream.

25. A bitstream, wherein, The bitstream is generated by bit encoding according to the information to be encoded; wherein, the information to be encoded at least includes: prediction residual, prediction mode of the current block, the identification parameter, the preset quantization parameter.

26. An encoder, the encoder includes a first determination unit; wherein The first determination part is configured to determine the first reconstruction value of the image component of the current block according to the prediction mode and prediction residual of the current block; Determine the reconstruction value of the image component corresponding to the current block based on the enhancement model; wherein, the input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a Transformer network structure; determine the identification parameter according to the first reconstruction value and the reconstruction value of the image component; write the decoding parameters of the current block into the bitstream; wherein, the decoding parameters include the prediction residual, the prediction mode of the current block, the identification parameter, the preset quantization parameter.

27. An encoder, the encoder includes a first memory and a first processor; wherein The first memory is used to store a computer program that can run on the first processor; The first processor is configured to execute the method according to any one of claims 14 to 24 when running the computer program.

28. A decoder, the decoder includes a second determination unit; wherein The second determination unit is configured to decode a bitstream and determine decoding parameters of a current block, where the decoding parameters include a prediction residual, a prediction mode of the current block, an identification parameter, and a preset quantization parameter; determine a first reconstruction value of an image component of the current block according to the prediction mode and the prediction residual; and based on the enhancement model, determine a reconstruction value of the image component of the current block when the identification parameter indicates that the enhancement model is used to decode the current block; where input parameters of the enhancement model include the first reconstruction value and the preset quantization parameter, and the enhancement model includes a transformer network structure.

29. A decoder, the decoder includes a second memory and a second processor; wherein, the second memory is configured to store a computer program capable of running on the second processor; the second processor is configured to execute the method according to any one of claims 1 to 13 when running the computer program.

30. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the method according to any one of claims 1 to 13, or implements the method according to any one of claims 14 to 24.

Citation Information

Patent Citations

  • Image compressed sensing reconstruction method based on Transform enhanced residual self-encoding network

    CN115984392A

  • Transform-based low-illumination image enhancement method

    CN116342409A

  • Server, method and computer program for monitoring wireless quality about 5g communal network

    KR1020230128990A

  • Image processing method and apparatus, device, system, and storage medium

    WO2023184088A1

Cited By

  • Battery system fault identification method and system based on time sequence contrast learning encoder

    CN120892801A

  • Image defect segmentation method, device and equipment and computer readable storage medium

    CN121661065A