Video decoding method, video encoding method, decoder, encoder, and computer-readable storage medium

By optimizing the number of feature channels and backbone network layers in the preset filtering network, the complexity of neural network loop filtering is reduced, the high complexity problem is solved, and the application on more terminals and high encoding and codec performance is achieved.

WO2025147877A1PCT designated stage expired Publication Date: 2025-07-17GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/071469
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The existing neural network-based loop filtering scheme is too complex, making it difficult to apply on ordinary devices or terminals that do not have neural processing units (NPU) chips, reducing its breadth.

Method used

Reduce the model complexity by reducing the number of feature channels of the preprocessed portion and/or reducing the number of layers of the backbone network in the preset filter network, including adjusting the number of feature channels of the convolution module and the depth of the backbone network layer.

Benefits of technology

On the basis of reducing the computational complexity, high encoding and codec performance is maintained, which improves the widespread application of neural network-based filtering solutions on more terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024071469_17072025_PF_FP_ABST
    Figure CN2024071469_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video decoding method, a video encoding method, a decoder, an encoder, and a computer-readable storage medium, capable of reducing the complexity of a neural network-based filtering scheme while ensuring the encoding and decoding compression performance. The method comprises: by means of a pre-processing portion in a preset filtering network, performing feature extraction on the basis of a reconstructed image block corresponding to a current block to determine an initial feature, the number of feature channels of the pre-processing portion being a preset value; by means of a preset number of backbone networks in the preset filtering network, performing feature extraction on the basis of the initial feature to determine an intermediate feature, the preset value being less than the number of feature channels for input data feature extraction in HOP.2 and LOP.2 filters, and / or the preset number being less than the number of backbone network layers in the HOP.2 and LOP.2 filters; and by means of an output portion in the preset filtering network, determining, on the basis of the intermediate feature, a filtered image block corresponding to the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding and decoding method, decoder, encoder and computer-readable storage medium Technical Field

[0001] The embodiments of the present application relate to video coding and decoding technology, and relate to but are not limited to a video coding and decoding method, a decoder, an encoder, and a computer-readable storage medium. Background Art

[0002] At present, the complexity of the neural network-based loop filtering solution is relatively high, which makes it only applicable to some high-performance terminals and difficult to implement on some ordinary devices or terminals that are not equipped with a neural network processing unit (NPU) chip.

[0003] Therefore, the current neural network-based loop filtering solution has a high complexity, which reduces the wide application of the neural network-based filtering solution.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a video encoding and decoding method, decoder, encoder and computer-readable storage medium, which can achieve higher encoding and decoding compression performance on the basis of reducing the complexity of loop filtering, thereby improving the wide application of neural network-based filters.

[0006] In a first aspect, an embodiment of the present application provides a video decoding method, the method comprising:

[0007] The preprocessing part in the preset filtering network is used to extract features based on the reconstructed image block corresponding to the current block to determine the initial features; the number of feature channels in the preprocessing part is a preset value;

[0008] Determining intermediate features by extracting features based on the initial features using a preset number of backbone networks in the preset filtering network; wherein the preset value is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters;

[0009] The output part of the preset filtering network is used to determine the filtered image block corresponding to the current block according to the intermediate features.

[0010] In a second aspect, an embodiment of the present application further provides an encoding method, the method comprising:

[0011] The preprocessing part in the preset filtering network is used to extract features based on the reconstructed image block corresponding to the current block to determine the initial features; the number of feature channels in the preprocessing part is a preset value;

[0012] Determining intermediate features by extracting features based on the initial features using a preset number of backbone networks in the preset filtering network; wherein the preset value is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters;

[0013] The output part of the preset filtering network is used to determine the filtered image block corresponding to the current block according to the intermediate features.

[0014] In a third aspect, an embodiment of the present application provides a decoder, including:

[0015] A first feature preprocessing part is configured to extract features based on the reconstructed image block corresponding to the current block through a preprocessing part in a preset filtering network to determine an initial feature; the number of feature channels of the preprocessing part is a preset value;

[0016] A first intermediate feature extraction section is configured to perform feature extraction based on the initial features using a preset number of backbone networks in the preset filtering network to determine intermediate features; the preset value is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters;

[0017] The first filtering output part is configured to determine the filtered image block corresponding to the current block according to the intermediate features through the output part in the preset filtering network.

[0018] In a fourth aspect, an embodiment of the present application provides an encoder, including:

[0019] A second feature preprocessing part is configured to extract features based on the reconstructed image block corresponding to the current block through a preprocessing part in a preset filtering network to determine an initial feature; the number of feature channels of the preprocessing part is a preset value;

[0020] A second intermediate feature extraction portion is configured to perform feature extraction based on the initial features using a preset number of backbone networks in the preset filtering network to determine intermediate features; the preset value is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters;

[0021] The second filtering output part is configured to determine the filtered image block corresponding to the current block according to the intermediate features through the output part in the preset filtering network.

[0022] In a fifth aspect, an embodiment of the present application further provides a decoder, including:

[0023] A first memory, configured to store executable instructions;

[0024] The first processor is configured to implement the video decoding method in the embodiment of the present application when executing the executable instructions stored in the first memory.

[0025] In a sixth aspect, an embodiment of the present application further provides an encoder, including:

[0026] a second memory for storing executable instructions;

[0027] The second processor is configured to implement the video encoding method provided in the embodiment of the present application when executing the executable instructions stored in the second memory.

[0028] An embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following:

[0029] The encoding information of the current image;

[0030] The coding information of the current image is determined by performing inter-frame prediction based on a historical filtered image block; the historical filtered image block is determined by filtering a historical block in the historical image through a preset filtering network, including:

[0031] By using a preprocessing part in a preset filtering network, feature extraction is performed based on the historical reconstructed image block corresponding to the historical block to determine the historical initial feature; the number of feature channels of the preprocessing part is a preset value;

[0032] Determining historical intermediate features by extracting features based on the historical initial features using a preset number of backbone networks in the preset filtering network; wherein the preset value is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters;

[0033] The historical filtered image block is determined according to the historical intermediate features through the output part in the preset filtering network.

[0034] The embodiments of the present application provide a video encoding and decoding method, a decoder, an encoder, and a computer-readable storage medium. In the preset filtering network, the number of feature channels of the preprocessing part is a preset value, which is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the number of backbone networks is a preset number, which is less than the number of backbone network layers in the HOP.2 and LOP.2 filters. In this way, by reducing the number of feature channels in the preprocessing part and / or reducing the depth of the backbone network layer, the model complexity is effectively reduced, which is conducive to the application of neural network-based loop filtering tools on more terminals, thereby improving the wide application of neural network-based filtering solutions. Moreover, it has been experimentally verified that the preset filtering network of the embodiment of the present application can still guarantee higher encoding and decoding performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] FIG1 is a block diagram of a current HOP.2 filter;

[0036] FIG2 is a block diagram of a current LOP.2 filter;

[0037] FIG3 is a schematic block diagram of an encoder provided in an embodiment of the present application;

[0038] FIG4 is a schematic block diagram of a decoder according to an embodiment of the present application;

[0039] FIG5 is a schematic diagram of a network architecture of a coding and decoding system provided in an embodiment of the present application;

[0040] FIG6 is a schematic diagram of an optional flow chart of an encoding method provided in an embodiment of the present application;

[0041] FIG7 is a schematic diagram of an optional structure of a preset filtering network provided in an embodiment of the present application;

[0042] FIG8 is a schematic diagram of an optional structure of a preset filtering network provided in an embodiment of the present application;

[0043] FIG9 is a schematic diagram of an optional structure of a preset filtering network provided in an embodiment of the present application;

[0044] FIG10 is a schematic diagram of an optional structure of a preset filtering network provided in an embodiment of the present application;

[0045] FIG11 is a schematic diagram of an optional structure of a backbone network module provided in an embodiment of the present application;

[0046] FIG12 is a schematic diagram of an optional structure of a preset filtering network provided in an embodiment of the present application;

[0047] FIG13 is a schematic diagram of an optional structure of a backbone network module provided in an embodiment of the present application;

[0048] FIG14 is a schematic diagram of an optional structure of a preset filtering network provided in an embodiment of the present application;

[0049] FIG15 is a schematic diagram of an optional structure of a first backbone network module or a second backbone network module provided in an embodiment of the present application;

[0050] FIG16 is a schematic diagram of an optional structure of a preset filtering network provided in an embodiment of the present application;

[0051] FIG17 is a schematic diagram of an optional structure of a first backbone network module provided in an embodiment of the present application;

[0052] FIG18 is a schematic diagram of an optional structure of a second backbone network module provided in an embodiment of the present application;

[0053] FIG19 is a schematic diagram of an optional structure of a preset filtering network provided in an embodiment of the present application;

[0054] FIG20 is a schematic diagram of an optional flow chart of a decoding method provided in an embodiment of the present application;

[0055] FIG21 is a structural diagram of a decoder provided in an embodiment of the present application;

[0056] FIG22 is a second structural diagram of a decoder provided in an embodiment of the present application;

[0057] FIG23 is a structural diagram 1 of an encoder provided in an embodiment of the present application;

[0058] FIG24 is a second structural diagram of an encoder provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to explain the related applications and are not intended to limit the applications. It should also be noted that for ease of description, only the parts relevant to the related applications are shown in the drawings.

[0060] It should be noted that the terms “first”, “second”, “third”, etc. mentioned throughout the specification are only used to distinguish different features and do not have the function of limiting priority, sequence, size relationship, etc.

[0061] The nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:

[0062] H.266 / VVC: Versatile Video Coding (VVC)

[0063] Joint Video Experts Team (JVET)

[0064] Neural-Network based Video Coding (NNVC)

[0065] Deblocking Filter (DBF)

[0066] Sample Adaptive Offset (SAO)

[0067] Adaptive Loop Filter (ALF)

[0068] Coding Unit (CU)

[0069] Coding Tree Unit (CTU)

[0070] Currently, common video codec standards (such as H.266 / VVC) all use a block-based hybrid coding framework. Each frame in the video is divided into square maximum coding units (LCUs) of the same size (such as 128x128, 64x64, etc.). Each LCU can be divided into rectangular coding units (CUs) according to rules. Coding units may also be divided into prediction units (PUs) and transform units (TUs). The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module includes intra-frame prediction and inter-frame prediction. Inter-frame prediction includes motion estimation and motion compensation. Since there is a strong correlation between adjacent pixels in a video frame, intra-frame prediction is used in video coding technology to eliminate spatial redundancy between adjacent pixels. Since there is a strong similarity between adjacent frames in a video, the inter-frame prediction method is used in video coding and decoding technology to eliminate the temporal redundancy between adjacent frames, thereby improving coding efficiency.

[0071] The basic process of a video codec is as follows. At the encoder, a frame is divided into blocks. Intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block for the current block. The prediction block is subtracted from the original image block to obtain a residual block. This residual block is transformed and quantized to obtain a quantization coefficient matrix. This quantization coefficient matrix is ​​entropy-encoded and output to the bitstream. At the decoder, intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block for the current block. The bitstream is then parsed to obtain a quantization coefficient matrix. This quantization coefficient matrix is ​​inversely quantized and inversely transformed to obtain a residual block. The prediction block and residual block are added together to obtain a reconstructed block. The reconstructed blocks form a reconstructed image, which is then subjected to image-based or block-based loop filtering to obtain a decoded image. The encoder also performs similar operations as the decoder to obtain a decoded image. The decoded image can serve as a reference frame for inter-frame prediction for subsequent frames. The block division information determined by the encoder, as well as information about the prediction, transform, quantization, entropy coding, loop filtering, and other modes or parameters, are output to the bitstream if necessary. The decoder determines the same block division information as the encoder by parsing and analyzing the existing information, as well as the mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, thereby ensuring that the decoded image obtained by the encoder is the same as the decoded image obtained by the decoder. The decoded image obtained by the encoder is also usually called a reconstructed image. During prediction, the current block can be divided into prediction units, and during transformation, the current block can be divided into transformation units. The division of prediction units and transformation units can be different. The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. The current block (current block) can be the current coding unit (CU) or the current prediction unit (PU), etc.

[0072] Deep learning and neural networks are currently hot topics across various industries. In computer vision, in particular, deep learning-based methods often hold an overwhelming advantage. Leveraging the powerful learning capabilities of neural networks, neural network-based coding tools often achieve highly efficient encoding. Among them, neural network-based loop filtering methods offer the most outstanding encoding performance, achieving over 10% improvement.

[0073] Neural network-based loop filtering solutions are used in loop filtering to improve the quality of reconstructed images. Currently, a neural network-based loop filtering model (High Operation Point, HOP.2) in JVET can be shown in Figure 1.

[0074] In Figure 1, the input part mainly includes reconstructed samples rec, predicted samples pred, boundary strength BS, mode information IPB and quantization parameter QP. The boundary strength here can be understood as the strength information detected based on the deblocking filter. The mode information indicates that the coding block where the sample is located is intra-frame prediction I, unidirectional inter-frame prediction P and bidirectional inter-frame prediction B. The final quantization parameter includes the basic quantization value BaseQP and the frame-level quantization value SliceQP. In Figure 1, CONV represents the convolution module; d1-d6 represents the number of feature channels of the corresponding convolution module, and 3×3 or 1×1 represents the convolution kernel size of the corresponding convolution module. PReLU represents the ReLU activation function of the learnable parameter P. Figure 1 includes N backbone network modules (Back Bone Block) connected in series. The network structure of each backbone network module can be shown in the dotted box in Figure 1. [C,h,w] represents the feature values ​​of the input backbone network module in three feature dimensions, C represents the number of feature channels, h represents height, and w represents width; C×C1, C×C 21 、C 21 ×C 22 、(C1+C 22 )×C, C×C 31 、C 31 ×C, etc., are the number of feature channels of the corresponding convolutional module in different feature dimensions. The number of feature channels of CONV↓2 is C. Figure 1 shows a single-network model, meaning that a single model processes different color components and different frame types. Therefore, the input information also includes the different color components of the aforementioned input, such as the reconstructed luminance component samples and chrominance component samples in the YUV domain; alternatively, it may also include predicted luminance component samples, predicted chrominance component samples, etc. For Figure 1, N is typically 24.

[0075] Regarding the output part of FIG1 , based on the description of the input part above, the input includes different color components, and the output part also includes different color components, namely, the luminance component filtered samples filteredRec and the chrominance component filtered samples filterRec.

[0076] At present, another neural network-based loop filter model (Low Operation Point, LOP.2) in JVET can be shown in Figure 2. The difference from Figure 1 is that the neural network loop filter in Figure 1 has no difference in the processing of luminance components and chrominance components in the main framework, that is, the Backbone Block of the same path is used for feature extraction and filtering. The main framework of the neural network loop filter in Figure 2 processes the luminance component and the chrominance component separately, that is, the complexity of the corresponding component can be increased or decreased according to the resource situation, so as to achieve the purpose of optimizing resource allocation. As shown in Figure 2, after input fusion, the information of the luminance component enters the luminance channel, and the number of Backbone Blocks of the luminance channel is N. Y Determined; similarly, the network depth of the chroma channel is determined by N UV The different complexity of the luma and chroma channels also biases the network's filtering capabilities towards brightness and chroma. For YUV420 video, the codec focuses more on the compression gain of the luma component, so this model is able to better achieve the benefits of the luma component at a lower complexity.

[0077] The most commonly used complexity metric for neural network tools is the number of multiplications performed by the model. Neural network-based tools used in video codecs also employ a similar metric. By expressing the complexity of multiplications per pixel, the complexity is expressed as KMAC / pixel. KMAC / pixel is used to measure the computational complexity of neural network-based tools during the codec process. KMAC stands for thousand multiplications accumulated, and KMAC / pixel represents the number of thousand multiplications required per pixel by the neural network tool.

[0078] The neural network-based loop filtering solutions in JVET shown in Figures 1 and 2 above are both highly complex. Using KMAC as the metric for complexity calculation, they require multiplications ranging from 17KMACs to 477KMACs. The NPU chips in current mainstream mobile phones have a processing capacity of between 10 and 20 TOPs, which translates to approximately 20 to 30KMACs / pixel. It's important to note that these NPU chips assume that one OP can complete one convolution multiplication. Therefore, these neural network tools can only be used on high-performance mobile phones. These neural network-based filtering tools are not applicable to standard devices or terminals without NPU chips.

[0079] In summary, the current neural network-based loop filtering solutions in JVET, such as HOP.2 or LOP.2 filters, are too complex, which reduces the wide application of neural network-based loop filters.

[0080] The embodiments of the present application provide a video encoding and decoding method, a decoder, an encoder, and a computer-readable storage medium, which can achieve higher encoding and decoding compression performance while reducing loop filtering complexity. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0081] Referring to Figure 3, which shows a schematic block diagram of the composition of an encoder provided in an embodiment of the present application. As shown in Figure 3, the encoder (specifically, a "video encoder") 50 may include a transform and quantization unit 501, an intra-frame estimation unit 502, an intra-frame prediction unit 503, an inter-frame prediction unit 504, a motion estimation unit 505, an inverse transform and inverse quantization unit 506, a filter control analysis unit 507, a filtering unit 508, an encoding unit 509, and a decoded image cache unit 510, etc., wherein the filtering unit 508 can implement deblocking filtering and sample adaptive offset (SAO) filtering, and the encoding unit 509 can implement header information encoding and context-based adaptive binary arithmetic coding (CABAC).For the input original video signal, a video coding block can be obtained by dividing the coding tree unit (CTU). Then, the residual pixel information obtained after intra-frame or inter-frame prediction is transformed by the transformation and quantization unit 501, including transforming the residual information from the pixel domain to the transform domain and quantizing the obtained transform coefficients to further reduce the bit rate; the intra-frame estimation unit 502 and the intra-frame prediction unit 503 are used to perform intra-frame prediction on the video coding block; specifically, the intra-frame estimation unit 502 and the intra-frame prediction unit 503 are used to determine the intra-frame prediction mode to be used to encode the video coding block; the inter-frame prediction unit 504 and the motion estimation unit 505 are used to perform inter-frame prediction coding of the received video coding block relative to one or more blocks in one or more reference frames to provide temporal prediction information; the motion estimation performed by the motion estimation unit 505 is the process of generating motion vectors, The motion vector can estimate the motion of the video coding block, and then the inter-frame prediction unit 504 performs motion compensation based on the motion vector determined by the motion estimation unit 505, so the inter-frame prediction unit 504 can also be called a motion compensation unit; after determining the intra-frame prediction mode, the intra-frame prediction unit 503 is also used to provide the selected intra-frame prediction data to the encoding unit 509, and the motion estimation unit 505 also sends the calculated and determined motion vector data to the encoding unit 509; in addition, the inverse transform and inverse quantization unit 506 is used to reconstruct the video coding block, reconstruct the residual block in the pixel domain, and the reconstruction The reconstructed residual block passes through the filter control analysis unit 507 and the filtering unit 508 to remove blocking artifacts. The reconstructed residual block is then added to a prediction block in the frame of the decoded image cache unit 510 to generate a reconstructed video coding block. The coding unit 509 is used to encode various coding parameters and quantized transform coefficients. In the CABAC-based coding algorithm, the context content can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-frame prediction mode and output the bitstream of the video signal. The decoded image cache unit 510 is used to store the reconstructed video coding block for prediction reference. As the video image encoding progresses, new reconstructed video coding blocks are continuously generated, and these reconstructed video coding blocks are all stored in the decoded image cache unit 510.

[0082] Referring to FIG4 , which shows a block diagram of a decoder according to an embodiment of the present application, the decoder (specifically, a “video decoder”) 60 includes a decoding unit 601, an inverse transform and inverse quantization unit 602, an intra-frame prediction unit 603, an inter-frame prediction unit 604, a filtering unit 605, and a decoded image buffer unit 606. The decoding unit 601 can implement header information decoding and CABAC decoding, and the filtering unit 605 can implement deblocking filtering and SAO filtering. After the input video signal is encoded as shown in FIG3 , a code stream of the video signal is output; the code stream is input to the decoder 60 and first passes through the decoding unit 601 to obtain the decoded transform coefficients; the transform coefficients are processed by the inverse transform and inverse quantization unit 602 to generate a residual block in the pixel domain; the intra-frame prediction unit 603 can be used to generate prediction data for the current video decoding block based on the determined intra-frame prediction mode and the data of the previously decoded block from the current frame or picture; the inter-frame prediction unit 604 is to determine the prediction information for the video decoding block by parsing the motion vector and other associated syntax elements, and use The prediction information is used to generate a prediction block for the video decoding block being decoded; a decoded video block is formed by summing the residual block from the inverse transform and inverse quantization unit 602 with the corresponding prediction block generated by the intra-frame prediction unit 603 or the inter-frame prediction unit 604; the decoded video signal passes through the filtering unit 605 to remove blocking artifacts, thereby improving video quality; the decoded video block is then stored in the decoded image buffer unit 606, which stores reference images used for subsequent intra-frame prediction or motion compensation, and is also used for outputting the video signal, thereby obtaining the restored original video signal.

[0083] Furthermore, the embodiment of the present application also provides a network architecture of a codec system including an encoder and a decoder. FIG5 is a schematic diagram of a network architecture of a codec system provided by the embodiment of the present application. As shown in FIG5 , the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01. During implementation, the electronic device can be various types of devices with video codec functions. For example, the electronic device can include a smart phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensing device, a server, etc., which is not specifically limited in the embodiment of the present application. Here, the decoder or encoder described in the embodiment of the present application can be the above-mentioned electronic device.

[0084] It should be noted that the method of the embodiment of the present application is mainly applied to the filtering unit 508 shown in Figure 3 and the filtering unit 605 shown in Figure 4. In other words, the embodiment of the present application can be applied to both the encoder and the decoder, and can even be applied to both the encoder and the decoder at the same time, but the embodiment of the present application does not specifically limit this.

[0085] It should also be noted that when applied to the filter unit 508 portion, the "current block" specifically refers to the current encoding block; when applied to the filter unit 605 portion, the "current block" specifically refers to the current decoding block. The filter unit 605 portion and the filter unit 605 portion may include a filter unit based on a neural network loop filtering algorithm.

[0086] In one embodiment of the present application, referring to FIG6 , a schematic flow chart of an encoding method provided by an embodiment of the present application is shown. The method may include:

[0087] S101. Perform feature extraction based on the reconstructed image block corresponding to the current block through a preprocessing part in a preset filtering network to determine initial features; the number of feature channels in the preprocessing part is a preset value.

[0088] In S101, the encoder predicts, transforms, quantizes, dequantizes, and de-transforms the current block to determine a reconstructed image block corresponding to the current block. In some embodiments, the current block can be a current coding unit (CU), a current transform unit (TU), a current prediction unit (PU), a current coding block (CB), etc., and is not specifically limited in this embodiment of the present application.

[0089] In S101, the filtering unit of the neural network-based loop filtering algorithm can be implemented as a preset filtering network. The preset filtering network includes a preprocessing portion, a preset number of backbone networks, and an output portion. The preprocessing portion is used to extract initial features from input data; the preset number of backbone networks are used to perform deep feature extraction based on the initial features and determine intermediate features; and the output portion is used to determine filtered image blocks based on the extracted intermediate features.

[0090] In some embodiments, the preprocessing portion includes at least one set of first convolutional components and a second convolutional component. The at least one set of first convolutional components corresponds to at least one input data of a preset filtering network. Each set of first convolutional components is configured to perform feature extraction on one type of input data to obtain input features corresponding to the input data. The at least one input data includes at least a reconstructed image block. The second convolutional component is configured to combine the input features corresponding to each type of input data, perform feature extraction, and determine initial features.

[0091] In some embodiments, the first convolution part includes a first convolution module; the second convolution part includes at least one set of second convolution modules and a second activation function. S101 can be implemented by the following process:

[0092] At least one set of first convolution parts is used to extract features from at least one input data to determine at least one input feature; at least one input feature is merged through the second convolution part, and at least one feature extraction is performed based on the merged input feature to determine the initial feature. Here, the above-mentioned preset value may include: the preset feature channel number of at least one first convolution module and / or the preset feature channel number of the second convolution module. In at least one first convolution module, the preset feature channel number of each convolution module is less than the feature channel number of the convolution module in the HOP.2 and LOP.2 filters. The preset feature channel number of the second convolution module is less than the feature channel number of the corresponding convolution module in the HOP.2 and LOP.2 filters. Exemplarily, the second convolution module may correspond to the convolution module (1×1, d6) and the convolution module (3×3, C) in the HOP.2 filter of Figure 1 or the LOP.2 filter of Figure 2; for the convolution module (1×1, d6), the preset number of feature channels of the second convolution module is less than d6; for the convolution module (3×3, C), the preset number of feature channels of the second convolution module is less than C, or the preset number of feature channels of the second convolution module is less than C. Y +C UV .

[0093] In some embodiments, at least one input data includes at least one of: a reconstructed image block, a predicted image block, a boundary strength, a quantization parameter, and mode information; accordingly, at least one first convolution module includes at least one of: a first convolution module corresponding to the reconstructed image block, a first convolution module corresponding to the predicted image block, a first convolution module corresponding to the boundary strength, a first convolution module corresponding to the quantization parameter, and a first convolution module corresponding to the mode information. For example, at least one group of first convolution modules may be shown as part 701 in FIG. 7 , and the second convolution module may be shown as part 702 in FIG. 7 . Alternatively, at least one group of first convolution modules may be shown as part 801 in FIG. 8 , and the second convolution module may be shown as part 802 in FIG. 8 . The first convolution module corresponding to the quantization parameter includes a first convolution module corresponding to the frame-level quantization value and a first convolution module corresponding to the base quantization value. The at least one group of first convolution parts in FIG. 7 may include a first convolution module and a first activation function (e.g., PReLU) module.

[0094] Based on FIG. 7 or FIG. 8 , the preset number of feature channels of the first convolution module corresponding to the reconstructed image block is a first value; the first value is less than the number of feature channels of the convolution module used for feature extraction of the reconstructed image block in the HOP.2 and LOP.2 filters. For example, the number of feature channels of the convolution module used for feature extraction of the reconstructed image block in the HOP.2 and LOP.2 filters can be the value of d1 in FIG. 1 or FIG. 2 .

[0095] And / or, the preset number of feature channels of the first convolution module corresponding to the predicted sample is a second value; the second value is less than the number of feature channels of the convolution module used to extract features from the predicted sample in the HOP.2 and LOP.2 filters. Exemplarily, the number of feature channels of the convolution module used to extract features from the predicted sample in the HOP.2 and LOP.2 filters can be the value of d2 in Figures 1 and 2.

[0096] And / or, the preset number of feature channels of the first convolution module corresponding to the boundary strength is a third value; the third value is less than the number of feature channels of the convolution module used for feature extraction of boundary strength in the HOP.2 and LOP.2 filters. Exemplarily, the number of feature channels of the convolution module used for feature extraction of boundary strength in the HOP.2 and LOP.2 filters can be the value of d3 in Figures 1 and 2.

[0097] And / or, the preset number of feature channels of the first convolution module corresponding to the quantization parameter is a fourth value; the fourth value is less than the number of feature channels of the convolution module used for feature extraction of the quantization parameter in the HOP.2 and LOP.2 filters. Exemplarily, the number of feature channels of the convolution module used for feature extraction of the quantization parameter in the HOP.2 and LOP.2 filters can be the value of d4 in Figures 1 and 2.

[0098] And / or, the number of feature channels of the first convolution module corresponding to the pattern information is a fifth value; the fifth value is less than the number of feature channels of the convolution module used to extract features from the pattern information in the HOP.2 and LOP.2 filters. Exemplarily, the number of feature channels of the convolution module used to extract features from the pattern information in the HOP.2 and LOP.2 filters can be the value of d5 in Figures 1 and 2.

[0099] It should be noted that the applicant has found through experiments that reducing the types of input information has little impact on the network model's capabilities. For example, removing boundary strength or predicted image blocks from at least one input data can help further reduce the complexity of the network model and lower the hardware implementation overhead.

[0100] S102: extracting features based on the initial features through a preset number of backbone networks in the preset filtering network to determine intermediate features.

[0101] In the embodiment of the present application, the preset number of backbone networks in the preset filtering network is equivalent to a preset number of layers of backbone networks. The input of the first backbone network in the preset number of backbone networks is connected to the output of the preprocessing part; the input of the i-th backbone network is connected to the output of the i-1-th backbone network (i is an integer greater than 1), and the output of the last backbone network is connected to the output part of the preset filtering network. In this way, through the preset number of backbone networks, layer-by-layer feature extraction is performed based on the initial features output by the preprocessing part, and deeper feature information is extracted from the initial features and determined as intermediate features.

[0102] In the embodiment of the present application, the preset value is smaller than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the preset number is smaller than the number of backbone network layers in the HOP.2 and LOP.2 filters. In other words, the number of feature channels in the preprocessing portion is smaller than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the number of backbone network layers is smaller than the number of backbone network layers in the HOP.2 and LOP.2 filters.

[0103] The applicant has found through experiments that for at least one first convolution module or second convolution part in the pre-processing part, the number of characteristic channels has little effect on the quality of the subsequent filtered image, so the number of characteristic channels of the pre-processing part can be optimized to reduce the complexity of the overall filtering network. And / or, based on Figures 1 and 2, it can be seen that the backbone network module contains multiple layers of convolution modules. Therefore, reducing the number of backbone network modules, that is, reducing the number of layers of the backbone network modules, can greatly reduce the complexity of the overall filtering network. In practical applications, the number of characteristic channels of the pre-processing part can be set to a preset value, or the number of backbone networks can be set to a preset number, or both the number of characteristic channels of the pre-processing part and the number of backbone networks can be set to a preset number, etc., to reduce the complexity of the overall filtering network. The specific selection is made according to the actual situation, and the embodiments of the present application are not limited thereto.

[0104] S103: Determine the filtered image block corresponding to the current block according to the intermediate features through the output part of the preset filtering network.

[0105] In S103, prediction is performed based on the intermediate features through the output part of the preset filtering network, thereby completing the filtering of the current block and determining the filtered image block corresponding to the current block.

[0106] In some embodiments, the output part may include a feature extraction part and an inference part. The feature extraction part further extracts features from the intermediate features, and makes predictions based on the further extracted features to determine the filtered image blocks.

[0107] For example, based on FIG7 , the output portion may be shown as portion 901 in FIG9 , and based on FIG8 , the output portion may be shown as portion 1001 in FIG10 . The output portion is used to determine the chroma component filtered image block and the luminance component filtered image block corresponding to the current block, and the filtered image block corresponding to the current block is determined by combining the chroma component filtered image block and the luminance component filtered image block.

[0108] In some embodiments, based on Figure 9, a 3×3 convolution module and a PReLU activation function may be connected between the output terminals of the preset number of backbone network modules and the input terminals of the output portion. Specific selections are made based on actual conditions and are not limited in this embodiment of the application.

[0109] In some embodiments, based on Figure 9, a 3×3 convolution module and a PReLU activation function module may be connected between the output terminals of a preset number of backbone network modules and the input terminals of the output portion. Specific selections are made based on actual conditions and are not limited in this embodiment of the application.

[0110] It is understandable that in the preset filtering network of the embodiment of the present application, the number of characteristic channels of the preprocessing part is a preset value, which is less than the number of characteristic channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the number of backbone networks is a preset number, which is less than the number of backbone network layers in the HOP.2 and LOP.2 filters. In this way, by reducing the number of characteristic channels of the preprocessing part and / or reducing the depth of the backbone network layer, the model complexity is effectively reduced, which is conducive to the application of neural network-based loop filtering tools on more terminals, thereby improving the wide application of neural network-based filtering solutions. Moreover, it has been experimentally verified that the preset filtering network of the embodiment of the present application can still guarantee high encoding and decoding performance under extremely low computational complexity (such as 5KMAC / pixel).

[0111] In some embodiments, the preset number of backbone networks includes: a preset number of backbone network modules connected in series. In the preset number of backbone network modules connected in series, each backbone network module includes: a first backbone network convolution module, a first separation convolution module, a third activation function, a second backbone network convolution module, and a second separation convolution module. The process of S102 can be implemented by the following process:

[0112] The first backbone network convolution module and the first separation convolution module are used to extract features of the initial features or the features output by the previous backbone network module respectively, and the features extracted by the first backbone network convolution module and the first separation convolution module are combined and nonlinearly operated through the third activation function, and the features determined by the nonlinear operation are converged through the second backbone network convolution module, and the converged features are extracted through the second separation convolution module to determine the intermediate features.

[0113] Among them, the number of feature channels of at least one convolution module among the first backbone network convolution module, the first separation convolution module, the second backbone network convolution module and the second separation convolution module is less than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0114] Exemplarily, based on Figure 7 or Figure 9, a network structure of the backbone network module can be as shown in Figure 11, including: a first backbone network convolution module 1101, a first separation convolution module 1102, a third activation function module 1103, a second backbone network convolution module 1104 and a second separation convolution module 1105. Among them, the number of feature channels corresponding to the first backbone network convolution module 1101 is determined based on the preset number of first feature channels; the first separation convolution module 1102 includes a first convolution kernel 11021 and a second convolution kernel 11022; the number of feature channels corresponding to the first convolution kernel 11021 is determined based on the preset number of second feature channels; the number of feature channels corresponding to the second convolution kernel 11022 is determined based on the preset number of second feature channels and the preset number of third feature channels; the number of feature channels corresponding to the second backbone network convolution module 1104 is determined based on the preset number of first feature channels and the preset number of third feature channels; the second separation convolution module 1105 includes a third convolution kernel 11051 and a fourth convolution kernel 11052; the number of feature channels corresponding to the third convolution kernel 11051 is determined based on the preset number of fourth feature channels; the number of feature channels corresponding to the fourth convolution kernel 11052 is determined based on the preset number of fourth feature channels. In Figure 11, [C, h, w] represents the feature values ​​of the input or output features corresponding to each backbone network module in three feature dimensions.

[0115] Among them, at least one of the preset first feature channel number, the preset second feature channel number, the preset third feature channel number and the preset fourth feature channel number is smaller than the feature channel number of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

[0116] Exemplarily, as shown in FIG12 , in the above-mentioned preprocessing part, the preset number of feature channels of the first convolution module corresponding to the reconstructed image block, i.e., the first value, is 4; the preset number of feature channels of the first convolution module corresponding to the predicted sample, i.e., the second value, is 2; the preset number of feature channels of the first convolution module corresponding to the boundary strength, i.e., the third value, is 2; the preset number of feature channels of the first convolution module corresponding to the quantization parameter, i.e., the fourth value, is 1; the preset number of feature channels of the first convolution module corresponding to the pattern information, i.e., the fifth value, is 1. The preset number of feature channels of the second convolution module is 8; the preset filtering network includes 16 backbone network modules connected in series, and the network structure of each of the 16 backbone network modules can be shown in FIG11 . In each backbone network module, the preset number of first feature channels is 4; the preset number of second feature channels is 2; the preset number of third feature channels is 4; and the preset number of fourth feature channels is 4.

[0117] As shown in Figure 12, in each backbone network module, the number of feature channels corresponding to the first backbone network convolution module is 8×4, where 8 represents the number of feature channels C of the input features (such as initial features or features output by the previous backbone network module) [C, h, w] corresponding to the backbone network module; 4 is the number of preset first feature channels corresponding to the first backbone network convolution module. The number of feature channels corresponding to the first convolution kernel is 8×2; where 8 represents the number of feature channels C of the input features [C, h, w] corresponding to the backbone network module, and 2 is the number of preset second feature channels corresponding to the first convolution kernel. The number of feature channels corresponding to the second convolution kernel is 2×4, where 2 represents the number of preset second feature channels corresponding to the second convolution kernel, and 4 represents the number of preset third feature channels corresponding to the second convolution kernel. The number of feature channels corresponding to the second backbone network convolution module is 8×8, where the first 8 represents the sum of the preset first feature channel number (4) and the preset third feature channel number (4), and the second 8 represents the number of feature channels C of the input features [C, h, w] corresponding to the backbone network module. The third convolution kernel corresponds to an 8×4 feature channel count, where 8 represents the number of feature channels C of the input features [C, h, w] corresponding to the backbone network module, and 4 represents the number of preset fourth feature channels. The fourth convolution kernel corresponds to an 8×8 feature channel count, where 4 represents the number of preset fourth feature channels, and 8 represents the number of feature channels C of the input features [C, h, w] corresponding to the backbone network module.

[0118] It can be understood that, based on the HOP.2 filter shown in FIG1 , the embodiment of the present application can reduce the complexity of the filter network and improve the wide application of the neural network filtering tool by reducing the number of feature channels of the convolution module in the preset processing part and / or the convolution module in the backbone network module. In addition, the values ​​of the first value, the second value, the third value, the fourth value, the fifth value, the preset number of feature channels of the second convolution module, and the preset number of first feature channels, the preset number of second feature channels, the preset number of third feature channels, and the preset number of fourth feature channels, etc., have been verified through experiments to be able to maintain codec compression performance while reducing complexity.

[0119] In some embodiments, based on FIG. 7 or FIG. 9 , the preset number of backbone networks includes: a preset number of backbone network modules connected in series. Each of the preset number of backbone network modules connected in series includes: a third separation convolution module, a fourth activation function module, and a fourth separation convolution module. The above S102 can be implemented by the following process:

[0120] The third separation convolution module is used to extract features from the initial features or the features output by the previous backbone network module, and the fourth activation function is used to perform a nonlinear operation on the features extracted by the third separation convolution module, and the fourth separation convolution module is used to extract features from the features determined by the nonlinear operation to determine the intermediate features.

[0121] Exemplarily, based on FIG. 7 or FIG. 9 , a network structure of the backbone network module may be as shown in FIG. 13 , including: a third separation convolution module 1301 , a fourth activation function module 1302 and a fourth separation convolution module 1303 .

[0122] Exemplarily, as shown in Figure 14, in the above-mentioned preprocessing part, the preset feature channel number of the first convolution module corresponding to the reconstructed image block, that is, the first value, is 4; the preset feature channel number of the first convolution module corresponding to the predicted sample, that is, the second value, is 2; the preset feature channel number of the first convolution module corresponding to the boundary strength, that is, the third value, is 2; the preset feature channel number of the first convolution module corresponding to the quantization parameter, that is, the fourth value, is 1; the feature channel number of the first convolution module corresponding to the pattern information, that is, the fifth value, is 1; the preset feature channel number of the second convolution module is 8; the preset filtering network includes 7 backbone network modules connected in series, and for each of the 7 backbone network modules, the preset feature channel number of the convolution kernel in the third separation convolution module and the fourth separation convolution module is determined based on the feature channel number of the initial feature and the sixth value, such as the feature channel number of the convolution kernel in the third separation convolution module and the fourth separation convolution module is 8×8, wherein the feature channel number of the initial feature is 8 and the sixth value is 8.

[0123] It is understandable that, due to the extremely low complexity, the performance gain brought by the bifurcated path may not be as good as that of the separated convolution. The embodiment of the present application is based on the HOP.2 filter, which removes the branch structure in the backbone network module and adopts a backbone network module with a single branch structure (for example, two pairs of 1x3 and 3x1 separated convolutions are used, and PReLU is used in the middle of the two convolution pairs for nonlinear operations to ensure the depth of the network model). In this way, the number of feature channels of the relevant convolution module is reduced, thereby reducing the complexity of the filtering network and improving the wide application of neural network filtering tools. In addition, the preset feature channel numbers of the above-mentioned first value, second value, third value, fourth value, fifth value, and second convolution module have been verified by experiments, and can maintain the codec compression performance on the basis of reducing complexity.

[0124] In some embodiments, based on FIG8 or FIG10, the preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: luminance intermediate features and chrominance intermediate features. The above S102 can be implemented by the following process:

[0125] A first preset number of first backbone network modules are used to extract brightness features based on the initial features to determine initial brightness features. A sixth separation convolution module on the first network branch is used to extract features from the initial brightness features to determine intermediate brightness features. A second preset number of second backbone network modules are used to extract chromaticity features from the initial features to determine initial chromaticity features. A seventh separation convolution module on the second network branch is used to extract features from the initial chromaticity features to determine intermediate chromaticity features.

[0126] Among them, the number of feature channels of at least one convolution module among the third backbone network convolution module, the fourth backbone network convolution module, the fifth separation convolution module, the sixth separation convolution module and the seventh separation convolution module is less than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0127] Among them, the first backbone network module includes: a third backbone network convolution module, a fourth backbone network convolution module, and a fifth separation convolution module; the second backbone network module includes: a third backbone network convolution module, a fourth backbone network convolution module, and a fifth separation convolution module; that is, the network structure of the first backbone network module and the second backbone network module are the same, and the number of feature channels of the same backbone network convolution module or separation convolution module is also the same. Exemplarily, based on Figure 8 or Figure 10, the structure of the first backbone network module or the second backbone network module can be as shown in Figure 15, including: a third backbone network convolution module 1501, a fourth backbone network convolution module 1502, and a fifth separation convolution module 1503. It can be seen that the first backbone network module or the second backbone network module can also include a 1×1 convolution module connected to the output end of the fifth separation convolution module 1503. The number of feature channels of this convolution module is determined by the number of feature channels C of the features input to the first backbone network module or the second backbone network module.

[0128] In some embodiments, the number of feature channels corresponding to the third backbone network convolution module 1501 is determined based on a preset first number of feature channels; the number of feature channels corresponding to the fourth backbone network convolution module 1502 is determined based on a preset first number of feature channels; the number of feature channels corresponding to the fifth separation convolution module 1503 is determined based on a preset second number of feature channels; the number of feature channels corresponding to the sixth separation convolution module is determined based on a preset fifth number of feature channels; and the number of feature channels corresponding to the seventh separation convolution module is determined based on a preset sixth number of feature channels. At least one of the preset first number of feature channels, the preset second number of feature channels, the preset fifth number of feature channels, and the preset sixth number of feature channels is less than the number of feature channels of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

[0129] For example, as shown in FIG16 , in the preprocessing portion, the preset number of feature channels of the first convolution module corresponding to the reconstructed image block, i.e., the first value, is 4; the preset number of feature channels of the first convolution module corresponding to the predicted sample, i.e., the second value, is 2; the preset number of feature channels of the first convolution module corresponding to the boundary strength, i.e., the third value, is 1; the preset number of feature channels of the first convolution module corresponding to the quantization parameter, i.e., the fourth value, is 1; and the preset number of feature channels of the first convolution module corresponding to the pattern information, i.e., the fifth value, is 1. The second convolution module includes a first merging convolution module 1601 and a first downsampling convolution module 1602. The preset number of feature channels corresponding to the first merging convolution module 1601 is 4; the preset number of feature channels corresponding to the first downsampling convolution module 1602 is 32. The preset number of feature channels corresponding to the first downsampling convolution module 1602 includes 16 feature channels corresponding to the luminance component and 16 feature channels corresponding to the chrominance component. The preset number of backbone networks includes: 8 first backbone network modules on the first network branch, and 1 second backbone network module on the second network branch. The first network branch is used to extract features based on the luminance components of the 16 feature channels, and the first network branch is used to extract features based on the chrominance components of the 16 feature channels.

[0130] For each first backbone network module or second backbone network module, the preset number of first characteristic channels is 32; the preset number of second characteristic channels is 16; the preset number of fifth characteristic channels is 32; and the preset number of sixth characteristic channels is 16.

[0131] As shown in Figure 16, the number of feature channels corresponding to the third backbone network convolution module is 16×32, where 16 represents the number of feature channels C of the input features (such as initial features or features output by the previous backbone network module) [C, h, w] corresponding to this backbone network module; and 32 is the preset number of first feature channels corresponding to the third backbone network convolution module. The number of feature channels corresponding to the fourth backbone network convolution module is 32×16, where 16 represents the number of feature channels C of the input features [C, h, w] corresponding to this backbone network module; and 32 is the preset number of first feature channels corresponding to the fourth backbone network convolution module. The number of feature channels corresponding to the fifth separation convolution module is 16×16, where one 16 represents the number of feature channels C of the input features [C, h, w] corresponding to this backbone network module, and the other 16 represents the preset number of second feature channels. The number of feature channels corresponding to the sixth separation convolution module 1603 is the preset number of fifth feature channels, i.e., 32. The number of feature channels corresponding to the seventh separation convolution module 1604 is the preset number of sixth feature channels, i.e., 16.

[0132] It can be understood that based on the LOP.2 filter shown in Figure 2, the embodiment of the present application can reduce the complexity of the filtering network and improve the wide application of the neural network filtering tool by reducing the number of characteristic channels of the convolution module in the preset processing part and / or the convolution module in the backbone network module. In addition, the above-mentioned first value, second value, third value, fourth value, fifth value, the preset number of characteristic channels of the second convolution module, and the values ​​of the preset first characteristic channel number, the preset second characteristic channel number, the preset fifth characteristic channel number, and the preset sixth characteristic channel number have been verified through experiments to be able to maintain the codec compression performance while reducing the complexity.

[0133] In some embodiments, based on FIG8 or FIG10, the preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: luminance intermediate features and chrominance intermediate features. The above S102 can be implemented by the following process:

[0134] Through a first preset number of first backbone network modules, brightness features are extracted based on the initial features to determine the initial brightness features; the first backbone network module includes an eighth separation convolution module; through the sixth separation convolution module on the first network branch, feature extraction is performed on the initial brightness features to determine the intermediate brightness features; through a second preset number of second backbone network modules, chroma features are extracted on the initial features to determine the initial chroma features; the second backbone network module includes a ninth separation convolution module; through the seventh separation convolution module on the second network branch, feature extraction is performed on the initial chroma features to determine the intermediate chroma features.

[0135] Exemplarily, based on Figure 8 or Figure 10, a network structure of the first backbone network module can be as shown in Figure 17, including an eighth separation convolution module 1701. A network structure of the second backbone network module can be as shown in Figure 18, including a ninth separation convolution module 1801. Among them, the module structure of the eighth separation convolution module and the ninth separation convolution module is the same, but the number of feature channels is different. It can be seen that in some embodiments, the output end of the eighth separation convolution module 1701 or the ninth separation convolution module 1801 can also be connected to a PReLU activation function module and a 1×1 convolution module, and the number of feature channels of the 1×1 convolution module is determined by the number of feature channels C of the features input to the backbone network module.

[0136] For example, as shown in FIG19 , in the preprocessing portion, the preset number of feature channels of the first convolution module corresponding to the reconstructed image block, i.e., the first value, is 8; the preset number of feature channels of the first convolution module corresponding to the predicted sample, i.e., the second value, is 4; the preset number of feature channels of the first convolution module corresponding to the boundary strength, i.e., the third value, is 1; the preset number of feature channels of the first convolution module corresponding to the quantization parameter, i.e., the fourth value, is 1; and the preset number of feature channels of the first convolution module corresponding to the pattern information, i.e., the fifth value, is 1. The second convolution module includes a merging convolution module 1901 and a downsampling convolution module 1902. The preset number of feature channels corresponding to the second merging convolution module 1901 is 8; the preset number of feature channels corresponding to the second downsampling convolution module 1902 is 32. The preset number of feature channels corresponding to the second downsampling convolution module 1902 includes 16 feature channels corresponding to the luminance component and 16 feature channels corresponding to the chrominance component. The first network branch includes 7 first backbone network modules, and the second network branch includes 2 second backbone network modules. The first network branch is used to extract features based on the luminance components of the 16 feature channels, and the first network branch is used to extract features based on the chrominance components of the 16 feature channels.

[0137] For each first backbone network module, the number of feature channels corresponding to the eighth separation convolution module is determined based on the number of feature channels of the initial features. For each second backbone network module, the number of feature channels corresponding to the ninth separation convolution module is less than the number of feature channels corresponding to the eighth separation convolution module. For example, as shown in FIG19 , the number of feature channels corresponding to the eighth separation convolution module is 32; and the number of feature channels corresponding to the ninth separation convolution module is 16.

[0138] As can be understood, based on the LOP.2 filter shown in Figure 2, backbone network modules with different numbers of feature channels are applied to the chroma and luma component pathways. The backbone network modules for the chroma and luma component pathways share the same primary structure, consisting of a pair of 1x3 and 3x1 separable convolutions followed by a PReLU nonlinear operation, and finally a 1x1 convolution for feature convergence. However, the number of feature channels in the backbone network modules for the chroma and luma component pathways differs. This ensures compression performance by using more feature channels for feature extraction in the luma component pathway, which has richer feature information, while reducing the complexity of the filtering network for the chroma component pathway, which has less feature information. This reduces the complexity of the filtering network while ensuring coding compression performance, thereby increasing the applicability of neural network filtering tools. Furthermore, the preset feature channel numbers for the first, second, third, fourth, fifth, and second convolution modules, as well as the preset feature channel numbers for the eighth and ninth separable convolution modules, have been experimentally verified to maintain codec compression performance while reducing complexity.

[0139] In some embodiments, when determining the filtered image block corresponding to the current block in S103, the encoder may further determine the filtered image corresponding to the current image based on the filtered image block corresponding to the current block; and implement inter-frame prediction decoding using the filtered image corresponding to the current image.

[0140] In some embodiments, the above-mentioned separation convolution module can also be replaced by a transformer network model.

[0141] Exemplarily, after prediction, transformation, quantization, inverse quantization, and inverse transformation, the encoder obtains a reconstructed image block and begins loop filtering to improve image quality. A flag for enabling the use of neural network-based loop filtering, sps_nnlf_enable_flag, is set to true. If the flag is set to true, the neural network-based loop filtering tool is enabled; otherwise, the neural network-based loop filtering tool is not enabled.

[0142] Step 1: If the neural network-based loop filtering tool allows the use of the flag bit is true, then execute step 2; otherwise, skip step 2 and directly execute step 3.

[0143] Step 2: Load and initialize the preset filter network according to pre-set parameters. The convolution and other parameters in the preset filter network have been trained and integer quantized. Obtain the reconstructed sample rec, predicted sample pred, boundary strength BS, mode information IPB, quantization information BaseQP, and SliceQP of the current coding tree unit area. This information is input into the preset filter network for inference calculation. The initial filtered image block filteredRec is obtained after the preset filter network.

[0144] If the current frame or current coding tree unit uses residual scaling, the residual scaling parameters are calculated based on the current block (i.e., the original image block), the reconstructed image block, and the filtered image block. The encoder calculates the rate-distortion cost of the solved parameters and the default parameters and selects the optimal parameters. The parameters are then multiplied by the residual of the initial filtered image block filteredRec and the reconstructed image block rec, and the scaled residual is added back to the reconstructed image block rec to obtain the filtered image block output. If the current frame or current coding tree unit does not use residual scaling, the initial filtered image block filteredRec is the filtered image block output.

[0145] After the filtering is completed, if all coding tree units of the current image use neural network-based loop filtering, the image-level flag sh_nnlf_flag is set to the first flag value and written into the bitstream. If some coding tree units in the current image use neural network-based loop filtering, the image-level flag sh_nnlf_flag is set to the second flag value, the coding tree units that use neural network loop filtering technology set ctb_nnlf_flag to true, and the coding tree units that do not use this technology set ctb_nnlf_flag to false. Finally, sh_nnlf_flag is written into the bitstream together with all ctb_nnlf_flags. If all coding tree units of the current image do not use neural network-based loop filtering technology, the image-level flag sh_nnlf_flag is set to the third flag value and written into the bitstream;

[0146] If sh_nnlf_flag is the first flag value or the second flag value, it is necessary to check the usage of the residual scaling technology. If the current frame uses the residual scaling technology, set the image-level flag scaleFlag to true and write it to the bitstream together with the corresponding residual scaling index scaleIdx; if the current image does not use the residual scaling technology, set the image-level flag scaleFlag to false and write it to the bitstream.

[0147] Step 3: The encoder continues to operate other loop filtering techniques.

[0148] Step 4: After executing all loop filtering tools, the final output image is obtained and the code stream information is output.

[0149] In some embodiments, the intermediate features include: at least two intermediate features, the at least two intermediate features corresponding to at least two initial features; the at least two initial features are determined by a preprocessing portion, based on the reconstructed image block, and feature extraction in combination with at least two candidate basic quantization parameters corresponding to the current block; the at least two candidate basic quantization parameters are determined by a basic quantization parameter corresponding to the current block and at least two candidate quantization offsets; the process of determining the filtered image block corresponding to the current block based on the intermediate features through the output portion of the preset filtering network may include:

[0150] Through the output part of the preset filtering network, at least two candidate filtered image blocks corresponding to the current block are determined according to at least two intermediate features; at least two first distortion costs corresponding to the at least two candidate filtered image blocks are determined; and based on the at least two first distortion costs, the filtered image block corresponding to the current block is determined.

[0151] In some embodiments, a target quantization bias among at least two candidate quantization biases is determined based on at least two first distortion costs, and first syntax identification information corresponding to the target quantization bias is determined; entropy encoding is performed on the first syntax identification information, and the obtained coded bits are written into the bitstream.

[0152] In other words, the neural network-based loop filter model uses a wider range of quantization parameters as input during training, resulting in better generalization capabilities. Different quantization values ​​can be selected as input during filtering on the encoder side. For example, the base quantization parameter BaseQP of the current frame is obtained, and the allowed offset Qpoffset {-10, -5, 5, 10} is selected on the encoder side. Different filtered images are obtained, the distortion cost of each filtered image is calculated, and the optimal quantization parameter is selected as input. This offset index Qpoffset Index is then written into the bitstream and transmitted to the decoder. The decoder, based on the parsed offset index, adds the corresponding Qpoffset to BaseQP as the network model input and filters the decoded and reconstructed image.

[0153] In some embodiments, the intermediate features include: at least two intermediate features; the at least two intermediate features correspond to at least two reconstructed image blocks; the at least two reconstructed image blocks are determined by transposing the reconstructed image blocks through at least two preset transposition angles;

[0154] Through the output part of the preset filtering network, according to the intermediate features, the filtered image block corresponding to the current block is determined, including:

[0155] Determining, by an output portion of a preset filtering network, at least two candidate filtered image blocks corresponding to the current block according to at least two intermediate features;

[0156] determining at least two second distortion costs corresponding to at least two candidate filtered image blocks;

[0157] A filtered image block corresponding to the current block is determined according to the at least two second distortion costs.

[0158] In some embodiments, a target transposition angle among at least two preset transposition angles is determined based on the at least two second distortion costs, and second syntax identification information corresponding to the target transposition angle is determined;

[0159] Perform entropy coding on the second syntax identification information, and write the obtained coded bits into the bitstream.

[0160] In some embodiments, the preset filtering network is determined by performing network training on the initial filtering network using a set of reconstructed sample image blocks including reconstructed sample image blocks of at least two preset transposition angles.

[0161] In other words, the encoder can try flipping the current block horizontally, vertically, rotating it 90 degrees forward, rotating it 90 degrees backward, transposing it 45 degrees, and transposing it 135 degrees. The encoder obtains filtered image blocks under different conditions, optimizes the rate-distortion cost for different transposition angles, and selects the optimal method to write the corresponding index GeoTransformIndex into the bitstream and transmit it to the decoder. The decoder then transforms the input and output accordingly based on the geometric transformation index GeoTransformIndex obtained by parsing to obtain the filtered reconstructed image.

[0162] In some embodiments, the intermediate feature includes: a first intermediate feature and a second intermediate feature; the first intermediate feature corresponds to a first reconstructed image block; the second intermediate feature corresponds to a second reconstructed image block; the first reconstructed image block includes an original reconstructed image block corresponding to the current block; the second reconstructed image block includes an initial filtered image block determined by performing Gaussian filtering on the original reconstructed image block;

[0163] Through the output part of the preset filtering network, according to the intermediate features, the filtered image block corresponding to the current block is determined, including:

[0164] Determining, through an output portion of a preset filtering network, a first candidate filtered image block corresponding to the current block based on the first intermediate feature; and determining, based on the second intermediate feature, a second candidate filtered image block corresponding to the current block;

[0165] Determining a third distortion cost corresponding to the first candidate filtered image block and a fourth distortion cost corresponding to the second candidate filtered image block;

[0166] A filtered image block corresponding to the current block is determined according to the third distortion cost and the fourth distortion cost.

[0167] In some embodiments, third syntax identification information is determined based on the third distortion cost and the fourth distortion cost; the third syntax identification information indicates whether Gaussian filtering is performed on the current block; the third syntax identification information is entropy encoded, and the obtained encoded bits are written into the bitstream.

[0168] In other words, current neural network-based loop filtering techniques all use unfiltered reconstructed images as input, resulting in relatively high noise levels. Instead, the reconstructed image can be filtered using a Gaussian filter before being used as input for the network model. After inference, the network model generates a new filtered image. The rate-distortion cost of this new image is calculated and compared with the unfiltered reconstructed image input. If the Gaussian filter yields the lower cost, the Gaussian filter flag is written into the bitstream. Otherwise, the Gaussian filter flag is written into the bitstream.

[0169] In some embodiments, the encoder can use at least two reconstructed sample image sets to perform network training on the initial filtering network and determine a preset filtering network; the at least two reconstructed sample image sets include: a reconstructed sample image set without using loop filtering, a reconstructed sample image set using deblocking filtering, a reconstructed sample image set using sample adaptive compensation filtering, and at least two of the reconstructed sample image sets using adaptive loop filtering.

[0170] In some embodiments, the process of training the initial filter network using at least two reconstructed sample image sets and determining a preset filter network includes: training the initial filter network using at least two reconstructed sample image sets and determining at least two candidate filter networks. Furthermore, the encoder may filter based on the reconstructed image block using the at least two candidate filter networks to determine at least two candidate filtered image blocks corresponding to the at least two candidate filter networks; determine at least two fourth distortion costs corresponding to the at least two candidate filtered images; and determine, based on the at least two fourth distortion costs, a filtered image block corresponding to the current block, and / or determine a preset filter network corresponding to the current image.

[0171] In other words, multiple models can be trained based on different data sets, and the encoding end can decide which model has better filtering effect for the current frame or current coding tree unit.

[0172] In one embodiment of the present application, referring to FIG20 , a flowchart of a decoding method provided by an embodiment of the present application is shown. The method may include:

[0173] S201. Perform feature extraction based on the reconstructed image block corresponding to the current block through a preprocessing part in a preset filtering network to determine initial features; the number of feature channels in the preprocessing part is a preset value.

[0174] In some embodiments, the current block may be a current coding unit (CU), a current transform unit (TU), a current prediction unit (PU), a current coding block (CB), etc., which is not specifically limited in the embodiments of the present application.

[0175] In S201 , the number of feature channels in the preprocessing part is a preset value.

[0176] Among them, the number of feature channels in the preprocessing part, that is, the preset value, is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters.

[0177] S202: extract features based on the initial features through a preset number of backbone networks in the preset filtering network to determine intermediate features.

[0178] S203: Determine the filtered image block corresponding to the current block according to the intermediate features through the output part of the preset filtering network.

[0179] The process in S201-S203 is consistent with the process description in S101-S103 in the encoder, and will not be repeated here.

[0180] In some embodiments, the preprocessing part includes at least one group of first convolution parts and a second convolution part, the at least one group of first convolution parts corresponds to at least one input data; the at least one input data includes at least the reconstructed image block; the first convolution part includes a first convolution module; the second convolution part includes at least one group of second convolution modules and a second activation function module; the preset value includes: a preset number of feature channels of at least one first convolution module and / or a preset number of feature channels of the second convolution module; the preprocessing part in the preset filtering network extracts features based on the reconstructed image block corresponding to the current block to determine the initial features, including:

[0181] performing feature extraction on at least one input data by using the at least one set of first convolution parts to determine at least one input feature;

[0182] Through the second convolution part, the at least one input feature is merged, and at least one feature extraction is performed based on the merged input feature to determine the initial feature; the preset number of feature channels of the first convolution module is smaller than the number of feature channels of the first convolution module in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the second convolution module is smaller than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0183] In some embodiments, the preset number of backbone networks includes: a preset number of backbone network modules connected in series; the backbone network modules include: a first backbone network convolution module, a first separation convolution module, a third activation function module, a second backbone network convolution module and a second separation convolution module; the feature extraction based on the initial features through the preset number of backbone networks in the preset filtering network to determine the intermediate features includes:

[0184] The first backbone network convolution module and the first separation convolution module are respectively used to extract features of the initial features or the features output by the previous backbone network module, and the third activation function module is used to merge and perform nonlinear operations on the features extracted by the first backbone network convolution module and the first separation convolution module, and the second backbone network convolution module is used to converge the features determined by the nonlinear operation, and the second separation convolution module is used to extract features of the converged features to determine the intermediate features; the number of feature channels of at least one convolution module among the first backbone network convolution module, the first separation convolution module, the second backbone network convolution module and the second separation convolution module is less than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0185] In some embodiments, the at least one input data includes: at least one of a reconstructed image block, a predicted image block, a boundary strength, a quantization parameter, and a mode information; the preset number of feature channels of the first convolution module corresponding to the reconstructed image block is a first value; the first value is less than the number of feature channels of the convolution module used for feature extraction of the reconstructed image block in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the first convolution module corresponding to the predicted sample is a second value; the second value is less than the number of feature channels of the convolution module used for feature extraction of the predicted sample in the HOP.2 and LOP.2 filters; and / or, the boundary strength is The preset number of feature channels of the first convolution module corresponding to the quantization parameter is a third value; the third value is less than the number of feature channels of the convolution module used for feature extraction of boundary strength in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the first convolution module corresponding to the quantization parameter is a fourth value; the fourth value is less than the number of feature channels of the convolution module used for feature extraction of quantization parameter in the HOP.2 and LOP.2 filters; and / or, the number of feature channels of the first convolution module corresponding to the pattern information is a fifth value; the fifth value is less than the number of feature channels of the convolution module used for feature extraction of pattern information in the HOP.2 and LOP.2 filters.

[0186] In some embodiments, the number of feature channels corresponding to the first backbone network convolution module is determined based on a preset number of first feature channels; the first separation convolution module includes a first convolution kernel and a second convolution kernel; the number of feature channels corresponding to the first convolution kernel is determined based on a preset number of second feature channels; the number of feature channels corresponding to the second convolution kernel is determined based on a preset number of second feature channels and a preset number of third feature channels; the number of feature channels corresponding to the second backbone network convolution module is determined based on a preset number of first feature channels and a preset number of third feature channels; the second separation convolution module includes a third convolution kernel and a fourth convolution kernel; the number of feature channels corresponding to the third convolution kernel is determined based on a preset number of fourth feature channels; the number of feature channels corresponding to the fourth convolution kernel is determined based on a preset number of fourth feature channels;

[0187] At least one of the preset first feature channel number, the preset second feature channel number, the preset third feature channel number and the preset fourth feature channel number is smaller than the feature channel number of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

[0188] In some embodiments, the first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; the preset filtering network includes 16 backbone network modules connected in series, and for each of the 16 backbone network modules, the preset number of first feature channels is 4; the preset number of second feature channels is 2; the preset number of third feature channels is 4; and the preset number of fourth feature channels is 4.

[0189] In some embodiments, the preset number of backbone networks includes: a preset number of backbone network modules connected in series; the backbone network module includes: a third separation convolution module, a fourth activation function module and a fourth separation convolution module; the predetermined number of backbone networks in the preset filtering network is used to extract features based on the initial features and determine intermediate features, including:

[0190] The third separation convolution module is used to extract features from the initial features or the features output by the previous backbone network module, and the fourth activation function module is used to perform a nonlinear operation on the features extracted by the third separation convolution module, and the fourth separation convolution module is used to extract features from the features determined by the nonlinear operation to determine the intermediate features.

[0191] In some embodiments, the first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; the preset filtering network includes 7 backbone network modules connected in series, and for each of the 7 backbone network modules, the preset number of feature channels of the convolution kernel in the third separation convolution module and the fourth separation convolution module is determined based on the sixth value.

[0192] In some embodiments, the preset number of backbone networks includes: a first preset number of first backbone network modules on a first network branch, and a second preset number of second backbone network modules on a second network branch; the intermediate features include: luminance intermediate features and chrominance intermediate features; the feature extraction based on the initial features through the preset number of backbone networks in the preset filtering network to determine the intermediate features includes:

[0193] Through the first preset number of first backbone network modules, brightness features are extracted based on the initial features to determine the initial brightness features; the first backbone network module includes a third backbone network convolution module, a fourth backbone network convolution module and a fifth separation convolution module; through the sixth separation convolution module on the first network branch, feature extraction is performed on the initial brightness features to determine the brightness intermediate features; through the second preset number of second backbone network modules, chroma features are extracted on the initial features to determine the initial chroma features; the second backbone network module includes: the third backbone network convolution module, the fourth backbone network convolution module and the fifth separation convolution module; through the seventh separation convolution module on the second network branch, feature extraction is performed on the initial chroma features to determine the chroma intermediate features; the number of feature channels of at least one convolution module among the third backbone network convolution module, the fourth backbone network convolution module, the fifth separation convolution module, the sixth separation convolution module and the seventh separation convolution module is less than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0194] In some embodiments, the number of feature channels corresponding to the third backbone network convolution module is determined based on a preset number of first feature channels;

[0195] The number of feature channels corresponding to the fourth backbone network convolution module is determined based on the preset first feature channel number; the number of feature channels corresponding to the fifth separation convolution module is determined based on the preset second feature channel number; the number of feature channels corresponding to the sixth separation convolution module is determined based on the preset fifth feature channel number; the number of feature channels corresponding to the seventh separation convolution module is determined based on the preset sixth feature channel number; at least one of the preset first feature channel number, the preset second feature channel number, the preset fifth feature channel number and the preset sixth feature channel number is smaller than the number of feature channels of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

[0196] In some embodiments, the first value is 4; the second value is 2; the third value is 1; the fourth value is 1; the fifth value is 1; the second convolution module includes a first merging convolution module and a first downsampling convolution module; the preset number of feature channels corresponding to the first merging convolution module is 4; the preset number of feature channels corresponding to the first downsampling convolution module is 32; the preset number of backbone networks includes: 8 first backbone network modules on the first network branch, and 1 second backbone network module on the second network branch; for the first backbone network module or the second backbone network module, the preset number of first feature channels is 32; the preset number of second feature channels is 16; the preset number of fifth feature channels is 32; and the preset number of sixth feature channels is 16.

[0197] In some embodiments, the preset number of backbone networks includes: a first preset number of first backbone network modules on a first network branch, and a second preset number of second backbone network modules on a second network branch; the intermediate features include: luminance intermediate features and chrominance intermediate features; the feature extraction based on the initial features through the preset number of backbone networks in the preset filtering network to determine the intermediate features includes:

[0198] Through the first preset number of first backbone network modules, brightness features are extracted based on the initial features to determine the initial brightness features; the first backbone network module includes an eighth separation convolution module; through the sixth separation convolution module on the first network branch, feature extraction is performed on the initial brightness features to determine the intermediate brightness features; through the second preset number of second backbone network modules, chroma features are extracted on the initial features to determine the initial chroma features; the second backbone network module includes a ninth separation convolution module; through the seventh separation convolution module on the second network branch, feature extraction is performed on the initial chroma features to determine the intermediate chroma features.

[0199] In some embodiments, the first value is 8; the second value is 4; the third value is 1; the fourth value is 1; the fifth value is 1; the second convolution module includes a second merging convolution module and a second downsampling convolution module; the number of preset feature channels corresponding to the second merging convolution module is 8; the number of preset feature channels corresponding to the second downsampling convolution module is 32; the first network branch includes 7 first backbone network modules, and the second network branch includes 2 second backbone network modules; the number of feature channels corresponding to the eighth separation convolution module is 32; the number of feature channels corresponding to the ninth separation convolution module is 16.

[0200] In some embodiments, based on the filtered image block corresponding to the current block, a filtered image corresponding to the current image is determined; and inter-frame prediction decoding is implemented using the filtered image corresponding to the current image.

[0201] In some embodiments, the decoding method further comprises:

[0202] Parse the bitstream to determine first syntax identification information; determine a target quantization bias from at least two candidate base quantization parameters based on the first syntax identification information; and update the quantization parameter corresponding to the current block using the target quantization bias to filter the reconstructed image block based on the updated quantization parameter through the preset filtering network to determine the filtered image block.

[0203] In some embodiments, the method further comprises:

[0204] Parse the bitstream to determine second syntax identification information; determine a target transposition angle from at least two preset transposition angles based on the second syntax identification information; transpose the reconstructed image block according to the target transposition angle, filter the reconstructed image block based on the transposed processing through the preset filtering network, and determine the filtered image block.

[0205] In some embodiments, the decoding method further comprises:

[0206] Parse the code stream to determine third syntax identification information; when the third syntax identification information indicates that Gaussian filtering is performed on the reconstructed image block, perform Gaussian filtering on the reconstructed image block to determine an initial filtered image block, and perform filtering based on the initial filtered image block through the preset filtering network to determine the filtered image block.

[0207] In some embodiments, the decoding method further comprises:

[0208] Using at least two reconstructed sample image sets, performing network training on an initial filtering network to determine the preset filtering network;

[0209] The at least two reconstructed sample image sets include:

[0210] At least two of a reconstructed sample image set without loop filtering, a reconstructed sample image set using deblocking filtering, a reconstructed sample image set using sample adaptive compensation filtering, and a reconstructed sample image set using adaptive loop filtering.

[0211] In some embodiments, the method of performing network training on an initial filtering network using at least two reconstructed sample image sets to determine the preset filtering network includes:

[0212] Using the at least two reconstructed sample image sets, performing network training on the initial filtering network to determine at least two candidate filtering networks; the decoding method further includes:

[0213] Filtering is performed based on the reconstructed image block using the at least two candidate filtering networks to determine at least two candidate filtered image blocks corresponding to the at least two candidate filtering networks; at least two fourth distortion costs corresponding to the at least two candidate filtered images are determined; based on the at least two fourth distortion costs, a filtered image block corresponding to the current block is determined, and / or a preset filtering network corresponding to the current image is determined.

[0214] For example, corresponding to the embodiment of the encoder described above, the decoder parses or obtains a flag indicating the permission to use neural network loop filtering, which is a sequence-level flag (sps_nnlf_enable_flag), indicating that the current decoder allows the use of neural network loop filtering technology. If sps_nnlf_enable_flag is true, the process starts from step 1; otherwise, the process starts from step 3.

[0215] Step 1: parse the code stream to obtain the frame-level usage flag sh_nnlf_flag of the neural network-based loop filtering technology.

[0216] If sh_nnlf_flag indicates that all coding tree units in the current frame use neural network-based loop filtering technology, then ctb_nnlf_flag of all coding tree units in the current frame is set to true.

[0217] If sn_nnlf_flag indicates that some coding tree units in the current frame use neural network-based loop filtering technology, the decoder parses the ctb_nnlf_flag flag of each coding tree unit.

[0218] If sh_nnlf_flag indicates that all coding tree units in the current frame do not use neural network-based loop filtering technology, then ctb_nnlf_flag of all coding tree units in the current frame is set to false.

[0219] If it is sh_case1 or sh_case2, the residual scaling flag scaleFlag of the current frame is parsed; otherwise, scaleFlag defaults to false.

[0220] If scaleFlag is true, the residual scaling parameter index scaleIdx is further parsed; if the parsed scaleIdx indicates that the code stream needs to be further parsed to obtain the residual scaling parameter, the code stream is parsed to obtain the residual scaling parameter scale of the color component of the current frame; otherwise, the preset residual scaling parameter scale is obtained according to the index.

[0221] Step 2: Load and initialize the preset filter network according to the preset parameters. The convolution and other parameters in the preset filter network have been obtained through training and integer quantization.

[0222] If the neural network-based loop filtering flag ctb_nnlf_flag for the current coding tree unit is true, the reconstructed image block rec, predicted image block pred, boundary strength BS, mode information IPB, quantization information BaseQP and SliceQP, and edge image of the current coding tree unit region are obtained and input into the preset filtering network for inference calculation. After inference by the preset filtering network, the initial filtered image block filteredRec is obtained. If scaleFlag is true, the residual scaling factor scale for each color component is obtained based on the parsed residual scaling parameter index scaleIdx. The residual of scale is multiplied by the initial filtered image block filteredRec and the reconstructed image block rec, and the scaled residual is added back to the initial filtered image block rec to obtain the filtered image block output. If the current image or current coding tree unit does not use residual scaling technology, the initial filtered image block filteredRec is the filtered image block output.

[0223] If the neural network-based loop filtering flag ctb_nnlf_flag of the current coding tree unit is false, the reconstructed image block rec is the filtered image block output.

[0224] Step 3: The decoder continues to operate other loop filtering techniques.

[0225] Step 4: After executing all loop filtering tools, the final output image is obtained.

[0226] It should be noted that the encoder and decoder should use a preset filter network with the same network structure.

[0227] In some embodiments, based on the video codec method described above, the filter unit in the codec was implemented using the preset filter network provided in the embodiments of the present application and tested under full-scale conditions. The test results are shown in Tables 1 and 2. Table 1 shows the test results based on the preset filter network of Figure 12, and Table 2 shows the test results based on the preset filter network of Figure 16.

[0228] Table 1

[0229] Table 2

[0230] In Tables 1 and 2, negative numbers represent performance gains, i.e., the bit rate savings ratio at the same quality. It can be seen that the preset filtering network of the embodiment of the present application can still ensure high codec compression performance while reducing network complexity.

[0231] Based on the implementation basis of the aforementioned embodiment, as shown in FIG21 , the embodiment of the present application provides a decoder 1, comprising:

[0232] The first feature preprocessing part 10 is configured to extract features based on the reconstructed image block corresponding to the current block through a preprocessing part in a preset filtering network to determine the initial features; the number of feature channels of the preprocessing part is a preset value;

[0233] A first intermediate feature extraction section 11 is configured to perform feature extraction based on the initial features using a preset number of backbone networks in the preset filtering network to determine intermediate features; the preset value is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters;

[0234] The first filtering output part 12 is configured to determine the filtered image block corresponding to the current block according to the intermediate features through the output part in the preset filtering network.

[0235] In some embodiments, the preprocessing part includes at least one group of first convolution parts and a second convolution part, the at least one group of first convolution parts corresponding to at least one input data; the at least one input data includes at least the reconstructed image block; the first convolution part includes a first convolution module; the second convolution part includes at least one group of second convolution modules and a second activation function module; the preset value includes: a preset number of feature channels of at least one first convolution module and / or a preset number of feature channels of the second convolution module; the first feature preprocessing part 10 is further configured to perform feature extraction on at least one input data through the at least one group of first convolution parts to determine at least one input feature; merge the at least one input feature through the second convolution part, and perform at least one feature extraction based on the merged input feature to determine the initial feature; the preset number of feature channels of the first convolution module is less than the number of feature channels of the first convolution module in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the second convolution module is less than the number of feature channels of the corresponding convolution module in the HOP.2 and LOP.2 filters.

[0236] In some embodiments, the preset number of backbone networks includes: a preset number of backbone network modules connected in series; the backbone network modules include: a first backbone network convolution module, a first separation convolution module, a third activation function module, a second backbone network convolution module and a second separation convolution module; the first intermediate feature extraction part 11 is also configured to respectively extract features of the initial features or the features output by the previous backbone network module through the first backbone network convolution module and the first separation convolution module, and perform feature merging and nonlinear operation on the features extracted by the first backbone network convolution module and the first separation convolution module through the third activation function module, and converge the features determined by the nonlinear operation through the second backbone network convolution module, and perform feature extraction on the converged features through the second separation convolution module to determine the intermediate features; the number of feature channels of at least one convolution module among the first backbone network convolution module, the first separation convolution module, the second backbone network convolution module and the second separation convolution module is less than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0237] In some embodiments, the at least one input data includes: at least one of a reconstructed image block, a predicted image block, a boundary strength, a quantization parameter, and a mode information; the preset number of feature channels of the first convolution module corresponding to the reconstructed image block is a first value; the first value is less than the number of feature channels of the convolution module used for feature extraction of the reconstructed image block in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the first convolution module corresponding to the predicted sample is a second value; the second value is less than the number of feature channels of the convolution module used for feature extraction of the predicted sample in the HOP.2 and LOP.2 filters; and / or, the boundary strength is The preset number of feature channels of the first convolution module corresponding to the quantization parameter is a third value; the third value is less than the number of feature channels of the convolution module used for feature extraction of boundary strength in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the first convolution module corresponding to the quantization parameter is a fourth value; the fourth value is less than the number of feature channels of the convolution module used for feature extraction of quantization parameter in the HOP.2 and LOP.2 filters; and / or, the number of feature channels of the first convolution module corresponding to the pattern information is a fifth value; the fifth value is less than the number of feature channels of the convolution module used for feature extraction of pattern information in the HOP.2 and LOP.2 filters.

[0238] In some embodiments, the number of feature channels corresponding to the first backbone network convolution module is determined based on a preset first feature channel number; the first separation convolution module includes a first convolution kernel and a second convolution kernel; the number of feature channels corresponding to the first convolution kernel is determined based on a preset second feature channel number; the number of feature channels corresponding to the second convolution kernel is determined based on a preset second feature channel number and a preset third feature channel number; the number of feature channels corresponding to the second backbone network convolution module is determined based on a preset first feature channel number and a preset third feature channel number; the second separation convolution module includes a third convolution kernel and a fourth convolution kernel; the number of feature channels corresponding to the third convolution kernel is determined based on a preset fourth feature channel number; the number of feature channels corresponding to the fourth convolution kernel is determined based on a preset fourth feature channel number; at least one of the preset first feature channel number, the preset second feature channel number, the preset third feature channel number and the preset fourth feature channel number is less than the number of feature channels of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

[0239] In some embodiments, the first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; the preset filtering network includes 16 backbone network modules connected in series, and for each of the 16 backbone network modules, the preset number of first feature channels is 4; the preset number of second feature channels is 2; the preset number of third feature channels is 4; and the preset number of fourth feature channels is 4.

[0240] In some embodiments, the preset number of backbone networks includes: a preset number of backbone network modules connected in series; the backbone network module includes: a third separation convolution module, a fourth activation function module and a fourth separation convolution module; the first intermediate feature extraction part 11 is also configured to perform feature extraction on the initial features or the features output by the previous backbone network module through the third separation convolution module, and perform nonlinear operations on the features extracted by the third separation convolution module through the fourth activation function module, and perform feature extraction on the features determined by the nonlinear operation through the fourth separation convolution module to determine the intermediate features.

[0241] In some embodiments, the first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; the preset filtering network includes 7 backbone network modules connected in series, and for each of the 7 backbone network modules, the preset number of feature channels of the convolution kernel in the third separation convolution module and the fourth separation convolution module is determined based on the sixth value.

[0242] In some embodiments, the preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: brightness intermediate features and chromaticity intermediate features; the first intermediate feature extraction part 11 is also configured to perform brightness feature extraction based on the initial features through the first preset number of first backbone network modules to determine the initial brightness features; the first backbone network module includes a third backbone network convolution module, a fourth backbone network convolution module and a fifth separation convolution module; the initial brightness features are extracted through the sixth separation convolution module on the first network branch to determine the Luminance intermediate features; through the second preset number of second backbone network modules, the initial features are subjected to chromaticity feature extraction to determine the initial chromaticity features; the second backbone network module includes: the third backbone network convolution module, the fourth backbone network convolution module and the fifth separation convolution module; through the seventh separation convolution module on the second network branch, the initial chromaticity features are subjected to feature extraction to determine the chromaticity intermediate features; the number of feature channels of at least one convolution module among the third backbone network convolution module, the fourth backbone network convolution module, the fifth separation convolution module, the sixth separation convolution module and the seventh separation convolution module is less than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0243] In some embodiments, the number of feature channels corresponding to the third backbone network convolution module is determined based on a preset first feature channel number; the number of feature channels corresponding to the fourth backbone network convolution module is determined based on a preset first feature channel number; the number of feature channels corresponding to the fifth separation convolution module is determined based on a preset second feature channel number; the number of feature channels corresponding to the sixth separation convolution module is determined based on a preset fifth feature channel number; the number of feature channels corresponding to the seventh separation convolution module is determined based on a preset sixth feature channel number; at least one of the preset first feature channel number, the preset second feature channel number, the preset fifth feature channel number and the preset sixth feature channel number is smaller than the number of feature channels of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

[0244] In some embodiments, the first value is 4; the second value is 2; the third value is 1; the fourth value is 1; the fifth value is 1; the second convolution module includes a first merging convolution module and a first downsampling convolution module; the preset number of feature channels corresponding to the first merging convolution module is 4; the preset number of feature channels corresponding to the first downsampling convolution module is 32; the preset number of backbone networks includes: 8 first backbone network modules on the first network branch, and 1 second backbone network module on the second network branch; for the first backbone network module or the second backbone network module, the preset number of first feature channels is 32; the preset number of second feature channels is 16; the preset number of fifth feature channels is 32; and the preset number of sixth feature channels is 16.

[0245] In some embodiments, the preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: luminance intermediate features and chrominance intermediate features; the first intermediate feature extraction part 11 is also configured to perform luminance feature extraction based on the initial features through the first preset number of first backbone network modules to determine the initial luminance features; the first backbone network module includes an eighth separation convolution module; through the sixth separation convolution module on the first network branch, feature extraction is performed on the initial luminance features to determine the luminance intermediate features; through the second preset number of second backbone network modules, chrominance feature extraction is performed on the initial features to determine the initial chrominance features; the second backbone network module includes a ninth separation convolution module; through the seventh separation convolution module on the second network branch, feature extraction is performed on the initial chrominance features to determine the chrominance intermediate features.

[0246] In some embodiments, the first value is 8; the second value is 4; the third value is 1; the fourth value is 1; the fifth value is 1; the second convolution module includes a second merging convolution module and a second downsampling convolution module; the number of preset feature channels corresponding to the second merging convolution module is 8; the number of preset feature channels corresponding to the second downsampling convolution module is 32; the first network branch includes 7 first backbone network modules, and the second network branch includes 2 second backbone network modules; the number of feature channels corresponding to the eighth separation convolution module is 32; the number of feature channels corresponding to the ninth separation convolution module is 16.

[0247] In some embodiments, the decoder 1 further includes a decoding part; the decoding part is configured to determine a filtered image corresponding to the current image based on a filtered image block corresponding to the current block; and implement inter-frame prediction decoding using the filtered image corresponding to the current image.

[0248] In some embodiments, the decoder 1 further includes a parsing part and a first parameter updating part; the parsing part is configured to parse the code stream and determine first syntax identification information; the first parameter updating part is configured to determine a target quantization bias from at least two candidate basic quantization parameters based on the first syntax identification information; the quantization parameter corresponding to the current block is updated using the target quantization bias, so as to filter based on the reconstructed image block and the updated quantization parameter through the preset filtering network to determine the filtered image block.

[0249] In some embodiments, the decoder 1 further includes a parsing part and a first transposition part; the parsing part is configured to parse the code stream to determine the second syntax identification information;

[0250] The first transposition part is configured to determine a target transposition angle from at least two preset transposition angles according to the second syntax identification information; transpose the reconstructed image block according to the target transposition angle, so as to filter the reconstructed image block based on the transposed processing through the preset filtering network to determine the filtered image block.

[0251] In some embodiments, the decoder 1 further includes a parsing part and a first Gaussian filtering part;

[0252] The parsing part is configured to parse the code stream and determine the third syntax identification information;

[0253] The first Gaussian filtering part is configured to perform Gaussian filtering on the reconstructed image block when the third syntax identification information represents that Gaussian filtering is performed on the reconstructed image block, determine an initial filtered image block, and perform filtering based on the initial filtered image block through the preset filtering network to determine the filtered image block.

[0254] In some embodiments, the decoder 1 further comprises a first training part;

[0255] The first training part is configured to perform network training on the initial filtering network using at least two reconstructed sample image sets to determine the preset filtering network; the at least two reconstructed sample image sets include:

[0256] At least two of a reconstructed sample image set without loop filtering, a reconstructed sample image set using deblocking filtering, a reconstructed sample image set using sample adaptive compensation filtering, and a reconstructed sample image set using adaptive loop filtering.

[0257] In some embodiments, the first training part is further configured to use the at least two reconstructed sample image sets to perform network training on the initial filtering network and determine at least two candidate filtering networks; the decoder also includes a first network determination part; the first network determination part is configured to perform filtering based on the reconstructed image block through the at least two candidate filtering networks, and determine at least two candidate filtered image blocks corresponding to the at least two candidate filtering networks; determine at least two fourth distortion costs corresponding to the at least two candidate filtered images; determine the filtered image block corresponding to the current block according to the at least two fourth distortion costs, and / or determine the preset filtering network corresponding to the current image.

[0258] In the actual application of the present application, as shown in FIG22 , the embodiment of the present application further provides a decoder, including:

[0259] A first memory 14 and a first processor 15;

[0260] The first memory 14 stores a computer program that can be run on the first processor 15. When the first processor 15 executes the program, the video decoding method provided in the embodiment of the present application is implemented.

[0261] Among them, the first processor 15 can be implemented by software, hardware, firmware or a combination thereof, and can use circuits, single or multiple application specific integrated circuits (ASICs), single or multiple general-purpose integrated circuits, single or multiple microprocessors, single or multiple programmable logic devices, or a combination of the aforementioned circuits or devices, or other suitable circuits or devices, so that the first processor 15 can execute the corresponding steps of the video decoding method provided in the embodiment of the present application.

[0262] It should be noted that the description of the above decoder embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the decoder embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding.

[0263] The embodiment of the present application provides an encoder 2, as shown in FIG23 , including:

[0264] The second feature preprocessing part 20 is configured to extract features based on the reconstructed image block corresponding to the current block through a preprocessing part in a preset filtering network to determine the initial features; the number of feature channels of the preprocessing part is a preset value;

[0265] A second intermediate feature extraction section 21 is configured to perform feature extraction based on the initial features using a preset number of backbone networks in the preset filtering network to determine intermediate features; the preset value is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters;

[0266] The second filtering output part 22 is configured to determine the filtered image block corresponding to the current block according to the intermediate features through the output part in the preset filtering network.

[0267] In some embodiments, the preprocessing part includes at least one group of first convolution parts and a second convolution part, the at least one group of first convolution parts corresponding to at least one input data; the at least one input data includes at least the reconstructed image block; the first convolution part includes a first convolution module; the second convolution part includes at least one group of second convolution modules and a second activation function module; the preset value includes: a preset number of feature channels of at least one first convolution module and / or a preset number of feature channels of the second convolution module; the second feature preprocessing part 20 is further configured to perform feature extraction on at least one input data through the at least one group of first convolution parts to determine at least one input feature; merge the at least one input feature through the second convolution part, and perform at least one feature extraction based on the merged input feature to determine the initial feature; the preset number of feature channels of the first convolution module is less than the number of feature channels of the first convolution module in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the second convolution module is less than the number of feature channels of the corresponding convolution module in the HOP.2 and LOP.2 filters.

[0268] In some embodiments, the preset number of backbone networks includes: a preset number of backbone network modules connected in series; the backbone network modules include: a first backbone network convolution module, a first separation convolution module, a third activation function module, a second backbone network convolution module and a second separation convolution module; the second intermediate feature extraction part 21 is also configured to respectively extract features of the initial features or the features output by the previous backbone network module through the first backbone network convolution module and the first separation convolution module, and perform feature merging and nonlinear operation on the features extracted by the first backbone network convolution module and the first separation convolution module through the third activation function module, and converge the features determined by the nonlinear operation through the second backbone network convolution module, and perform feature extraction on the converged features through the second separation convolution module to determine the intermediate features; the number of feature channels of at least one convolution module among the first backbone network convolution module, the first separation convolution module, the second backbone network convolution module and the second separation convolution module is less than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0269] In some embodiments, the at least one input data includes: at least one of a reconstructed image block, a predicted image block, a boundary strength, a quantization parameter, and a mode information; the preset number of feature channels of the first convolution module corresponding to the reconstructed image block is a first value; the first value is less than the number of feature channels of the convolution module used for feature extraction of the reconstructed image block in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the first convolution module corresponding to the predicted sample is a second value; the second value is less than the number of feature channels of the convolution module used for feature extraction of the predicted sample in the HOP.2 and LOP.2 filters; and / or, the boundary strength corresponds to The preset number of feature channels of the first convolution module is a third value; the third value is smaller than the number of feature channels of the convolution module used for feature extraction of boundary strength in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the first convolution module corresponding to the quantization parameter is a fourth value; the fourth value is smaller than the number of feature channels of the convolution module used for feature extraction of quantization parameter in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the first convolution module corresponding to the pattern information is a fifth value; the fifth value is smaller than the number of feature channels of the convolution module used for feature extraction of pattern information in the HOP.2 and LOP.2 filters.

[0270] In some embodiments, the number of feature channels corresponding to the first backbone network convolution module is determined based on a preset first feature channel number; the first separation convolution module includes a first convolution kernel and a second convolution kernel; the number of feature channels corresponding to the first convolution kernel is determined based on a preset second feature channel number; the number of feature channels corresponding to the second convolution kernel is determined based on a preset second feature channel number and a preset third feature channel number; the number of feature channels corresponding to the second backbone network convolution module is determined based on a preset first feature channel number and a preset third feature channel number; the second separation convolution module includes a third convolution kernel and a fourth convolution kernel; the number of feature channels corresponding to the third convolution kernel is determined based on a preset fourth feature channel number; the number of feature channels corresponding to the fourth convolution kernel is determined based on a preset fourth feature channel number; at least one of the preset first feature channel number, the preset second feature channel number, the preset third feature channel number and the preset fourth feature channel number is smaller than the number of feature channels of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

[0271] In some embodiments, the first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; the preset filtering network includes 16 backbone network modules connected in series, and for each of the 16 backbone network modules, the preset number of first feature channels is 4; the preset number of second feature channels is 2; the preset number of third feature channels is 4; and the preset number of fourth feature channels is 4.

[0272] In some embodiments, the preset number of backbone networks includes: a preset number of backbone network modules connected in series; the backbone network module includes: a third separation convolution module, a fourth activation function module and a fourth separation convolution module; the second intermediate feature extraction part 21 is also configured to perform feature extraction on the initial features or the features output by the previous backbone network module through the third separation convolution module, and perform nonlinear operations on the features extracted by the third separation convolution module through the fourth activation function module, and perform feature extraction on the features determined by the nonlinear operation through the fourth separation convolution module to determine the intermediate features.

[0273] In some embodiments, the first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; the preset filtering network includes 7 backbone network modules connected in series, and for each of the 7 backbone network modules, the number of feature channels of the convolution kernel in the third separation convolution module and the fourth separation convolution module is determined based on the number of feature channels of the initial feature.

[0274] In some embodiments, the preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: brightness intermediate features and chromaticity intermediate features; the second intermediate feature extraction part 21 is also configured to perform brightness feature extraction based on the initial features through the first preset number of first backbone network modules to determine the initial brightness features; the first backbone network module includes a third backbone network convolution module, a fourth backbone network convolution module and a fifth separation convolution module; the initial brightness features are extracted through the sixth separation convolution module on the first network branch to determine the Luminance intermediate features; through the second preset number of second backbone network modules, the initial features are subjected to chromaticity feature extraction to determine the initial chromaticity features; the second backbone network module includes: the third backbone network convolution module, the fourth backbone network convolution module and the fifth separation convolution module; through the seventh separation convolution module on the second network branch, the initial chromaticity features are subjected to feature extraction to determine the chromaticity intermediate features; the number of feature channels of at least one convolution module among the third backbone network convolution module, the fourth backbone network convolution module, the fifth separation convolution module, the sixth separation convolution module and the seventh separation convolution module is less than the number of feature channels of the corresponding convolution modules in the HOP.2 and LOP.2 filters.

[0275] In some embodiments, the number of feature channels corresponding to the third backbone network convolution module is determined based on a preset first feature channel number; the number of feature channels corresponding to the fourth backbone network convolution module is determined based on a preset first feature channel number; the number of feature channels corresponding to the fifth separation convolution module is determined based on a preset second feature channel number; the number of feature channels corresponding to the sixth separation convolution module is determined based on a preset fifth feature channel number; the number of feature channels corresponding to the seventh separation convolution module is determined based on a preset sixth feature channel number; at least one of the preset first feature channel number, the preset second feature channel number, the preset fifth feature channel number and the preset sixth feature channel number is smaller than the number of feature channels of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

[0276] In some embodiments, the first value is 4; the second value is 2; the third value is 1; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 4; the preset number of backbone networks includes: 8 first backbone network modules on the first network branch, and 8 second backbone network modules on the second network branch; for the first backbone network module or the second backbone network module, the preset number of first feature channels is 32; the preset number of second feature channels is 16; the preset number of fifth feature channels is 32; and the preset number of sixth feature channels is 16.

[0277] In some embodiments, the preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: luminance intermediate features and chrominance intermediate features; the second intermediate feature extraction part 21 is also configured to perform luminance feature extraction based on the initial features through the first preset number of first backbone network modules to determine the initial luminance features; the first backbone network module includes an eighth separation convolution module; through the sixth separation convolution module on the first network branch, feature extraction is performed on the initial luminance features to determine the luminance intermediate features; through the second preset number of second backbone network modules, chrominance feature extraction is performed on the initial features to determine the initial chrominance features; the second backbone network module includes a ninth separation convolution module; through the seventh separation convolution module on the second network branch, feature extraction is performed on the initial chrominance features to determine the chrominance intermediate features.

[0278] In some embodiments, the first value is 8; the second value is 4; the third value is 1; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; the first network branch includes 7 first backbone network modules, and the second network branch includes 7 second backbone network modules; the number of feature channels corresponding to the eighth separation convolution module is determined based on the number of feature channels of the initial feature; the number of feature channels corresponding to the ninth separation convolution module is less than the number of feature channels corresponding to the eighth separation convolution module.

[0279] In some embodiments, the encoder 2 further includes: an encoding part; the encoding part is configured to determine a filtered image corresponding to the current image based on the filtered image block corresponding to the current block; and implement inter-frame prediction encoding using the filtered image corresponding to the current image.

[0280] In some embodiments, the intermediate features include: at least two intermediate features, the at least two intermediate features corresponding to at least two initial features; the at least two initial features are determined by the preprocessing portion performing feature extraction based on the reconstructed image block and in combination with at least two candidate basic quantization parameters corresponding to the current block; the at least two candidate basic quantization parameters are determined by the basic quantization parameter corresponding to the current block and at least two candidate quantization offsets;

[0281] The second filtering output part 22 is further configured to determine, through the output part in the preset filtering network, at least two candidate filtered image blocks corresponding to the current block based on the at least two intermediate features; determine at least two first distortion costs corresponding to the at least two candidate filtered image blocks; and determine the filtered image block corresponding to the current block based on the at least two first distortion costs.

[0282] In some embodiments, the encoder 2 also includes an encoding part, which is configured to determine a target quantization bias among the at least two candidate quantization biases based on the at least two first distortion costs, and determine a first syntax identification information corresponding to the target quantization bias; perform entropy encoding on the first syntax identification information, and write the obtained encoding bits into the bitstream.

[0283] In some embodiments, the intermediate features include: at least two intermediate features; the at least two intermediate features correspond to at least two reconstructed image blocks; the at least two reconstructed image blocks are determined by transposing the reconstructed image blocks using at least two preset transposition angles;

[0284] The second filtering output part 22 is further configured to determine, through the output part in the preset filtering network, at least two candidate filtered image blocks corresponding to the current block based on the at least two intermediate features; determine at least two second distortion costs corresponding to the at least two candidate filtered image blocks; and determine the filtered image block corresponding to the current block based on the at least two second distortion costs.

[0285] In some embodiments, the encoder 2 also includes an encoding part, which is configured to determine a target transposition angle among the at least two preset transposition angles based on the at least two second distortion costs, and determine a second syntax identification information corresponding to the target transposition angle; perform entropy encoding on the second syntax identification information, and write the obtained encoding bits into the bitstream.

[0286] In some embodiments, the encoder 2 also includes a second training part; the second training part is configured to perform network training on the initial filtering network through a set of reconstructed sample image blocks including at least two preset transposition angles to determine the preset filtering network.

[0287] In some embodiments, the intermediate feature includes: a first intermediate feature and a second intermediate feature; the first intermediate feature corresponds to a first reconstructed image block; the second intermediate feature corresponds to a second reconstructed image block; the first reconstructed image block includes an original reconstructed image block corresponding to the current block; the second reconstructed image block includes an initial filtered image block determined by performing Gaussian filtering on the original reconstructed image block;

[0288] The second filtering output part 22 is also configured to determine, through the output part in the preset filtering network, the first candidate filtering image block corresponding to the current block according to the first intermediate feature; and determine the second candidate filtering image block corresponding to the current block according to the second intermediate feature; determine the third distortion cost corresponding to the first candidate filtering image block, and the fourth distortion cost corresponding to the second candidate filtering image block; and determine the filtering image block corresponding to the current block according to the third distortion cost and the fourth distortion cost.

[0289] In some embodiments, the encoder 2 also includes an encoding part, which is configured to determine third syntax identification information based on the third distortion cost and the fourth distortion cost; the third syntax identification information indicates whether Gaussian filtering is performed on the current block; the third syntax identification information is entropy encoded, and the obtained encoded bits are written into the bitstream.

[0290] In some embodiments, the encoder 2 further includes a second training part; the second training part is configured to perform network training on the initial filtering network using at least two reconstructed sample image sets to determine the preset filtering network; the at least two reconstructed sample image sets include:

[0291] At least two of a reconstructed sample image set without loop filtering, a reconstructed sample image set using deblocking filtering, a reconstructed sample image set using sample adaptive compensation filtering, and a reconstructed sample image set using adaptive loop filtering.

[0292] In some embodiments, the second training part is further configured to perform network training on the initial filtering network using the at least two reconstructed sample image sets to determine at least two candidate filtering networks;

[0293] The encoder 2 also includes a second network determination part; the second network determination part is configured to perform filtering based on the reconstructed image block through the at least two candidate filtering networks, determine at least two candidate filtered image blocks corresponding to the at least two candidate filtering networks; determine at least two fourth distortion costs corresponding to the at least two candidate filtered images; determine the filtered image block corresponding to the current block corresponding to the current block based on the at least two fourth distortion costs, and / or determine the preset filtering network corresponding to the current image.

[0294] In practical applications, as shown in FIG24 , an embodiment of the present application further provides an encoder, including:

[0295] a second memory 25 and a second processor 26;

[0296] The second memory 25 stores a computer program that can be run on the second processor 26. When the second processor 26 executes the program, the video encoding method provided in the embodiment of the present application is implemented.

[0297] It should be noted that the description of the above encoder embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the encoder embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding.

[0298] An embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following:

[0299] The encoding information of the current image;

[0300] The coding information of the current image is determined by performing inter-frame prediction based on a historical filtered image block; the historical filtered image block is determined by filtering a historical block in the historical image through a preset filtering network, including:

[0301] By using a preprocessing part in a preset filtering network, feature extraction is performed based on the historical reconstructed image block corresponding to the historical block to determine the historical initial feature; the number of feature channels of the preprocessing part is a preset value;

[0302] Determining historical intermediate features by extracting features based on the historical initial features using a preset number of backbone networks in the preset filtering network; wherein the preset value is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters;

[0303] The historical filtered image block is determined according to the historical intermediate features through the output part in the preset filtering network.

[0304] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a first processor, the video decoding method provided by the embodiment of the present application is implemented; or, when the computer program is executed by a second processor, the video encoding method provided by the embodiment of the present application is implemented.

[0305] The various components in the embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of software functional modules.

[0306] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned computer-readable storage medium includes: ferromagnetic random access memory (FRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface storage, optical disk, or compact disc read-only memory (CD-ROM), etc. Various media that can store program codes are not limited in the embodiments of the present disclosure.

[0307] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability

[0308] The embodiments of the present application provide a video encoding and decoding method, a decoder, an encoder, and a computer-readable storage medium. In the preset filtering network, the number of feature channels of the preprocessing part is a preset value, which is less than the number of feature channels used for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the number of backbone networks is a preset number, which is less than the number of backbone network layers in the HOP.2 and LOP.2 filters. In this way, by reducing the number of feature channels in the preprocessing part and / or reducing the depth of the backbone network layer, the model complexity is effectively reduced, which is conducive to the application of neural network-based loop filtering tools on more terminals, thereby improving the wide application of neural network-based filtering solutions. Moreover, it has been experimentally verified that the preset filtering network of the embodiment of the present application can still guarantee higher encoding and decoding performance.

Claims

1. A video decoding method, applied to a decoder, comprising: Performing feature extraction based on a reconstructed image block corresponding to a current block through a preprocessing part in a preset filtering network to determine initial features; The number of feature channels of the preprocessing part is a preset value; Performing feature extraction based on the initial features through a preset number of backbone networks in the preset filtering network to determine intermediate features; The preset value is less than the number of feature channels for input data feature extraction in the HOP.2 and LOP.2 filters; And / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters; Determining a filtered image block corresponding to the current block according to the intermediate features through an output part in the preset filtering network.

2. The method according to claim 1, wherein, The preprocessing part includes at least one group of first convolution parts and a second convolution part. The at least one group of first convolution parts corresponds to at least one type of input data. The at least one type of input data includes at least the reconstructed image block. The first convolution part includes a first convolution module. The second convolution part includes at least one group of second convolution modules and a second activation function module. The preset value includes: the preset feature channel number of at least one first convolution module and / or the preset feature channel number of the second convolution module. The performing feature extraction based on a reconstructed image block corresponding to a current block through a preprocessing part in a preset filtering network to determine initial features includes: Performing feature extraction on at least one type of input data through the at least one group of first convolution parts to determine at least one input feature; Merging the at least one input feature through the second convolution part and performing at least one feature extraction based on the merged input features to determine the initial features; The preset feature channel number of the first convolution module is less than the number of feature channels of this first convolution module in the HOP.2 and LOP.2 filters; and / or, the preset feature channel number of the second convolution module is less than the number of feature channels of the corresponding convolution module in the HOP.2 and LOP.2 filters.

3. The method according to claim 1 or 2, wherein The preset number of backbone networks includes: a preset number of backbone network modules connected in series. The backbone network module includes: a first backbone network convolution module, a first separable convolution module, a third activation function module, a second backbone network convolution module, and a second separable convolution module. The performing feature extraction based on the initial features through a preset number of backbone networks in the preset filtering network to determine intermediate features includes: Respectively performing feature extraction on the initial features or the features output by the previous backbone network module through the first backbone network convolution module and the first separable convolution module, and performing feature merging and non-linear operation on the features extracted by the first backbone network convolution module and the first separable convolution module through the third activation function module, and converging the features determined by the non-linear operation through the second backbone network convolution module, and performing feature extraction on the converged features through the second separable convolution module to determine the intermediate features; The number of feature channels of at least one convolutional module among the first backbone network convolutional module, the first separable convolutional module, the second backbone network convolutional module, and the second separable convolutional module is less than the number of feature channels of the corresponding convolutional module in the HOP.2 and LOP.2 filters.

4. The method according to claim 2, wherein The at least one type of input data includes at least one of a reconstructed image block, a predicted image block, a boundary strength, a quantization parameter, and mode information; The preset number of feature channels of the first convolutional module corresponding to the reconstructed image block is a first value; the first value is less than the number of feature channels of the convolutional module in the HOP.2 and LOP.2 filters for feature extraction of the reconstructed image block; and / or, The preset number of feature channels of the first convolutional module corresponding to the predicted sample is a second value; the second value is less than the number of feature channels of the convolutional module in the HOP.2 and LOP.2 filters for feature extraction of the predicted sample; and / or, The preset number of feature channels of the first convolutional module corresponding to the boundary strength is a third value; the third value is less than the number of feature channels of the convolutional module in the HOP.2 and LOP.2 filters for feature extraction of the boundary strength; and / or, The preset number of feature channels of the first convolutional module corresponding to the quantization parameter is a fourth value; the fourth value is less than the number of feature channels of the convolutional module in the HOP.2 and LOP.2 filters for feature extraction of the quantization parameter; and / or, The number of feature channels of the first convolutional module corresponding to the mode information is a fifth value; the fifth value is less than the number of feature channels of the convolutional module in the HOP.2 and LOP.2 filters for feature extraction of the mode information.

5. The method according to claim 3, wherein, The number of feature channels corresponding to the first backbone network convolutional module is determined based on a preset first number of feature channels; The first separable convolutional module includes a first convolutional kernel and a second convolutional kernel; the number of feature channels corresponding to the first convolutional kernel is determined based on a preset second number of feature channels; the number of feature channels corresponding to the second convolutional kernel is determined based on the preset second number of feature channels and a preset third number of feature channels; The number of feature channels corresponding to the second backbone network convolutional module is determined based on the preset first number of feature channels and the preset third number of feature channels; The second separable convolutional module includes a third convolutional kernel and a fourth convolutional kernel; the number of feature channels corresponding to the third convolutional kernel is determined based on a preset fourth number of feature channels; the number of feature channels corresponding to the fourth convolutional kernel is determined based on the preset fourth number of feature channels; At least one of the preset first number of feature channels, the preset second number of feature channels, the preset third number of feature channels, and the preset fourth number of feature channels is less than the number of feature channels of the corresponding convolutional module in the backbone network module of the HOP.2 and LOP.2 filters.

6. The method according to claim 4 or 5, wherein, The first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolutional module is 8; The preset filtering network includes 16 backbone network modules connected in series. For each of the 16 backbone network modules, the preset number of first feature channels is 4; the preset number of second feature channels is 2; the preset number of third feature channels is 4; the preset number of fourth feature channels is 4.

7. The method according to claim 1 or 2, wherein The preset number of backbone networks includes: a preset number of backbone network modules connected in series; the backbone network module includes: a third separable convolution module, a fourth activation function module, and a fourth separable convolution module; the feature extraction based on the initial feature through the preset number of backbone networks in the preset filtering network to determine the intermediate feature includes: Through the third separable convolution module, feature extraction is performed on the initial feature or the feature output by the previous backbone network module, and through the fourth activation function module, a non-linear operation is performed on the feature extracted by the third separable convolution module, and through the fourth separable convolution module, feature extraction is performed on the feature determined by the non-linear operation to determine the intermediate feature.

8. The method according to claim 7, wherein, The first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; The preset filtering network includes 7 backbone network modules connected in series. For each of the 7 backbone network modules, the preset number of feature channels of the convolution kernels in the third separable convolution module and the fourth separable convolution module is determined based on the sixth value.

9. The method according to claim 1 or 2, wherein, The preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch and a second preset number of second backbone network modules on the second network branch; the intermediate feature includes: a luminance intermediate feature and a chrominance intermediate feature; the feature extraction based on the initial feature through the preset number of backbone networks in the preset filtering network to determine the intermediate feature includes: Through the first preset number of first backbone network modules, luminance feature extraction is performed based on the initial feature to determine the initial luminance feature; the first backbone network module includes a third backbone network convolution module, a fourth backbone network convolution module, and a fifth separable convolution module; Through the sixth separable convolution module on the first network branch, feature extraction is performed on the initial luminance feature to determine the luminance intermediate feature; Through the second preset number of second backbone network modules, chrominance feature extraction is performed on the initial feature to determine the initial chrominance feature; the second backbone network module includes: the third backbone network convolution module, the fourth backbone network convolution module, and the fifth separable convolution module; Through the seventh separable convolution module on the second network branch, feature extraction is performed on the initial chrominance feature to determine the chrominance intermediate feature; The number of feature channels of at least one of the third backbone network convolution module, the fourth backbone network convolution module, the fifth separable convolution module, the sixth separable convolution module, and the seventh separable convolution module is less than the number of feature channels of the corresponding convolution module in the HOP.2 and LOP.2 filters.

10. The method according to claim 9, wherein, The number of feature channels corresponding to the third backbone network convolution module is determined based on a preset first number of feature channels; The number of feature channels corresponding to the fourth backbone network convolution module is determined based on a preset first number of feature channels; The number of feature channels corresponding to the fifth separable convolution module is determined based on a preset second number of feature channels; The number of feature channels corresponding to the sixth separable convolution module is determined based on a preset fifth number of feature channels; The number of feature channels corresponding to the seventh separable convolution module is determined based on a preset sixth number of feature channels; At least one of the preset first number of feature channels, the preset second number of feature channels, the preset fifth number of feature channels, and the preset sixth number of feature channels is less than the number of feature channels of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

11. The method according to claim 4 or 10, wherein, The first value is 4; the second value is 2; the third value is 1; the fourth value is 1; the fifth value is 1; the second convolution module includes a first merging convolution module and a first downsampling convolution module; the preset number of feature channels corresponding to the first merging convolution module is 4; The preset number of feature channels corresponding to the first downsampling convolution module is 32; The preset number of backbone networks includes: 8 first backbone network modules on the first network branch, and 1 second backbone network module on the second network branch; For the first backbone network module or the second backbone network module, the preset first number of feature channels is 32; the preset second number of feature channels is 16; the preset fifth number of feature channels is 32; the preset sixth number of feature channels is 16.

12. The method according to claim 1 or 2, wherein The preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: a luminance intermediate feature and a chrominance intermediate feature; the method of determining the intermediate features based on the initial features by passing through the preset number of backbone networks in the preset filtering network includes: Performing luminance feature extraction on the initial features through the first preset number of first backbone network modules to determine an initial luminance feature; the first backbone network module includes an eighth separable convolution module; Performing feature extraction on the initial luminance feature through the sixth separable convolution module on the first network branch to determine the luminance intermediate feature; Performing chrominance feature extraction on the initial features through the second preset number of second backbone network modules to determine an initial chrominance feature; the second backbone network module includes a ninth separable convolution module; Performing feature extraction on the initial chrominance feature through the seventh separable convolution module on the second network branch to determine the chrominance intermediate feature.

13. The method according to claim 12, wherein, The first value is 8; the second value is 4; the third value is 1; the fourth value is 1; the fifth value is 1; the second convolution module includes a second merging convolution module and a second downsampling convolution module; the preset number of feature channels corresponding to the second merging convolution module is 8; the preset number of feature channels corresponding to the second downsampling convolution module is 32; The first network branch includes 7 first backbone network modules, and the second network branch includes 2 second backbone network modules; the number of feature channels corresponding to the eighth separable convolution module is 32; the number of feature channels corresponding to the ninth separable convolution module is 16.

14. The method according to any one of claims 1, 2, 4, 5, 8, 10, 13, wherein The method further includes: Based on the filtered image block corresponding to the current block, determining the filtered image corresponding to the current image; Performing inter-frame prediction decoding by using the filtered image corresponding to the current image.

15. The method according to any one of claims 1, 2, 4, 5, 8, 10, 13, wherein The method further includes: Parsing the bitstream to determine first syntax identification information; Determining a target quantization bias from at least two candidate base quantization parameters according to the first syntax identification information; Updating the quantization parameter corresponding to the current block by using the target quantization bias, so as to perform filtering based on the reconstructed image block and the updated quantization parameter through the preset filtering network to determine the filtered image block.

16. The method according to any one of claims 1, 2, 4, 5, 8, 10, 13, wherein The method further includes: Parsing the bitstream to determine second syntax identification information; Determining a target transpose angle from at least two preset transpose angles according to the second syntax identification information; Performing a transpose process on the reconstructed image block according to the target transpose angle, so as to perform filtering based on the transposed reconstructed image block through the preset filtering network to determine the filtered image block.

17. The method according to any one of claims 1, 2, 4, 5, 8, 10, 13, wherein The method further includes: Parsing the bitstream to determine third syntax identification information; When the third syntax identification information indicates that Gaussian filtering is to be performed on the reconstructed image block, performing Gaussian filtering on the reconstructed image block to determine an initial filtered image block, so as to perform filtering based on the initial filtered image block through the preset filtering network to determine the filtered image block.

18. The method according to any one of claims 1, 2, 4, 5, 8, 10, 13, wherein, The method further includes: Using at least two sets of reconstructed sample images to perform network training on an initial filtering network to determine the preset filtering network; The at least two sets of reconstructed sample images include: At least two of a set of reconstructed sample images without using loop filtering, a set of reconstructed sample images using deblocking filtering, a set of reconstructed sample images using sample adaptive compensation filtering, and a set of reconstructed sample images using adaptive loop filtering.

19. The method according to claim 18, wherein, The using at least two sets of reconstructed sample images to perform network training on an initial filtering network to determine the preset filtering network includes: Using the at least two sets of reconstructed sample images to perform network training on an initial filtering network to determine at least two candidate filtering networks; The method further includes: Performing filtering based on the reconstructed image block through the at least two candidate filtering networks to determine at least two candidate filtered image blocks corresponding to the at least two candidate filtering networks; Determining at least two fourth distortion costs corresponding to the at least two candidate filtered images; Determining the filtered image block corresponding to the current block corresponding to the current block according to the at least two fourth distortion costs, and / or determining the preset filtering network corresponding to the current image.

20. A video coding method applied to an encoder, including: Performing feature extraction based on the reconstructed image block corresponding to the current block through a preprocessing part in a preset filtering network to determine initial features; The number of feature channels of the preprocessing part is a preset value; Feature extraction is performed on the initial features based on a preset number of backbone networks in the preset filtering network to determine intermediate features; The preset value is less than the number of feature channels for input data feature extraction in the HOP.2 and LOP.2 filters; And / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters; The filtering image block corresponding to the current block is determined based on the intermediate features through the output part in the preset filtering network.

21. The method according to claim 20, wherein, The preprocessing part includes at least one set of first convolution parts and a second convolution part. The at least one set of first convolution parts corresponds to at least one type of input data. The at least one type of input data includes at least the reconstructed image block. The first convolution part includes a first convolution module. The second convolution part includes at least one set of second convolution modules and a second activation function module. The preset value includes: the preset number of feature channels of at least one first convolution module and / or the preset number of feature channels of the second convolution module. The feature extraction is performed on the reconstructed image block corresponding to the current block through the preprocessing part in the preset filtering network to determine the initial features, including: Feature extraction is performed on at least one type of input data through the at least one set of first convolution parts to determine at least one input feature; The at least one input feature is merged through the second convolution part, and at least one feature extraction is performed based on the merged input features to determine the initial features; The preset number of feature channels of the first convolution module is less than the number of feature channels of this first convolution module in the HOP.2 and LOP.2 filters; and / or, the preset number of feature channels of the second convolution module is less than the number of feature channels of the corresponding convolution module in the HOP.2 and LOP.2 filters.

22. The method according to claim 20 or 21, wherein, The preset number of backbone networks includes: a preset number of backbone network modules connected in series. The backbone network module includes: a first backbone network convolution module, a first separable convolution module, a third activation function module, a second backbone network convolution module, and a second separable convolution module. The feature extraction is performed on the initial features based on the preset number of backbone networks in the preset filtering network to determine the intermediate features, including: Feature extraction is respectively performed on the initial features or the features output by the previous backbone network module through the first backbone network convolution module and the first separable convolution module, and the features extracted by the first backbone network convolution module and the first separable convolution module are feature-merged and non-linearly operated through the third activation function module, and the features determined by the non-linear operation are converged through the second backbone network convolution module, and feature extraction is performed on the converged features through the second separable convolution module to determine the intermediate features; The number of feature channels of at least one convolution module among the first backbone network convolution module, the first separable convolution module, the second backbone network convolution module, and the second separable convolution module is less than the number of feature channels of the corresponding convolution module in the HOP.2 and LOP.2 filters.

23. The method according to claim 21, wherein The at least one input data includes at least one of: a reconstructed image block, a predicted image block, a boundary strength, a quantization parameter, and mode information; The preset number of feature channels of the first convolutional module corresponding to the reconstructed image block is a first value; the first value is less than the number of feature channels of the convolutional module for feature extraction of the reconstructed image block in the HOP.2 and LOP.2 filters; and / or, The preset number of feature channels of the first convolutional module corresponding to the prediction sample is a second value; the second value is less than the number of feature channels of the convolutional module for feature extraction of the prediction sample in the HOP.2 and LOP.2 filters; and / or, The preset number of feature channels of the first convolutional module corresponding to the boundary strength is a third value; the third value is less than the number of feature channels of the convolutional module for feature extraction of the boundary strength in the HOP.2 and LOP.2 filters; and / or, The preset number of feature channels of the first convolutional module corresponding to the quantization parameter is a fourth value; the fourth value is less than the number of feature channels of the convolutional module for feature extraction of the quantization parameter in the HOP.2 and LOP.2 filters; and / or, The preset number of feature channels of the first convolutional module corresponding to the mode information is a fifth value; the fifth value is less than the number of feature channels of the convolutional module for feature extraction of the mode information in the HOP.2 and LOP.2 filters.

24. The method according to claim 22, wherein, The number of feature channels corresponding to the first backbone network convolutional module is determined based on a preset first number of feature channels; The first separable convolutional module includes a first convolution kernel and a second convolution kernel; the number of feature channels corresponding to the first convolution kernel is determined based on a preset second number of feature channels; the number of feature channels corresponding to the second convolution kernel is determined based on the preset second number of feature channels and a preset third number of feature channels; The number of feature channels corresponding to the second backbone network convolutional module is determined based on the preset first number of feature channels and the preset third number of feature channels; The second separable convolutional module includes a third convolution kernel and a fourth convolution kernel; the number of feature channels corresponding to the third convolution kernel is determined based on a preset fourth number of feature channels; the number of feature channels corresponding to the fourth convolution kernel is determined based on the preset fourth number of feature channels; At least one of the preset first number of feature channels, the preset second number of feature channels, the preset third number of feature channels, and the preset fourth number of feature channels is less than the number of feature channels of the corresponding convolutional module in the backbone network module of the HOP.2 and LOP.2 filters.

25. The method according to claim 23 or 24, wherein The first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolutional module is 8; The preset filtering network includes 16 backbone network modules connected in series. For each of the 16 backbone network modules, the preset first number of feature channels is 4; the preset second number of feature channels is 2; the preset third number of feature channels is 4; the preset fourth number of feature channels is 4.

26. The method according to claim 20 or 21, wherein The preset number of backbone networks includes: a preset number of backbone network modules connected in series; the backbone network module includes: a third separable convolution module, a fourth activation function module, and a fourth separable convolution module; the process of performing feature extraction on the initial feature based on the preset number of backbone networks in the preset filtering network to determine intermediate features includes: Performing feature extraction on the initial feature or the feature output by the previous backbone network module through the third separable convolution module, performing a non-linear operation on the feature extracted by the third separable convolution module through the fourth activation function module, and performing feature extraction on the feature determined by the non-linear operation through the fourth separable convolution module to determine the intermediate feature.

27. The method according to claim 26, wherein The first value is 4; the second value is 2; the third value is 2; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; The preset filtering network includes 7 backbone network modules connected in series. For each backbone network module among the 7 backbone network modules, the number of feature channels of the convolution kernels in the third separable convolution module and the fourth separable convolution module is determined based on the number of feature channels of the initial feature.

28. The method according to claim 20 or 21, wherein, The preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: a luminance intermediate feature and a chrominance intermediate feature; the process of performing feature extraction on the initial feature based on the preset number of backbone networks in the preset filtering network to determine intermediate features includes: Performing luminance feature extraction on the initial feature through the first preset number of first backbone network modules to determine an initial luminance feature; the first backbone network module includes a third backbone network convolution module, a fourth backbone network convolution module, and a fifth separable convolution module; Performing feature extraction on the initial luminance feature through the sixth separable convolution module on the first network branch to determine the luminance intermediate feature; Performing chrominance feature extraction on the initial feature through the second preset number of second backbone network modules to determine an initial chrominance feature; the second backbone network module includes: the third backbone network convolution module, the fourth backbone network convolution module, and the fifth separable convolution module; Performing feature extraction on the initial chrominance feature through the seventh separable convolution module on the second network branch to determine the chrominance intermediate feature; The number of feature channels of at least one of the third backbone network convolution module, the fourth backbone network convolution module, the fifth separable convolution module, the sixth separable convolution module, and the seventh separable convolution module is less than the number of feature channels of the corresponding convolution module in the HOP.2 and LOP.2 filters.

29. According to the method of claim 28, wherein, The number of feature channels corresponding to the third backbone network convolution module is determined based on a preset first number of feature channels; The number of feature channels corresponding to the fourth backbone network convolution module is determined based on a preset first number of feature channels; The number of feature channels corresponding to the fifth separable convolution module is determined based on a preset second number of feature channels; The number of feature channels corresponding to the sixth separable convolution module is determined based on a preset fifth number of feature channels; The number of feature channels corresponding to the seventh separable convolution module is determined based on a preset sixth number of feature channels; At least one of the preset first number of feature channels, the preset second number of feature channels, the preset fifth number of feature channels, and the preset sixth number of feature channels is less than the number of feature channels of the corresponding convolution module in the backbone network module of the HOP.2 and LOP.2 filters.

30. The method according to claim 23 or 29, wherein The first value is 4; the second value is 2; the third value is 1; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 4; The preset number of backbone networks includes: 8 first backbone network modules on the first network branch, and 8 second backbone network modules on the second network branch; for the first backbone network module or the second backbone network module, the preset first number of feature channels is 32; the preset second number of feature channels is 16; the preset fifth number of feature channels is 32; the preset sixth number of feature channels is 16.

31. The method according to claim 20 or 21, wherein, The preset number of backbone networks includes: a first preset number of first backbone network modules on the first network branch, and a second preset number of second backbone network modules on the second network branch; the intermediate features include: a luminance intermediate feature and a chrominance intermediate feature; the method of determining the intermediate features based on the initial features through the preset number of backbone networks in the preset filter network includes: Performing luminance feature extraction on the initial features through the first preset number of first backbone network modules to determine an initial luminance feature; the first backbone network module includes an eighth separable convolution module; Performing feature extraction on the initial luminance feature through the sixth separable convolution module on the first network branch to determine the luminance intermediate feature; Performing chrominance feature extraction on the initial features through the second preset number of second backbone network modules to determine an initial chrominance feature; the second backbone network module includes a ninth separable convolution module; Performing feature extraction on the initial chrominance feature through the seventh separable convolution module on the second network branch to determine the chrominance intermediate feature.

32. The method according to claim 31, wherein The first value is 8; the second value is 4; the third value is 1; the fourth value is 1; the fifth value is 1; the preset number of feature channels of the second convolution module is 8; The first network branch includes 7 first backbone network modules, and the second network branch includes 7 second backbone network modules; the number of feature channels corresponding to the eighth separable convolution module is determined based on the number of feature channels of the initial features; the number of feature channels corresponding to the ninth separable convolution module is less than the number of feature channels corresponding to the eighth separable convolution module.

33. The method according to any one of claims 20, 21, 23, 24, 27, 29, 32, wherein, The method further includes: Determining a filtered image corresponding to the current image based on the filtered image block corresponding to the current block; Performing inter-frame prediction coding using the filtered image corresponding to the current image.

34. The method according to any one of claims 20, 21, 23, 24, 27, 29, 32, wherein, The intermediate features include: at least two intermediate features, where the at least two intermediate features correspond to at least two initial features; the at least two initial features are determined by feature extraction through the preprocessing part based on the reconstructed image block and in combination with at least two candidate basic quantization parameters corresponding to the current block respectively; the at least two candidate basic quantization parameters are determined by the basic quantization parameter corresponding to the current block and at least two candidate quantization biases. Determining, by an output part in the preset filtering network, a filtered image block corresponding to the current block according to the intermediate features includes: Determining, by the output part in the preset filtering network, at least two candidate filtered image blocks corresponding to the current block according to the at least two intermediate features; Determining at least two first distortion costs corresponding to the at least two candidate filtered image blocks; Determining the filtered image block corresponding to the current block according to the at least two first distortion costs.

35. The method according to claim 34, wherein, The method further includes: Determining a target quantization bias in the at least two candidate quantization biases according to the at least two first distortion costs, and determining first syntax identification information corresponding to the target quantization bias; Performing entropy coding on the first syntax identification information and writing the obtained coded bits into a bitstream.

36. The method according to any one of claims 20, 21, 23, 24, 27, 29, 32, wherein The intermediate features include: at least two intermediate features; the at least two intermediate features correspond to at least two types of reconstructed image blocks; the at least two types of reconstructed image blocks are determined by transposing the reconstructed image block through at least two preset transpose angles. Determining, by an output part in the preset filtering network, a filtered image block corresponding to the current block according to the intermediate features includes: Determining, by the output part in the preset filtering network, at least two candidate filtered image blocks corresponding to the current block according to the at least two intermediate features; Determining at least two second distortion costs corresponding to the at least two candidate filtered image blocks; Determining the filtered image block corresponding to the current block according to the at least two second distortion costs.

37. The method according to claim 36, wherein The method further includes: Determining a target transpose angle in the at least two preset transpose angles according to the at least two second distortion costs, and determining second syntax identification information corresponding to the target transpose angle; Performing entropy coding on the second syntax identification information and writing the obtained coded bits into a bitstream.

38. The method according to claim 36, wherein, The method further includes: Training a network of an initial filtering network through a set of reconstructed sample image blocks including at least two preset transpose angles of reconstructed sample image blocks to determine the preset filtering network.

39. The method according to claim 33, wherein, The intermediate features include: a first intermediate feature and a second intermediate feature; the first intermediate feature corresponds to a first reconstructed image block; the second intermediate feature corresponds to a second reconstructed image block; the first reconstructed image block includes the original reconstructed image block corresponding to the current block; the second reconstructed image block includes an initial filtered image block determined by performing Gaussian filtering on the original reconstructed image block. Determining, by an output part in the preset filtering network, a filtered image block corresponding to the current block according to the intermediate features includes: Through the output part in the preset filtering network, determine a first candidate filtered image block corresponding to the current block according to the first intermediate feature; and determine a second candidate filtered image block corresponding to the current block according to the second intermediate feature; Determine a third distortion cost corresponding to the first candidate filtered image block and a fourth distortion cost corresponding to the second candidate filtered image block; Determine a filtered image block corresponding to the current block according to the third distortion cost and the fourth distortion cost.

40. The method according to claim 39, wherein The method further includes: Determine third syntax identification information according to the third distortion cost and the fourth distortion cost; the third syntax identification information indicates whether Gaussian filtering is performed on the current block; Perform entropy coding on the third syntax identification information and write the obtained coded bits into the bitstream.

41. The method according to any one of claims 20, 21, 23, 24, 27, 29, 32, wherein, The method further includes: Use at least two reconstructed sample image sets to perform network training on an initial filtering network to determine the preset filtering network; The at least two reconstructed sample image sets include: At least two of a reconstructed sample image set without using loop filtering, a reconstructed sample image set using deblocking filtering, a reconstructed sample image set using sample adaptive compensation filtering, and a reconstructed sample image set using adaptive loop filtering.

42. The method according to claim 41, wherein, The using at least two reconstructed sample image sets to perform network training on an initial filtering network to determine the preset filtering network includes: Use the at least two reconstructed sample image sets to perform network training on an initial filtering network to determine at least two candidate filtering networks; The method further includes: Through the at least two candidate filtering networks, perform filtering based on the reconstructed image block to determine at least two candidate filtered image blocks corresponding to the at least two candidate filtering networks; Determine at least two fourth distortion costs corresponding to the at least two candidate filtered images; According to the at least two fourth distortion costs, determine a filtered image block corresponding to the current block corresponding to the current block, and / or determine a preset filtering network corresponding to the current image.

43. A decoder, comprising: A first feature preprocessing part configured to perform feature extraction based on a reconstructed image block corresponding to a current block through a preprocessing part in a preset filtering network to determine an initial feature; The number of feature channels of the preprocessing part is a preset value; A first intermediate feature extraction part configured to perform feature extraction based on the initial feature through a preset number of backbone networks in the preset filtering network to determine an intermediate feature; The preset value is less than the number of feature channels for input data feature extraction in the HOP.2 and LOP.2 filters; and / or the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters; A first filtering output part configured to determine a filtered image block corresponding to the current block according to the intermediate feature through an output part in the preset filtering network.

44. An encoder, comprising: A second feature preprocessing part configured to perform feature extraction based on a reconstructed image block corresponding to a current block through a preprocessing part in a preset filtering network to determine an initial feature; The number of feature channels of the preprocessing part is a preset value; The second intermediate feature extraction part is configured to extract features based on the initial features through a preset number of backbone networks in the preset filtering network to determine intermediate features; The preset value is less than the number of feature channels for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters; The second filtering output part is configured to determine the filtered image block corresponding to the current block according to the intermediate features through the output part in the preset filtering network.

45. A decoder, the decoder includes a first memory and a first processor; wherein, The first memory is used to store a computer program that can run on the first processor; The first processor is configured to execute the method according to any one of claims 1 to 19 when running the computer program.

46. An encoder, the encoder includes a second memory and a second processor; wherein, The second memory is used to store a computer program that can run on the second processor; The second processor is configured to execute the method according to any one of claims 20 to 42 when running the computer program.

47. A bitstream, which is generated by performing bit encoding according to information to be encoded; wherein, The information to be encoded at least includes at least one of the following: The encoding information of the current image; Wherein, the encoding information of the current image is determined by inter-frame prediction based on historical filtered image blocks; the historical filtered image blocks are determined by filtering historical blocks in a historical image through a preset filtering network, including: Through the preprocessing part in the preset filtering network, feature extraction is performed based on the historical reconstructed image block corresponding to the historical block to determine historical initial features; the number of feature channels of the preprocessing part is a preset value; Through a preset number of backbone networks in the preset filtering network, feature extraction is performed based on the historical initial features to determine historical intermediate features; the preset value is less than the number of feature channels for input data feature extraction in the HOP.2 and LOP.2 filters; and / or, the preset number is less than the number of backbone network layers in the HOP.2 and LOP.2 filters; Through the output part in the preset filtering network, the historical filtered image block is determined according to the historical intermediate features.

48. A storage medium, wherein, The storage medium stores a computer program, and when the computer program is executed, it implements the method according to any one of claims 1 to 19, or implements the method according to any one of claims 20 to 42.

Citation Information

Patent Citations

  • Image coding method, image decoding method, coder and decoder

    CN114025164A

  • Filtering, encoding and decoding method and device, computer readable medium and electronic equipment

    CN115883842A

  • Neural network loop filtering method and device for video coding

    CN115914654A

  • Loop filtering decision-making method of deep neural network based on coding information

    CN116347108A

  • Hybrid neural network based end-to-end image and video coding method

    US20230096567A1