Image encoding method, image decoding method, encoder, and decoder
By combining the reconstruction image of the image to be encoded with different encoding information, input it into the neural network for filtering, and selecting the target encoding information based on the different values, the problem of limited input of the neural network in the prior art is solved, and the reconstruction quality of the image is improved.
Patent Information
- Application Number
- CN202111162709.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In the existing video encoding and decoding methods, the input of the neural network is usually the currently reconstructed image, and the encoding and decoding information is not fully utilized, resulting in the reconstruction quality after image filtering cannot be further improved.
By obtaining the reconstruction image of the image to be encoded, combining different encoding information as different network inputs, inputting it into the preset neural network for filtering, obtaining the filtered reconstruction image, and selecting the target encoding information based on the different value size, and encoding it into the code stream of the image to be encoded.
Add coding information to the neural network filtering process to improve the generalization of the neural network and strengthen the reconstruction quality.
Smart Images

Figure CN114025164B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video encoding and decoding technologies, and particularly to an image encoding method, an image decoding method, an encoder, and a decoder. Background Art
[0002] The amount of video image data is relatively large. Usually, it is necessary to compress the video pixel data (RGB, YUV, etc.). The compressed data is called a video bitstream. The video bitstream is transmitted to the user side through a wired or wireless network and then decoded for viewing. The entire video encoding process includes processes such as block partitioning, prediction, transformation, quantization, and encoding. Subsequently, various filtering processes can be added to make the image look more natural.
[0003] However, in the existing methods of using a neural network for filtering in the encoding and decoding process, the input of the neural network is usually the current reconstructed image, without combining too much encoding and decoding information, and the utilization of the information generated by the encoding and decoding itself is limited, resulting in the inability to further improve the reconstruction quality of the image after filtering. Summary of the Invention
[0004] This application provides an image encoding method, an image decoding method, an encoder, and a decoder.
[0005] To solve the above technical problems, a technical solution adopted by this application is to provide an image encoding method, and the image encoding method includes:
[0006] Obtain the reconstructed image of the image to be encoded;
[0007] Respectively combine the reconstructed image with different encoding information as different network inputs, and respectively input them into a preset neural network, so that the preset neural network filters each network input to obtain a filtered reconstructed image;
[0008] Obtain the difference value between each filtered reconstructed image and the image to be encoded;
[0009] Based on the size of the difference value, select the target encoding information corresponding to the filtered reconstructed image, and encode the target encoding information into the bitstream of the image to be encoded.
[0010] Among them, the encoding information includes one or a combination of multiple of partitioning information, prediction information, residual information, and quantization information.
[0011] Among them, the step of respectively combining the reconstructed image with different encoding information as different network inputs, and respectively inputting them into a preset neural network, so that the preset neural network filters each network input to obtain a filtered reconstructed image includes:
[0012] Divide the reconstructed image into a number of reconstructed sub - blocks;
[0013] Combine each reconstructed sub - block with different coding information as different network inputs respectively, and input them into a preset neural network respectively to obtain the filtered reconstructed sub - blocks output by the preset neural network;
[0014] Combine all the filtered reconstructed sub - blocks to obtain a filtered reconstructed image;
[0015] Among them, the preset neural network is used to replace the traditional filtering module in the filtering network and / or be added to the filtering network. The functions of the traditional filtering module include one or more of de - blocking filtering, sample - adaptive compensation filtering, adaptive loop filtering, and cross - component adaptive loop filtering.
[0016] Among them, the reconstructed image includes a first - component reconstructed image, a second - component reconstructed image, and a third - component reconstructed image; the image coding method further includes:
[0017] Input the first - component reconstructed image, the second - component reconstructed image, and the third - component reconstructed image into the preset neural network simultaneously;
[0018] Obtain the first - component reconstructed image, the second - component reconstructed image, and the third - component reconstructed image after being filtered by the preset neural network, and combine them into a filtered reconstructed image.
[0019] Among them, the reconstructed image includes a first - component reconstructed image, a second - component reconstructed image, and a third - component reconstructed image; the image coding method further includes:
[0020] Input the first - component reconstructed image, and one or two of the second - component reconstructed image and the third - component reconstructed image into the preset neural network;
[0021] Obtain the first - component reconstructed image after being filtered by the preset neural network.
[0022] Among them, the preset neural network arranges a convolutional layer, an activation layer, an attention module, and a convolutional layer in sequence from the input end to the output end. The preset neural network also includes a residual connection from the input end to the output end;
[0023] Among them, the attention module includes a channel - domain attention module and / or a spatial - domain attention module.
[0024] Among them, the preset neural network arranges a convolutional layer, an activation layer, a residual block, and a convolutional layer in sequence from the input end to the output end. The preset neural network also includes a residual connection from the input end to the output end, and an attention module inserted as a network branch between the output of the activation layer and the output of the subsequent convolutional layer.
[0025] Among them, the preset neural network arranges multiple dense connection modules in sequence from the input end to the output end. Among them, the dense connection module includes multiple convolutional layers, and the input of the convolutional block is connected to its own output and / or the output of other convolutional blocks.
[0026] Among them, the image encoding method further includes:
[0027] Encoding a number of encoding and decoding markers in the bitstream based on the encoding and decoding results, where the encoding and decoding markers include a neural network tool switch syntax, a frame-level switch syntax, a block-level switch syntax, and / or a reconstructed pixel value adjustment syntax.
[0028] To solve the above technical problems, a technical solution adopted by this application is to provide an image decoding method, which includes:
[0029] Obtaining the above bitstream and its encoding and decoding information;
[0030] Decoding the bitstream based on the encoding and decoding method in the encoding and decoding information to obtain a decoded reconstructed image;
[0031] Inputting the decoded reconstructed image and the encoding information into the preset neural network, so that the preset neural network filters the decoded reconstructed image based on the encoding information to obtain a final reconstructed image.
[0032] Among them, the image decoding method further includes:
[0033] Obtaining the encoding and decoding markers in the bitstream;
[0034] Judging whether to use the preset neural network for filtering based on the encoding and decoding markers;
[0035] If so, inputting the decoded reconstructed image and the encoding and decoding information into the preset neural network.
[0036] To solve the above technical problems, a technical solution adopted by this application is to provide an encoder, which includes a processor and a memory; a computer program is stored in the memory, and the processor is used to execute the computer program to implement the steps of the above image encoding method.
[0037] To solve the above technical problems, a technical solution adopted by this application is to provide a decoder, which includes a processor and a memory; a computer program is stored in the memory, and the processor is used to execute the computer program to implement the steps of the above image encoding method.
[0038] To solve the above technical problems, a technical solution adopted by this application is to provide a computer storage medium, which stores a computer program, and when the computer program is executed, the steps of the above image encoding method and / or image decoding method are implemented.
[0039] Different from the prior art, the beneficial effect of this application is that the encoder obtains the reconstructed image of the image to be encoded; combines the reconstructed image with different encoding information respectively as different network inputs, and inputs them into a preset neural network respectively, so that the preset neural network filters each network input to obtain a filtered reconstructed image; obtains the difference value between each filtered reconstructed image and the image to be encoded; based on the size of the difference value, selects the target encoding information corresponding to the filtered reconstructed image, and encodes the target encoding information into the bitstream of the image to be encoded. Through the above image encoding method, this application can add encoding information during neural network filtering, improve the generalization of the neural network, and enhance the reconstruction quality of the neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic diagram of the framework of the first embodiment of the neural network based on the attention mechanism provided by this application;
[0042] Figure 2 It is a schematic diagram of the framework of the second embodiment of the neural network based on the attention mechanism provided by this application;
[0043] Figure 3 It is a schematic diagram of the framework of the third embodiment of the neural network based on the attention mechanism provided by this application;
[0044] Figure 4 It is a schematic diagram of the framework of the channel domain attention module provided by this application;
[0045] Figure 5 It is a schematic diagram of the framework of the spatial domain attention module provided by this application;
[0046] Figure 6 It is a schematic diagram of the framework of the combination of the channel domain attention module and the spatial domain attention module provided by this application;
[0047] Figure 7 It is a schematic diagram of the framework of the fourth embodiment of the neural network based on the attention mechanism provided by this application;
[0048] Figure 8 It is a schematic framework diagram of the first embodiment of the attention module provided by this application;
[0049] Figure 9 It is a schematic framework diagram of the fifth embodiment of the neural network based on the attention mechanism provided by this application;
[0050] Figure 10 It is a schematic framework diagram of the second embodiment of the attention module provided by this application;
[0051] Figure 11 It is a schematic framework diagram of the sixth embodiment of the neural network based on the attention mechanism provided by this application;
[0052] Figure 12 It is a schematic framework diagram of the third embodiment of the attention module provided by this application;
[0053] Figure 13 It is a schematic framework diagram of the seventh embodiment of the neural network based on the attention mechanism provided by this application;
[0054] Figure 14 It is a schematic framework diagram of the first embodiment of the neural network based on dense connections provided by this application;
[0055] Figure 15 It is a schematic framework diagram of the second embodiment of the neural network based on dense connections provided by this application;
[0056] Figure 16 It is a schematic framework diagram of the third embodiment of the neural network based on dense connections provided by this application;
[0057] Figure 17 It is a schematic framework diagram of the fourth embodiment of the neural network based on dense connections provided by this application;
[0058] Figure 18 It is a schematic flowchart of an embodiment of the image encoding method provided by this application;
[0059] Figure 19 It is a schematic diagram of the application process of the neural network provided by this application;
[0060] Figure 20 It is a schematic flowchart of an embodiment of the image decoding method provided by this application;
[0061] Figure 21 It is a schematic structural diagram of an embodiment of the encoder provided by this application;
[0062] Figure 22 It is a schematic structural diagram of an embodiment of the decoder provided by this application;
[0063] Figure 23Schematic diagram of a structure of an embodiment of a computer storage medium provided by the present application. Detailed implementation manners
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0065] The present application proposes a convolutional neural network for image filtering, and an application method of the neural network in an encoding and decoding standard. Specifically, the neural network provided by the present application can be used to replace one or more traditional filtering modules in an encoder and / or a decoder, or can be inserted into a filtering system as an additional filtering module to improve the filtering effect of the filtering system in combination with traditional filtering modules.
[0066] A traditional filtering model can perform loop filtering on a reconstructed image, that is, after the entire frame of the image is reconstructed, the pixel values in the reconstructed image are adjusted. The loop filtering is arranged in the front and back processes: deblocking filtering (DBF), sample adaptive offset (SAO), adaptive loop filtering (ALF), and cross-component adaptive loop filtering (CCALF).
[0067] Specifically, in loop filtering, deblocking filtering is performed first. The deblocking filtering (DBF) technology mainly filters the block boundaries in the block coding process to remove blocking artifacts and greatly improve the subjective quality of the image. Then, sample adaptive offset (SAO) is performed. This technology further improves the image quality by classifying pixels and adding specific offset values to each class of pixels, and can solve problems such as color shift and loss of high-frequency information in the image. Then, adaptive loop filtering (ALF) is performed. The ALF technology uses a diamond filter at the encoding end and obtains the filtering coefficients by Wiener filtering to filter the luminance and chrominance components to reduce image distortion. The cross-component adaptive loop filtering (CCALF) technology uses the luminance component after Wiener filtering as an adjustment value to further adjust the chrominance component after ALF.
[0068] Therefore, the neural network provided by the present application needs to be able to implement the filtering functions in one or more of the above loop filterings. The neural network provided by the present application will be introduced below:
[0069] The present application proposes a neural network for loop filtering based on an attention mechanism. The neural network includes as Figures 1 - 3One or more modules shown in the figure, including an attention module and a residual connection from the input end to the output end.
[0070] Among them, the combination method of the attention module in the neural network includes at least the following three types:
[0071] The first type, as Figure 1 shown, the attention module combines with a residual block as a component in the neural network structure, and directly uses the weights generated by the attention module to adjust the input of the attention module. Specifically, the residual block contains several convolutional layers, several activation layers, and a residual connection from the input end to the output end. For example, Figure 1 in the neural network structure from the input end to the output end, there are a convolutional layer, an activation layer, an attention module, and a convolutional layer arranged in sequence.
[0072] The second type, as Figure 2 shown, the attention module is directly set in the network backbone. By adding several attention modules, the weights generated by the attention module are directly used to adjust the input of the attention module. Among them, the multiple attention modules in the network backbone can be the same or different. For example, Figure 2 in the neural network structure from the input end to the output end, there are a convolutional layer, an activation layer, an attention module, a residual block, an attention module, and a convolutional layer arranged in sequence.
[0073] The third type, as Figure 3 shown, the attention module is used as a network branch of the neural network, and the weights obtained by the attention module act on the feature map obtained after the input of the attention module passes through several layers of the network backbone, including convolutional layers and activation layers, etc. Further, the neural network can be provided with multiple network branches, and the attention modules in each network branch can be the same or different. For example, Figure 3 in the neural network structure from the input end to the output end, there are a convolutional layer, an activation layer, a residual block, a convolutional layer, a residual block, a convolutional layer, and a convolutional layer arranged in sequence. Among them, the input of the attention module is connected to the input of the activation layer, and the output of the attention module is connected to the output of the second convolutional layer.
[0074] Next, continue to introduce the specific structure of the attention module:
[0075] The attention module provided in this application can be divided into three types:
[0076] (1) Channel domain attention mechanism module, as Figure 4As shown, the feature map undergoes Pooling operations, including maxPooling, average Pooling, etc., convolution operations, and / or fully connected operations, etc., to obtain a weight feature vector with the same length as the number of channels. After being processed by the Sigmoid activation module, its weight values are normalized to [0, 1], and the resulting feature map is used as the output of the attention module. Subsequently, it is multiplied corresponding to each feature point in each channel. The corresponding method is that the first value of the weight vector is multiplied by all the values of the first-channel feature map of the feature map to obtain the output of the first channel, and so on.
[0077] (2) Spatial domain attention mechanism module, such as Figure 5 As shown, the feature map undergoes Pooling operations, including maxPooling, average Pooling, etc., convolution operations, and / or fully connected operations, etc., to obtain a weight feature matrix with the same resolution size as the feature map. After being processed by the Sigmoid module, its weight values are normalized to [0, 1], and the resulting feature map is used as the output of the attention module. Subsequently, it is multiplied corresponding to each feature point in each channel. The corresponding method is that the weight value at a certain position of the weight matrix is multiplied by the values at this position in all channels of the feature map to obtain the outputs at this position in all channels respectively, and so on to obtain the weighted feature map.
[0078] (3) The combination of the channel domain attention mechanism module and the spatial domain attention mechanism module has the following three specific combination methods:
[0079] (31) The feature map first passes through the channel domain attention mechanism module, is weighted by the weights obtained, and then passes through the spatial domain attention mechanism module for a second weighting.
[0080] (32) The feature map first passes through the spatial domain attention mechanism module, is weighted by the weights obtained, and then passes through the channel domain attention mechanism module for a second weighting.
[0081] (33) As shown in Figure 6 As shown, the feature map passes through the channel domain attention mechanism module to obtain channel domain weights, and at the same time passes through the spatial domain attention mechanism module to obtain spatial domain weights. The channel domain weights are multiplied by the spatial domain weights, that is, each value of the channel domain weights is multiplied by all the spatial domain weights to obtain weights corresponding to each point of the feature map, and these weights are used to multiply the feature map correspondingly.
[0082] For example, the following provides an example of a neural network for filtering based on the attention mechanism.
[0083] The overall structure of the neural network is as shown in Figure 7As shown, it includes a residual connection from input to output. In the network backbone, the image first passes through a convolutional layer and a ReLU (Rectified Linear Unit) activation in the input, then passes through several attention residual structures, and finally passes through another convolutional layer. After adding through the residual connection, the output is obtained. Among them, the attention residual structure is as shown in Figure 7 shown on the right in the figure. The network backbone includes a convolution, a ReLU activation, and an attention module, and finally connects to a convolutional layer. The entire attention residual structure also includes a residual connection from input to output.
[0084] Among them, Figure 7 for the specific attention module in the figure, please continue to refer to Figure 8 . The feature map first passes through the channel domain attention mechanism module. After being weighted by the weights obtained, it passes through the spatial domain attention mechanism module for a second weighting. Among them, the channel domain attention mechanism module obtains the weight vector through the Max Pooling operation and the Sigmoid operation. Among them, the spatial domain attention mechanism module respectively obtains the feature map tensor with 2 channels through the Max Pooling operation and the Avg Pooling operation, and then obtains the weights through convolution and Sigmoid.
[0085] Another example of a neural network for filtering based on the attention mechanism is provided below.
[0086] The overall structure of the neural network is as shown in Figure 9 . It includes a residual connection from input to output. In the network backbone, the image passes through several convolutional layers and activation layers from the input, then an attention structure, repeats such connections several times, and finally, after connecting to another convolutional layer, the output of the network backbone is obtained. The obtained output of the network backbone is added with a residual connection from input to output to obtain the network output.
[0087] Among them, Figure 9 for the specific attention structure in the figure, please continue to refer to Figure 10 . The input feature map passes through the channel domain attention mechanism module to obtain the weights in the channel domain. At the same time, the input feature map passes through the spatial domain attention mechanism module to obtain the weights in the spatial domain. Multiply the two weights to obtain a weight tensor of the same size as the input feature map. Finally, multiply the weight tensor and the input feature map tensor element by element to obtain the output of the attention module. Among them, the channel domain attention mechanism module obtains the weight vector through the Max Pooling operation and the Sigmoid operation. Among them, the spatial domain attention mechanism module respectively obtains the feature map tensor with 2 channels through the Max Pooling operation and the Avg Pooling operation, and then obtains the weights through convolution and Sigmoid.
[0088] Another example of a neural network for filtering based on the attention mechanism is provided below.
[0089] The overall structure of the neural network is as Figure 11 shown, including a residual connection from the input to the output. In the network backbone, two different attention structure branches are connected.
[0090] Among them, Figure 11 For the specific attention structure in Figure 12 , please continue to refer to
[0091] Another example of a neural network for filtering based on the attention mechanism is provided below.
[0092] The overall structure of the neural network is as Figure 13 shown, including a residual connection from the input to the output. In the network backbone, two different attention structure branches are connected. Specifically, at the input end of the network backbone, a channel domain attention structure branch is first connected, then a spatial domain attention structure branch is connected, and finally several convolutional layers and several activation layers are connected.
[0093] In the embodiments of the present application, by proposing a neural network for loop filtering based on the attention mechanism, attention modules can be added at different positions in the network, and channel-based and / or space-based attention modules can be added. Different weights are added according to the importance of different channels in the network and / or different positions of the image, enhancing the network filtering performance.
[0094] The present application also proposes a neural network for loop filtering based on dense connections. The neural network includes one or more modules as shown in Figures 14 - 16 and a residual connection from the input end to the output end.
[0095] Among them, the combination methods of dense connections in the neural network include at least the following three:
[0096] The first one, as shown in Figure 14As shown, the inputs and outputs of all the feature maps within the dense connection module are connected to the inputs and outputs of all the feature maps.
[0097] Second, as Figure 15 shown, the inputs and outputs of all the feature maps within the dense connection module are connected to the final output.
[0098] Third, as Figure 16 shown, the input within the dense connection module is connected to the inputs and outputs of all the feature maps.
[0099] For example, an example of a neural network based on a dense connection block is provided below.
[0100] As Figure 17 shown, on the right side of the figure is the overall structure of the neural network. The neural network consists of a network backbone and a residual connection from the input to the output. The backbone of the neural network is composed of a convolutional layer, a ReLU activation layer, and several dense connection modules DenseBlock. The structure of the dense connection module DenseBlock is shown on the left side of the figure. The inputs and outputs of all the feature maps within the dense connection module DenseBlock are connected to the final output.
[0101] In the embodiments of the present application, by proposing a neural network for loop filtering based on dense connections, the learning ability of each layer of the network is better enhanced through several different dense connections, and the network filtering performance is enhanced.
[0102] The present application also proposes an image coding method based on the above neural network. For details, please refer to Figure 18 and Figure 19 , Figure 18 which is a schematic flowchart of an embodiment of the image coding method provided by the present application, Figure 19 and Figure 19 is a schematic diagram of the application process of the neural network provided by the present application. Among them, the image coding method of the embodiments of the present application is applied to an encoder.
[0103] As Figure 18 shown, the image coding method of this embodiment specifically includes the following steps:
[0104] Step S11: Obtain a reconstructed image of the image to be encoded.
[0105] In the embodiments of the present application, as Figure 19 shown, the neural network shown in the above embodiments can act on the encoding end.
[0106] First, the encoding end processes the image to be encoded using a preset encoding method to obtain a reconstructed image. Among them, the preset encoding method includes, but is not limited to: partitioning processing, prediction processing, residual processing (transformation processing), quantization processing, etc. During this encoding process, the encoding end obtains the encoding information generated during the encoding process, such as partitioning information, prediction information, residual information, and quantization information, etc., based on the preset encoding method and its reconstructed image, for subsequent input to the neural network.
[0107] Step S12: Combine the reconstructed image with different encoding information respectively as different network inputs, and input them into the preset neural network respectively, so that the preset neural network filters each network input to obtain a filtered reconstructed image.
[0108] In the embodiment of the present application, the encoding end inputs the reconstructed image into the neural network and adds encoding information that can be obtained by the decoding end, such as partitioning information, prediction information, residual information, and quantization information. Among them, during the encoding process and before the filtering process, the above information can be obtained.
[0109] Specifically, the encoding end can input the reconstructed image combined with a single type of encoding information into the neural network, or input the reconstructed image combined with two or more types of encoding information into the neural network.
[0110] The following details the methods for adding the above information:
[0111] (a) Add partitioning information to the neural network input, that is, add a channel with the same resolution as the image to be filtered, that is, the reconstructed image. The value of each point in this channel is either 1 or 0. When the current point position is on the partitioning boundary, the value is 1, otherwise it is 0, or vice versa. Among them, the partitioning boundary includes one or more of coding unit partitioning, prediction block partitioning, and transform block partitioning.
[0112] (b) Add prediction information to the neural network input, that is, add a channel with the same resolution as the image to be filtered, that is, the reconstructed image. The value of each point in this channel is the predicted value obtained through intra-frame prediction, inter-frame prediction, or other prediction processes at this position.
[0113] (c) Add residual information to the neural network input, that is, add a channel with the same resolution as the image to be filtered, that is, the reconstructed image. The value of each point in this channel is the reconstructed residual value obtained after transformation, quantization, inverse quantization, and inverse transformation.
[0114] (d) Add quantization information to the neural network input, that is, add a channel with the same resolution as the image to be filtered, i.e., the reconstructed image. The value of each point in this channel can be the sliceQP of this frame, or the Q_step value, or, when the rate control is enabled, the QP (Quantizer Parameter) value of the image block corresponding to each pixel point.
[0115] Furthermore, since the reconstructed image generally has three color components of YUV, through network configuration, the neural network can be set to filter the three color components simultaneously, or filter the three color components separately. Specifically, the embodiments of the present application provide the following two methods for different color components:
[0116] (a) Input at least 2 components together and output the filtered and reconstructed images of at least 2 components.
[0117] (b) Input at least 2 components together and output the filtered and reconstructed image of 1 component.
[0118] For example, in the encoding and decoding process with three components of YUV, train a separate model for the Y component and the same model for the UV components.
[0119] The output of the model for the Y component is the reconstructed image of the Y component after being filtered by the neural network, and the input is:
[0120] (1) The reconstructed image of the Y component before being filtered by the neural network;
[0121] (2) The reconstructed image of the U component before being filtered by the neural network;
[0122] (3) The reconstructed image of the V component before being filtered by the neural network;
[0123] (4) The feature map with the same resolution generated by the partitioning information of the Y component;
[0124] (5) The feature map with the same resolution generated by the residual information of the Y component;
[0125] (6) The feature map with the same resolution generated by the sliceQP of each frame in the quantization information of the Y component.
[0126] The output of the model for the U component is the reconstructed image of the U component after being filtered by the neural network, and the input is:
[0127] (1) The reconstructed image of the Y component before being filtered by the neural network;
[0128] (2) The reconstructed image of the U component before being filtered by the neural network;
[0129] (4) Feature maps with the same resolution generated from the partitioning information of the U component;
[0130] (5) Feature maps with the same resolution generated from the residual information of the U component;
[0131] (6) Feature maps with the same resolution generated by sliceQP of each frame in the quantization information of the U component.
[0132] The output of the V component model is the reconstructed image of the V component after being filtered by the neural network, and the input is:
[0133] (1) The reconstructed image of the Y component before being filtered by the neural network;
[0134] (2) The reconstructed image of the V component before being filtered by the neural network;
[0135] (3) Feature maps with the same resolution generated from the partitioning information of the V component;
[0136] (4) Feature maps with the same resolution generated from the residual information of the V component;
[0137] (5) Feature maps with the same resolution generated by sliceQP of each frame in the quantization information of the V component.
[0138] Furthermore, the encoding end can also divide the reconstructed image into several reconstructed sub-blocks; combine each reconstructed sub-block with different coding information as different network inputs, and input them into the preset neural network respectively to obtain the filtered reconstructed sub-blocks output by the preset neural network; combine all the filtered reconstructed sub-blocks to obtain the filtered reconstructed image. The process is the same as the above content and will not be elaborated here.
[0139] Step S13: Obtain the difference value between each filtered reconstructed image and the image to be encoded.
[0140] In the embodiment of the present application, the main idea of adjusting the reconstructed value is: divide the input of the neural network into blocks, input each block X into the neural network for filtering to obtain the reconstructed image block Y after being filtered by the neural network, obtain the original image block Y_org corresponding to the position of the reconstructed block, and obtain an appropriate scaling factor scale such that under the set metric D, D((Y - X)·scale + X, Y_org) is closest to 0.
[0141] The encoding and decoding ends can adopt frame-level or block-level scaling factors. If a certain image increases the frame-level scaling factor, the residual image needs to be multiplied by the frame-level scaling factor and then added to the network input image to obtain the final output image; if a certain image increases the block-level scaling factor, that is, the scaling factor of each block can be different, at this time, each residual image block needs to be multiplied by its respective scaling factor and then added to the network input image to obtain the final output image.
[0142] The encoding end combines different encoding information or combinations of different encoding information with the reconstructed image and inputs them into the neural network to obtain corresponding filtered reconstructed images respectively. Then, the encoding end calculates the difference value between each filtered reconstructed image and the image to be encoded, and uses the difference value to characterize the filtering effect after the neural network adds the encoding information. Among them, the difference value can be calculated from the difference between the pixel values of the filtered reconstructed image and the image to be encoded.
[0143] Step S14: Based on the magnitude of the difference value, select the target encoding information corresponding to the filtered reconstructed image, and encode the target encoding information into the bitstream of the image to be encoded.
[0144] In the embodiment of the present application, the encoding end can select the encoding information corresponding to the minimum value among all the difference values between the images to be encoded and the filtered reconstructed images as the target encoding information, and write the target encoding information into the encoded bitstream of the image to be encoded, so that the decoding end can decode the encoded bitstream according to the encoding information and perform image filtering reconstruction.
[0145] Furthermore, after obtaining the encoded bitstream of the image to be encoded, the encoding end also needs to configure relevant syntax elements in the encoded bitstream. The specific syntax elements and their introductions are as follows:
[0146] Neural network tool switch syntax:
[0147] When applying the neural network, a neural network tool switch syntax can be set to indicate whether neural network filtering is used for the encoding and decoding of the current sequence. If not used, there is no need to transmit other syntax.
[0148] Frame-level and block-level switch syntax:
[0149] If the neural network tool switch is turned on, when applying the neural network, a switch syntax can be set for each frame to indicate whether neural network filtering is used for that frame.
[0150] If the neural network filtering switch of a certain frame is turned on, when applying the neural network, a switch is provided for each p-coded block, including the case where the coded block is the largest coding unit, to control whether the current frame or the current coded block can use the neural network. This switch also needs to be transmitted as a syntax element.
[0151] Reconstructed pixel value adjustment syntax:
[0152] If the neural network tool switch is turned on and the neural network filtering switch of the current frame is turned on, the syntax for reconstructed pixel value adjustment needs to be transmitted. If a frame-level scaling factor is added to a certain image, the frame-level scaling factor syntax needs to be transmitted; if a block-level scaling factor is added to a certain image, that is, the scaling factor of each block can be different, the block-level scaling factor needs to be transmitted as a syntax element.
[0153] In an embodiment of the present application, an encoding end obtains a reconstructed image of an image to be encoded; combines the reconstructed image with different encoding information respectively as different network inputs, and inputs them into a preset neural network respectively, so that the preset neural network filters each network input to obtain a filtered reconstructed image; obtains the difference value between each filtered reconstructed image and the image to be encoded; based on the size of the difference value, selects the target encoding information corresponding to the filtered reconstructed image, and encodes the target encoding information into the bitstream of the image to be encoded. Through the above image encoding method, the present application can add encoding information during neural network filtering, improve the generalization of the neural network, and enhance the reconstruction quality of the neural network.
[0154] Please continue to refer to Figure 20 , Figure 20 which is a schematic flowchart of an embodiment of the image decoding method provided by the present application. In an embodiment of the present application, as Figure 19 shown, the neural network shown in the above embodiment can act on the decoding end.
[0155] As Figure 20 shown, the image decoding method of this embodiment specifically includes the following steps:
[0156] Step S21: Obtain a bitstream and its encoding information.
[0157] Step S22: Decode the bitstream to obtain a decoded reconstructed image.
[0158] In an embodiment of the present application, a decoder decodes relevant syntax elements to obtain the opening situation of the neural network filtering switch and confirm the range of neural network filtering.
[0159] Step S23: Input the decoded reconstructed image and the encoding information into a preset neural network, so that the preset neural network filters the decoded reconstructed image based on the encoding information to obtain a final reconstructed image.
[0160] In an embodiment of the present application, a decoder constructs an input of a neural network with the decoded reconstructed image and the encoding information, and performs filtering to obtain a final reconstructed image. Among them, the encoding information is the target encoding information obtained by the encoder by comparing different encoding information during the encoding stage, or a preset optimal encoding information.
[0161] Specifically, the decoder divides the network input into blocks, inputs the sub-blocks with the filtering state turned on into the neural network for filtering to obtain the reconstructed image blocks after network filtering, and reorganizes them to obtain the final reconstructed image.
[0162] In this application, a method of combining more information in the encoding and decoding in the network input is proposed, enabling the network to better filter according to the characteristics of the current sequence, improving the generalization ability during the network learning, and further enhancing the reconstruction quality after network filtering. A method of further adjusting the reconstructed image after neural network filtering is proposed, including adjusting the reconstructed values and determining whether to use the neural network filtering switch for blocks, further utilizing the adaptability of the encoder and better improving the quality of the reconstructed image.
[0163] To implement the image encoding method of the above embodiment, this application proposes an encoder. For details, please refer to Figure 21 , Figure 21 which is a schematic structural diagram of an embodiment of the encoder provided in this application.
[0164] The encoder 300 includes a memory 31 and a processor 32. Among them, the memory 31 is coupled to the processor 32.
[0165] The memory 31 is used to store computer programs, and the processor 32 is used to execute the computer programs to implement the image encoding method of the above embodiment.
[0166] In this embodiment, the processor 32 can also be referred to as a CPU (Central Processing Unit). The processor 32 may be an integrated circuit chip with signal processing capabilities. The processor 32 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor, or the processor 32 can also be any conventional processor, etc.
[0167] To implement the image decoding method of the above embodiment, this application proposes a decoder. For details, please refer to Figure 22 , Figure 22 which is a schematic structural diagram of an embodiment of the decoder provided in this application.
[0168] The decoder 400 includes a memory 41 and a processor 42. Among them, the memory 41 is coupled to the processor 42.
[0169] The memory 41 is used to store computer programs, and the processor 42 is used to execute the computer programs to implement the image decoding method of the above embodiment.
[0170] In this embodiment, the processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with the ability to process signals. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 42 may also be any conventional processor, etc.
[0171] The present application also provides a computer storage medium. Please continue to refer to Figure 23 , Figure 23 FIG. is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. The computer storage medium 600 stores a computer program 61. When the computer program 61 is executed by a processor, it is used to implement the image encoding method and / or the image decoding method of the above embodiment.
[0172] When the embodiments of the present application are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0173] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. An image encoding method, characterized in that, The described image encoding method includes: Obtaining a reconstructed image of the image to be encoded; Combining the reconstructed image of the image to be encoded with different encoding information respectively as different network inputs, and inputting them into a preset neural network respectively, so that the preset neural network filters each network input to obtain a filtered reconstructed image, where the encoding information includes a combination of one or more of partitioning information, prediction information, residual information, and quantization information; Obtaining the difference value between each filtered reconstructed image and the reconstructed image of the image to be encoded; Selecting the encoding information of the filtered reconstructed image corresponding to the minimum value of the difference value as the target encoding information, and encoding the target encoding information into the bitstream of the image to be encoded.
2. The image encoding method according to claim 1, characterized in that, The step of combining the reconstructed image of the image to be encoded with different encoding information respectively as different network inputs, and inputting them into a preset neural network respectively, so that the preset neural network filters each network input to obtain a filtered reconstructed image includes: Dividing the reconstructed image of the image to be encoded into several reconstructed sub-blocks; Combining each reconstructed sub-block with different encoding information respectively as different network inputs, and inputting them into a preset neural network respectively to obtain the filtered reconstructed sub-blocks output by the preset neural network; Combining all the filtered reconstructed sub-blocks to obtain a filtered reconstructed image; Wherein, the preset neural network is used to replace the traditional filtering module in the filtering network and / or be added to the filtering network, and the functions of the traditional filtering module include one or more of deblocking filtering, sample adaptive offset filtering, adaptive loop filtering, and cross-component adaptive loop filtering.
3. The image encoding method according to claim 1, characterized in that, The reconstructed image of the image to be encoded includes a first component reconstructed image, a second component reconstructed image, and a third component reconstructed image; the image encoding method further includes: Inputting the first component reconstructed image, the second component reconstructed image, and the third component reconstructed image into the preset neural network simultaneously; Obtaining the first component reconstructed image, the second component reconstructed image, and the third component reconstructed image after being filtered by the preset neural network, and combining them into a filtered reconstructed image.
4. The image encoding method according to claim 1, characterized in that, The reconstructed image of the image to be encoded includes a first component reconstructed image, a second component reconstructed image, and a third component reconstructed image; the image encoding method further includes: Inputting the first component reconstructed image and one or two of the second component reconstructed image and the third component reconstructed image into the preset neural network; Obtaining the first component reconstructed image after being filtered by the preset neural network.
5. The image encoding method according to claim 1, characterized in that, The preset neural network arranges a convolutional layer, an activation layer, an attention module, and a convolutional layer in sequence from the input end to the output end, and the preset neural network further includes a residual connection from the input end to the output end; Wherein, the attention module includes a channel domain attention module and / or a spatial domain attention module.
6. The image encoding method according to claim 1, characterized in that, The preset neural network arranges a convolutional layer, an activation layer, a residual block, and a convolutional layer in sequence from the input end to the output end. The preset neural network further includes a residual connection from the input end to the output end, and an attention module inserted as a network branch between the output of the activation layer and the output of the subsequent convolutional layer.
7. The image encoding method according to claim 1, characterized in that, The preset neural network arranges multiple groups of densely connected modules in sequence from the input end to the output end. Among them, each densely connected module includes multiple convolutional layers, and the input of each convolutional layer is connected to its own output and / or the output of other convolutional layers.
8. The image encoding method according to claim 1, characterized in that, The image encoding method further includes: Encoding a plurality of encoding and decoding markers in the bitstream based on the encoding and decoding results. The encoding and decoding markers include a neural network tool switch syntax, a frame-level switch syntax, a block-level switch syntax, and / or a reconstructed pixel value adjustment syntax.
9. An image decoding method, characterized in that, The image decoding method includes: Obtaining the bitstream described in any one of claims 1-8 and its encoding information; Decoding the bitstream to obtain a decoded reconstructed image; Inputting the decoded reconstructed image and the encoding information into the preset neural network, so that the preset neural network filters the decoded reconstructed image based on the encoding information to obtain a final reconstructed image.
10. The image decoding method according to claim 9,It is characterized in that wherein the bitstream includes encoding and decoding markers; The image decoding method further includes: Obtaining the encoding and decoding markers in the bitstream, where the encoding and decoding markers include a neural network tool switch syntax; Judging whether to use the preset neural network for filtering based on the encoding and decoding markers; If so, inputting the decoded reconstructed image and the encoding and decoding markers into the preset neural network.
11. An encoder, characterized in that, The encoder includes a processor and a memory; a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the steps of the image encoding method described in any one of claims 1-8.
12. A decoder, characterized in that, The decoder includes a processor and a memory; a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the steps of the image decoding method described in any one of claims 9-10.
13. A computer storage medium, characterized in that, The computer storage medium stores a computer program, and when the computer program is executed by a computer, it implements the steps of the image encoding method described in any one of claims 1-8 and / or the image decoding method described in any one of claims 9-10.
Citation Information
Patent Citations
Method and system for realizing filtration in adversarial generative network based video coding and decoding loop
CN108174225A
Loop filtering method, device and equipment in video encoding and decoding, and storage medium
CN111711824A