Image Filtering Method, Apparatus, Device, and Storage Medium
By determining the target filtering order in the neural network filter to improve the filtering effect, the problem of poor generalization of chroma component filtering in the prior art is solved, and better filtering effect and generalization are achieved.
Patent Information
- Application Number
- CN202310430930.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-04-14
AI Technical Summary
When existing neural network filters filter chromaticity components, their generalization is poor and the filtering effect is poor.
The target filtering sequence is determined by decoding the code stream or filtering costs based on the N filtering orders, and the first chrominance component and the second chrominance component of the image block are input to the neural network filter for filtering.
It improves the filtering effect, improves the generalization of neural network filters, and enhances the image encoding and codec performance.
Smart Images

Figure CN116405701B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of image coding and decoding technologies, and in particular, to an image filtering method, apparatus, device, and storage medium. Background Art
[0002] With the development of video technology, the amount of data included in video data is large. To facilitate the transmission of video data, video devices perform video compression technology to make the video data more effectively transmitted or stored. In video compression, both the encoding end and the decoding end need to perform operations such as inverse quantization and inverse transformation to obtain a reconstructed image. Since losses are introduced in video compression, the reconstructed image is filtered to reduce the compression loss of the image.
[0003] With the rapid development of neural network technology, neural network filters have been widely used in video processing. However, currently, when filtering the chrominance component, the neural network filter has problems of poor generalization and poor filtering effect. Summary of the Invention
[0004] The present application provides an image filtering method, apparatus, device, and storage medium to improve the filtering effect of an image and enhance the generalization of a neural network filter.
[0005] In a first aspect, the present application provides an image filtering method, including:
[0006] Decoding the bitstream of the current image to obtain the residual value of the current image, and determining the reconstructed image of the current image based on the residual value;
[0007] For the current image block to be filtered in the reconstructed image, determining the target filtering order of the first chrominance component and the second chrominance component of the current image block, where the target filtering order is determined by decoding the bitstream or based on the filtering cost of N filtering orders, and N is a positive integer greater than 1;
[0008] Based on the target filtering order, inputting the first chrominance component and the second chrominance component of the current image block into a neural network filter for filtering to obtain the chrominance filtering block of the current image block.
[0009] In a second aspect, the present application provides an image filtering method, including:
[0010] Encoding the current image to obtain the reconstructed image of the current image;
[0011] For the current image block to be filtered in the reconstructed image, determining the target filtering order of the first chrominance component and the second chrominance component of the current image block, where the target filtering order is determined based on the filtering cost of N filtering orders, and N is a positive integer greater than 1;
[0012] Based on the target filtering order, input the first chrominance component and the second chrominance component of the current image block into a neural network filter for filtering to obtain the filtered image block of the current image block.
[0013] In a third aspect, the present application provides an image filtering device, including:
[0014] A decoding unit, configured to decode the bitstream of the current image to obtain the residual value of the current image, and determine the reconstructed image of the current image based on the residual value;
[0015] An order determination unit, configured to determine the target filtering order of the first chrominance component and the second chrominance component of the current image block to be filtered in the reconstructed image, where the target filtering order is determined by decoding the bitstream or based on the filtering cost of N filtering orders, and N is a positive integer greater than 1;
[0016] A filtering unit, configured to input the first chrominance component and the second chrominance component of the current image block into a neural network filter for filtering based on the target filtering order to obtain the chrominance filtered block of the current image block.
[0017] In a fourth aspect, the present application provides an image filtering device, including:
[0018] An encoding unit, configured to encode the current image to obtain the reconstructed image of the current image;
[0019] An order determination unit, configured to determine the target filtering order of the first chrominance component and the second chrominance component of the current image block to be filtered in the reconstructed image, where the target filtering order is determined based on the filtering cost of N filtering orders, and N is a positive integer greater than 1;
[0020] A filtering unit, configured to input the first chrominance component and the second chrominance component of the current image block into a neural network filter for filtering based on the target filtering order to obtain the filtered image block of the current image block.
[0021] In a fifth aspect, a decoder is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementation manners.
[0022] In a sixth aspect, an encoder is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the second aspect or its various implementation manners.
[0023] In a seventh aspect, a chip is provided for implementing the method in any one of the first aspect to the second aspect or its various implementation manners above. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device installed with the chip executes the method in any one of the first aspect to the second aspect or its various implementation manners above.
[0024] In an eighth aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program enables a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners above.
[0025] In a ninth aspect, a computer program product is provided, including computer program instructions, and the computer program instructions enable a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners above.
[0026] In a tenth aspect, a computer program is provided, which when running on a computer, enables the computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners above.
[0027] In summary, the present application determines a reconstructed image of a current image; for a current image block to be filtered in the reconstructed image, determines a target filtering order of a first chrominance component and a second chrominance component of the current image block, where the target filtering order is determined by decoding a bitstream or based on filtering costs of N filtering orders, and N is a positive integer greater than 1; based on the target filtering order, inputs the first chrominance component and the second chrominance component of the current image block into a neural network filter for filtering to obtain a chrominance filtered block of the current image block. That is to say, the embodiment of the present application determines the target filtering order based on the filtering costs of N filtering orders, improves the selection accuracy of the target filtering order, and when inputting the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering based on the accurately determined target filtering order, can improve the filtering effect, thereby improving the generalization of the neural network filter and enhancing the image coding and decoding performance. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0029] Figure 1 It is a schematic block diagram of a video coding and decoding system according to an embodiment of the present application;
[0030] Figure 2 Schematic diagram of the encoding framework provided by the embodiment of the present application;
[0031] Figure 3 Schematic diagram of the decoding framework provided by the embodiment of the present application;
[0032] Figure 4 Schematic diagram of a coding unit;
[0033] Figure 5 Schematic diagram of the filtering process of a neural network filter;
[0034] Figure 6A Schematic diagram of a chrominance training sequence of a neural network filter;
[0035] Figure 6B Schematic diagram of a filtering sequence of a neural network filter;
[0036] Figure 7 Schematic flowchart of the image filtering method provided by an embodiment of the present application;
[0037] Figures 8A to 8C Schematic diagram of the current image block;
[0038] Figure 9 Schematic diagram of the filtered area around the current image block;
[0039] Figure 10 Schematic diagram of the determination of a target filtering sequence;
[0040] Figures 11A to 12B Schematic diagram of the determination of a target filtering sequence;
[0041] Figure 13 Schematic flowchart of the image filtering method provided by an embodiment of the present application;
[0042] Figure 14 Schematic diagram of the determination of a target filtering sequence;
[0043] Figure 15A and Figure 15B Schematic diagram of the determination of a target filtering sequence;
[0044] Figure 16 Schematic block diagram of the image filtering device provided by an embodiment of the present application;
[0045] Figure 17 Schematic block diagram of the image filtering device provided by an embodiment of the present application;
[0046] Figure 18 Schematic block diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0047] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two.
[0049] The present application can be applied to the fields of image coding and decoding, video coding and decoding, hardware video coding and decoding, dedicated circuit video coding and decoding, real-time video coding and decoding, etc. For example, the solution of the present application can be combined with the standard of end-to-end image coding based on deep learning, such as JPEG AI. Or, the solution of the present application can be combined with other proprietary or industry standards for operation, and the standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions. It should be understood that the technology of the present application is not limited to any specific coding and decoding standard or technology.
[0050] For ease of understanding, first, in combination with Figure 1 a video coding and decoding system involved in the embodiments of the present application will be introduced.
[0051] Figure 1Schematic block diagram of a video encoding and decoding system according to an embodiment of the present application. It should be noted that Figure 1 This is only an example. The video encoding and decoding system of the embodiments of the present application includes but is not limited to Figure 1 as shown. As Figure 1 shown, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) video data to generate a bitstream and transmit the bitstream to the decoding device. The decoding device decodes the bitstream generated by the encoding device to obtain the decoded video data.
[0052] The encoding device 110 of the embodiments of the present application can be understood as a device with video encoding function, and the decoding device 120 can be understood as a device with video decoding function. That is, the embodiments of the present application include a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebooks (e.g., laptops) computers, tablet computers, set-top boxes, TVs, cameras, display devices, digital media players, video game consoles, in-vehicle computers, etc.
[0053] In some embodiments, the encoding device 110 can transmit the encoded video data (such as a bitstream) to the decoding device 120 via a channel 130. The channel 130 can include one or more media and / or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.
[0054] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to directly transmit the encoded video data to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media can also include wired communication media, such as one or more physical transmission lines.
[0055] In another example, the channel 130 includes a storage medium that can store the encoded video data of the encoding device 110. The storage medium includes various locally accessible data storage media, such as optical discs, DVDs, flash memories, etc. In this example, the decoding device 120 can obtain the encoded video data from the storage medium.
[0056] In another example, the channel 130 may include a storage server that can store the video data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded video data from the storage server. Optionally, the storage server can store the encoded video data and can transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.
[0057] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. Among them, the output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0058] In some embodiments, in addition to the video encoder 112 and the input interface 113, the encoding device 110 may further include a video source 111.
[0059] The video source 111 may include at least one of a video capture device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system. Among them, the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.
[0060] The video encoder 112 encodes the video data from the video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the encoding information of the pictures or the sequence of pictures in the form of a bitstream. The encoding information may include encoded image data and associated data. The associated data may include a sequence parameter set (SPS for short), a picture parameter set (PPS for short), and other syntax structures. The SPS may contain parameters applied to one or more sequences. The PPS may contain parameters applied to one or more pictures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.
[0061] The video encoder 112 directly transmits the encoded video data to the decoding device 120 via the output interface 113. The encoded video data may also be stored on a storage medium or a storage server for subsequent reading by the decoding device 120.
[0062] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122.
[0063] In some embodiments, in addition to the input interface 121 and the video decoder 122, the decoding device 120 may further include a display device 123.
[0064] Among them, the input interface 121 includes a receiver and / or a modem. The input interface 121 can receive the encoded video data through the channel 130.
[0065] The video decoder 122 is used to decode the encoded video data to obtain the decoded video data, and transmit the decoded video data to the display device 123.
[0066] The display device 123 displays the decoded video data. The display device 123 can be integrated with the decoding device 120 or outside the decoding device 120. The display device 123 can include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0067] In addition, Figure 1 merely for example, the technical solutions of the embodiments of the present application are not limited to Figure 1 , for example, the technology of the present application can also be applied to unilateral video encoding or unilateral video decoding.
[0068] The video encoding framework related to the embodiments of the present application will be introduced below.
[0069] Figure 2 It is a schematic block diagram of a video encoder according to an embodiment of the present application. It should be understood that the video encoder 200 can be used for lossy compression of images or lossless compression of images. The lossless compression can be visually lossless compression or mathematically lossless compression.
[0070] The video encoder 200 can be applied to image data in the luminance chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4. Y represents luminance, Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance used to describe color and saturation. For example, in the color format, 4:2:0 means that for every 4 pixels, there are 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that for every 4 pixels, there are 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 means full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0071] For example, the video encoder 200 reads video data. For each frame of the video data, a frame of the image is divided into a number of coding tree units (CTUs). In some examples, a CTU can be referred to as a "tree block", "Largest Coding Unit" (LCU for short), or "coding tree block" (CTB for short). Each CTU can be associated with a pixel block of equal size within the image. Each pixel can correspond to one luminance sample and two chrominance samples. Therefore, each CTU can be associated with a luminance sample block and two chrominance sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU can be further divided into a number of Coding Units (CUs) for encoding. A CU can be a rectangular block or a square block. A CU can be further divided into a prediction Unit (PU for short) and a transform unit (TU for short), thus separating encoding, prediction, and transformation, making the processing more flexible. In one example, a CTU is divided into CUs in a quadtree manner, and a CU is divided into TUs and PUs in a quadtree manner.
[0072] Video encoders and video decoders can support various PU sizes. Assuming that the size of a specific CU is 2N×2N, video encoders and video decoders can support PU sizes of 2N×2N or N×N for intra prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter prediction. Video encoders and video decoders can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.
[0073] In some embodiments, such asFigure 2 As shown, the video encoder 200 may include: a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image buffer 270, and an entropy encoding unit 280. It should be noted that the video encoder 200 may include more, fewer, or different functional components.
[0074] Optionally, in this application, a current block may be referred to as a current coding unit (CU) or a current prediction unit (PU), etc. A prediction block may also be referred to as a predicted image block or an image prediction block, and a reconstructed image block may also be referred to as a reconstruction block or an image reconstruction block.
[0075] In some embodiments, the prediction unit 210 includes an inter-frame prediction unit 211 and an intra-frame prediction unit 212. Since there is a strong correlation between adjacent pixels in a frame of a video, the intra-frame prediction method is used in video coding and decoding technologies to eliminate the spatial redundancy between adjacent pixels. Since there is a strong similarity between adjacent frames in a video, the inter-frame prediction method is used in video coding and decoding technologies to eliminate the temporal redundancy between adjacent frames, thereby improving the coding efficiency.
[0076] The inter-frame prediction unit 211 can be used for inter-frame prediction, which can include motion estimation and motion compensation. Motion estimation can search for reference pictures in the list of reference pictures to find the reference block of the image block to be encoded. Motion estimation can generate an index indicating the reference block and a motion vector indicating the spatial displacement between the image block to be encoded and the reference block. Motion estimation can output the index of the reference block and the motion vector as the motion information of the image block to be encoded. Motion compensation can obtain the prediction information of the image block to be encoded based on the motion information of the image block to be encoded. Inter-frame prediction can refer to the image information of different frames. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block to eliminate temporal redundancy. The frames used for inter-frame prediction can be P frames and / or B frames. A P frame refers to a forward prediction frame, and a B frame refers to a bidirectional prediction frame. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block. The motion information includes the reference frame list where the reference frame is located, the reference frame index, and the motion vector. The motion vector can be an integer pixel or a fractional pixel. If the motion vector is a fractional pixel, then an interpolation filter needs to be used in the reference frame to create the required fractional pixel block. Here, the integer pixel or fractional pixel block found in the reference frame according to the motion vector is called the reference block. Some techniques directly use the reference block as the prediction block, and some techniques further process the reference block to generate the prediction block. Further processing the reference block to generate the prediction block can also be understood as using the reference block as the prediction block and then further processing the prediction block to generate a new prediction block.
[0077] The intra-frame prediction unit 212 only refers to the information of the same frame image and predicts the pixel information within the currently encoded image block to eliminate spatial redundancy. The frame used for intra-frame prediction can be an I frame.
[0078] There are various intra-frame prediction modes. Taking the international digital video coding standard H series as an example, the H.264 / AVC standard has 8 angular prediction modes and 1 non-angular prediction mode. The H.265 / HEVC has been extended to 33 angular prediction modes and 2 non-angular prediction modes. The intra-frame prediction modes used in HEVC include the Planar mode, DC, and 33 angular modes, for a total of 35 prediction modes. The intra-frame modes used in VVC include Planar, DC, and 65 angular modes, for a total of 67 prediction modes.
[0079] It should be noted that with the increase in the number of angular modes, intra-frame prediction will be more accurate and more in line with the requirements for the development of high-definition and ultra-high-definition digital videos.
[0080] Residual unit 220 can generate a residual block of a CU based on the pixel block of the CU and the prediction block of the PU of the CU. For example, residual unit 220 can generate a residual block of the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the PU of the CU.
[0081] Transform / quantization unit 230 can quantize transform coefficients. Transform / quantization unit 230 can quantize the transform coefficients associated with the TU of the CU based on the quantization parameter (QP) value associated with the CU. Video encoder 200 can adjust the quantization degree applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU. Exemplarily, the residual video signal undergoes transform operations such as DFT, DCT, etc., to convert the signal into the transform domain, which is called transform coefficients. The signal in the transform domain further undergoes a lossy quantization operation, losing certain information, such that the quantized signal is conducive to compressed representation. In some video coding standards, there may be more than one transform method to choose from. Therefore, the encoding end also needs to select one of them for the currently encoded CU and inform the decoding end. The fineness of quantization is usually determined by the quantization parameter (QP). A larger QP value means that coefficients in a larger value range will be quantized to the same output, so it usually brings greater distortion and a lower bit rate; on the contrary, a smaller QP value means that coefficients in a smaller value range will be quantized to the same output, so it usually brings less distortion and a corresponding higher bit rate.
[0082] Inverse transform / quantization unit 240 can apply inverse quantization and inverse transform to the quantized transform coefficients respectively to reconstruct the residual block from the quantized transform coefficients.
[0083] Reconstruction unit 250 can add the samples of the reconstructed residual block to the corresponding samples of one or more prediction blocks generated by prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of the CU in this way, video encoder 200 can reconstruct the pixel block of the CU.
[0084] The loop filter unit 260 is used to process the pixels after inverse transformation and inverse quantization, compensate for the distorted information, and provide a better reference for the subsequent encoded pixels. For example, it can perform deblocking filtering operations to reduce the blocking effect of the pixel blocks associated with the CU. As described above, for an already encoded image, through the operations of inverse quantization, inverse transformation, and prediction compensation, a reconstructed decoded image can be obtained. Compared with the original image, due to the influence of quantization, some information is different from the original image, resulting in distortion. Filtering operations on the reconstructed image, such as deblocking filter (DBF), sample adaptive offset (SAO), or adaptive loop filter (ALF), can effectively reduce the degree of distortion caused by quantization. Since these filtered reconstructed images will be used as references for subsequent encoded images to predict future signals, the above-mentioned filtering operations are also called loop filtering, that is, filtering operations within the encoding loop.
[0085] The decoded image buffer 270 can store the reconstructed pixel blocks. The inter prediction unit 211 can use the reference image containing the reconstructed pixel blocks to perform inter prediction on the PUs of other images. Additionally, the intra prediction unit 212 can use the reconstructed pixel blocks in the decoded image buffer 270 to perform intra prediction on the PUs of other images in the same image as the CU.
[0086] The entropy coding unit 280 can receive the quantized transform coefficients from the transform / quantization unit 230. The entropy coding unit 280 can perform one or more entropy coding operations on the quantized transform coefficients to generate entropy-coded data. Exemplarily, the quantized transform-domain signal will be statistically compressed and encoded according to the frequency of each value, and finally a binary (0 or 1) compressed bitstream will be output. At the same time, other information generated during encoding, such as the selected mode, motion vector, etc., also needs to be entropy-coded to reduce the bit rate. In one example, statistical coding is a lossless coding method that can effectively reduce the bit rate required to represent the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).
[0087] Figure 3 It is a schematic block diagram of the video decoder involved in the embodiments of the present application.
[0088] As Figure 3As shown, video decoder 300 includes: entropy decoding unit 310, prediction unit 320, inverse quantization / transformation unit 330, reconstruction unit 340, loop filtering unit 350, and decoded picture buffer 360. It should be noted that video decoder 300 may include more, fewer, or different functional components.
[0089] Video decoder 300 may receive a bitstream. Entropy decoding unit 310 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, entropy decoding unit 310 may parse the entropy-coded syntax elements in the bitstream. Prediction unit 320, inverse quantization / transformation unit 330, reconstruction unit 340, and loop filtering unit 350 may decode video data according to the syntax elements extracted from the bitstream, that is, generate decoded video data.
[0090] In some embodiments, prediction unit 320 includes intra prediction unit 322 and inter prediction unit 321.
[0091] Intra prediction unit 322 may perform intra prediction to generate a predicted block of a PU. Intra prediction unit 322 may use an intra prediction mode to generate a predicted block of a PU based on pixel blocks of spatially adjacent PUs. Intra prediction unit 322 may also determine the intra prediction mode of a PU according to one or more syntax elements parsed from the bitstream.
[0092] Inter prediction unit 321 may construct a first reference picture list (list 0) and a second reference picture list (list 1) according to the syntax elements parsed from the bitstream. In addition, if a PU is encoded using inter prediction, entropy decoding unit 310 may parse the motion information of the PU. Inter prediction unit 321 may determine one or more reference blocks of the PU according to the motion information of the PU. Inter prediction unit 321 may generate a predicted block of the PU according to one or more reference blocks of the PU.
[0093] Inverse quantization / transformation unit 330 may inverse quantize (i.e., dequantize) the transform coefficients associated with a TU. Inverse quantization / transformation unit 330 may use the QP value associated with the CU of the TU to determine the degree of quantization.
[0094] After inverse quantizing the transform coefficients, inverse quantization / transformation unit 330 may apply one or more inverse transforms to the inverse quantized transform coefficients to generate a residual block associated with the TU.
[0095] Reconstruction unit 340 uses the residual block associated with the TU of the CU and the predicted block of the PU of the CU to reconstruct the pixel block of the CU. For example, reconstruction unit 340 may add the samples of the residual block to the corresponding samples of the predicted block to reconstruct the pixel block of the CU, obtaining a reconstructed picture block.
[0096] The loop filter unit 350 can perform deblocking filtering operations to reduce the blocking artifacts of the pixel blocks associated with the CU.
[0097] The video decoder 300 can store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 can use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
[0098] The basic process of video coding and decoding is as follows: At the encoding end, a frame of image is divided into blocks. For the current block, the prediction unit 210 generates a predicted block of the current block using intra-frame prediction or inter-frame prediction. The residual unit 220 can calculate a residual block based on the predicted block and the original block of the current block, that is, the difference between the predicted block and the original block of the current block. This residual block can also be referred to as residual information. The residual block can go through processes such as transformation and quantization by the transform / quantization unit 230 to remove information that is insensitive to the human eye, so as to eliminate visual redundancy. Optionally, the residual block before being transformed and quantized by the transform / quantization unit 230 can be referred to as a temporal residual block, and the temporal residual block after being transformed and quantized by the transform / quantization unit 230 can be referred to as a frequency residual block or a frequency-domain residual block. The entropy coding unit 280 receives the quantized transform coefficients output by the transform quantization unit 230, and can perform entropy coding on the quantized transform coefficients to output a bitstream. For example, the entropy coding unit 280 can eliminate character redundancy according to the target context model and the probability information of the binary bitstream.
[0099] At the decoding end, the entropy decoding unit 310 can parse the bitstream to obtain prediction information, a quantization coefficient matrix, etc. of the current block. The prediction unit 320 generates a predicted block of the current block using intra-frame prediction or inter-frame prediction based on the prediction information. The inverse quantization / transformation unit 330 uses the quantization coefficient matrix obtained from the bitstream to perform inverse quantization and inverse transformation on the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the predicted block and the residual block to obtain a reconstructed block. The reconstructed blocks form a reconstructed image. The loop filter unit 350 performs loop filtering on the reconstructed image based on the image or based on the block to obtain a decoded image. The encoding end also needs to perform operations similar to those of the decoding end to obtain a decoded image. This decoded image can also be referred to as a reconstructed image, and the reconstructed image can be used as a reference frame for inter-frame prediction for subsequent frames.
[0100] It should be noted that the block partitioning information determined at the encoding end, as well as mode information or parameter information such as prediction, transformation, quantization, entropy coding, and loop filtering, etc. are carried in the bitstream when necessary. The decoding end determines the same block partitioning information, prediction, transformation, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoding end by parsing the bitstream and analyzing based on the existing information, so as to ensure that the decoded image obtained at the encoding end is the same as the decoded image obtained at the decoding end.
[0101] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. The present application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.
[0102] In the existing hybrid coding framework, each frame of a video is often divided into units of a certain size before the subsequent encoding and decoding process. Figure 4 As shown in the figure, the largest coding unit (CTU) is the basic coding unit in the hybrid coding framework, which usually contains two parts: luminance Y and chrominance UV. Since the U component and V component in chrominance have similar characteristics, they are usually processed in the order of U first and then V using the same coding parameters, and the corresponding coding results of U and V are obtained.
[0103] The existing hybrid coding framework uses traditional loop filters to suppress the distortion of reconstructed images, improve the quality of reconstructed images, and expects to restore the encoded reconstructed images to the original images. However, traditional loop filters are based on manual design and are difficult to effectively reduce the distortion of reconstructed images, leaving a large room for optimization. Due to the excellent performance of deep learning tools in image processing, deep learning-based loop filters are applied to the loop filter module.
[0104] The main technology involved in this application is the filter based on neural network loop filter (NNLF). Figure 5 As shown, by inputting the image to be filtered before filtering into the trained filter, the enhanced image after filtering can be obtained.
[0105] During the training process, neural networks usually use loss functions to constrain the filtered image so that it can be restored to the original image as much as possible. The loss function measures the difference between the filtered value and the true value. The larger the loss value, the greater the difference, and the goal of training is to reduce the loss. For deep learning-based encoding tools, exemplary and commonly used loss functions are: L1 norm loss function, L2 norm loss function and smooth L1 loss function.
[0106] In the training process of the neural network filter, the UV component of the chroma is often input in units of the maximum coding unit, and the filtering results of the U component and the V component are output accordingly. Then, the loss function is used to constrain the filtering results of the U component and the V component respectively, prompting them to be restored to the original image. After the training process is completed, the parameters of the neural network filter are fixed. In order to maintain the consistency of training and testing, the testing process of the neural network filter usually adopts the same chroma filtering order as the training process. For example, Figure 6AAs shown, during the training process of the neural network filter, first input the U component of the maximum coding unit into the neural network filter, and then input the V component of the maximum coding unit into the neural network filter to obtain the filtered value of the U component of the maximum coding unit and the filtered value of the V component of the maximum coding unit. Then, based on the filtered value of the U component of the maximum coding unit and the original value of the U component of the maximum coding unit, determine the loss of the U component, and based on the filtered value of the V component of the maximum coding unit and the original value of the V component of the maximum coding unit, determine the loss of the V component. Finally, based on the loss of the U component and the loss of the V component, adjust the parameters of the neural network filter to achieve the training of the neural network filter. Correspondingly, as Figure 6B shown, during the testing process of the neural network filter, also first input the U component of the maximum coding unit into the neural network filter, and then input the V component of the maximum coding unit into the neural network filter for filtering to obtain the filtered value of the U component of the maximum coding unit and the filtered value of the V component of the maximum coding unit.
[0107] However, since the U component and the V component in chrominance have relatively similar own characteristics, there is a certain similarity in the learning of the neural network filter for chrominance component filtering. Strictly specifying that the chrominance filtering order during the testing process is consistent with the training process may limit the generalization of the neural network filter, and there is room for optimization.
[0108] To solve the above technical problems, in the embodiment of the present application when decoding the current image, first determine the reconstructed image of the current image. For the current image block to be filtered in the reconstructed image, determine the target filtering order of the first chrominance component and the second chrominance component of the current image block, where the target filtering order is determined based on the filtering costs of N filtering orders, and N is a positive integer greater than 1; then, based on the target filtering order, input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering to obtain the filtered image block of the current image block. That is to say, in the embodiment of the present application, the target filtering order of the first chrominance component and the second chrominance component of the current image block input into the neural network filter is determined based on the filtering costs of N filtering orders, rather than defaulting to the training order. This can improve the selection accuracy of the target filtering order. When inputting the first chrominance component and the second chrominance component of the current image block into the neural network filter based on the accurately determined target filtering order for filtering, the filtering effect can be improved, thereby enhancing the generalization of the neural network filter and improving the decoding performance.
[0109] The technical solutions of the embodiments of the present application will be described in detail through some embodiments below. These several embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0110] First, take the decoding end as an example to introduce the image filtering method provided by the embodiments of the present application.
[0111] Figure 7 It is a schematic flowchart of the image filtering method provided by an embodiment of the present application. The embodiments of the present application are applied to Figure 1 or Figure 3 the decoder or decoding device shown in. As Figure 7 shown, the method of the embodiments of the present application includes:
[0112] S101. Decode the bitstream of the current image to obtain the residual value of the current image, and determine the reconstructed image of the current image based on the residual value.
[0113] In the embodiments of the present application, when the encoding end encodes the current image, the current image is divided into encoding blocks, and each block is encoded one by one with the encoding block as the encoding unit. For example, for the current block to be encoded in the current image, first obtain the predicted value of the current block through inter-frame and / or intra-frame prediction methods. Then, based on the predicted value and the current block of the current block, obtain the residual value of the current block. The encoding end transforms the residual value of the current block to obtain transform coefficients. In one example, the encoding end does not quantize the transform coefficients of the current block, but directly encodes the transform coefficients to obtain a bitstream. In another example, the encoding end quantizes the transform coefficients of the current block to obtain quantization coefficients, and then encodes the quantization coefficients to obtain a bitstream.
[0114] During the encoding process, as Figure 2 shown, the encoding end also inverse-transforms the transform coefficients to obtain the residual value, and adds the residual value and the predicted value to obtain the reconstructed value of the current block. Based on the above steps, the reconstructed values of each encoding block in the current image can be obtained, and these reconstructed values constitute the reconstructed image of the current image. Then, in order to further improve the quality of the reconstructed image, the reconstructed image is filtered to obtain the decoded image of the current image. In one example, the decoded image can be stored in the decoding buffer for prediction of subsequent images.
[0115] As Figure 3As shown, for each block to be decoded in the current image, such as the current block, after the decoding end obtains the bitstream, it decodes the bitstream to obtain the transform coefficients of the current block. In one example, if the encoding end quantizes the transform coefficients and then encodes them, the decoding end decodes the bitstream to obtain the quantization coefficients of the current block, and then inverse-quantizes the quantization coefficients to obtain the transform coefficients of the current image. Next, the decoding end performs an inverse transform on the transform coefficients of the current block to obtain the residual value of the current block. At the same time, the decoding end uses an inter-frame and / or intra-frame prediction method to predict the predicted value of the current block. In this way, the predicted value and the residual value of the current block are added together to obtain the reconstructed value of the current block. Based on the above steps, the decoding end can decode and determine the reconstructed values of each block to be decoded in the current image, and these reconstructed values constitute the reconstructed image of the current image. Next, in order to further improve the quality of the reconstructed image, the decoding end filters the reconstructed image to obtain the decoded image of the current image. In one example, the decoding end can store the decoded image in a decoding buffer for prediction of subsequent images. In one example, the decoding end can output the decoded image to a display device for display.
[0116] In some embodiments, the image filtering method proposed in the embodiments of the present application can be used to filter at least one frame of image in a video. That is, the above current image is an image in a video.
[0117] In some embodiments, the image filtering method proposed in the embodiments of the present application can be used to decode a single image. That is, the above current image is a single image, such as an image generated by an electronic device.
[0118] After the decoding end obtains the reconstructed image of the current image based on the above steps, it performs the following steps of S102.
[0119] S102. For the current image block to be filtered in the reconstructed image, determine the target filtering order of the first chrominance component and the second chrominance component of the current image block.
[0120] Among them, the target filtering order is determined by decoding the bitstream or based on the filtering cost of N filtering orders, and N is a positive integer greater than 1.
[0121] In the embodiments of the present application, in order to improve the quality of the reconstructed image, the reconstructed image is filtered. Specifically, a neural network filter is used to filter the reconstructed image. When filtering the reconstructed image, the reconstructed image is divided into at least one image block, and each image block is filtered separately. Among them, the process of the decoding end using the neural network filter to filter each image block in the reconstructed image is basically the same. For the convenience of description, here, the filtering of the current image block in the reconstructed image is taken as an example for illustration.
[0122] It should be noted that the embodiments of the present application do not limit the size and shape of the above-mentioned current image block.
[0123] In a possible implementation manner, the above-mentioned current image block to be filtered is at least one CTU of the reconstructed image. That is, at least one CTU of the reconstructed image is divided into an image block and input into the neural network filter for filtering.
[0124] In some examples, such as Figure 8A shown, the current image block is one CTU of the reconstructed image, that is, one CTU of the reconstructed image is used as the input image block of a neural network filter.
[0125] In another example, such as Figure 8B shown, the current image block is 4 CTUs of the reconstructed image, that is, 4 CTUs of the reconstructed image are used as the input image block of a neural network filter.
[0126] In one example, multiple CTUs such as 2 CTUs or 3 CTUs of the reconstructed image can also be used as the input image block of a neural network filter. These multiple CTUs can be multiple CTUs in the horizontal direction, or multiple CTUs in the vertical direction. Optionally, these multiple CTUs can be adjacent, or not adjacent, or partially adjacent and partially not adjacent.
[0127] In another possible implementation manner, the above-mentioned current image block to be filtered is a preset image region of the reconstructed image. That is, a preset image region of the reconstructed image is used as the input image block of a neural network filter.
[0128] The embodiments of the present application do not limit the specific shape and size of this preset image region.
[0129] In one example, such as Figure 8C shown, the preset image region includes at least one CTU with a defective quantity of the reconstructed image, that is, at least one CTU with a defective quantity of the reconstructed image is used as the input image block of the neural network filter.
[0130] In some embodiments, the above-mentioned preset image region is a fixed region. For example, during each filtering, according to this preset image region, the image block to be currently filtered in the reconstructed image is obtained, and this image block is used as the input image block of the neural network filter. At this time, the size and shape of the image block input into the neural network filter each time are the same, both being the preset image region.
[0131] In some embodiments, the above-mentioned preset image region is a variation value. For example, during the first filtering, according to the first preset image region, an image block to be filtered in the reconstructed image is obtained and used as an input image block to be input into the neural network filter for filtering. During the second filtering, according to the second preset image region, an image block to be filtered in the reconstructed image is obtained and used as an input image block to be input into the neural network filter for filtering, and so on. In an example of this embodiment, the decoding end can divide the reconstructed image into several image blocks to be filtered, and the shapes and sizes of these several image blocks to be filtered can be the same, different, or partially the same and partially different.
[0132] After the decoding end determines the current image block to be filtered in the reconstructed image based on the above steps, the neural network filter is used to filter the current image block.
[0133] The current image block includes a luminance component and a chrominance component, and the chrominance component includes a U component and a V component. Since the characteristics of the U component and the V component in chrominance are relatively close, usually the same neural network filter is used for filtering.
[0134] As can be seen from the above, currently when filtering the chrominance component, the filtering order of the chrominance component is fixed and is usually the same as the input order of the chrominance component during the training of the neural network filter. For example, during training, the U component and the V component are input into the neural network filter in the order of the U component first and the V component second to train the neural network filter. In this way, during the actual filtering process, the U component and the V component are also input into the neural network filter in the filtering order of the U component first and the V component second for filtering. However, when filtering by keeping the filtering order of the chrominance component consistent with the training order in this way, it will lead to poor filtering effect and reduce the generalization of the neural network filter.
[0135] To solve this technical problem, in the embodiment of the present application, when filtering the chrominance component of the current image block, the target filtering order of the first chrominance component and the second chrominance component of the current image block is first determined. The target filtering order is determined based on the filtering cost of N filtering orders, which can improve the accuracy of determining the target filtering order. In this way, when filtering the chrominance component of the current image block based on the accurately determined target filtering order, the filtering effect of the chrominance component can be effectively improved, and thus the generalization of the neural network filter can be enhanced.
[0136] The filtering cost in the embodiments of the present application includes at least one of the computational cost and the distortion cost. That is to say, in some embodiments, the filtering cost of the filtering order in the embodiments of the present application includes the computational cost of the filtering order. For example, when the computational time and / or computational complexity are higher, it means that the computational cost is greater. In some embodiments, the filtering cost of the filtering order in the embodiments of the present application includes the distortion cost of the filtering order. For example, when the distortion degree is higher, it means that the distortion cost is greater. In some embodiments, the filtering cost of the filtering order in the embodiments of the present application includes the computational cost and the distortion cost of the filtering order. For example, when the sum of the distortion cost and the computational cost is higher, it means that the filtering cost is greater.
[0137] The following introduces the specific process of the decoding end determining the target filtering order of the first chrominance component and the second chrominance component of the current image block.
[0138] In the embodiments of the present application, the specific ways for the decoding end to determine the target filtering order include but are not limited to the following several types:
[0139] Method 1: The decoding end obtains the target filtering order by decoding the bitstream. At this time, determining the target filtering order of the first chrominance component and the second chrominance component of the current image block in S102 above includes the following steps of S102-A1 and S102-A2:
[0140] S102-A1: Decode the bitstream to obtain first information, where the first information is used to indicate the target filtering order;
[0141] S102-A2: Based on the first information, obtain the target filtering order.
[0142] In this method 1, the encoding end determines the target filtering order of the first chrominance component and the second chrominance component of the current image block, and then writes the first information in the bitstream, and indicates the target filtering order through the first information. In this way, the decoding end decodes the bitstream to obtain the first information, and then based on the first information, obtains the target filtering order.
[0143] The embodiments of the present application do not limit the specific manifestation form of the first information, as long as it is any syntax field that can indicate the target filtering order.
[0144] In some embodiments, the above first information includes a first flag, and the target filtering order is indicated by different values of the first flag.
[0145] Exemplarily, the corresponding relationship between the values of the first flag and the filtering order of the chrominance component is shown in Table 1:
[0146] Table 1
[0147] Value of the first flag Filtering order A1 U before V A2 V before U …… ……
[0148] The embodiments of the present application do not limit the specific values of A1, A2, etc. For example, A1 is equal to 0, A2 is equal to 1, or A1 is equal to 1, A2 is equal to 0, etc.
[0149] The embodiments of the present application do not limit the specific type of the filtering order of the chrominance components. For example, in addition to the two filtering orders of U first and V second, and V first and U second shown in Table 1 above, it may at least include the following several types:
[0150] Example 1, the filtering order of the chrominance components includes: dividing the U component into multiple sub-U components, and these multiple sub-U components and the V component form multiple filtering orders.
[0151] For example, divide the U component into a first sub-U component and a second sub-U component. In this way, the filtering orders formed by the first sub-U component, the second sub-U component, and the V component include: first the first sub-U component, then the V component, and then the second sub-U component; first the second sub-U component, then the V component, and then the first sub-U component; first the second sub-U component, then the first sub-U component, and then the V component; first the V component, then the second sub-U component, and then the first sub-U component, etc.
[0152] Example 2, the filtering order of the chrominance components includes: dividing the V component into multiple sub-V components, and these multiple sub-V components and the U component form multiple filtering orders.
[0153] For example, divide the V component into a first sub-V component and a second sub-V component. In this way, the filtering orders formed by the first sub-V component, the second sub-V component, and the U component include: first the first sub-V component, then the U component, and then the second sub-V component; first the second sub-V component, then the U component, and then the first sub-V component; first the second sub-V component, then the first sub-V component, and then the U component; first the U component, then the second sub-V component, and then the first sub-V component, etc.
[0154] Example 3, the filtering order of the chrominance components includes: dividing the U component into multiple sub-U components and dividing the V component into multiple sub-V components, and these multiple sub-U components and multiple sub-V components form multiple filtering orders.
[0155] For example, divide the U component into a first sub-U component and a second sub-U component, and divide the V component into a first sub-V component and a second sub-V component. In this way, the filtering orders formed by the first sub-U component, the second sub-U component, the first sub-V component, and the second sub-V component include: first the first sub-U component, then the first sub-V component, then the second sub-U component, and then the second sub-V component; first the first sub-V component, then the first sub-U component, then the second sub-V component, and then the second sub-U component; first the second sub-V component, then the first sub-U component, then the first sub-V component, and then the second sub-U component, etc.
[0156] In the first method, the encoding end can determine the value of the first flag corresponding to the target filtering order based on Table 1 above. After that, the first flag is set to this value and then written into the bitstream. In this way, the decoding end decodes the bitstream to obtain the first flag, and then according to the value of the first flag, by querying Table 1 above, the target filtering orders of the first chrominance component and the second chrominance component of the current image block are obtained.
[0157] In one example, assume that the first chrominance component is the U component and the second chrominance component is the V component. If the encoding end determines that the target filtering orders of the first chrominance component and the second chrominance component of the current image block are that the first chrominance component is before the second chrominance component, based on Table 1 above, it can be determined that the value of the first flag is A1. After that, the first flag is set to A1 and then written into the bitstream. In this way, the decoding end decodes the bitstream to obtain the first flag, and then based on the value of the first flag, by querying Table 1 above, it is determined that the target filtering orders of the first chrominance component and the second chrominance component of the current image block are that the first chrominance component is before the second chrominance component.
[0158] In another example, assume that the first chrominance component is the U component and the second chrominance component is the V component. If the encoding end determines that the target filtering orders of the first chrominance component and the second chrominance component input into the neural network filter of the current image block are that the second chrominance component is before the first chrominance component, based on Table 1 above, it can be determined that the value of the first flag is A2. After that, the first flag is set to A2 and then written into the bitstream. In this way, the decoding end decodes the bitstream to obtain the first flag, and then based on the value of the first flag, by querying Table 1 above, it is determined that the target filtering orders of the first chrominance component and the second chrominance component of the current image block are that the second chrominance component is before the first chrominance component.
[0159] In some embodiments, the above filtering cost includes a first filtering cost, and the above target filtering order is determined based on the first filtering cost of each of the N filtering orders, where the first filtering cost of a filtering order is the filtering cost determined when the first chrominance component and the second chrominance component of the current image block are input into the neural network filter according to this filtering order.
[0160] That is to say, in this embodiment, for each of the N filtering orders at the encoding end, according to this filtering order, the first chrominance component and the second chrominance component of the current image block are input into the neural network filter for filtering, and the filtering cost corresponding to this filtering order is determined, and this filtering cost is recorded as the first filtering cost. The specific types of the above N filtering orders are not limited in the embodiments of the present application. For example, the N filtering orders include some or all of the various filtering orders shown in Table 1 above.
[0161] Exemplarily, for the j-th filtering order among N filtering orders, the encoding end inputs the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order, and obtains the filtered value of the first chrominance component of the current image block and the filtered value of the second chrominance component. For the convenience of description, the filtered value of the chrominance component of the current image block is denoted as the j-th filtered value. At this time, the j-th filtered value includes the filtered value of the first chrominance component of the current image block and the filtered value of the second chrominance component of the current image block under the j-th filtering order.
[0162] Next, the encoding end determines the first filtering cost corresponding to the j-th filtering order based on the j-th filtered value of the current image block and the original image block of the current image block. For example, based on the filtered value of the first chrominance component of the current image block and the first chrominance component of the original image block of the current image block, and the filtered value of the second chrominance component of the current image block and the second chrominance component of the original image block of the current image block under the j-th filtering order, the first filtering cost corresponding to the j-th filtering order is determined.
[0163] The embodiments of the present application do not limit the specific calculation method of the above first filtering cost. For example, the above first filtering cost may be a rate-distortion cost (RDO), or may also be an approximate cost, such as SSD, STAD or SAD, etc.
[0164] Based on the above steps, the encoding end can determine the first filtering cost corresponding to each of the N filtering orders, and then determine the target filtering order from the N filtering orders based on the first filtering cost corresponding to each filtering order.
[0165] In some embodiments, the above target filtering order is the filtering order with the smallest first filtering cost among the N filtering orders. That is to say, the encoding end determines the filtering order with the smallest first filtering cost among the N filtering orders as the target filtering order of the chrominance component of the current image block, and then indicates the target filtering order to the decoding end.
[0166] As can be seen from the above, in the first method, the decoding end can quickly obtain the target filtering order of the first chrominance component and the second chrominance component of the current image block by decoding the bitstream, and thus can improve the filtering speed of the image.
[0167] In addition to using the method described in the first method above to determine the target filtering order, the decoding end can also use the following method of the second method to determine the target filtering order.
[0168] In the second method, the decoding end determines the target filtering order by itself. At this time, determining the target filtering orders of the first chrominance component and the second chrominance component of the current image block in S102 above includes the following steps from S102-B1 to S102-B3:
[0169] S102-B1. Determine the already-filtered region around the current image block;
[0170] S102-B2. For the i-th filtering order among the N filtering orders, input the first chrominance component and the second chrominance component of the already-filtered region around into the neural network filter for filtering according to the i-th filtering order, and determine the i-th second filtering cost of the already-filtered region under the i-th filtering order, where i is a positive integer less than or equal to N;
[0171] S102-B3. Based on the second filtering costs respectively corresponding to the N filtering orders, determine the target filtering order from the N filtering orders.
[0172] In this second method, the decoding end determines the target filtering order from the N filtering orders based on the already-filtered region around the current image block.
[0173] The embodiments of the present application do not limit the size and shape of the already-filtered region around the current image block.
[0174] In some embodiments, the already-filtered region around the current image block is the already-filtered region adjacent to the current image block around the current image block.
[0175] In some embodiments, as Figure 9 shown, the already-filtered region around the current image block includes the already-filtered region above the current image block and the already-filtered region on the left side.
[0176] In some embodiments, the already-filtered region around the current image block includes the template region of the current image.
[0177] After the decoding end determines the already-filtered region around the current image block in the current image, it executes the steps of S102-B2 above. According to each of the N filtering orders, input the first chrominance component and the second chrominance component of the already-filtered region around into the neural network filter for filtering, determine the filtering cost corresponding to each of the N filtering orders, and record this filtering cost as the second filtering cost. In this embodiment, the specific process of the decoding end determining the second filtering cost corresponding to each of the N filtering orders is basically the same. Taking the i-th filtering order among the N filtering orders as an example for illustration. That is, according to the i-th filtering order, input the first chrominance component and the second chrominance component of the already-filtered region around into the neural network filter for filtering, and determine the i-th second filtering cost of the already-filtered region under the i-th filtering order.
[0178] The embodiments of the present application do not limit the specific manner of determining the i-th second filtering cost of the surrounding filtered region in the i-th filtering order.
[0179] In a possible implementation manner, the decoding end inputs the first chrominance component and the second chrominance component of the surrounding filtered region into a neural network filter for filtering according to the i-th filtering order, to obtain the filtered value of the first chrominance component of the surrounding filtered region in the i-th filtering order, and the filtered value of the second chrominance component in the i-th filtering order. Then, based on the filtered value of the first chrominance component of the surrounding filtered region in the i-th filtering order and the filtered value of the second chrominance component in the i-th filtering order, the second filtering cost corresponding to the i-th filtering order is determined. For example, if the second filtering cost includes a calculation cost, the decoding end determines the calculation cost when filtering the first chrominance component and the second chrominance component of the surrounding filtered region through the neural network filter in the i-th filtering order, and then determines the second filtering cost corresponding to the i-th filtering order based on this calculation cost.
[0180] In a possible implementation manner, the above S102-B2 includes the following steps S102-B21 and S102-B22:
[0181] S102-B21: Input the first chrominance component and the second chrominance component of the surrounding filtered region into a neural network filter for filtering according to the i-th filtering order, to obtain the i-th filtered value of the surrounding filtered region;
[0182] S102-B22: Determine the i-th second filtering cost based on the i-th filtered value and the surrounding filtered region
[0183] In this implementation manner, the decoding end inputs the first chrominance component and the second chrominance component of the surrounding filtered region into a neural network filter for filtering according to the i-th filtering order, to obtain the filtered value of the first chrominance component of the surrounding filtered region in the i-th filtering order and the filtered value of the second chrominance component in the i-th filtering order. For the convenience of description, the filtered value of the first chrominance component of the surrounding filtered region in the i-th filtering order and the filtered value of the second chrominance component in the i-th filtering order are denoted as the i-th filtered value of the surrounding filtered region.
[0184] Then, based on the i-th filtered value and the surrounding filtered region, the second filtering cost corresponding to the i-th filtering order is determined.
[0185] For example, if the second filtering cost includes a distortion cost, the decoding end determines the second filtering cost corresponding to the i-th filtering order based on the filtered value and the unfiltered value of the first chrominance component in the already filtered surrounding area, and the filtered value and the unfiltered value of the second chrominance component in the already filtered surrounding area under the i-th filtering order.
[0186] The embodiments of the present application do not limit the specific calculation method of the above-mentioned second filtering cost. For example, the above-mentioned second filtering cost can be a rate-distortion cost (RDO), or an approximate cost, such as SSD, STAD, or SAD, etc.
[0187] For another example, if the second filtering cost includes a calculation cost and a distortion cost, the decoding end determines the distortion cost corresponding to the i-th filtering order based on the filtered value and the unfiltered value of the first chrominance component in the already filtered surrounding area, and the filtered value and the unfiltered value of the second chrominance component in the already filtered surrounding area under the i-th filtering order. Meanwhile, the calculation cost when filtering the first chrominance component and the second chrominance component in the already filtered surrounding area through the neural network filter is determined under the i-th filtering order. In this way, the second filtering cost corresponding to the i-th filtering order is determined according to the distortion cost and the calculation cost corresponding to the i-th filtering order. For example, the sum or weighted sum of the distortion cost and the calculation cost corresponding to the i-th filtering order is determined as the second filtering cost corresponding to the i-th filtering order.
[0188] Based on the above steps, the decoding end can determine the second filtering cost of the already filtered surrounding area of the current image block under each of the N filtering orders. Then, based on the second filtering costs corresponding to the N filtering orders respectively, the target filtering order is determined from the N filtering orders.
[0189] The embodiments of the present application do not limit the specific method for the decoding end to determine the target filtering order from the N filtering orders based on the second filtering costs corresponding to the N filtering orders respectively.
[0190] For example, the filtering order with the minimum second filtering cost among the N filtering orders is determined as the target filtering order.
[0191] For another example, any filtering order with a second filtering cost less than a preset value among the N filtering orders is determined as the target filtering order.
[0192] The embodiments of the present application do not limit the specific types of the above-mentioned N filtering orders.
[0193] In one example, the N filtering orders are shown in Table 2:
[0194] Table 2
[0195]
[0196]
[0197] In some embodiments, the N filtering orders include a first filtering order and a second filtering order. In the first filtering order, the first chrominance component in the input of the neural network filter comes first and the second chrominance component comes second. That is, when the first chrominance component and the second chrominance component are input into the neural network filter, the first chrominance component is before the second chrominance component. In the second filtering order, the second chrominance component in the input of the neural network filter comes first and the first chrominance component comes second. That is, when the first chrominance component and the second chrominance component are input into the neural network filter, the second chrominance component is before the first chrominance component.
[0198] At this time, as Figure 10 shown, the decoding end inputs the first chrominance component (e.g., U component) and the second chrominance component (e.g., V component) of the surrounding filtered region into the neural network filter for filtering according to the first filtering order, and determines the second filtering cost 1 corresponding to the first filtering order. At the same time, the decoding end inputs the first chrominance component and the second chrominance component of the surrounding filtered region into the neural network filter for filtering according to the second filtering order, and determines the second filtering cost 2 corresponding to the second filtering order. Finally, according to the second filtering cost 1 corresponding to the first filtering order and the second filtering cost 2 corresponding to the second filtering order, the target filtering order is selected from the first filtering order and the second filtering order. For example, the filtering order with the minimum second filtering cost among the first filtering order and the second filtering order is determined as the target filtering order.
[0199] As can be seen from the above, in the embodiment of the present application, when determining the target filtering order of the chrominance components of the current image block, the training order of the chrominance components of the neural network filter is not considered. That is to say, regardless of the training order of the chrominance components of the neural network filter, the decoding end determines the target filtering order of the chrominance components of the current chrominance block according to the above steps.
[0200] The embodiment of the present application does not limit the training method and related training parameters of the neural network filter.
[0201] In some embodiments, the above neural network filter is trained with at least one CTU as a training unit. That is to say, at least one CTU of the training image is divided into a training unit and input into the neural network filter to train the neural network filter.
[0202] In some embodiments, the neural network filter is trained with a preset image region as a training unit. That is, the preset image region of the training image is divided into a training unit and input into the neural network filter for training. For the relevant description of the above preset image region, reference may be made to the relevant description of the above preset image region, which will not be elaborated here.
[0203] In some embodiments, the training order of the chrominance components of the neural network filter is any one of N training orders. Optionally, the above N training orders may be the same as, different from, or partially the same and partially different from the above N filtering orders.
[0204] In some embodiments, the above N training orders include a first training order and a second training order. In the first training order, the first chrominance component is in the front and the second chrominance component is in the back in the input of the neural network filter. In the second training order, the second chrominance component is in the front and the first chrominance component is in the back in the input of the neural network filter.
[0205] In one example, during the training process, the chrominance training order of the neural network filter is the first training order. For example, the U component and the V component are input successively, and the loss function is used to constrain the filtered results of the U component and the V component to realize the training of the neural network filter. When using the trained neural network filter to filter the first chrominance component and the second chrominance component of the current image block, the target filtering order of the first chrominance component and the second chrominance component is determined. Specifically, as Figure 11A shown, the decoding end inputs the U component and the V component of the filtered region around the current image block into the neural network filter NNLF for filtering according to the first filtering order, for example, the filtering order with the U component in the front and the V component in the back, and outputs the filtered value of the U component and the filtered value of the V component of the filtered region around. Then, based on the filtered value of the U component and the filtered value of the V component of the filtered region around under the first filtering order, and the filtered region around, the second filtering cost 1 corresponding to the first filtering order is determined. Similarly, as Figure 11A shown, the decoding end inputs the U component and the V component of the filtered region around the current image block into the neural network filter NNLF for filtering according to the second filtering order, for example, the filtering order with the V component in the front and the U component in the back, and outputs the filtered value of the U component and the filtered value of the V component of the filtered region around. Then, based on the filtered value of the U component and the filtered value of the V component of the filtered region around under the second filtering order, and the filtered region around, the second filtering cost 2 corresponding to the second filtering order is determined. Finally, according to the second filtering costs corresponding to the first filtering order and the second filtering order respectively, the target filtering order of the chrominance components of the current image block is determined from the first filtering order and the second filtering order.
[0206] In another example, during the training process, the chrominance training order of the neural network filter is the first training order. For example, the V component and the U component are input successively, and the loss function is used to constrain the filtered results of the V component and the U component respectively to implement the training of the neural network filter. When using the trained neural network filter to filter the first chrominance component and the second chrominance component of the current image block, the target filtering order of the first chrominance component and the second chrominance component is determined. Specifically, as Figure 11B shown, at the decoding end, according to the first filtering order, for example, the filtering order with the U component first and the V component second, the U component and the V component of the already filtered region around the current image block are input into the neural network filter NNLF for filtering, and the filtered value of the U component and the filtered value of the V component of the already filtered region around are output. Then, based on the filtered value of the U component and the filtered value of the V component of the already filtered region around under the first filtering order, and the already filtered region around, the second filtering cost 1 corresponding to the first filtering order is determined. Similarly, as Figure 11B shown, at the decoding end, according to the second filtering order, for example, the filtering order with the V component first and the U component second, the U component and the V component of the already filtered region around the current image block are input into the neural network filter NNLF for filtering, and the filtered value of the U component and the filtered value of the V component of the already filtered region around are output. Then, based on the filtered value of the U component and the filtered value of the V component of the already filtered region around under the second filtering order, and the already filtered region around, the second filtering cost 2 corresponding to the second filtering order is determined. Finally, according to the second filtering costs corresponding to the first filtering order and the second filtering order respectively, the target filtering order of the chrominance components of the current image block is determined from the first filtering order and the second filtering order.
[0207] As can be seen from the above, the determination process of the target filtering order in the embodiments of the present application is independent of the chrominance training order of the neural network filter, thereby improving the selection flexibility and accuracy of the target filtering training.
[0208] It should be noted that the above takes the N filtering orders including the first filtering order and the second filtering order as an example to introduce the process of determining the target filtering order. If the N filtering orders further include other filtering orders, for example, when including other filtering orders shown in Table 2 above, the same method as the above first filtering order and second filtering order is adopted to determine the second filtering cost of the already decoded region under other filtering orders, so that the second filtering cost corresponding to each of the N filtering orders can be determined, and then based on the second filtering cost corresponding to each filtering order, the target filtering order is determined from the N filtering orders.
[0209] After determining the target filtering order of the first chrominance component and the second chrominance component of the current image block based on the above steps, the decoding end performs the following steps of S103.
[0210] S103: Input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering based on the target filtering order, to obtain the chrominance filtering block of the current image block.
[0211] After determining the target filtering order of the first chrominance component and the second chrominance component of the current image block based on the above steps, according to the target filtering order, input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering, to obtain the filtered image block of the current image block.
[0212] In one example, as Figure 12A shown, if the above target filtering order is that the first chrominance component is before the second chrominance component, then splice the first chrominance component and the second chrominance component of the current image block, with the first chrominance component before the second chrominance component during splicing. Then, input the spliced first chrominance component and second chrominance component into the neural network filter for filtering, to obtain the filtering value of the first chrominance component and the filtering value of the second chrominance component of the current image block, where the filtering value of the first chrominance component and the filtering value of the second chrominance component of the current image block form the chrominance filtering block of the current image block.
[0213] In one example, as Figure 12B shown, if the above target filtering order is that the second chrominance component is before the first chrominance component, then splice the first chrominance component and the second chrominance component of the current image block, with the first chrominance component after the second chrominance component during splicing. Then, input the spliced first chrominance component and second chrominance component into the neural network filter for filtering, to obtain the filtering value of the first chrominance component and the filtering value of the second chrominance component of the current image block, where the filtering value of the first chrominance component and the filtering value of the second chrominance component of the current image block form the chrominance filtering block of the current image block.
[0214] The above introduces the filtering process of the current image block in the reconstructed image. The filtering process of other image blocks to be filtered in the reconstructed image can refer to the filtering process of the current image block above, and finally obtain the filtered reconstructed image.
[0215] In some embodiments, the above neural network filter is used as a loop filter. At this time, the output of the neural network filter affects video decoding. For example, at the decoding end, by the above method, the neural network filter is used to filter the reconstructed image of the current image to obtain a filtered reconstructed image, and the filtered reconstructed image is stored in the decoding cache as a decoded image for subsequent image filtering. Since the embodiments of the present application provide the accuracy of determining the filtering order, the filtering quality of the reconstructed image is improved. In this way, when subsequent decoding is performed based on the reconstructed image with better quality, the decoding effect of the video can be enhanced.
[0216] In some embodiments, the above neural network filter is used for post-processing, that is, to filter and optimize the decoded video. At this time, the output of the neural network filter does not affect video decoding. For example, the decoding end decodes the video stream to obtain the decoded video. Then, the neural network filter is used to filter at least one image in the decoded video. At this time, each image in the at least one image can be recorded as a reconstructed image. Then, by the above method, the decoding end uses the neural network filter to filter the reconstructed image to obtain a filtered reconstructed image. The filtered reconstructed image is not cached in the decoding cache for subsequent decoding, but is directly stored elsewhere or directly output for display. At this time, the image filtering method of the embodiments of the present application is used for post-processing of the decoded video.
[0217] In some embodiments, the image filtering method proposed by the embodiments of the present application, in addition to being applied in the field of video or image decoding, can also be used in conventional image filtering. For example, for the current image block in the image to be filtered, the target filtering order of the first chrominance component and the second chrominance component of the current image block is determined. The target filtering order is determined based on the filtering costs of N filtering orders, where N is a positive integer greater than 1; based on the target filtering order, the first chrominance component and the second chrominance component of the current image block are input into the neural network filter for filtering to obtain the chrominance filtering block of the current image block.
[0218] The image filtering method provided by the embodiments of the present application is as follows: when the decoding end decodes the current image, it first decodes the bitstream of the current image to obtain the residual value of the current image, and determines the reconstructed image of the current image based on the residual value; for the current image block to be filtered in the reconstructed image, it determines the target filtering order of the first chrominance component and the second chrominance component of the current image block, and the target filtering order is determined by decoding the bitstream or based on the filtering cost of N filtering orders, where N is a positive integer greater than 1; based on the target filtering order, it inputs the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering to obtain the chrominance filtering block of the current image block. That is to say, the embodiments of the present application determine the target filtering order based on the filtering cost of N filtering orders, improving the selection accuracy of the target filtering order. When filtering the first chrominance component and the second chrominance component of the current image block based on the accurately determined target filtering order by inputting them into the neural network filter, the filtering effect can be improved, thereby enhancing the generalization of the neural network filter and improving the decoding performance.
[0219] The above takes the decoding end as an example to introduce the image filtering method of the embodiments of the present application. Next, the encoding end is taken as an example to introduce the image filtering method of the embodiments of the present application.
[0220] Figure 13 is a schematic flowchart of the image filtering method provided by an embodiment of the present application, and the embodiments of the present application are applied to Figure 1 or Figure 2 the encoder shown. As Figure 13 shown, the method of the embodiments of the present application includes:
[0221] S201. Encode the current image to obtain the reconstructed image of the current image.
[0222] In the embodiments of the present application, when the encoding end encodes the current image, it divides the current image into coding blocks and performs block-by-block encoding with the coding blocks as the encoding units. For example, for the current block to be encoded in the current image, first, the predicted value of the current block is obtained through inter-frame and / or intra-frame prediction methods. Then, based on the predicted value and the current block of the current block, the residual value of the current block is obtained. The encoding end transforms the residual value of the current block to obtain transform coefficients. In one example, the encoding end does not quantize the transform coefficients of the current block and directly encodes the transform coefficients to obtain a bitstream. In another example, the encoding end quantizes the transform coefficients of the current block to obtain quantization coefficients, and then encodes the quantization coefficients to obtain a bitstream.
[0223] During the encoding process, as Figure 2As shown, the encoding end also performs an inverse transformation on the transform coefficients to obtain a residual value, and adds the residual value and the predicted value to obtain the reconstructed value of the current block. Based on the above steps, the reconstructed values of each encoding block in the current image can be obtained, and these reconstructed values constitute the reconstructed image of the current image. Then, in order to further improve the quality of the reconstructed image, the reconstructed image is filtered to obtain the decoded image of the current image. In one example, the decoded image can be stored in a decoding cache for prediction of subsequent images.
[0224] In some embodiments, the image filtering method proposed in the embodiments of the present application can be used to filter at least one frame of an image in a video. That is, the above-mentioned current image is an image in the video.
[0225] In some embodiments, the image filtering method proposed in the embodiments of the present application can be used to decode a single image. That is, the above-mentioned current image is a single image, for example, an image generated by an electronic device.
[0226] After obtaining the reconstructed image of the current image based on the above steps, the decoding end performs the following steps of S202.
[0227] S202: For the current image block to be filtered in the reconstructed image, determine the target filtering order of the first chrominance component and the second chrominance component of the current image block.
[0228] Wherein, the target filtering order is determined based on the filtering costs of N filtering orders, and N is a positive integer greater than 1.
[0229] In the embodiments of the present application, in order to improve the quality of the reconstructed image, the reconstructed image is filtered. Specifically, a neural network filter is used to filter the reconstructed image. When filtering the reconstructed image, the reconstructed image is divided into at least one image block, and each image block is filtered separately. Among them, the process of the decoding end using the neural network filter to filter each image block in the reconstructed image is basically the same. For ease of description, here, taking the filtering of the current image block in the reconstructed image as an example for illustration.
[0230] It should be noted that the embodiments of the present application do not limit the size and shape of the above-mentioned current image block.
[0231] In one possible implementation, the above-mentioned current image block to be filtered is at least one CTU of the reconstructed image. That is, at least one CTU of the reconstructed image is divided into an image block and input into the neural network filter for filtering.
[0232] In some examples, such as Figure 8AAs shown, the current image block is a CTU of the reconstructed image, that is, a CTU of the reconstructed image is used as the input image block of a neural network filter.
[0233] In another example, as Figure 8B shown, the current image block is 4 CTUs of the reconstructed image, that is, 4 CTUs of the reconstructed image are used as the input image block of a neural network filter.
[0234] In one example, multiple CTUs such as 2 CTUs or 3 CTUs of the reconstructed image can also be used as the input image block of a neural network filter. These multiple CTUs can be multiple CTUs in the horizontal direction or multiple CTUs in the vertical direction. Optionally, these multiple CTUs can be adjacent, or not adjacent, or partially adjacent and partially non - adjacent.
[0235] In another possible implementation, the current image block to be filtered above is a preset image area of the reconstructed image. That is, the encoding end uses a preset image area of the reconstructed image as the input image block of a neural network filter.
[0236] The embodiments of the present application do not limit the specific shape and size of this preset image area.
[0237] In one example, as Figure 8C shown, the preset image area includes at least one CTU with a defective quantity of the reconstructed image, that is, at least one CTU with a defective quantity of the reconstructed image is used as the input image block of the neural network filter.
[0238] In some embodiments, the above - mentioned preset image area is a fixed area. For example, each time filtering is performed, according to this preset image area, the image block to be currently filtered in the reconstructed image is obtained and used as the input image block of the neural network filter. At this time, the size and shape of the image block input to the neural network filter each time are the same, both being the preset image area.
[0239] In some embodiments, the above - mentioned preset image area is a variable value. For example, when filtering for the first time, according to the first preset image area, an image block to be filtered in the reconstructed image is obtained and used as the input image block to be input into the neural network filter for filtering. When filtering for the second time, according to the second preset image area, an image block to be filtered in the reconstructed image is obtained and used as the input image block to be input into the neural network filter for filtering, and so on. In one example of this embodiment, the decoding end can divide the reconstructed image into several image blocks to be filtered, and the shapes and sizes of these several image blocks to be filtered can be the same, or different, or partially the same and partially different.
[0240] After determining the current image block to be filtered in the reconstructed image based on the above steps, the encoding end filters the current image block using a neural network filter.
[0241] The current image block includes a luminance component and chrominance components. The chrominance components include a U component and a V component. Since the characteristics of the U component and the V component in chrominance are relatively close, usually the same neural network filter is used for filtering.
[0242] As can be seen from the above, currently when filtering chrominance components, the filtering order of the chrominance components is fixed and is usually the same as the input order of the chrominance components during the training of the neural network filter. For example, during training, the U component and the V component are input into the neural network filter in the order of the U component first and the V component second to train the neural network filter. In this way, during the actual filtering process, the U component and the V component are also input into the neural network filter in the filtering order of the U component first and the V component second for filtering. However, when filtering with the filtering order of the chrominance components being the same as the training order like this, it will lead to poor filtering effects and reduce the generalization of the neural network filter.
[0243] To solve this technical problem, in the embodiments of the present application, when filtering the chrominance components of the current image block, first determine the target filtering order of the first chrominance component and the second chrominance component of the current image block. The target filtering order is determined based on the filtering cost of N filtering orders, which can improve the accuracy of determining the target filtering order. In this way, based on the accurately determined target filtering order, when filtering the chrominance components of the current image block, the filtering effect of the chrominance components can be effectively improved, thereby enhancing the generalization of the neural network filter.
[0244] The filtering cost in the embodiments of the present application includes at least one of a computational cost and a distortion cost. That is to say, in some embodiments, the filtering cost of the filtering order in the embodiments of the present application includes the computational cost of the filtering order. For example, when the computational time and / or computational complexity is higher, it indicates that the computational cost is greater. In some embodiments, the filtering cost of the filtering order in the embodiments of the present application includes the distortion cost of the filtering order. For example, when the distortion degree is higher, it indicates that the distortion cost is greater. In some embodiments, the filtering cost of the filtering order in the embodiments of the present application includes the computational cost and the distortion cost of the filtering order. For example, when the sum of the distortion cost and the computational cost is greater, it indicates that the filtering cost is greater.
[0245] The following introduces the specific process of the encoding end determining the target filtering order of the first chrominance component and the second chrominance component of the current image block.
[0246] In the embodiments of the present application, the specific ways for the encoding end to determine the target filtering order include but are not limited to the following several:
[0247] In Method 1, the encoding end determines the target filtering order based on the current image block. At this time, the steps of determining the target filtering orders of the first chrominance component and the second chrominance component of the current image block in S202 above include the following steps of S202-A1 and S202-A2:
[0248] S202-A1. For the j-th filtering order among the N filtering orders, input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order, and determine the j-th first filtering cost of the current image block under the j-th filtering order, where j is a positive integer less than or equal to N;
[0249] S202-A2. Based on the first filtering costs respectively corresponding to the N filtering orders, determine the target filtering order from the N filtering orders.
[0250] In this Method 1, the encoding end inputs the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to each of the N filtering orders, determines the filtering cost corresponding to each of the N filtering orders, and records this filtering cost as the first filtering cost. In this embodiment, the specific processes of the encoding end determining the first filtering costs corresponding to each of the N filtering orders are basically the same. Taking the j-th filtering order among the N filtering orders as an example for illustration. That is, input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order, and determine the j-th first filtering cost of the current image block under the j-th filtering order.
[0251] The embodiments of the present application do not limit the specific manner of determining the j-th first filtering cost of the current image block under the j-th filtering order.
[0252] In a possible implementation manner, the encoding end inputs the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order, and obtains the filtered values of the first chrominance component of the current image block under the j-th filtering order and the filtered values of the second chrominance component under the j-th filtering order. Furthermore, based on the filtered values of the first chrominance component of the current image block under the j-th filtering order and the filtered values of the second chrominance component under the j-th filtering order, determine the first filtering cost corresponding to the j-th filtering order. For example, if the first filtering cost includes a calculation cost, the encoding end determines the calculation cost when filtering the first chrominance component and the second chrominance component of the current image block through the neural network filter in the j-th filtering order, and then determines the first filtering cost corresponding to the j-th filtering order based on this calculation cost.
[0253] In a possible implementation, the above S202-B2 includes the following steps S202-B21 and S202-B22:
[0254] S202-B21. Input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order, to obtain the j-th filtered image block of the current image block;
[0255] S202-B22. Determine the first filtering cost corresponding to the j-th filtering order based on the j-th filtered image block and the original image block of the current image block
[0256] In this implementation, at the encoding end, the first chrominance component and the second chrominance component of the current image block are input into the neural network filter for filtering according to the j-th filtering order, to obtain the filtered value of the first chrominance component of the current image block under the j-th filtering order, and the filtered value of the second chrominance component of the current image block under the j-th filtering order. For the sake of description, the filtered value of the first chrominance component of the current image block under the j-th filtering order, and the filtered value of the second chrominance component of the current image block under the j-th filtering order are denoted as the j-th filtered image block of the current image block.
[0257] Then, based on the j-th filtered image block and the original image block of the current image block, determine the first filtering cost corresponding to the j-th filtering order.
[0258] For example, if the first filtering cost includes a distortion cost, the encoding end determines the first filtering cost corresponding to the j-th filtering order based on the filtered value of the first chrominance component of the current image block and the first chrominance component of the original image block of the current image block, and the filtered value of the second chrominance component of the current image block and the second chrominance component of the original image block of the current image block under the j-th filtering order.
[0259] The embodiments of the present application do not limit the specific calculation method of the above first filtering cost. For example, the above first filtering cost may be a rate-distortion cost (RDO), or may also be an approximate cost, such as SSD, STAD or SAD, etc.
[0260] For another example, if the first filtering cost includes a computational cost and a distortion cost, then at the encoding end, based on the filtering order j, the filtered value of the first chrominance component of the current image block and the first chrominance component of the original image block of the current image block, as well as the filtered value of the second chrominance component of the current image block and the second chrominance component of the original image block of the current image block, the distortion cost corresponding to the j-th filtering order is determined. At the same time, the computational cost when filtering the first chrominance component and the second chrominance component of the current image block through the neural network filter in the j-th filtering order is determined. In this way, according to the distortion cost and the computational cost corresponding to the j-th filtering order, the first filtering cost corresponding to the j-th filtering order is determined. For example, the sum or weighted sum of the distortion cost and the computational cost corresponding to the j-th filtering order is determined as the first filtering cost corresponding to the j-th filtering order.
[0261] Based on the above steps, the encoding end can determine the first filtering cost of the current image block under each of the N filtering orders. Then, based on the first filtering costs corresponding to the N filtering orders respectively, the target filtering order is determined from the N filtering orders.
[0262] The embodiments of the present application do not limit the specific manner in which the encoding end determines the target filtering order from the N filtering orders based on the first filtering costs corresponding to the N filtering orders respectively.
[0263] For example, the filtering order with the smallest first filtering cost among the N filtering orders is determined as the target filtering order.
[0264] For another example, any filtering order among the N filtering orders with a first filtering cost less than a preset value is determined as the target filtering order.
[0265] The embodiments of the present application do not limit the specific types of the above N filtering orders.
[0266] In one example, the N filtering orders are shown in Table 2.
[0267] In some embodiments, the N filtering orders include a first filtering order and a second filtering order. In the first filtering order, the first chrominance component is in front and the second chrominance component is behind in the input of the neural network filter. In the second filtering order, the second chrominance component is in front and the first chrominance component is behind in the input of the neural network filter.
[0268] At this time, as Figure 14As shown, the encoding end inputs the first chrominance component (e.g., U component) and the second chrominance component (e.g., V component) of the current image block into the neural network filter for filtering according to the first filtering order, and determines the first filtering cost 1 corresponding to the first filtering order. At the same time, the encoding end inputs the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the second filtering order, and determines the first filtering cost 2 corresponding to the first filtering order. Finally, according to the first filtering cost 1 corresponding to the first filtering order and the first filtering cost 2 corresponding to the second filtering order, the target filtering order is selected from the first filtering order and the second filtering order. For example, the filtering order with the minimum first filtering cost among the first filtering order and the second filtering order is determined as the target filtering order.
[0269] As can be seen from the above, in the embodiment of the present application, when determining the target filtering order of the chrominance components of the current image block, the training order of the chrominance components of the neural network filter is not considered. That is to say, regardless of the training order of the chrominance components of the neural network filter, the encoding end determines the target filtering order of the chrominance components of the current chrominance block according to the above steps.
[0270] The embodiment of the present application does not limit the training method and related training parameters of the neural network filter.
[0271] In some embodiments, the above neural network filter is trained with at least one CTU as a training unit. That is to say, at least one CTU of the training image is divided into a training unit and input into the neural network filter to train the neural network filter.
[0272] In some embodiments, the neural network filter is trained with a preset image area as a training unit. That is to say, the preset image area of the training image is divided into a training unit and input into the neural network filter to train the neural network filter. For the relevant description of the above preset image area, reference can be made to the relevant description of the above preset image area, which will not be elaborated here.
[0273] In some embodiments, the training order of the chrominance components of the neural network filter is any one of N training orders. Optionally, the above N training orders may be the same as, different from, or partially the same and partially different from the above N filtering orders.
[0274] In some embodiments, the above N training orders include a first training order and a second training order. In the first training order, the first chrominance component is in front and the second chrominance component is behind in the input of the neural network filter. In the second training order, the second chrominance component is in front and the first chrominance component is behind in the input of the neural network filter.
[0275] In one example, during the training process, the chrominance training order of the neural network filter is the first training order. For example, the U component and the V component are input successively, and the loss function is used to constrain the filtered results of the U component and the V component respectively to implement the training of the neural network filter. When using the trained neural network filter to filter the first chrominance component and the second chrominance component of the current image block, the target filtering order of the first chrominance component and the second chrominance component is determined. Specifically, as Figure 15A shown, at the encoding end, according to the first filtering order, for example, the filtering order with the U component first and the V component second, the U component and the V component of the current image block are input into the neural network filter NNLF for filtering, and the filtered value of the U component of the current image block and the filtered value of the V component are output. Then, based on the filtered value of the U component and the filtered value of the V component of the current image block under the first filtering order, and the original image block of the current image block, the first filtering cost 1 corresponding to the first filtering order is determined. Similarly, as Figure 15A shown, at the encoding end, according to the second filtering order, for example, the filtering order with the V component first and the U component second, the U component and the V component of the current image block are input into the neural network filter NNLF for filtering, and the filtered value of the U component of the current image block and the filtered value of the V component are output. Then, based on the filtered value of the U component and the filtered value of the V component of the current image block under the second filtering order, and the original image block of the current image block, the first filtering cost 2 corresponding to the second filtering order is determined. Finally, according to the first filtering costs corresponding to the first filtering order and the second filtering order respectively, the target filtering order of the chrominance components of the current image block is determined from the first filtering order and the second filtering order.
[0276] In another example, during the training process, the chrominance training order of the neural network filter is the first training order. For example, the V component and the U component are input successively, and the loss function is used to constrain the filtered results of the V component and the U component respectively to implement the training of the neural network filter. When using the trained neural network filter to filter the first chrominance component and the second chrominance component of the current image block, the target filtering order of the first chrominance component and the second chrominance component is determined. Specifically, as Figure 15B shown, at the encoding end, according to the first filtering order, for example, the filtering order with the U component first and the V component second, the U component and the V component of the current image block are input into the neural network filter NNLF for filtering, and the filtered value of the U component of the current image block and the filtered value of the V component are output. Then, based on the filtered value of the U component and the filtered value of the V component of the current image block under the first filtering order, and the original image block of the current image block, the first filtering cost 1 corresponding to the first filtering order is determined. Similarly, as Figure 15BAs shown, the encoding end inputs the U and V components of the current image block into the neural network filter NNLF for filtering according to the second filtering order, for example, the filtering order with the V component first and the U component second, and outputs the filtered value of the U component and the filtered value of the V component of the current image block. Then, based on the filtered value of the U component and the filtered value of the V component of the current image block in the second filtering order, and the original image block of the current image block, the first filtering cost 2 corresponding to the second filtering order is determined. Finally, according to the first filtering costs corresponding to the first filtering order and the second filtering order respectively, the target filtering order of the chrominance component of the current image block is determined from the first filtering order and the second filtering order.
[0277] As can be seen from the above, the determination process of the target filtering order in the embodiments of the present application is independent of the chrominance training order of the neural network filter, thereby improving the selection flexibility and accuracy of the target filtering training.
[0278] It should be noted that the above takes N filtering orders including the first filtering order and the second filtering order as an example to introduce the process of determining the target filtering order. If the N filtering orders also include other filtering orders, for example, other filtering orders shown in Table 2 above, then the same method as the above first filtering order and the second filtering order is used to determine the first filtering cost of the surrounding encoded region in other filtering orders, so that the first filtering cost corresponding to each filtering order in the N filtering orders can be determined, and then based on the first filtering cost corresponding to each filtering order, the target filtering order is determined from the N filtering orders.
[0279] In some embodiments, after the encoding end determines the target filtering order based on the above steps, the first information is written into the code stream, and the first information indicates the target filtering order. In this way, the decoding end decodes the code stream to obtain the first information, and then based on the first information, obtains the target filtering order.
[0280] The embodiments of the present application do not limit the specific form of the first information, and any syntax field that can indicate the target filtering order is acceptable.
[0281] In some embodiments, the above first information includes a first flag, and the target filtering order is indicated by different values of the first flag.
[0282] Exemplarily, the correspondence between the values of the first flag and the filtering order of the chrominance component is shown in Table 1.
[0283] In the first method, the encoding end can determine the value of the first flag corresponding to the target filtering order based on Table 1 above. Then, after setting the first flag to this value, it is written into the bitstream. In this way, the decoding end decodes the bitstream to obtain the first flag, and then according to the value of the first flag, by querying Table 1 above, it obtains the target filtering orders of the first chrominance component and the second chrominance component of the current image block.
[0284] In addition to determining the target filtering order by using the method described in the first method above, the encoding end can also use the following method of the second method to determine the target filtering order.
[0285] Second method: The encoding end determines the target filtering order by itself. At this time, determining the target filtering orders of the first chrominance component and the second chrominance component of the current image block in S202 above includes the following steps of S202-B1 to S202-B3:
[0286] S202-B1: Determine the already filtered area around the current image block;
[0287] S202-B2: For the i-th filtering order among the N filtering orders, input the first chrominance component and the second chrominance component of the already filtered area around according to the i-th filtering order into the neural network filter for filtering, and determine the i-th second filtering cost of the already filtered area under the i-th filtering order, where i is a positive integer less than or equal to N;
[0288] S202-B3: Based on the second filtering costs respectively corresponding to the N filtering orders, determine the target filtering order from the N filtering orders.
[0289] In this second method, the encoding end determines the target filtering order from the N filtering orders based on the already filtered area around the current image block.
[0290] The embodiments of the present application do not limit the size and shape of the already filtered area around the current image block.
[0291] In some embodiments, the already filtered area around the current image block is the already filtered area adjacent to the current image block around the current image block.
[0292] In some embodiments, as Figure 9 shown, the already filtered area around the current image block includes the already filtered area above the current image block and the already filtered area on the left side.
[0293] In some embodiments, the already filtered area around the current image block includes the template area of the current image.
[0294] After the encoding end determines the filtered area around the current image block in the current image, it executes the steps of S202-B2 above. According to each of the N filtering orders, the first chrominance component and the second chrominance component of the filtered area around are input into the neural network filter for filtering, and the filtering cost corresponding to each of the N filtering orders is determined, and this filtering cost is recorded as the second filtering cost. In this embodiment, the specific process of the encoding end determining the second filtering cost corresponding to each of the N filtering orders is basically the same. Taking the i-th filtering order among the N filtering orders as an example for illustration. That is, according to the i-th filtering order, the first chrominance component and the second chrominance component of the filtered area around are input into the neural network filter for filtering, and the i-th second filtering cost of the filtered area around under the i-th filtering order is determined.
[0295] The embodiment of the present application does not limit the specific manner of determining the i-th second filtering cost of the filtered area around under the i-th filtering order.
[0296] In a possible implementation manner, the encoding end inputs the first chrominance component and the second chrominance component of the filtered area around into the neural network filter for filtering according to the i-th filtering order, and obtains the filtered value of the first chrominance component of the filtered area around under the i-th filtering order, and the filtered value of the second chrominance component under the i-th filtering order. Furthermore, based on the filtered value of the first chrominance component of the filtered area around under the i-th filtering order, and the filtered value of the second chrominance component under the i-th filtering order, the second filtering cost corresponding to the i-th filtering order is determined. For example, if the second filtering cost includes a calculation cost, the encoding end determines the calculation cost when filtering the first chrominance component and the second chrominance component of the filtered area around through the neural network filter in the i-th filtering order, and then determines the second filtering cost corresponding to the i-th filtering order based on this calculation cost.
[0297] In a possible implementation manner, the above S202-B2 includes the following steps of S202-B21 and S202-B22:
[0298] S202-B21: Input the first chrominance component and the second chrominance component of the filtered area around into the neural network filter for filtering according to the i-th filtering order, and obtain the i-th filtered value of the filtered area around;
[0299] S202-B22: Based on the i-th filtered value and the filtered area around, determine the i-th second filtering cost
[0300] In this implementation, the encoding end inputs the first chrominance component and the second chrominance component of the surrounding filtered region into the neural network filter for filtering according to the i-th filtering order, and obtains the filtered value of the first chrominance component of the surrounding filtered region under the i-th filtering order, and the filtered value of the second chrominance component under the i-th filtering order. For the sake of description, the filtered value of the first chrominance component of the surrounding filtered region under the i-th filtering order and the filtered value of the second chrominance component under the i-th filtering order are denoted as the i-th filtered value of the surrounding filtered region.
[0301] Next, based on the i-th filtered value and the surrounding filtered region, the second filtering cost corresponding to the i-th filtering order is determined.
[0302] For example, if the second filtering cost includes a distortion cost, the encoding end determines the second filtering cost corresponding to the i-th filtering order based on the filtered value of the first chrominance component of the surrounding filtered region and the first chrominance component of the surrounding filtered region, and the filtered value of the second chrominance component of the surrounding filtered region and the second chrominance component of the surrounding filtered region under the i-th filtering order.
[0303] The embodiments of the present application do not limit the specific calculation method of the above-mentioned second filtering cost. For example, the above-mentioned second filtering cost can be a rate-distortion cost (RDO), or an approximate cost, such as SSD, STAD, or SAD, etc.
[0304] For another example, if the second filtering cost includes a calculation cost and a distortion cost, the encoding end determines the distortion cost corresponding to the i-th filtering order based on the filtered value of the first chrominance component of the surrounding filtered region and the first chrominance component of the surrounding filtered region, and the filtered value of the second chrominance component of the surrounding filtered region and the second chrominance component of the surrounding filtered region under the i-th filtering order. At the same time, the calculation cost when filtering the first chrominance component and the second chrominance component of the surrounding filtered region by the neural network filter in the i-th filtering order is determined. In this way, according to the distortion cost and the calculation cost corresponding to the i-th filtering order, the second filtering cost corresponding to the i-th filtering order is determined. For example, the sum or weighted sum of the distortion cost and the calculation cost corresponding to the i-th filtering order is determined as the second filtering cost corresponding to the i-th filtering order.
[0305] Based on the above steps, the encoding end can determine the second filtering cost of the surrounding filtered region of the current image block under each of the N filtering orders. Then, based on the second filtering costs corresponding to the N filtering orders respectively, the target filtering order is determined from the N filtering orders.
[0306] The embodiments of the present application do not limit the specific manner of determining the target filtering order from the second filtering costs corresponding to N filtering orders at the encoding end.
[0307] For example, one filtering order with the minimum second filtering cost among the N filtering orders is determined as the target filtering order.
[0308] For another example, any filtering order with a second filtering cost less than a preset value among the N filtering orders is determined as the target filtering order.
[0309] After the encoding end determines the target filtering orders of the first chrominance component and the second chrominance component of the current image block based on the above steps, it executes the following step S203.
[0310] S203: Based on the target filtering order, input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering to obtain the chrominance filtered block of the current image block.
[0311] After the encoding end determines the target filtering orders of the first chrominance component and the second chrominance component of the current image block based on the above steps, according to the target filtering order, input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering to obtain the filtered image block of the current image block.
[0312] In one example, as Figure 12A shown, if the above target filtering order is that the first chrominance component is before and the second chrominance component is after, then the first chrominance component and the second chrominance component of the current image block are concatenated, and the first chrominance component is before the second chrominance component during concatenation. Then, the concatenated first chrominance component and second chrominance component are input into the neural network filter for filtering to obtain the filtered value of the first chrominance component and the filtered value of the second chrominance component of the current image block, where the filtered value of the first chrominance component and the filtered value of the second chrominance component of the current image block form the chrominance filtered block of the current image block.
[0313] In one example, as Figure 12B shown, if the above target filtering order is that the second chrominance component is before and the first chrominance component is after, then the first chrominance component and the second chrominance component of the current image block are concatenated, and the first chrominance component is after the second chrominance component during concatenation. Then, the concatenated first chrominance component and second chrominance component are input into the neural network filter for filtering to obtain the filtered value of the first chrominance component and the filtered value of the second chrominance component of the current image block, where the filtered value of the first chrominance component and the filtered value of the second chrominance component of the current image block form the chrominance filtered block of the current image block.
[0314] The above describes the filtering process of the current image block in the reconstructed image. The filtering processes of other image blocks to be filtered in the reconstructed image can refer to the filtering process of the current image block above, and finally the filtered reconstructed image is obtained.
[0315] In some embodiments, the above neural network filter is used as a loop filter. At this time, the output of the neural network filter affects video coding. For example, at the encoding end, by the above method, the neural network filter is used to filter the reconstructed image of the current image to obtain the filtered reconstructed image, and the filtered reconstructed image is stored in the encoding cache as the encoded image for subsequent image filtering. Since the embodiments of the present application provide the accuracy of determining the filtering order, the filtering quality of the reconstructed image is improved. When subsequent encoding is performed based on the reconstructed image with better quality, the encoding effect of the video can be enhanced.
[0316] In some embodiments, the above neural network filter is used for post-processing, that is, filtering and optimizing the decoded video. At this time, the output of the neural network filter does not affect video coding.
[0317] In some embodiments, the image filtering method proposed in the embodiments of the present application can be used not only in the field of video or image coding, but also in conventional image filtering. For example, for the current image block in the image to be filtered, the target filtering order of the first chrominance component and the second chrominance component of the current image block is determined. The target filtering order is determined based on the filtering costs of N filtering orders, where N is a positive integer greater than 1; based on the target filtering order, the first chrominance component and the second chrominance component of the current image block are input into the neural network filter for filtering to obtain the chrominance filtering block of the current image block.
[0318] In the image filtering method provided by the embodiments of the present application, when the encoding end encodes the current image, it first encodes the current image to obtain the reconstructed image of the current image; for the current image block to be filtered in the reconstructed image, the target filtering order of the first chrominance component and the second chrominance component of the current image block is determined. The target filtering order is determined based on the filtering costs of N filtering orders, where N is a positive integer greater than 1; based on the target filtering order, the first chrominance component and the second chrominance component of the current image block are input into the neural network filter for filtering to obtain the chrominance filtering block of the current image block. That is to say, the embodiments of the present application determine the target filtering order based on the filtering costs of N filtering orders, improving the selection accuracy of the target filtering order. When the first chrominance component and the second chrominance component of the current image block are input into the neural network filter for filtering based on the accurately determined target filtering order, the filtering effect can be improved, thereby enhancing the generalization of the neural network filter and improving the encoding performance.
[0319] The preferred embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all fall within the protection scope of the present application. For example, in the various specific technical features described in the above specific embodiments, they can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the present application will not separately describe various possible combination methods. For another example, any combination can be made between various different embodiments of the present application, as long as it does not violate the idea of the present application, it should also be regarded as the content disclosed in the present application.
[0320] It should also be understood that in various method embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0321] As described above in conjunction with Figures 7 to 15B , the method embodiments of the present application have been described in detail. Below, in conjunction with Figures 16 to 17 , the apparatus embodiments of the present application will be described in detail.
[0322] Figure 16 FIG. is a schematic block diagram of an image filtering apparatus provided by an embodiment of the present application.
[0323] As Figure 10 shown, the image filtering apparatus 10 may include:
[0324] A decoding unit 11, configured to decode the bitstream of the current image to obtain the residual value of the current image, and determine the reconstructed image of the current image based on the residual value;
[0325] An order determination unit 12, configured to determine, for a current image block to be filtered in the reconstructed image, the target filtering order of the first chrominance component and the second chrominance component of the current image block, where the target filtering order is determined by decoding the bitstream or based on the filtering cost of N filtering orders, and N is a positive integer greater than 1;
[0326] A filtering unit 13, configured to input the first chrominance component and the second chrominance component of the current image block into a neural network filter for filtering based on the target filtering order to obtain the chrominance filtering block of the current image block.
[0327] In some embodiments, the order determination unit 12 is specifically configured to decode the bitstream to obtain first information, where the first information is used to indicate the target filtering order; and obtain the target filtering order based on the first information.
[0328] In some embodiments, the filtering cost includes a first filtering cost, and the target filtering order is determined based on the first filtering cost of each of the N filtering orders. The first filtering cost of a filtering order is the filtering cost determined when the first chrominance component and the second chrominance component of the current image block are input into the neural network filter according to the filtering order.
[0329] In some embodiments, the target filtering order is the one with the minimum first filtering cost among the N filtering orders.
[0330] In some embodiments, the filtering cost includes a second filtering cost. The order determination unit 12 is specifically configured to determine the already-filtered region around the current image block; for the i-th filtering order among the N filtering orders, input the first chrominance component and the second chrominance component of the already-filtered region into the neural network filter for filtering according to the i-th filtering order, and determine the i-th second filtering cost of the already-filtered region under the i-th filtering order, where i is a positive integer less than or equal to N; and determine the target filtering order from the N filtering orders based on the second filtering costs respectively corresponding to the N filtering orders.
[0331] In some embodiments, the order determination unit 12 is specifically configured to input the first chrominance component and the second chrominance component of the already-filtered region into the neural network filter for filtering according to the i-th filtering order to obtain the i-th filtered value of the already-filtered region; and determine the i-th second filtering cost based on the i-th filtered value and the already-filtered region.
[0332] In some embodiments, the order determination unit 12 is specifically configured to determine the filtering order with the minimum second filtering cost among the N filtering orders as the target filtering order.
[0333] In some embodiments, the N filtering orders include a first filtering order and a second filtering order. In the first filtering order, the first chrominance component is before the second chrominance component in the input of the neural network filter, and in the second filtering order, the second chrominance component is before the first chrominance component in the input of the neural network filter.
[0334] In some embodiments, the current image block is at least one CTU of the reconstructed image, or the current image block is a preset image region of the reconstructed image.
[0335] In some embodiments, the neural network filter is trained with at least one CTU as a training unit, or the neural network filter is trained with a preset image region as a training unit.
[0336] In some embodiments, the training order of the chrominance components of the neural network filter is any one of N training orders.
[0337] In some embodiments, the N training orders include a first training order and a second training order. In the first training order, the first chrominance component is before the second chrominance component in the input of the neural network filter, and in the second training order, the second chrominance component is before the first chrominance component in the input of the neural network filter.
[0338] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they are not elaborated here. Specifically, Figure 16 the illustrated device can execute the above Figure 7 illustrated method embodiments, and the foregoing and other operations and / or functions of each module in the device respectively implement the corresponding method embodiments of the decoder. For the sake of brevity, they are not elaborated here.
[0339] Figure 17 is a schematic block diagram of an image filtering device provided by an embodiment of the present application.
[0340] As Figure 17 shown, the image filtering device 20 may include:
[0341] An encoding unit 21, configured to encode a current image to obtain a reconstructed image of the current image;
[0342] An order determination unit 22, configured to determine, for a current image block to be filtered in the reconstructed image, a target filtering order of a first chrominance component and a second chrominance component of the current image block, where the target filtering order is determined based on a filtering cost of N filtering orders, and N is a positive integer greater than 1;
[0343] A filtering unit 23, configured to input the first chrominance component and the second chrominance component of the current image block into a neural network filter for filtering based on the target filtering order, to obtain a filtered image block of the current image block.
[0344] In some embodiments, the order determination unit 22 is specifically configured to, for the j-th filtering order among the N filtering orders, input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order, and determine the j-th first filtering cost of the current image block in the j-th filtering order, where j is a positive integer less than or equal to N; and determine the target filtering order from the N filtering orders based on the first filtering costs respectively corresponding to the N filtering orders.
[0345] In some embodiments, the order determination unit 22 is specifically configured to input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order, to obtain the j-th filtered image block of the current image block; and determine the j-th first filtering cost based on the j-th filtered image block and the original image block of the current image block.
[0346] In some embodiments, the order determination unit 22 is specifically configured to determine the filtering order with the smallest first filtering cost among the N filtering orders as the target filtering order.
[0347] In some embodiments, the encoding unit 21 is further configured to write first information into the bitstream, where the first information is used to indicate the target filtering order.
[0348] In some embodiments, the order determination unit 22 is specifically configured to determine the already filtered region around the current image block; for the i-th filtering order among the N filtering orders, input the first chrominance component and the second chrominance component of the already filtered region around the current image block into the neural network filter for filtering according to the i-th filtering order, and determine the i-th second filtering cost of the already filtered region around the current image block in the i-th filtering order, where i is a positive integer less than or equal to N; and determine the target filtering order from the N filtering orders based on the second filtering costs respectively corresponding to the N filtering orders.
[0349] In some embodiments, the order determination unit 22 is specifically configured to input the first chrominance component and the second chrominance component of the already filtered region around the current image block into the neural network filter for filtering according to the i-th filtering order, to obtain the i-th filtered value of the already filtered region around the current image block; and determine the i-th second filtering cost based on the i-th filtered value and the already filtered region around the current image block.
[0350] In some embodiments, the order determination unit 22 is specifically configured to determine the filtering order with the smallest second filtering cost among the N filtering orders as the target filtering order.
[0351] In some embodiments, the N filtering orders include a first filtering order and a second filtering order. In the first filtering order, the first chrominance component is before the second chrominance component in the input of the neural network filter. In the second filtering order, the second chrominance component is before the first chrominance component in the input of the neural network filter.
[0352] In some embodiments, the current image block is at least one CTU of the reconstructed image, or the current image block is a preset image region of the reconstructed image.
[0353] In some embodiments, the neural network filter is trained with at least one CTU as a training unit, or the neural network filter is trained with a preset image region as a training unit.
[0354] In some embodiments, the training order of the chrominance components of the neural network filter is any one of N training orders.
[0355] In some embodiments, the N training orders include a first training order and a second training order. In the first training order, the first chrominance component is before the second chrominance component in the input of the neural network filter. In the second training order, the second chrominance component is before the first chrominance component in the input of the neural network filter.
[0356] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, details are not described herein again. Specifically, Figure 17 The illustrated apparatus can execute the embodiments of the above method, and the foregoing and other operations and / or functions of each module in the apparatus respectively implement the corresponding method embodiments of the encoder. For the sake of brevity, details are not described herein again.
[0357] The apparatus of the embodiments of the present application has been described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be implemented in hardware form, can also be implemented by instructions in software form, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0358] Figure 18 is a schematic block diagram of an electronic device provided by an embodiment of the present application. Figure 18 The electronic device can be the above-mentioned encoder or decoder.
[0359] As Figure 18 shown, the electronic device 30 may include:
[0360] A memory 31 and a processor 32. The memory 31 is used to store a computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiment of the present application.
[0361] For example, the processor 32 can be used to execute the steps in the above method 200 according to the instructions in the computer program 33.
[0362] In some embodiments of the present application, the processor 32 may include but is not limited to:
[0363] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.
[0364] In some embodiments of the present application, the memory 31 includes but is not limited to:
[0365] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0366] In some embodiments of the present application, the computer program 33 can be divided into one or more modules, and the one or more modules are stored in the memory 31 and executed by the processor 32 to complete the method for recording a page provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the electronic device 900.
[0367] As Figure 18 shown, the electronic device 30 may further include:
[0368] A transceiver 34, which can be connected to the processor 32 or the memory 31.
[0369] Among them, the processor 32 can control the transceiver 34 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 34 can include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas can be one or more.
[0370] It should be understood that the various components in the electronic device 30 are connected through a bus system. Among them, the bus system includes, in addition to a data bus, a power bus, a control bus, and a status signal bus.
[0371] According to one aspect of the present application, there is provided a computer storage medium having a computer program stored thereon. When the computer program is executed by a computer, the computer is enabled to execute the method of the above method embodiment. Or rather, the embodiment of the present application further provides a computer program product containing instructions. When the instructions are executed by a computer, the computer is enabled to execute the method of the above method embodiment.
[0372] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the method of the above method embodiment.
[0373] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0374] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0375] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or module can be in an electrical, mechanical, or other form.
[0376] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment. For example, in each embodiment of this application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0377] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. An image filtering method, characterized in that, Including: Decoding the bitstream of the current image to obtain the residual value of the current image, and determining the reconstructed image of the current image based on the residual value; For the current image block to be filtered in the reconstructed image, determining the target filtering order of the first chrominance component and the second chrominance component of the current image block, where the target filtering order is determined by decoding the bitstream or based on the filtering cost of N filtering orders, and N is a positive integer greater than 1; Based on the target filtering order, inputting the first chrominance component and the second chrominance component of the current image block into a neural network filter for filtering to obtain the chrominance filtering block of the current image block.
2. The method according to claim 1, characterized in that, The determining the target filtering order of the first chrominance component and the second chrominance component of the current image block includes: Decoding the bitstream to obtain first information, where the first information is used to indicate the target filtering order; Based on the first information, obtaining the target filtering order.
3. The method according to claim 2, wherein The filtering cost includes a first filtering cost, and the target filtering order is determined based on the first filtering cost of each of the N filtering orders. The first filtering cost of a filtering order is the filtering cost determined when inputting the first chrominance component and the second chrominance component of the current image block into the neural network filter according to the filtering order; Wherein, the target filtering order is the filtering order with the smallest first filtering cost among the N filtering orders.
4. The method according to claim 1, characterized in that, The filtering cost includes a second filtering cost, and the determining the target filtering order of the first chrominance component and the second chrominance component of the current image block includes: Determining the already filtered region around the current image block; For the i-th filtering order among the N filtering orders, inputting the first chrominance component and the second chrominance component of the already filtered region around into the neural network filter for filtering according to the i-th filtering order, and determining the i-th second filtering cost of the already filtered region under the i-th filtering order, where i is a positive integer less than or equal to N; Based on the second filtering costs respectively corresponding to the N filtering orders, determining the target filtering order from the N filtering orders.
5. The method according to claim 4, characterized in that, The inputting the first chrominance component and the second chrominance component of the already filtered region around into the neural network filter for filtering according to the i-th filtering order, and determining the i-th second filtering cost of the already filtered region under the i-th filtering order includes: Inputting the first chrominance component and the second chrominance component of the already filtered region around into the neural network filter for filtering according to the i-th filtering order to obtain the i-th filtered value of the already filtered region; Based on the i-th filtered value and the already filtered region, determining the i-th second filtering cost.
6. The method according to claim 4, wherein The based on the second filtering costs respectively corresponding to the N filtering orders, determining the target filtering order from the N filtering orders includes: Determining the filtering order with the smallest second filtering cost among the N filtering orders as the target filtering order.
7. The method according to any one of claims 1-6, characterized in that, The N filtering orders include a first filtering order and a second filtering order. In the first filtering order, the first chrominance component is before the second chrominance component in the input of the neural network filter. In the second filtering order, the second chrominance component is before the first chrominance component in the input of the neural network filter.
8. The method according to any one of claims 1-6, characterized in that, The current image block is at least one CTU of the reconstructed image, or the current image block is a preset image region of the reconstructed image.
9. The method according to any one of claims 1-6, characterized in that, The neural network filter is trained with at least one CTU as one training unit, or the neural network filter is trained with a preset image region as a training unit.
10. The method according to any one of claims 1-6, characterized in that, The training order of the chrominance components of the neural network filter is any one of N training orders.
11. The method according to claim 10, wherein The N training orders include a first training order and a second training order. In the first training order, the first chrominance component is before the second chrominance component in the input of the neural network filter. In the second training order, the second chrominance component is before the first chrominance component in the input of the neural network filter.
12. An image filtering method, characterized in that, including: Encoding the current image to obtain the reconstructed image of the current image; For the current image block to be filtered in the reconstructed image, determining the target filtering order of the first chrominance component and the second chrominance component of the current image block, where the target filtering order is determined based on the filtering cost of N filtering orders, and N is a positive integer greater than 1; Based on the target filtering order, inputting the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering to obtain the filtered image block of the current image block.
13. The method according to claim 12, characterized in that, The determining the target filtering order of the first chrominance component and the second chrominance component of the current image block includes: For the j-th filtering order among the N filtering orders, inputting the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order, and determining the j-th first filtering cost of the current image block in the j-th filtering order, where j is a positive integer less than or equal to N; Based on the first filtering costs respectively corresponding to the N filtering orders, determining the target filtering order from the N filtering orders.
14. The method according to claim 13, wherein The inputting the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order and determining the j-th first filtering cost of the current image block in the j-th filtering order includes: Inputting the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering according to the j-th filtering order to obtain the j-th filtered image block of the current image block; Based on the j-th filtered image block and the original image block of the current image block, determining the j-th first filtering cost; The determining the target filtering order from the N filtering orders based on the first filtering costs respectively corresponding to the N filtering orders includes: Determine the filtering order with the smallest first filtering cost among the N filtering orders as the target filtering order.
15. The method according to claim 13, wherein The method further includes: Write first information into the bitstream, where the first information is used to indicate the target filtering order.
16. The method according to claim 12, wherein The determining the target filtering orders of the first chrominance component and the second chrominance component of the current image block includes: Determine the already-filtered region around the current image block; For the i-th filtering order among the N filtering orders, input the first chrominance component and the second chrominance component of the already-filtered region around the current image block into the neural network filter for filtering according to the i-th filtering order, and determine the i-th second filtering cost of the already-filtered region under the i-th filtering order, where i is a positive integer less than or equal to N; Based on the second filtering costs respectively corresponding to the N filtering orders, determine the target filtering order from the N filtering orders.
17. An image filtering device, characterized in that, Includes: A decoding unit, configured to decode the bitstream of the current image to obtain the residual value of the current image, and determine the reconstructed image of the current image based on the residual value; An order determination unit, configured to determine the target filtering orders of the first chrominance component and the second chrominance component of the current image block to be filtered in the reconstructed image, where the target filtering order is determined by decoding the bitstream or based on the filtering costs of N filtering orders, and N is a positive integer greater than 1; A filtering unit, configured to input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering based on the target filtering order to obtain the chrominance filtering block of the current image block.
18. An image filtering device, characterized in that, Includes: An encoding unit, configured to encode the current image to obtain the reconstructed image of the current image; An order determination unit, configured to determine the target filtering orders of the first chrominance component and the second chrominance component of the current image block to be filtered in the reconstructed image, where the target filtering order is determined based on the filtering costs of N filtering orders, and N is a positive integer greater than 1; A filtering unit, configured to input the first chrominance component and the second chrominance component of the current image block into the neural network filter for filtering based on the target filtering order to obtain the filtered image block of the current image block.
19. An electronic device, including a processor and a memory; The memory is used to store a computer program; The processor is configured to execute the computer program to implement the method according to any one of claims 1 to 11 or 12 to 16 above.
20. A computer-readable storage medium, characterized in that, For storing a computer program; The computer program causes the computer to execute the method according to any one of claims 1 to 11 or 12 to 16 above.
Citation Information
Patent Citations
Image encoding / decoding method and apparatus
CN101507277A
Multi-view video coding strong filtering realization method for array structure
CN105847839A