Filtering method and device, electronic equipment and storage medium

By using neural network filters in video encoding and decoding technology, and combining time domain and airspace information, the problem of poor filtering effect in the prior art is solved, and more efficient image quality recovery is achieved.

CN119996677APending Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311497940.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing video encoding and decoding technology, the filtering effect has not yet been achieved, which makes it difficult to completely reduce the compression loss of the reconstructed image.

Method used

The filtering information is filtered by a neural network filter, and by determining at least one reference image in the reference image set, the time domain and airspace information are used to perform auxiliary filtering to improve the filtering performance.

Benefits of technology

Through the use of neural network filters, the filtering effect of the image to be filtered can be significantly improved, the compression loss of the reconstructed image can be reduced, and the performance of video encoding and decoding can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996677A_ABST
    Figure CN119996677A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a filtering method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring first to-be-filtered information; the first to-be-filtered information comprises a first to-be-filtered unit obtained by decoding; when a neural network filter is utilized to filter a first to-be-filtered unit in the first to-be-filtered information, determining at least one reference image in a reference image set; the reference image set comprises at least one of the following items: a reference image in a reference image list of the first to-be-filtered unit and a to-be-filtered image to which the first to-be-filtered unit belongs; and based on the first to-be-filtered information and the at least one reference image, filtering the first to-be-filtered unit by using the neural network filter, and obtaining a filtered unit of the first to-be-filtered unit. The filtering performance of the neural network filter is improved by enriching the selection range of the reference image, so that the filtering effect of the to-be-filtered image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video coding and decoding technology, and more specifically, to a filtering method, device, electronic device and storage medium. Background Art

[0002] With the development of video technology, the amount of data included in video data is large. In order to facilitate the transmission of video data, video devices perform video compression technology to make video data more efficiently transmitted or stored. In video compression, both the encoder and the decoder need to obtain a reconstructed image through operations such as inverse quantization and inverse transformation. Since loss is introduced in video compression, after obtaining the reconstructed image, it is necessary to further filter the reconstructed image to reduce the compression loss of the reconstructed image.

[0003] Typically, a deblocking filter (DBF), sample adaptive offset (SAO), and adaptive loop filter (ALF) may be used to filter the reconstructed image.

[0004] However, with the development of technology, there is still a need to pursue better filtering technology to improve the filtering effect. Summary of the invention

[0005] The embodiments of the present application provide a filtering method, device, electronic device and storage medium, which can improve the filtering effect.

[0006] In a first aspect, an embodiment of the present application provides a filtering method, including:

[0007] Acquire first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by decoding;

[0008] When filtering the first unit to be filtered in the first information to be filtered using a neural network filter, at least one reference image in a reference image set is determined; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs;

[0009] Based on the first information to be filtered and the at least one reference image, the first unit to be filtered is filtered using the neural network filter to obtain a filtered unit of the first unit to be filtered.

[0010] In a second aspect, an embodiment of the present application provides a filtering method, including:

[0011] Acquire first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by encoding;

[0012] When filtering the first unit to be filtered in the first information to be filtered using a neural network filter, at least one reference image in a reference image set is determined; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs;

[0013] Based on the first information to be filtered and the at least one reference image, the first unit to be filtered is filtered using the neural network filter to obtain a filtered unit of the first unit to be filtered.

[0014] In a third aspect, an embodiment of the present application provides a filtering device, including:

[0015] An acquisition unit, used for acquiring first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by decoding;

[0016] A determination unit, configured to determine at least one reference image in a reference image set when filtering a first unit to be filtered in the first information to be filtered using a neural network filter; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs;

[0017] A filtering unit is used to filter the first unit to be filtered using the neural network filter based on the first information to be filtered and the at least one reference image, and obtain a filtered unit of the first unit to be filtered.

[0018] In a fourth aspect, an embodiment of the present application provides a filtering device, including:

[0019] An acquisition unit, used for acquiring first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by encoding;

[0020] A determination unit, configured to determine at least one reference image in a reference image set when filtering a first unit to be filtered in the first information to be filtered using a neural network filter; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs;

[0021] A filtering unit is used to filter the first unit to be filtered using the neural network filter based on the first information to be filtered and the at least one reference image, and obtain a filtered unit of the first unit to be filtered.

[0022] In a fifth aspect, an embodiment of the present application provides a filtering device, including:

[0023] a processor adapted to implement computer instructions; and,

[0024] A computer-readable storage medium stores computer instructions, wherein the computer instructions are suitable for being loaded by a processor and executing the filtering method in the first aspect or its various implementation modes involved above.

[0025] In one implementation, the number of the processor is one or more, and the number of the memory is one or more.

[0026] In one implementation, the computer-readable storage medium may be integrated with the processor, or the computer-readable storage medium may be disposed separately from the processor.

[0027] In a sixth aspect, an embodiment of the present application provides a filtering device, including:

[0028] a processor adapted to implement computer instructions; and,

[0029] A computer-readable storage medium stores computer instructions, wherein the computer instructions are suitable for being loaded by a processor and executing the filtering method in the second aspect or its various implementation modes involved above.

[0030] In one implementation, the number of the processor is one or more, and the number of the memory is one or more.

[0031] In one implementation, the computer-readable storage medium may be integrated with the processor, or the computer-readable storage medium may be disposed separately from the processor.

[0032] In the seventh aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are read and executed by a processor of a computer device, the computer device executes the filtering method involved in the first aspect mentioned above or the filtering method involved in the second aspect mentioned above.

[0033] In an eighth aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including a computer instruction, the computer instruction being stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the filtering method involved in the first aspect mentioned above or the filtering method involved in the second aspect mentioned above.

[0034] In a ninth aspect, an embodiment of the present application provides a code stream, which may be a code stream decoded using the filtering method described in the first aspect mentioned above, or the code stream may be a code stream generated using the filtering method described in the second aspect mentioned above.

[0035] For the filtering method provided in the present application, the method includes: obtaining first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by decoding; when filtering the first unit to be filtered in the first information to be filtered using a neural network filter, determining at least one reference image in a reference image set; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, an image to be filtered to which the first unit to be filtered belongs; based on the first information to be filtered and the at least one reference image, filtering the first unit to be filtered using the neural network filter, and obtaining a filtered unit of the first unit to be filtered. That is to say, when filtering the first unit to be filtered in the first information to be filtered using a neural network filter, it is necessary to first determine at least one reference image in a reference image set, and then based on the first information to be filtered and the at least one reference image, filtering the first unit to be filtered using the neural network filter, and obtaining a filtered unit of the first unit to be filtered. Since the reference image set includes at least one of the following: the reference image in the reference image list of the first unit to be filtered, the image to be filtered to which the first unit to be filtered belongs, and the reference image in the reference image list of the first unit to be filtered can provide the first unit to be filtered with information in the time domain, and the image to be filtered to which the first unit to be filtered belongs can provide information in the spatial domain for the first filtering unit, therefore, based on the first information to be filtered and the at least one reference image, the first unit to be filtered is filtered using the neural network filter, which is equivalent to using at least one of the time domain information and the spatial domain information to assist the neural network filter in filtering the first unit to be filtered, that is, the filtering performance of the neural network filter is improved by enriching the selection range of the reference image, thereby improving the filtering effect of the image to be filtered. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0037] Figure 1 A schematic block diagram of a video encoding and decoding system involved in an embodiment of the present application.

[0038] Figure 2 It is a schematic block diagram of a video encoder involved in an embodiment of the present application.

[0039] Figure 3 It is a schematic structural diagram of the relationship between the coding tree unit and the coding unit provided in the present application.

[0040] Figure 4 It is a schematic block diagram of a video decoder involved in an embodiment of the present application.

[0041] Figure 5 It is another example of the video encoder involved in the embodiments of the present application.

[0042] Figure 6 It is a schematic structural diagram of the CTU provided in this application.

[0043] Figure 7 It is a schematic diagram of the filtering principle of the NNLF involved in the embodiment of the present application.

[0044] Figure 8 It is a schematic diagram of a filtering principle of a time-domain NNLF involved in the present application.

[0045] Fig. 9 It is a schematic flow chart of the filtering method provided in the embodiment of the present application.

[0046] Fig.10 This is an example of an image block to be filtered provided in an embodiment of the present application.

[0047] Fig.11 This is an example of at least one reference image provided in an embodiment of the present application.

[0048] Fig.12 is another example of at least one reference image provided by an embodiment of the present application.

[0049] Fig.13 is another schematic flow chart of the filtering method provided in an embodiment of the present application.

[0050] Fig.14 It is a schematic block diagram of a filtering device provided in an embodiment of the present application.

[0051] Fig.15 It is another schematic block diagram of the filtering device provided in an embodiment of the present application.

[0052] Fig.16 It is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments provided by this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0054] To facilitate understanding of the technical solution provided by this application, the relevant terms are explained below.

[0055] It should be noted that the terms used in the implementation method section of the present application are only used to explain the embodiments of the present application and are not intended to limit the present application.

[0056] For example, the term "and / or" in this article is only a way to describe the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" is only a way to describe the combination relationship of listed objects, indicating that one or more items may exist. For example, at least one of the following: A, B, C can mean the following combinations: A exists alone, B exists alone, C exists alone, A and B exist at the same time, A and C exist at the same time, B and C exist at the same time, and A, B, and C exist at the same time. The term "multiple" means two or more. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.

[0057] For another example, the term "corresponding" may indicate that there is a direct or indirect correspondence between the two, or that there is an association relationship between the two, or that there is an indication and being indicated, configuration and being configured, etc. The term "indication" may be a direct indication, an indirect indication, or an indication of an association relationship. For example, A indicates B, which may indicate that A directly indicates B, such as B can be obtained through A; it may also indicate that A indirectly indicates B, such as A indicates C, B can be obtained through C; it may also indicate that there is an association relationship between A and B. The term "predefined" or "preconfigured" may refer to the pre-storage of corresponding codes, tables or other relevant information that can be used for indication in the device, or it may refer to an agreement by protocol. "Protocol" may refer to a standard protocol in this field. The term "when..." may be interpreted as "if" or "if" or "when..." or "in response to" and other similar descriptions. Similarly, depending on the context, the phrase "if determined" or "if (stated condition or event) is detected" may be interpreted as "when determined" or "in response to determining" or "when (stated condition or event) is detected" or "in response to detecting (stated condition or event)" and similar descriptions. The terms "first", "second", "third", "fourth", "Ath", "Bth", etc. are used to distinguish different objects rather than to describe a specific order. The terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. Among them, digital video compression technology mainly compresses huge digital image video data for easy transmission and storage.

[0058] The solution provided by this application relates to the field of digital compression technology.

[0059] Among them, digital video compression technology is mainly used to compress huge digital image video data for easy transmission and storage.

[0060] The solution provided by the present application can be applied to the field of digital video encoding technology.

[0061] Among them, the field of digital video coding technology includes but is not limited to at least one of the following: image coding field, video coding field, hardware video coding field, dedicated circuit video coding field and real-time video coding field. In addition, the scheme provided in this application can be combined with the following standards: Audio Video Coding Standard (AVS), second-generation AVS standard (AVS2) or third-generation AVS standard (AVS3). For example, including but not limited to: H.264 / Audio Video Coding (AVC) standard, H.265 / High Efficiency Video Coding (HEVC) standard and H.266 / Versatile Video Coding (VVC) standard. In addition, the scheme provided in this application can be used for lossy compression of images, and can also be used for lossless compression of images. Among them, the lossless compression can be visually lossless compression or mathematically lossless compression.

[0062] To facilitate understanding, first combine Figure 1 The video encoding and decoding system involved in the embodiments of the present application is introduced.

[0063] Figure 1 A schematic block diagram of a video encoding and decoding system involved in an embodiment of the present application.

[0064] like Figure 1 As shown, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120 .

[0065] The encoding device 110 is used to encode (which can be understood as compressing) the video data to generate a code stream, and transmit the code stream to the decoding device 120. The decoding device 120 decodes the code stream generated by the encoding device 110 to obtain decoded video data.

[0066] The encoding device 110 can be understood as a device with a video encoding function, and the decoding device 120 can be understood as a device with a video decoding function, that is, the embodiments of the present application include a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, vehicle-mounted computers, etc.

[0067] The encoding device 110 may transmit the encoded video data (eg, a bitstream) to the decoding device 120 via the channel 130 .

[0068] Channel 130 may include one or more media and / or devices capable of transmitting encoded video data from encoding device 110 to decoding device 120 .

[0069] The channel 130 may include one or more communication media that enable the encoding device 110 to transmit the encoded video data directly to the decoding device 120 in real time. The encoding device 110 may modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media may include wireless communication media, such as radio frequency spectrum. The communication media may also include wired communication media, such as one or more physical transmission lines.

[0070] The channel 130 may include a storage medium that can store the video data encoded by the encoding device 110. The storage medium includes a variety of locally accessible data storage media, such as an optical disk, a DVD, a flash memory, etc. In this example, the decoding device 120 may obtain the encoded video data from the storage medium.

[0071] The channel 130 may include a storage server that can store the video data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded video data from the storage server. Alternatively, the storage server can store the encoded video data and transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0072] The encoding device 110 includes a video encoder 112 and an output interface 113 .

[0073] The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter. The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data may also be stored in a storage medium or a storage server for subsequent reading by the decoding device 120.

[0074] The encoding device 110 may include a video source 111 in addition to the video encoder 112 and the input interface 113 .

[0075] The video source 111 may include at least one of a video acquisition device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data. The video encoder 112 encodes the video data from the video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the encoding information of the picture or the sequence of pictures in the form of a bitstream. The encoding information may include the encoded picture data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. The SPS may contain parameters applied to one or more sequences. The PPS may contain parameters applied to one or more pictures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.

[0076] The decoding device 120 includes an input interface 121 and a video decoder 122. The input interface 121 may include a receiver and / or a modem.

[0077] The decoding device 120 may include a display device 123 in addition to the input interface 121 and the video decoder 122 .

[0078] The input interface 121 may receive the encoded video data through the channel 130. The video decoder 122 is used to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to the display device 123. The display device 123 displays the decoded video data. The display device 123 may be integrated with the decoding device 120 or outside the decoding device 120. The display device 123 may include a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0079] It should be understood that Figure 1 This is only an example of the present application and should not be understood as a display of the present application. In other words, the technical solution of the embodiment of the present application is not limited to Figure 1 The system framework shown, for example, the technology of the present application can also be applied to single-sided video encoding or single-sided video decoding.

[0080] The following is an introduction to the video encoding framework involved in the embodiments of the present application.

[0081] Figure 2 It is a schematic block diagram of a video encoder 200 involved in an embodiment of the present application.

[0082] It should be understood that the video encoder 200 can be applied to image data in luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2 or 4:4:4, Y represents brightness (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) for describing color and saturation. For example, in color format, 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr). Of course, it can also be applied to image data in red-green-blue (RGB) format, and this application does not make specific limitations on this.

[0083] After the video encoder 200 reads the video stream, each frame of the video stream may be divided into a number of coding tree units (CTUs). In some examples, a CTU may be referred to as a tree block, a largest coding unit (LCU), or a coding tree block (CTB). A CTU size may be, for example, 128×128, 64×64, 32×32, etc.

[0084] Figure 3 It is a schematic structural diagram of the relationship between the coding tree unit and the coding unit provided in the present application.

[0085] like Figure 3 As shown, a CTU can be further divided into several coding units (Coding Unit, CU) for encoding. CU can be a rectangular block or a square block. CU can be further divided into prediction unit (prediction Unit, PU) and transform unit (transform unit, TU), so that encoding, prediction, and transformation are separated, which is more flexible when processing. In one example, CTU is divided into CU in a tree manner (such as a quadtree), and CU is divided into TU and PU in a tree manner (such as a quadtree).

[0086] The video encoder and the video decoder may support various PU sizes.

[0087] Assuming that the size of a particular CU is 2N×2N, the video encoder and the video decoder may support a PU size of 2N×2N or N×N for intra prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N or similar sizes for inter prediction. The video encoder and the video decoder may also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.

[0088] like Figure 2 As shown, the video encoder 200 may include: a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image cache 270 and an entropy coding unit 280. It should be noted that the video encoder 200 may include more, fewer or different functional components. In the present application, the current block may be referred to as a current coding unit (CU) or a current prediction unit (PU), etc. The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or an image reconstructed image block.

[0089] The prediction unit 210 includes an inter prediction unit 211 and an intra prediction unit 212. Since there is a strong correlation between adjacent pixels in an image in a video, an intra prediction method is used to eliminate spatial redundancy between adjacent pixels in video coding and decoding technology. Since there is a strong similarity between adjacent images in a video, an inter prediction method is used to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency.

[0090] The inter prediction unit 211 can be used for inter prediction, which may include motion estimation and motion compensation. It may refer to the image information of different frames. Inter prediction uses motion information to find a reference block from a reference frame, and generates a prediction block based on the reference block to eliminate temporal redundancy. The reference frame may be a P frame and / or a B frame. A P frame refers to a forward prediction frame, and a B frame refers to a bidirectional prediction frame. After inter prediction uses motion information to find a reference block, a prediction block is generated based on the reference block. The motion information includes a frame list, a frame index, and a motion vector to which the reference frame belongs. The motion vector may be an integer pixel or a sub-pixel. If the motion vector is a sub-pixel, an interpolation filter is required in the reference frame to make the required sub-pixel block. The reference block is the integer pixel or sub-pixel block found based on the motion vector. Some technologies directly use the reference block as a prediction block, while some technologies further process the reference block to generate a prediction block. Reprocessing the reference block to generate a prediction block can also be understood as using the reference block as a prediction block and then processing the prediction block to generate a new prediction block.

[0091] The intra prediction unit 212 only refers to the information of the same frame image to predict the pixel information in the current code image block to eliminate spatial redundancy. The reference frame used for intra prediction can be an I frame.

[0092] Intra prediction has multiple prediction modes. The image block to be encoded can be predicted with the help of angle prediction mode and non-angle prediction mode to obtain the prediction block. According to the prediction block and the image block to be encoded, the rate distortion information is calculated to select the optimal prediction mode of the image block to be encoded, and the prediction mode is written into the bitstream for transmission to the decoder. The decoder parses the prediction mode, predicts the prediction block of the target decoding block and superimposes the time domain residual block obtained based on the bitstream to obtain the reconstructed block.

[0093] Taking the H series of international digital video coding standards as an example, the H.264 / AVC standard has 8 angle prediction modes and 1 non-angle prediction mode, and H.265 / HEVC is expanded to 33 angle prediction modes and 2 non-angle prediction modes. The intra prediction modes used by HEVC are planar mode, direct current (DC) and 33 angle modes, a total of 35 prediction modes. The intra-frame modes used by VVC are Planar, DC and 65 angle modes, a total of 67 prediction modes, which include traditional prediction modes and non-traditional prediction modes. Non-traditional prediction modes may include matrix weighted intra-frame prediction (MIP) mode. Traditional prediction modes include: planar mode with mode number 0, DC mode with mode number 1, and angle prediction modes with mode numbers 2 to 66. It should be noted that with the increase of angle modes, the prediction results of intra prediction will be more accurate and more in line with the needs of the development of high-definition and ultra-high-definition digital video. The above intra prediction modes are only examples of this application and should not limit this application.

[0094] The residual unit 220 may generate a residual block of the CU based on the pixel blocks of the CU and the prediction blocks of the PUs of the CU. For example, the residual unit 220 may generate a residual block of the CU so that each sample in the residual block has a value equal to the difference between the following two: a sample in the pixel blocks of the CU and a corresponding sample in the prediction blocks of the PUs of the CU.

[0095] The transform / quantization unit 230 may quantize the transform coefficients. The transform / quantization unit 230 may quantize the transform coefficients associated with the TUs of the CU based on a quantization parameter (QP) value associated with the CU. The video encoder 200 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.

[0096] The inverse transform / quantization unit 240 may apply inverse quantization and inverse transform to the quantized transform coefficients, respectively, to reconstruct a residual block from the quantized transform coefficients.

[0097] The reconstruction unit 250 may add the samples of the reconstructed residual block to the corresponding samples of one or more prediction blocks generated by the prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of the CU in this manner, the video encoder 200 may reconstruct the pixel blocks of the CU.

[0098] The loop filter unit 260 is used to process the inverse transformed and inverse quantized pixels, compensate for the distortion information, and provide a better reference for the subsequent coded pixels. For example, a deblocking filter operation can be performed to reduce the blocking effect of the pixel blocks associated with the CU. In some embodiments, the loop filter unit 260 includes: a deblocking filter (DBF) unit and a sample adaptive compensation / adaptive loop filter (SAO / ALF) unit, wherein the DBF unit is used to remove the block effect, and the SAO / ALF unit is used to remove the ringing effect.

[0099] The decoded image buffer 270 may store the reconstructed pixel blocks.

[0100] The inter prediction unit 211 may use the reference image containing the reconstructed pixel block in the decoded image buffer 270 to perform inter prediction on the PU of other images. In addition, the intra prediction unit 212 may use the reconstructed pixel block in the decoded image buffer 270 to perform intra prediction on other PUs in the same image as the CU.

[0101] The entropy encoding unit 280 may receive the quantized transform coefficients from the transform / quantization unit 230. The entropy encoding unit 280 may perform one or more entropy encoding operations on the quantized transform coefficients to generate entropy-encoded data.

[0102] Figure 4 It is a schematic block diagram of a video decoder involved in an embodiment of the present application.

[0103] like Figure 4 As shown, the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transformation unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded image buffer 360. It should be noted that the video decoder 300 may include more, fewer, or different functional components.

[0104] The video decoder 300 may receive a bitstream. The entropy decoding unit 310 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding unit 310 may parse the syntax elements in the bitstream that have been entropy encoded. The prediction unit 320, the inverse quantization / transformation unit 330, the reconstruction unit 340, and the loop filter unit 350 may decode the video data according to the syntax elements extracted from the bitstream, that is, generate decoded video data.

[0105] The prediction unit 320 includes an intra prediction unit 322 and an inter prediction unit 321 .

[0106] The intra prediction unit 322 may perform intra prediction to generate a prediction block of the PU. The intra prediction unit 322 may use an intra prediction mode to generate a prediction block of the PU based on a pixel block of a spatially neighboring PU. The intra prediction unit 322 may also determine the intra prediction mode of the PU according to one or more syntax elements parsed from the code stream.

[0107] The inter prediction unit 321 may construct a first reference image list (list 0) and a second reference image list (list 1) according to the syntax elements parsed from the code stream. In addition, if the PU is encoded using inter prediction, the entropy decoding unit 310 may parse the motion information of the PU. The inter prediction unit 321 may determine one or more reference blocks of the PU according to the motion information of the PU. The inter prediction unit 321 may generate a prediction block of the PU according to one or more reference blocks of the PU.

[0108] The inverse quantization / transform unit 330 may inversely quantize (i.e., dequantize) the transform coefficients associated with the TU. The inverse quantization / transform unit 330 may use the QP value associated with the CU of the TU to determine the degree of quantization. After inverse quantizing the transform coefficients, the inverse quantization / transform unit 330 may apply one or more inverse transforms to the inverse quantized transform coefficients to generate a residual block associated with the TU.

[0109] The reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, the reconstruction unit 340 may add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.

[0110] The loop filtering unit 350 may perform a deblocking filtering operation to reduce blocking effects of pixel blocks associated with a CU.

[0111] The video decoder 300 may store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.

[0112] Combination Figure 2 and Figure 4 For example, the basic process of video encoding and decoding is as follows:

[0113] At the encoding end, a frame of image is divided into image blocks. For the current block, the prediction unit 210 uses intra prediction or inter prediction to predict the prediction block of the current block (i.e., the block to be encoded). The residual unit 220 can calculate the residual block based on the original block of the prediction block and the current block (i.e., the block to be encoded), that is, the difference between the prediction block and the original block, and the residual block can also be called residual information. The residual block can remove information that is not sensitive to the human eye through the transformation and quantization process of the transformation / quantization unit 230 to eliminate visual redundancy. Optionally, the residual block before transformation and quantization by the transformation / quantization unit 230 can be called a time domain residual block, and the time domain residual block after transformation and quantization by the transformation / quantization unit 230 can be called a frequency residual block or a frequency domain residual block. The entropy coding unit 280 receives the quantized change coefficient output by the change quantization unit 230, and can entropy encode the quantized change coefficient and output a code stream. For example, the entropy coding unit 280 can eliminate character redundancy according to the target context model and the probability information of the binary code stream.

[0114] At the decoding end, the entropy decoding unit 310 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block (i.e., the block to be decoded). The prediction unit 320 uses intra prediction or inter prediction based on the prediction information to predict the prediction block of the current block (i.e., the block to be decoded). The inverse quantization / transformation unit 330 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed image, and the loop filtering unit 350 performs loop filtering on the reconstructed image based on the image or on the block to obtain a decoded image. It is worth noting that the encoding end also needs to use operations similar to those of the decoder to obtain a decoded image. The decoded image can also be called a reconstructed image, and the reconstructed image can be a subsequent image, which serves as a reference image for inter prediction.

[0115] In addition, the block division information determined by the encoder, as well as the mode information or parameter information such as prediction, transformation, quantization, entropy coding, loop filtering, etc., are carried in the bitstream when necessary. The decoder parses the bitstream and determines the same block division information, prediction, transformation, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder by analyzing the existing information, thereby ensuring that the decoded image obtained by the encoder is the same as the decoded image obtained by the decoder.

[0116] It should be noted that, due to the need for parallel processing, an image can be divided into slices, etc., and slices in the same image can be processed in parallel, that is, there is no data dependency between them. The term "frame" can be understood as an image or a slice, etc.

[0117] In addition, the above is a basic process of a video codec under a block-based codec framework. With the development of technology, some modules or steps of the framework or process may be optimized, that is, the present application is not limited to the framework and process.

[0118] Figure 5 is another example of the video encoder provided by this application.

[0119] like Figure 5 As shown, the original image signal s k [x,y] and predicted image signal Perform difference operation to obtain the residual signal u k [x,y], residual signal u k [x,y] is transformed and quantized to obtain the quantized coefficients. The quantized coefficients are entropy coded to obtain the encoded bit stream, and the reconstructed residual signal u' is obtained by inverse quantization and inverse transformation. k [x,y], predicted image signal and reconstructed residual signal u' k [x,y] superposition generates reconstructed image signal For reconstructing the image signal On the one hand, it is input to the intra-frame mode decision module and the intra-frame prediction module for intra-frame prediction processing, and on the other hand, it is filtered through loop filtering and outputs the filtered image signal s' k [x,y], filtered image signal s' k [x,y] can be used as a reference image for the next frame to perform motion estimation and motion compensation prediction for the next frame. Further, the result s' can be used based on the motion compensation prediction. r [x+m x ,y+m y ] and intra prediction results Predict the next frame and get the predicted image signal of the next frame And continue to repeat the above process until the encoding is completed.

[0120] To facilitate understanding of the technical solution provided by this application, the relevant contents are explained below.

[0121] (1) An overview of video coding technology.

[0122] From the perspective of signal acquisition, video signals can be divided into two types: those captured by cameras and those generated by computers. Due to different statistical characteristics, the corresponding compression encoding methods may also be different.

[0123] Taking the international video coding standards HEVC, VVC, and China's national video coding standard AVS as examples, a hybrid coding framework is adopted. The hybrid coding framework can perform the following series of operations and processing on the input original video signal:

[0124] 1. Block partition structure.

[0125] The input image is divided into several non-overlapping processing units according to the size of a, and each processing unit will perform similar compression operations. This processing unit is called CTU, or LCU. Below CTU, it can be further divided into more detailed divisions to obtain one or more basic coding units, called CU. Each CU is the most basic element in the encoding process. The following describes the various encoding methods that may be used for each CU.

[0126] 2. Predictive Coding

[0127] Predictive coding includes intra-frame prediction and inter-frame prediction. The original video signal is predicted by the selected reconstructed video signal to obtain a residual video signal. The encoder selects the most suitable one among many possible predictive coding modes for the current CU and informs the decoder.

[0128] a. Intra-frame prediction: The predicted signal comes from the area that has been coded and reconstructed in the same image. The frame encoded using intra-frame prediction is an I-frame. I-frame is also called intra-coded frame.

[0129] b. Inter-frame prediction: The predicted signal comes from other images that have been encoded and are different from the current image (called reference images). The frames encoded using inter-frame prediction are P frames and B frames. In other words, P frames refer to images that are encoded using only unidirectional reference frames, while B frames refer to images that are encoded using bidirectional reference frames (forward and backward). P frames are also called forward predictive-frames or inter-frame predictive-frames, and B frames are also called bi-directional interpolated prediction frames or bi-directional predictive-frames.

[0130] 3. Transform coding and quantization (Transform&Quantization).

[0131] The residual video signal is transformed into the transform domain through discrete Fourier transform (DFT), discrete cosine transform (DCT) and other transform operations, which are called transform coefficients. The signal in the transform domain is further subjected to lossy quantization operations, which discards certain information, making the quantized signal conducive to compression expression. In some video coding standards, there may be more than one transform mode to choose from, so the encoder also needs to select one of the transforms for the current coding CU and inform the decoder. The degree of quantization is usually determined by the quantization parameter (QP). A larger QP value means that coefficients with a larger value range will be quantized to the same output, which usually results in greater distortion and a lower bit rate. On the contrary, a smaller QP value means that coefficients with a smaller value range will be quantized to the same output, which usually results in less distortion and a higher bit rate.

[0132] 4. Entropy Coding or Statistical Coding.

[0133] The quantized transform domain signal will be statistically compressed and encoded according to the frequency of occurrence of each value, and finally output as a binary (0 or 1) compressed bit stream. At the same time, the encoding generates other information, such as the selected mode, motion vector, etc., which also needs to be entropy encoded to reduce the bit rate. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Statistical coding methods usually include: variable length coding (VLC), context-based binarization (i.e. image binarization) arithmetic coding (CABAC).

[0134] 5. In-Loop Filtering.

[0135] The encoded image can be reconstructed into a decoded image after inverse quantization, inverse transformation and prediction compensation. Compared with the original image, the reconstructed image has some information different from the original image due to the influence of quantization, and there is distortion. Filtering operations are performed on the reconstructed image, such as deblocking filter (DBF), sample adaptive offset (SAO) or adaptive loop filter (ALF), which can effectively reduce the distortion caused by quantization. Since these filtered reconstructed images will be stored in the decoder picture buffer (DPB) as a reference for subsequent encoded images and used to predict future signals, the above filtering operation is also called loop filtering, that is, filtering operations within the encoding loop.

[0136] According to the above encoding process, for each CU, after the decoder obtains the compressed code stream, the decoding end first performs entropy decoding to obtain various mode information and quantized transform coefficients. Each coefficient is inversely quantized and inversely transformed to obtain a residual signal. On the other hand, based on the known encoding mode information, the prediction signal corresponding to the CU can be obtained. After adding the two, the reconstructed signal can be obtained. Finally, the reconstructed value of the decoded image needs to undergo a loop filtering operation to generate the final output signal. In this series of encoding processes, the encoding framework mainly evaluates and selects the optimal encoding parameters and results based on the rate-distortion optimization criterion (RDO).

[0137] (2) Coding tree unit (CTU).

[0138] Each frame of the video is often divided into units of a certain size before the subsequent encoding and decoding process.

[0139] Figure 6 It is a schematic structural diagram of the CTU provided in this application.

[0140] like Figure 6As shown, CTU is the basic coding unit in the hybrid coding framework, which usually includes two parts: luminance Y and chrominance UV. That is, each CTU is composed of a luminance (Luma) coding tree block (Coding Tree Block, CTB), two chrominance (Chroma) coding tree blocks and syntax elements (Syntax Element). For example, taking an image with a YUV ratio of 4:2:0 as an example, the image is divided into multiple CTUs. Assuming that the size of CTU is 16×16, it contains 1 16×16 luminance coding tree block and 2 8×8 chrominance coding tree blocks. In other words, since the image adopts a 4:2:0 sampling method, the size of the luminance coding tree block is four times that of the chrominance coding tree block.

[0141] (3) Neural Network in-Loop Filter (NNLF).

[0142] Typically, the hybrid coding framework uses a traditional loop filter to suppress the distortion of the reconstructed image, improve the quality of the reconstructed image, and expects to restore the encoded reconstructed image to the original image. However, the loop filter is based on manual design, which is difficult to effectively reduce the distortion of the reconstructed image and there is a large room for optimization. Due to the excellent performance of deep learning tools in image processing, deep learning-based loop filters are applied to loop filters.

[0143] The present application mainly relates to the technology of filtering using NNLF.

[0144] Figure 7 It is a schematic diagram of the filtering principle of the NNLF involved in the embodiment of the present application.

[0145] like Figure 7 As shown, by inputting the image to be filtered before filtering into the trained NNLF, the enhanced image after filtering, that is, the filtered image, can be obtained.

[0146] NNLF usually uses a loss function to constrain the filtered image so that it can be restored to the original image as much as possible. The loss function measures the difference between the predicted value and the true value. The larger the loss value of NNLF, the greater the difference, and the goal of training is to reduce the loss value of NNLF. For encoding tools based on deep learning, commonly used loss functions are: L1 norm loss function, L2 norm loss function and smooth L1 loss function.

[0147] In addition, by considering the use of temporal information (such as reference images), a temporal NNLF (Temporal NNLF) can be constructed.

[0148] Figure 8It is a schematic diagram of a filtering principle of a time domain NNLF involved in an embodiment of the present application.

[0149] like Figure 8 As shown, the image to be filtered and N reference images (i.e., reference image 1 to reference image N) before filtering are input into the trained time domain NNLF, and the enhanced image after filtering, i.e., the filtered image, can be obtained. Wherein, N is a positive integer greater than or equal to 1.

[0150] (4) Image evaluation indicators.

[0151] Image evaluation indicators usually include Peak Signal-to-Noise Ratio (PSNR) and Multi Scale Structural Similarity Index Measure (MS-SSIM).

[0152] Among them, PNSR is a commonly used objective standard for evaluating images, which is used to calculate the visual error between two images. For example, in video coding, PSNR is used to measure the pixel difference between the compressed image and the original image. The larger the PNSR value, the better.

[0153] For example, given a denoised image I and a noisy image K of size m×n, the PSNR calculation formula is as follows:

[0154]

[0155] Among them, MAX represents the maximum pixel value in the image, and MSE is defined as:

[0156]

[0157] MS-SSIM is a widely used image quality evaluation standard that is used to measure the structural similarity of images, including differences in brightness, contrast, and structure. Its evaluation results are closer to the subjective quality evaluation results of the human eye.

[0158] For example, for samples x and y, the calculation formula of MS-SSIM is as follows:

[0159]

[0160] Among them, c j (x,y) and s j (x, y) represents the contrast measure and structure measure at different scales, and l(x, y) represents the brightness measure at the last scale. M , βj , γ j Used to adjust the importance of different components. M represents the number of scales.

[0161] The technical problem to be solved by this application is described below.

[0162] Typically, a reference image may be selected as input for the time domain NNLF to provide additional reference information for the image to be filtered.

[0163] Specifically, since there is a certain temporal correlation between consecutive images of a video, in video coding, the decoded image can be used as a reference image to provide a reference for the predictive coding of the current coded image. In neural network-based video coding, the image in the decoder picture buffer (DPB) can be input into the neural network filter to provide an effective reference for the quality enhancement of the reconstructed image to be filtered.

[0164] For example, a reference image may be selected from a reference image list that caches multiple decoded images. Specifically, for a P frame, the first image in the default order in its unidirectional reference image list may be used as a reference image; for a B frame, the first image in the default order in its forward reference image list and the first image in the backward reference image list may be used as a reference image.

[0165] However, due to the existence of motion relationship, there are often certain content differences between the reference image and the current image to be filtered. Therefore, although similar / identical areas in the reference image can provide effective filtering references for the image to be filtered, different areas in the reference image will introduce unnecessary interference or noise. In addition, reference images with higher quality levels (larger PSNR) can often provide more sufficient information for quality enhancement of the image to be filtered. Conversely, reference images with lower quality levels (smaller PSNR) are not conducive to filtering the image to be filtered. Since the reference images in the DPB are all compressed decoded images, different reference images have different quality levels (PSNR sizes). Therefore, directly / randomly selecting one or more images in the reference image list as the input of the neural network filter cannot obtain the optimal solution for time domain filtering.

[0166] In view of this, embodiments of the present application provide a filtering method, device, electronic device, and storage medium, which can improve the filtering effect.

[0167] Specifically, the embodiment of the present application provides an optimal reference image selection method based on considerations of the content characteristics and quality levels of the reference images, which can adaptively select the optimal reference image of the embodiment from a reference image list and a reference image set formed by the image to be filtered, and use it as the input of the neural network filter, thereby providing a more effective reference for quality enhancement of the image to be filtered, thereby improving the filtering effect of the image to be filtered and the performance of video encoding based on neural networks.

[0168] The filtering method provided by the present application is described in detail below.

[0169] It should be understood that the filtering method provided in the present application can be applied to video encoders, codecs, or video pre- and post-processing products.

[0170] For example, if the filtering method provided by the present application can be applied to a video encoder or a video decoder, a general neural network filter can be used as a loop filter (whose filter output affects video encoding and decoding).

[0171] If applied to a video coding scheme, when the encoder encodes the current image, the current image is first divided into coding blocks, and the coding blocks are used as coding units for block-by-block coding. For example, for the current block to be encoded in the current image, the prediction value of the current block is first obtained by inter-frame and / or intra-frame prediction. Then, based on the prediction value of the current block and the current block, the residual value of the current block is obtained. The encoder transforms the residual value of the current block to obtain a transformation coefficient. In one example, the encoder does not quantize the transformation coefficient of the current block, but directly encodes the transformation coefficient to obtain a code stream. In another example, the encoder quantizes the transformation coefficient of the current block to obtain a quantization coefficient, and then encodes the quantization coefficient to obtain a code stream. In addition, the encoder also dequantizes the transformation coefficient (dequantization may not be performed when quantization is not performed during the encoding process) and de-transforms it to obtain a residual value, and adds the residual value to the prediction value to obtain a reconstruction value of the current block. Based on the above steps, the reconstruction value of each coding block in the current image can be obtained, and these reconstruction values ​​constitute the reconstructed image of the current image. Then, in order to further improve the quality of the reconstructed image, the encoding end can filter the reconstructed image or the reconstructed image block in the reconstructed image using the filtering method provided by the present application to obtain a decoded image of the current image. Of course, the encoder can also store the decoded image in a decoding cache for prediction of subsequent images.

[0172] If applied to the video decoding scheme, for each block to be decoded in the current image, such as the current block, after the decoder obtains the code stream, it decodes the code stream to obtain the transform coefficient of the current block. Then, the decoder performs inverse quantization (inverse quantization may not be performed when quantization is not performed during the encoding process) and inverse transformation on the transform coefficient of the current block to obtain the residual value of the current block. At the same time, the decoder uses inter-frame and / or intra-frame prediction methods to predict the predicted value of the current block. In this way, the predicted value and the residual value of the current block are added to obtain the reconstructed value of the current block. Based on the above steps, the decoder can decode and determine the reconstructed value of each block to be decoded in the current image, and these reconstructed values ​​constitute the reconstructed image of the current image. Then, in order to further improve the quality of the reconstructed image, the decoder can use the filtering method provided by the present application to filter the reconstructed image or the reconstructed image block in the reconstructed image to obtain the decoded image of the current image. Of course, the decoder can store the decoded image in the decoding cache for prediction of subsequent images. In addition, the decoder can output the decoded image to a display device for display.

[0173] For another example, if the filtering method provided in the present application can be applied to a video post-processing product, filtering optimization can be performed on the decoded video (the filter output does not affect the video decoding).

[0174] In addition, the filtering method provided by the present application can also be used for any module using a neural network in a neural network-based video encoding, such as a neural network super-resolution module. In addition, the filtering method provided by the present application can be used not only in the reasoning process of a neural network-based video encoding, but also in the training process of a neural network filter. That is, the optimal one or more reference images can be used for training during the training process of the neural network filter, and the optimal one or more reference images can still be used for testing during the testing process to ensure the consistency of the training process and the testing process, thereby ensuring the filtering performance of the neural network filter.

[0175] Fig. 9 It is a schematic flow chart of the filtering method 400 provided in the present application.

[0176] It should be understood that the filtering method 400 may be performed by a filtering device. For example, the filtering method 400 may be performed by Figure 1 The filtering method 400 may be performed by the filtering device in the decoding device 120 or the decoder 122 shown in FIG. Figure 4 The filtering device in the video decoder 300 is shown to perform. For ease of description, the filtering method 400 is described below by taking the filtering device as an example.

[0177] like Fig. 9 As shown, the filtering method 400 may include some or all of the following:

[0178] S410, obtaining first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by decoding.

[0179] Exemplarily, the first unit to be filtered may be any unit of any size or type obtained by decoding the decoder and requiring filtering, which may include an image, an image block, an audio block, or other basic units of signal processing. These units are usually compressed and encoded during the encoding process, and the decoder needs to decode and decompress them, and then perform subsequent processing such as filtering.

[0180] Exemplarily, the first unit to be filtered may be an image to be filtered or an image block to be filtered. The image to be filtered may be a reconstructed image, and the image block to be filtered may be a partial area in the image to be filtered.

[0181] Fig.10 This is an example of an image block to be filtered provided in an embodiment of the present application.

[0182] like Fig.10 As shown, the image block to be filtered may include L CTUs (eg, L is a positive integer) or a certain area of ​​a random size. Fig.10 As shown in (a) in FIG. , the image block to be filtered may include 4 CTUs. Fig.10 As shown in (b) in FIG. 1 , the image block to be filtered may be a region whose upper left corner overlaps with the upper left corner of a CTU and whose size is larger than a CTU. Fig.10 As shown in (c) in FIG. 8 , the image block to be filtered may be a region whose upper right corner overlaps with the upper right corner of a CTU and whose size is larger than that of a CTU.

[0183] S420, when using a neural network filter to filter the first unit to be filtered in the first information to be filtered, determine at least one reference image in a reference image set; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs.

[0184] Exemplarily, if the image to be filtered to which the first unit to be filtered belongs is a B frame, the reference image list may include at least one of the following: a forward reference image list and a backward reference image list. If the image to be filtered to which the first unit to be filtered belongs is a P frame, the reference image list may include a unidirectional reference image list, which may also be referred to as a forward reference image list. Wherein, when the image to be filtered to which the first unit to be filtered belongs is a B frame, the reference image list includes a forward reference image list, which may be the same list as the forward reference image list when the image to be filtered to which the first unit to be filtered belongs is a P frame, or may be independent lists, and the present application does not limit this.

[0185] Exemplarily, when the reference image set includes the image to be filtered, it indicates that the image to be filtered can be used as the reference image of the first reference unit.

[0186] Exemplarily, the number of the at least one reference image may be a value agreed upon by a protocol, or may be a value notified to the decoding end by the encoding end (for example, through one or more flags).

[0187] S430, based on the first information to be filtered and the at least one reference image, use the neural network filter to filter the first unit to be filtered, and obtain a filtered unit of the first unit to be filtered.

[0188] Exemplarily, the filtering device may use the first information to be filtered and a reference unit in the at least one reference image that matches the first filtering unit as inputs of the neural network filter, and output a filtered unit of the first unit to be filtered.

[0189] In this embodiment, when the filtering device uses a neural network filter to filter the first unit to be filtered in the first information to be filtered, it is necessary to first determine at least one reference image in the reference image set, and then use the neural network filter to filter the first unit to be filtered based on the first information to be filtered and the at least one reference image, and obtain the filtered unit of the first unit to be filtered. Since the reference image set includes at least one of the following: the reference image in the reference image list of the first unit to be filtered, the image to be filtered to which the first unit to be filtered belongs, and the reference image in the reference image list of the first unit to be filtered can provide the first unit to be filtered with information in the time domain, and the image to be filtered to which the first unit to be filtered belongs can provide the first filtering unit with information in the spatial domain, therefore, based on the first information to be filtered and the at least one reference image, the first unit to be filtered is filtered using the neural network filter, which is equivalent to using at least one of the time domain information and the spatial domain information to assist the neural network filter in filtering the first unit to be filtered, that is, by enriching the selection range of the reference image, the filtering performance of the neural network filter is improved, thereby improving the filtering effect of the image to be filtered.

[0190] In some embodiments, when the at least one reference image constitutes a first reference image combination, the first reference image combination is a reference image combination with the smallest filtering cost among multiple reference image combinations, and the multiple reference image combinations include a reference image combination obtained by combining reference images in the reference image set.

[0191] Exemplarily, the multiple reference image combinations include reference image combinations obtained by combining all reference images in the reference image set. For example, the multiple reference image combinations may include all reference image combinations obtained by combining all reference images in the reference image set. For another example, the multiple reference image combinations may include some reference image combinations among all reference image combinations obtained by combining all reference images in the reference image set.

[0192] Exemplarily, the multiple reference image combinations include reference image combinations obtained by combining some of the reference images ranked top in the reference image set. For example, the multiple reference image combinations include all reference image combinations obtained by combining some of the reference images ranked top in the reference image set. For another example, the multiple reference image combinations include some reference combinations among all reference image combinations obtained by combining some of the reference images ranked top in the reference image set.

[0193] Of course, no matter what form the multiple reference image combinations are in, the encoder and decoder need to have the same understanding of the multiple reference image combinations. For example, the multiple reference image combinations may be combinations agreed upon by a protocol, or the multiple reference image combinations may be information that the encoder informs the decoder (for example, through one or more flags).

[0194] In this embodiment, the reference image included in the reference image combination with the smallest filtering cost among multiple reference image combinations is used as the at least one reference image, which is equivalent to selecting the optimal reference image combination from the reference image set as the input of the neural network filter, so that the neural network filter can better exert its generalization, improve the filtering performance of the neural network filter, and further improve the filtering effect of the image to be filtered.

[0195] In some embodiments, the filtering cost is determined based on at least one of the following indicators: image evaluation indicators generally include peak signal-to-noise ratio (PSNR), multi-scale structural similarity (MS-SSIM), and rate distortion optimization (RDO) cost.

[0196] In this embodiment, the filtering cost is determined by PSNR or MS-SSIM, which is equivalent to taking into account the differences in content characteristics and quality levels of different reference images. The filtering cost is determined by RDO, which is equivalent to being able to adaptively select the optimal reference image group from the reference image set as the input of the neural network filter; thereby, the neural network filter can better exert its generalization, improve the filtering performance of the neural network filter, and further improve the filtering effect of the image to be filtered.

[0197] Exemplarily, the filtering cost is an RDO cost.

[0198] Fig.11 This is an example of at least one reference image provided in an embodiment of the present application.

[0199] like Fig.11 As shown, assuming that the first filtering unit is an image to be filtered, its frame type is a B frame, and the reference image set includes a forward reference image list and a backward reference image list of the image to be filtered. If the at least one reference image is two reference images, the two reference images constitute a first reference image combination as a reference image combination with the minimum RDO cost among multiple reference image combinations. Among them, the multiple reference image combinations include reference image combination 1 composed of forward reference image 1 in the forward reference image list and backward reference image 1 in the backward reference image list, reference image combination 2 composed of forward reference image 2 in the forward reference image list and backward reference image 2 in the backward reference image list, and reference image combination 3 composed of forward reference image 1 in the forward reference image list and backward reference image 2 in the backward reference image list, that is, the at least one reference image can include reference images included in reference image combination 1, reference image combination 2, and reference image combination 3 with the minimum RDO cost. Specifically, the encoder uses reference image combination 1, reference image combination 2 and reference image combination 3 as reference images of the image to be filtered, respectively, to obtain three filtered images (i.e., filtered image 1, filtered image 2 and filtered image 3), then calculates the RDO cost of filtered image 1, the RDO cost of filtered image 2 and the RDO cost of filtered image 3, respectively, and determines the reference image combination used by the filtered image with the smallest RDO cost among the three filtered images as the optimal reference image combination, and determines the reference image in the optimal reference image combination as the at least one reference image.

[0200] Fig.12 is another example of at least one reference image provided by an embodiment of the present application.

[0201] like Fig.12As shown, it is assumed that the first filtering unit is an image to be filtered, and its frame type is a P frame, and the reference image set includes a unidirectional reference image list of the image to be filtered. If the at least one reference image is a reference image, then this reference image is the reference image with the smallest RDO cost among the top three reference images (i.e., unidirectional reference image 1, unidirectional reference image 2, and unidirectional reference image 3) in the unidirectional reference image list. Specifically, the encoder uses unidirectional reference image 1, unidirectional reference image 2, and unidirectional reference image 3 as reference images of the image to be filtered, respectively, to obtain three filtered images (i.e., filtered image 1, filtered image 2, and filtered image 3), then calculates the RDO cost of filtered image 1, the RDO cost of filtered image 2, and the RDO cost of filtered image 3, respectively, and determines the reference image used by the filtered image with the smallest RDO cost among the three filtered images as the optimal reference image, and determines the optimal reference image as the at least one reference image.

[0202] Of course, in other alternative embodiments, Fig.12 and Fig.13 The RDO cost in may also be replaced by other types of costs, for example, a cost determined based on at least one of PSNR and MS-SSIM, which is not limited in the present application.

[0203] Exemplarily, the filtering cost is determined based on PSNR and MS-SSIM.

[0204] In this embodiment, the filtering cost is determined based on PSNR and MS-SSIM, that is, the filtering cost of each reference image combination can be determined based on PSNR and MS-SSIM, and then the optimal reference image combination can be determined based on the filtering cost of each reference image combination. In this way, the complex RDO screening process to determine the optimal reference image combination can be skipped, thereby improving the filtering efficiency.

[0205] Exemplarily, the filtering cost may be determined by the following formula:

[0206] Score i =PSNR i +λ*MS-SSIM i .

[0207] Where i represents the i-th reference image in the reference image set, PSNR i represents the quality level of the i-th reference image in the reference image set, MS-SSIM represents the content difference between the i-th reference image in the reference image set and the image to be filtered to which the first filtering unit belongs, and λ is the balanced PSNR i and MS-SSIM i The weight factor of .

[0208] Of course, in other alternative embodiments, PSNR or MS-SSIM may also be replaced by other evaluation indicators, for example, including but not limited to: Sum of Absolute Differences (SAD), Sum of Absolute Transformed Difference (SATD), Sum of Squares for Error (SSE), Mean Square Error (MSE). The filtering cost may also be determined only by PSNR or MS-SSIM or other indicators, which is not limited in this application.

[0209] In some embodiments, S420 may include:

[0210] At least one first flag is obtained by decoding the code stream; wherein the at least one first flag is used to indicate the at least one reference image in the reference image set.

[0211] Exemplarily, when the filtering device uses the at least one reference image by default to filter the first unit to be filtered, it can obtain the at least one first flag by decoding the code stream; and then determine the reference image indicated by the at least one first flag as the at least one reference image.

[0212] Exemplarily, the value of the at least one first flag may be an index or subscript of the at least one reference image in the reference image set.

[0213] Exemplarily, the first flag is used to indicate one or more reference images in the reference image set. The first flag is used to indicate a reference image in the reference image set, indicating that the at least one first flag corresponds to the at least one reference image in a one-to-one relationship. The first flag is used to indicate multiple reference images in the reference image set, indicating that the at least one first flag and the at least one reference image are in a one-to-many relationship.

[0214] From the perspective of encoding and decoding, the encoder can sequentially obtain each reference image combination from the reference image set as additional input information, and send it to the neural network filter to generate the corresponding filtering result. Then, the filtering cost (such as rate-distortion cost) of each reference image combination is determined, and the reference image combination with the smallest filtering cost is selected based on the filtering cost of each reference image combination, which can be used as the optimal reference image combination, that is, the at least one reference image. Finally, the encoder can indicate the at least one reference image in the reference image set through the at least one first flag, and write the at least one first flag into the bitstream, so that the decoder can obtain the at least one reference image by decoding the at least one first flag.

[0215] Exemplarily, assuming that the first filtering unit is an image to be filtered, its frame type is a B frame, the reference image set includes a forward reference image list and a backward reference image list of the image to be filtered, and the reference image input to the neural network filter includes a reference image. In this case, the at least one first flag is a first flag, and its possible values ​​are shown in the following table:

[0216] Table 1

[0217]

[0218] As shown in Table 1, assuming that the value of the first flag is 00, the forward reference image 1 can be determined as the at least one reference image, that is, the reference image input to the neural network filter only includes the forward reference image 1. Similarly, assuming that the value of the first flag is 11, the backward reference image 2 can be determined as the at least one reference image, that is, the reference image input to the neural network filter only includes the backward reference image 2. Of course, the values ​​of the first flag in the above Table 1 and the corresponding relationship between the values ​​and the reference images are only exemplary and should not be understood as a limitation on the present application.

[0219] Exemplarily, assuming that the first filtering unit is an image to be filtered, whose frame type is a B frame, the reference image set includes a forward reference image list and a backward reference image list of the image to be filtered, and the reference image input to the neural network filter includes two reference images. In this case, the at least one first flag is two first flags, and its possible values ​​are shown in the following table:

[0220] Table 2

[0221]

[0222] As shown in Table 2, assuming that the value of the first first flag is 00 and the value of the second first flag is 01, the forward reference image 1 and the forward reference image 2 can be determined as the at least one reference image, that is, the reference image input to the neural network filter includes the forward reference image 1 and the forward reference image 2. Similarly, assuming that the value of the first first flag is 01 and the value of the second first flag is 10, the forward reference image 2 and the backward reference image 1 can be determined as the at least one reference image, that is, the reference image input to the neural network filter includes the forward reference image 2 and the backward reference image 1. Of course, the values ​​of the first flags in the above Table 2 and the corresponding relationship between the values ​​and the reference images are only exemplary and should not be understood as limitations on the present application.

[0223] Exemplarily, assuming that the first filtering unit is an image to be filtered, its frame type is a P frame, the reference image set includes a forward reference image list of the image to be filtered, and the reference image input to the neural network filter includes a reference image. In this case, the at least one first flag is a first flag, and its possible values ​​are shown in the following table:

[0224] Table 3

[0225]

[0226] As shown in Table 3, assuming that the value of the first flag is 00, the forward reference image 1 can be determined as the at least one reference image, that is, the reference image input to the neural network filter only includes the forward reference image 1. Similarly, assuming that the value of the first flag is 01, the backward reference image 2 can be determined as the at least one reference image, that is, the reference image input to the neural network filter only includes the backward reference image 2. Of course, the value of the first flag in the above Table 3 and the corresponding relationship between the value and the reference image are only exemplary and should not be understood as a limitation on the present application.

[0227] Exemplarily, assuming that the first filtering unit is an image to be filtered, its frame type is a P frame, the reference image set includes a forward reference image list of the image to be filtered, and the reference image input to the neural network filter includes two reference images. In this case, the at least one first flag is two first flags, and its possible values ​​are shown in the following table:

[0228] Table 4

[0229]

[0230] As shown in Table 4, assuming that the value of the first first flag is 00 and the value of the second first flag is 01, forward reference image 1 and forward reference image 2 can be determined as the at least one reference image, that is, the reference image input to the neural network filter includes forward reference image 1 and forward reference image 2. Similarly, assuming that the value of the first first flag is 01 and the value of the second first flag is 100, forward reference image 2 and forward reference image 3 can be determined as the at least one reference image, that is, the reference image input to the neural network filter includes forward reference image 2 and forward reference image 3. Of course, the values ​​of the first flags in the above Table 4 and the corresponding relationship between the values ​​and the reference images are only exemplary and should not be understood as limitations on the present application.

[0231] In some embodiments, when the first unit to be filtered is an image to be filtered, the i-th first mark in the at least one first mark is used to indicate the i-th reference image of the image to be filtered in the reference image set.

[0232] Exemplarily, when the first unit to be filtered is an image to be filtered, the i-th first flag among the at least one first flag is used to indicate the i-th reference image of the image to be filtered in the reference image set; equivalently, the at least one first flag is a frame-level flag (also called an image-level flag), which can simplify the complexity of the at least one first flag, thereby improving decoding efficiency.

[0233] In some embodiments, when the first unit to be filtered is an image block to be filtered, the i-th first flag in the at least one first flag is used to indicate in the reference image set the reference image to which the i-th reference image block of the image block to be filtered belongs.

[0234] Exemplarily, when the first unit to be filtered is an image block to be filtered, the i-th first flag among the at least one first flag is used to indicate the reference image to which the i-th reference image block of the image block to be filtered belongs in the reference image set; this is equivalent to the at least one first flag being a block-level flag, and indicating image-level information through the block-level flag, thereby avoiding directly indicating the reference image block of the image block to be filtered, simplifying the complexity of the at least one first flag, and thereby improving decoding efficiency.

[0235] Of course, whether the i-th first mark in the at least one first mark is used to indicate the i-th reference image of the image to be filtered in the reference image set, or the i-th first mark in the at least one first mark is used to indicate the reference image to which the i-th reference image block of the image block to be filtered belongs in the reference image set, the at least one first mark and the at least one reference image are in a one-to-one correspondence, but the present application is not limited thereto. For example, in other alternative embodiments, the at least one first mark and the at least one reference image are in a one-to-many relationship. In one example, the at least one first mark may also be a first mark, and the first mark is used to indicate a first reference image combination. At this time, the filtering device may determine the reference image in the first reference image combination indicated by the first mark as the at least one reference image. In another example, part of the first marks in the at least one first mark may be used to indicate multiple reference images in the reference image set, and another part of the first marks in the at least one first mark may be used to indicate a reference image in the reference image set. The specific implementation may be adjusted based on actual needs, and the present application is not limited thereto.

[0236] In some embodiments, the filtering device obtaining at least one first flag may be implemented as follows:

[0237] The second flag is obtained by decoding the code stream; if the second flag indicates that the first unit to be filtered is filtered using a reference image, the at least one first flag is obtained by decoding the code stream.

[0238] Exemplarily, the filtering device first analyzes the second flag to indicate whether to select the reference image as the input of the neural network filter. If it is determined to select the reference image as the input, the filtering device continues to analyze at least one subsequent first flag to indicate which image or images are selected as the reference image.

[0239] From the perspective of encoding and decoding, the encoder can sequentially obtain each reference image combination from the reference image set as additional input information, and send it to the neural network filter to generate the corresponding filtering result. Then, the filtering cost of each reference image combination is determined, and the reference image combination with the minimum filtering cost is selected based on the filtering cost of each reference image combination, and it is determined whether to use the reference image to filter the first unit to be filtered, and the second flag used to indicate whether the reference image is used to filter the first unit to be filtered is written into the code stream. If it is determined that the reference image is used to filter the first unit to be filtered, the reference image combination with the minimum filtering cost can be selected from the filtering costs of each reference image combination as the optimal reference image combination, that is, the at least one reference image. Further, the encoder can write the at least one first flag into the code stream to indicate the reference image used in the reference image set, that is, the at least one reference image. The decoder can determine whether to use the reference image to filter the first unit to be filtered by decoding the second flag, and when it is determined that the reference image is used to filter the first unit to be filtered, the at least one first flag is decoded to obtain the at least one reference image.

[0240] Exemplarily, assuming that the first filtering unit is an image to be filtered, whose frame type is a B frame, the reference image set includes a forward reference image list and a backward reference image list of the image to be filtered, and the reference image input to the neural network filter includes a reference image. In this case, the at least one first flag is a first flag, and the possible values ​​of the second flag and the at least one first flag (when the at least one first flag exists or the reference image needs to be used to filter the image to be filtered) are shown in the following table:

[0241] Table 5

[0242]

[0243] As shown in Table 5, the decoding device first decodes the second flag, and if the value of the second flag is 1, it continues to decode a subsequent first flag to indicate which specific image is selected as the reference image. Assuming that the value of this first flag is 00, the forward reference image 1 can be determined as the at least one reference image, that is, the reference image input to the neural network filter only includes the forward reference image 1. Similarly, assuming that the value of this first flag is 11, the backward reference image 2 can be determined as the at least one reference image, that is, the reference image input to the neural network filter only includes the backward reference image 2. Of course, the values ​​of the first flag in the above Table 5 and the correspondence between the values ​​and the reference images are only exemplary and should not be understood as limitations on the present application.

[0244] Exemplarily, assuming that the first filtering unit is an image to be filtered, whose frame type is a B frame, the reference image set includes a forward reference image list and a backward reference image list of the image to be filtered, and the reference image input to the neural network filter includes two reference images. In this case, the at least one first flag is two first flags, and the possible values ​​of the second flag and the at least one first flag (when the at least one first flag exists or the reference image needs to be used to filter the image to be filtered) are shown in the following table:

[0245] Table 6

[0246]

[0247] As shown in Table 6, the decoding device first decodes the second flag, and if the value of the second flag is 1, it continues to decode the subsequent two first flags to indicate which two images are specifically selected as reference images. Assuming that the value of the first first flag is 00 and the value of the second first flag is 01, the forward reference image 1 and the forward reference image 2 can be determined as the at least one reference image, that is, the reference image input to the neural network filter includes the forward reference image 1 and the forward reference image 2. Similarly, assuming that the value of the first first flag is 01 and the value of the second first flag is 10, the forward reference image 2 and the backward reference image 1 can be determined as the at least one reference image, that is, the reference image input to the neural network filter includes the forward reference image 2 and the backward reference image 1. Of course, the values ​​of the first flags in the above Table 6 and the corresponding relationship between the values ​​and the reference images are only exemplary and should not be understood as limitations on the present application.

[0248] Exemplarily, assuming that the first filtering unit is an image to be filtered, whose frame type is a P frame, the reference image set includes a forward reference image list of the image to be filtered, and the reference image input to the neural network filter includes a reference image. In this case, the at least one first flag is a first flag, and the possible values ​​of the second flag and the at least one first flag (when the at least one first flag exists or the reference image needs to be used to filter the image to be filtered) are shown in the following table:

[0249] Table 7

[0250]

[0251] As shown in Table 7, the decoding device first decodes the second flag, and if the value of the second flag is 1, it continues to decode a subsequent first flag to indicate which specific image is selected as the reference image. Assuming that the value of this first flag is 00, the forward reference image 1 can be determined as the at least one reference image, that is, the reference image input to the neural network filter only includes the forward reference image 1. Similarly, assuming that the value of this first flag is 01, the backward reference image 2 can be determined as the at least one reference image, that is, the reference image input to the neural network filter only includes the backward reference image 2. Of course, the values ​​of the first flag in the above Table 7 and the correspondence between the values ​​and the reference images are only exemplary and should not be understood as limitations on the present application.

[0252] Exemplarily, assuming that the first filtering unit is an image to be filtered, its frame type is a P frame, the reference image set includes a forward reference image list of the image to be filtered, and the reference image input to the neural network filter includes two reference images. In this case, the at least one first flag is two first flags, and the possible values ​​of the second flag and the at least one first flag (when the at least one first flag exists or the reference image needs to be used to filter the image to be filtered) are shown in the following table:

[0253] Table 8

[0254]

[0255] As shown in Table 8, the decoding device first decodes the second flag, and if the value of the second flag is 1, it continues to decode the subsequent two first flags to indicate which two images are specifically selected as reference images. Assuming that the value of the first first flag is 00 and the value of the second first flag is 01, forward reference image 1 and forward reference image 2 can be determined as the at least one reference image, that is, the reference image input to the neural network filter includes forward reference image 1 and forward reference image 2. Similarly, assuming that the value of the first first flag is 01 and the value of the second first flag is 100, forward reference image 2 and forward reference image 3 can be determined as the at least one reference image, that is, the reference image input to the neural network filter includes forward reference image 2 and forward reference image 3. Of course, the values ​​of the first flags in the above Table 8 and the correspondence between the values ​​and the reference images are only exemplary and should not be understood as limitations on the present application.

[0256] Of course, in other alternative embodiments, the filtering device may select a reference image as the input of the neural network filter by default, that is, the filtering device may directly parse the at least one first flag to indicate which image or images are specifically selected as reference images.

[0257] From the perspective of encoding and decoding, the encoder can use the reference image by default to filter the first unit to be filtered, that is, the encoder sequentially obtains each reference image combination from the reference image set as additional input information, and sends it to the neural network filter to generate the corresponding filtering result. Then, the encoder determines the filtering cost of each reference image combination, and selects the reference image combination with the smallest filtering cost as the optimal reference image combination, that is, the at least one reference image. Furthermore, the encoder can write the at least one first flag into the bitstream to indicate the reference image to be used in the reference image set, that is, the at least one reference image. The decoder can use the reference image by default to filter the first unit to be filtered, that is, the at least one reference image can be directly obtained by decoding the at least one first flag.

[0258] Of course, in other alternative embodiments, the second flag and the at least one first flag can be combined into a third flag, for example, by decoding the code stream to obtain the third flag; if the third flag indicates whether to use the reference image to filter the first unit to be filtered. Further, when the third flag indicates to use the reference image to filter the first unit to be filtered, it is also used to indicate the at least one reference image, that is, to indicate the first reference image combination in multiple reference image combinations.

[0259] In some embodiments, when the number of the at least one reference image is greater than an integer of 1, the reference images in the at least one reference image are different from each other.

[0260] Exemplarily, when the number of the at least one reference image is N and N>1, the reference images in the at least one reference image are different from each other. In other words, when the number of the at least one reference image is N and N>1, the kth reference image in the at least one reference image is different from any reference image in the first k-1 reference images. For example, the last reference image in the at least one reference image is different from any reference image in the first N-1 reference images.

[0261] In this embodiment, when the number of the at least one reference image is greater than an integer of 1, the reference images in the at least one reference image are different from each other, so that the at least one reference image can provide sufficient reference information for the neural network filter to filter the first unit to be filtered as much as possible, thereby improving the filtering performance of the neural network filter and the filtering effect of the first unit to be filtered.

[0262] Of course, in other alternative embodiments, when the number of the at least one reference image is greater than an integer of 1, the reference images in the at least one reference image may also be partially identical or completely identical, and the present application does not make any specific limitation on this.

[0263] In some embodiments, S430 may include:

[0264] When the first unit to be filtered is an image to be filtered, the first information to be filtered and the at least one reference image are used as input, and the neural network filter is used to filter the first unit to be filtered to obtain a filtered unit of the first unit to be filtered; or, when the first unit to be filtered is an image block to be filtered, the first information to be filtered and at least one reference image block in the at least one reference image that matches the image block to be filtered are used as input, and the neural network filter is used to filter the first unit to be filtered to obtain a filtered unit of the first unit to be filtered.

[0265] Exemplarily, if the first unit to be filtered is an image to be filtered, when the filtering device uses the at least one reference image to filter the first image to be filtered, it can first select N reference images, and then use these N reference images and the first information to be filtered as input into the neural network filter, that is, the information input into the NNLF includes these N reference images and the first information to be filtered.

[0266] Exemplarily, if the first unit to be filtered is an image block to be filtered, when the filtering device uses the at least one reference image to filter the first image to be filtered, it can first select N reference images, and then determine N reference image blocks that match the image block to be filtered in the at least one reference image, and then use these N reference image blocks and the first information to be filtered as input to the neural network filter, that is, the information input into the NNLF includes these N reference images and the first information to be filtered.

[0267] In some embodiments, the reference image block that matches the image block to be filtered in any one of the at least one reference images includes at least one of the following: an image block in the any one reference image that has the same position as the first unit to be filtered, an image block in the any one reference image that is adjacent to the first unit to be filtered, and an image block in the any one reference image that is a preset distance away from the first unit to be filtered.

[0268] Exemplarily, the reference image block in any one of the reference images that matches the image block to be filtered may include: an image block in the any one of the reference images that has the same position as the first unit to be filtered, and the image block in the any one of the reference images that has the same position as the first unit to be filtered may refer to: an image block in the any one of the reference images that has the same position as the first unit to be filtered in the image to be filtered to which the first unit to be filtered belongs. In other words, the image block in the any one of the reference images that has the same position as the first unit to be filtered may refer to: a co-located reference block in the any one of the reference images. The co-located reference block may be an image block in the any one of the reference images that has the same position as the first unit to be filtered in the image to be filtered to which the first unit to be filtered belongs.

[0269] Exemplarily, the reference image blocks in any one of the reference images that match the image block to be filtered may include: image blocks in any one of the reference images that are adjacent to the first unit to be filtered, and the image blocks in any one of the reference images that are adjacent to the first unit to be filtered may refer to: image blocks in any one of the reference images that are adjacent to the co-located reference blocks of the first unit to be filtered. The co-located reference block may be an image block in any one of the reference images that has the same position as the first unit to be filtered in the image to be filtered to which the first unit to be filtered belongs. Image blocks in any one of the reference images that are adjacent to the first unit to be filtered include, but are not limited to: image blocks in any one of the reference images that have common points with the first unit to be filtered, image blocks in any one of the reference images that are co-linear with the first unit to be filtered, and image blocks in any one of the reference images that are coplanar with the first unit to be filtered.

[0270] Exemplarily, the reference image block in any one of the reference images that matches the image block to be filtered may include: an image block in any one of the reference images whose distance from the first unit to be filtered is a preset distance, and the image block in any one of the reference images whose distance from the first unit to be filtered is a preset distance may refer to: an image block in any one of the reference images whose distance from the co-located reference block of the first unit to be filtered is a preset distance. The co-located reference block may be an image block in any one of the reference images whose position is the same as the position of the first unit to be filtered in the image to be filtered to which the first unit to be filtered belongs. The preset distance may be a distance agreed upon by a protocol, or the preset distance may be information on a storage medium in a storage filtering device or a device in which the filtering device is located. The preset distance may be any value greater than zero, such as an integer. The distance from the first unit to be filtered in any one of the reference images may be a Euclidean distance, a Morton distance, a Hilbert distance, or the like.

[0271] In some embodiments, the first information to be filtered further includes at least one of the following:

[0272] The prediction unit corresponding to the first unit to be filtered, the quantization parameter of the image to be filtered to which the first unit to be filtered belongs, and the frame type of the image to be filtered to which the first unit to be filtered belongs.

[0273] Exemplarily, the prediction unit corresponding to the first unit to be filtered may be a predicted image used to determine the first unit to be filtered. For example, when the first unit to be filtered is a reconstructed image to be filtered, and the reconstructed image to be filtered is a reconstructed image determined based on a predicted image of an original image and a residual image of the original image, the prediction unit corresponding to the first unit to be filtered may be the predicted image of the original image.

[0274] Exemplarily, the quantization parameter of the image to be filtered to which the first unit to be filtered belongs may include at least one of the following: a base quantization parameter (base QP), a slice-level quantization parameter (slice QP). Of course, the quantization parameter of the image to be filtered to which the first unit to be filtered belongs may also be a quantization parameter of other granularity, such as a block level. This application does not limit this.

[0275] Exemplarily, the frame type of the image to be filtered to which the first unit to be filtered belongs includes at least one of the following: P frame, B frame, I frame. The frame type of the image to be filtered to which the first unit to be filtered belongs can also be referred to as the slice type of the image to be filtered to which the first unit to be filtered belongs.

[0276] In some embodiments, S430 may include:

[0277] Based on the first information to be filtered and the at least one reference image, the neural network filter is used to filter at least one of the following items of the first unit to be filtered, and a filtered unit of the first unit to be filtered is obtained: the brightness component of the first unit to be filtered and the chrominance component of the first unit to be filtered.

[0278] In this embodiment, the at least one reference image can be used to filter at least one of the brightness component (Y) and the chrominance component (UV) of the first unit to be filtered, which is equivalent to using at least one of the time domain information and the spatial domain information to assist the neural network filter in filtering at least one of the brightness component (Y) and the chrominance component (UV) of the first unit to be filtered, that is, by enriching the selection range of reference images to improve the filtering performance of the neural network filter, thereby improving the filtering effect of the image to be filtered on each component.

[0279] The following will be combined Fig.13 The filtering method 500 according to the embodiment of the present application is described from the perspective of encoding.

[0280] It should be understood that the filtering method 500 may be performed by a filtering device in an encoder. Figure 1 The filtering device in the encoding device 110 or the encoder 112 shown in the figure is executed. For example, the filtering method 500 can be performed by Figure 2 The filtering device in the video encoder 200 is shown to perform.

[0281] like Fig.13 As shown, the filtering method 500 may include:

[0282] S510, obtaining first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by encoding;

[0283] S520, when filtering the first unit to be filtered in the first information to be filtered using a neural network filter, determining at least one reference image in a reference image set; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs;

[0284] S530, based on the first information to be filtered and the at least one reference image, use the neural network filter to filter the first unit to be filtered, and obtain a filtered unit of the first unit to be filtered.

[0285] In some embodiments, S520 may include:

[0286] Acquire a plurality of reference image combinations; the plurality of reference image combinations include reference image combinations obtained by combining reference images in the reference image set;

[0287] Determining a filtering cost using each reference image combination of the plurality of reference image combinations;

[0288] Based on the filtering cost of each reference image combination, the reference image included in the reference image combination with the smallest filtering cost is determined as the at least one reference image.

[0289] In some embodiments, the filtering cost is determined based on at least one of the following indicators: peak signal-to-noise ratio PSNR, multi-scale structural similarity MS-SSIM, and rate-distortion optimization RDO cost.

[0290] In some embodiments, the method 500 may further include:

[0291] encoding at least one first flag;

[0292] The at least one first flag is used to indicate the at least one reference image in the reference image set.

[0293] In some embodiments, when the first unit to be filtered is an image to be filtered, the i-th first flag in the at least one first flag is used to indicate the i-th reference image of the image to be filtered in the reference image list.

[0294] In some embodiments, when the first unit to be filtered is an image block to be filtered, the i-th first flag in the at least one first flag is used to indicate in the reference image list the reference image to which the i-th reference image block of the image block to be filtered belongs.

[0295] In some embodiments, the method 500 may further include:

[0296] Encode the second flag;

[0297] The second flag indicates that the first unit to be filtered is filtered using a reference image.

[0298] In some embodiments, when the number of the at least one reference image is greater than an integer of 1, the reference images in the at least one reference image are different from each other.

[0299] In some embodiments, S530 may include:

[0300] When the first unit to be filtered is an image to be filtered, the first information to be filtered and the at least one reference image are used as input, and the first unit to be filtered is filtered by using the neural network filter to obtain a filtered unit of the first unit to be filtered; or

[0301] When the first unit to be filtered is an image block to be filtered, the first information to be filtered and at least one reference image block in the at least one reference image that matches the image block to be filtered are used as input, and the first unit to be filtered is filtered using the neural network filter to obtain a filtered unit of the first unit to be filtered.

[0302] In some embodiments, the reference image block that matches the image block to be filtered in any one of the at least one reference images includes at least one of the following: an image block in the any one reference image that has the same position as the first unit to be filtered, an image block in the any one reference image that is adjacent to the first unit to be filtered, and an image block in the any one reference image that is a preset distance away from the first unit to be filtered.

[0303] In some embodiments, the first information to be filtered further includes at least one of the following:

[0304] The prediction unit corresponding to the first unit to be filtered, the quantization parameter of the image to be filtered to which the first unit to be filtered belongs, and the frame type of the image to be filtered to which the first unit to be filtered belongs.

[0305] In some embodiments, S530 may include:

[0306] Based on the first information to be filtered and the at least one reference image, the neural network filter is used to filter at least one of the following items of the first unit to be filtered, and a filtered unit of the first unit to be filtered is obtained:

[0307] The brightness component of the first unit to be filtered and the chrominance component of the first unit to be filtered.

[0308] It should be understood that filtering method 400 can be understood as the filtering method used in the decoding method, and filtering method 500 can be understood as the filtering method used in the encoding method. The specific principles are similar, and the difference is that filtering method 400 is used to filter the unit to be filtered obtained by decoding, and filtering method 500 is used to filter the unit to be filtered obtained by encoding. Therefore, the specific scheme of filtering method 500 can refer to the relevant content of filtering method 400. For the sake of ease of description, it will not be repeated here.

[0309] The preferred embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the embodiments mentioned above. Within the technical concept of the present application, the technical solution of the present application can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the present application. For example, the various specific technical features described in the specific embodiments mentioned above can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.

[0310] It should also be understood that in the various method embodiments of the present application, the size of the serial numbers of the processes involved above does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0311] The following is a detailed description of the device embodiments of the present application in conjunction with the accompanying drawings.

[0312] Fig.14 It is a schematic block diagram of the filtering device 600 provided in this application.

[0313] like Fig.14 As shown, the filtering device 600 may include:

[0314] An acquisition unit 610 is used to acquire first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by decoding;

[0315] The determining unit 620 is configured to determine at least one reference image in a reference image set when filtering the first unit to be filtered in the first information to be filtered using a neural network filter; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs;

[0316] The filtering unit 630 is used to filter the first unit to be filtered using the neural network filter based on the first information to be filtered and the at least one reference image, and obtain a filtered unit of the first unit to be filtered.

[0317] In some embodiments, when the at least one reference image constitutes a first reference image combination, the first reference image combination is a reference image combination with the smallest filtering cost among multiple reference image combinations, and the multiple reference image combinations include a reference image combination obtained by combining reference images in the reference image set.

[0318] In some embodiments, the filtering cost is determined based on at least one of the following indicators: peak signal-to-noise ratio PSNR, multi-scale structural similarity MS-SSIM, and rate-distortion optimization RDO cost.

[0319] In some embodiments, the determining unit 620 is specifically configured to:

[0320] At least one first flag is obtained by decoding the code stream; wherein the at least one first flag is used to indicate the at least one reference image in the reference image set.

[0321] In some embodiments, when the first unit to be filtered is an image to be filtered, the i-th first mark in the at least one first mark is used to indicate the i-th reference image of the image to be filtered in the reference image set.

[0322] In some embodiments, when the first unit to be filtered is an image block to be filtered, the i-th first flag in the at least one first flag is used to indicate in the reference image set the reference image to which the i-th reference image block of the image block to be filtered belongs.

[0323] In some embodiments, the determining unit 620 is specifically configured to:

[0324] The second flag is obtained by decoding the code stream; if the second flag indicates that the first unit to be filtered is filtered using a reference image, the at least one first flag is obtained by decoding the code stream.

[0325] In some embodiments, when the number of the at least one reference image is greater than an integer of 1, the reference images in the at least one reference image are different from each other.

[0326] In some embodiments, the filtering unit 630 is specifically used for:

[0327] When the first unit to be filtered is an image to be filtered, the first information to be filtered and the at least one reference image are used as input, and the neural network filter is used to filter the first unit to be filtered to obtain a filtered unit of the first unit to be filtered; or, when the first unit to be filtered is an image block to be filtered, the first information to be filtered and at least one reference image block in the at least one reference image that matches the image block to be filtered are used as input, and the neural network filter is used to filter the first unit to be filtered to obtain a filtered unit of the first unit to be filtered.

[0328] In some embodiments, the reference image block that matches the image block to be filtered in any one of the at least one reference images includes at least one of the following: an image block in the any one reference image that has the same position as the first unit to be filtered, an image block in the any one reference image that is adjacent to the first unit to be filtered, and an image block in the any one reference image that is a preset distance away from the first unit to be filtered.

[0329] In some embodiments, the first information to be filtered further includes at least one of the following:

[0330] The prediction unit corresponding to the first unit to be filtered, the quantization parameter of the image to be filtered to which the first unit to be filtered belongs, and the frame type of the image to be filtered to which the first unit to be filtered belongs.

[0331] In some embodiments, the filtering unit 630 is specifically used for:

[0332] Based on the first information to be filtered and the at least one reference image, the neural network filter is used to filter at least one of the following items of the first unit to be filtered, and a filtered unit of the first unit to be filtered is obtained: the brightness component of the first unit to be filtered and the chrominance component of the first unit to be filtered.

[0333] It should be understood that the device embodiment and the method embodiment of the filtering method may correspond to each other, and similar descriptions may refer to the method embodiment. Specifically, the filtering device 600 may correspond to the corresponding subject in the filtering method 400 of the embodiment of the present application, and the aforementioned and other operations and / or functions of each unit in the filtering device 600 are respectively for implementing the corresponding process in the filtering method 400. To avoid repetition, it will not be repeated here.

[0334] Fig.15 It is a schematic block diagram of the filtering device 700 provided in this application.

[0335] like Fig.15 As shown, the filtering device 700 may include:

[0336] An acquisition unit 710 is used to acquire first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by encoding;

[0337] The determining unit 720 is configured to determine at least one reference image in a reference image set when filtering the first unit to be filtered in the first information to be filtered using a neural network filter; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs;

[0338] The filtering unit 730 is used to filter the first unit to be filtered using the neural network filter based on the first information to be filtered and the at least one reference image, and obtain a filtered unit of the first unit to be filtered.

[0339] In some embodiments, the determining unit 720 is specifically configured to:

[0340] Acquire a plurality of reference image combinations; the plurality of reference image combinations include reference image combinations obtained by combining reference images in the reference image set;

[0341] Determining a filtering cost using each reference image combination of the plurality of reference image combinations;

[0342] Based on the filtering cost of each reference image combination, the reference image included in the reference image combination with the smallest filtering cost is determined as the at least one reference image.

[0343] In some embodiments, the filtering cost is determined based on at least one of the following indicators: peak signal-to-noise ratio PSNR, multi-scale structural similarity MS-SSIM, and rate-distortion optimization RDO cost.

[0344] In some embodiments, the determining unit 720 is further configured to:

[0345] encoding at least one first flag;

[0346] The at least one first flag is used to indicate the at least one reference image in the reference image set.

[0347] In some embodiments, when the first unit to be filtered is an image to be filtered, the i-th first flag in the at least one first flag is used to indicate the i-th reference image of the image to be filtered in the reference image list.

[0348] In some embodiments, when the first unit to be filtered is an image block to be filtered, the i-th first flag in the at least one first flag is used to indicate in the reference image list the reference image to which the i-th reference image block of the image block to be filtered belongs.

[0349] In some embodiments, the determining unit 720 is specifically configured to:

[0350] Encode the second flag;

[0351] The second flag indicates that the first unit to be filtered is filtered using a reference image.

[0352] In some embodiments, when the number of the at least one reference image is greater than an integer of 1, the reference images in the at least one reference image are different from each other.

[0353] In some embodiments, the filtering unit 730 is specifically used for:

[0354] When the first unit to be filtered is an image to be filtered, the first information to be filtered and the at least one reference image are used as input, and the first unit to be filtered is filtered by using the neural network filter to obtain a filtered unit of the first unit to be filtered; or

[0355] When the first unit to be filtered is an image block to be filtered, the first information to be filtered and at least one reference image block in the at least one reference image that matches the image block to be filtered are used as input, and the first unit to be filtered is filtered using the neural network filter to obtain a filtered unit of the first unit to be filtered.

[0356] In some embodiments, the reference image block that matches the image block to be filtered in any one of the at least one reference images includes at least one of the following: an image block in the any one reference image that has the same position as the first unit to be filtered, an image block in the any one reference image that is adjacent to the first unit to be filtered, and an image block in the any one reference image that is a preset distance away from the first unit to be filtered.

[0357] In some embodiments, the first information to be filtered further includes at least one of the following:

[0358] The prediction unit corresponding to the first unit to be filtered, the quantization parameter of the image to be filtered to which the first unit to be filtered belongs, and the frame type of the image to be filtered to which the first unit to be filtered belongs.

[0359] In some embodiments, the filtering unit 730 is specifically used for:

[0360] Based on the first information to be filtered and the at least one reference image, the neural network filter is used to filter at least one of the following items of the first unit to be filtered, and a filtered unit of the first unit to be filtered is obtained:

[0361] The brightness component of the first unit to be filtered and the chrominance component of the first unit to be filtered.

[0362] It should be understood that the device embodiment of the encoder and the method embodiment of the filtering method may correspond to each other, and similar descriptions may refer to the method embodiment. The filtering device 700 may correspond to the corresponding subject in the filtering method 500 of the embodiment of the present application, and the aforementioned and other operations and / or functions of each unit in the filtering device 700 are respectively to implement the corresponding processes in each method such as the filtering method 500. To avoid repetition, it will not be repeated here.

[0363] It should also be understood that the various units in the filter device 600 or filter device 700 involved in the embodiment of the present application are divided based on logical functions. In practical applications, the function of a unit can also be implemented by multiple units, or the functions of multiple units are implemented by one unit, and even these functions can also be implemented with the assistance of one or more other units. For example, part or all of the filter device 600 or filter device 700 are merged into one or several other units. For another example, a certain (some) unit in the filter device 600 or filter device 700 can also be split into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effect of the embodiment of the present application. For another example, the filter device 600 or filter device 700 can also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0364] According to another embodiment of the present application, the filtering device 600 or filtering device 700 involved in the embodiment of the present application can be constructed by running a computer program (including program code) capable of executing each step involved in the corresponding method on a general computing device of a general-purpose computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), and the filtering method of the embodiment of the present application can be implemented. Wherein, the computer program can be recorded on, for example, a computer-readable storage medium, and loaded into an electronic device through a computer-readable storage medium, so that the computer program executes the corresponding method of the embodiment of the present application when it is run in the electronic device.

[0365] In other words, the units mentioned above can be implemented in the form of hardware, can be implemented by software instructions, or can be implemented in the form of a combination of hardware and software.

[0366] Specifically, each step of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor, or a combination of hardware and software in the decoding processor. Optionally, the software can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The processor reads the information in the storage medium and completes the steps in the method embodiment mentioned above in combination with its hardware.

[0367] Fig.16 It is a schematic structural diagram of an electronic device 800 provided in this application.

[0368] like Fig.16As shown, the electronic device 800 at least includes a processor 810 and a computer-readable storage medium 820. The processor 810 and the computer-readable storage medium 820 may be connected via a bus or other means. The computer-readable storage medium 820 is used to store a computer program 821, which includes computer instructions, and the processor 810 is used to execute the computer instructions stored in the computer-readable storage medium 820. The processor 810 is the computing core and control core of the electronic device 800, which is suitable for implementing one or more computer instructions, and is specifically suitable for loading and executing one or more computer instructions to implement the corresponding method flow or corresponding function.

[0369] Exemplarily, the processor 810 may also be referred to as a central processing unit (CPU). The processor 810 may include, but is not limited to, a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, discrete hardware components, and the like.

[0370] Exemplarily, the computer-readable storage medium 820 may be a high-speed RAM memory, or a non-volatile memory (Non-VolatileMemory), such as at least one disk memory; optionally, it may also be at least one computer-readable storage medium located away from the aforementioned processor 810. Specifically, the computer-readable storage medium 820 includes, but is not limited to: a volatile memory and / or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0371] Exemplarily, the electronic device 800 may be a filtering device involved in an embodiment of the present application; the computer-readable storage medium 820 stores computer instructions; the processor 810 loads and executes the computer instructions stored in the computer-readable storage medium 820 to implement the corresponding steps in the filtering method provided in the present application; in other words, the computer instructions in the computer-readable storage medium 820 may be loaded by the processor 810 and the corresponding steps may be executed. To avoid repetition, they will not be repeated here.

[0372] According to another aspect of the present application, the present application further provides a coding method or an encoder, which can be encoded using the filtering method provided by the present application, and the encoder can be encoded using the filtering method provided by the present application.

[0373] According to another aspect of the present application, the present application also provides a decoding method or a decoder, which can use the filtering method provided by the present application for encoding, and the decoder can use the filtering method provided by the present application for decoding.

[0374] According to another aspect of the present application, the present application also provides a coding and decoding system, including the encoder and decoder mentioned above, the encoder can use the filtering method provided by the present application for encoding, and the decoder can use the filtering method provided by the present application for decoding.

[0375] According to another aspect of the present application, the present application also provides a computer-readable storage medium (Memory), which stores computer instructions. When the computer instructions are read and executed by a processor of a computer device, the computer device executes the filtering method involved above.

[0376] Among them, the computer-readable storage medium is a memory device in a decoder or encoder, which is used to store programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in electronic devices and, of course, extended storage media supported by electronic devices. The computer-readable storage medium can be used to provide storage space, and the storage space can store the operating system of the electronic device. In addition, one or more computer instructions suitable for being loaded and executed by the processor are also stored in the storage space, for example, one or more computer instructions for executing the filtering method involved above are stored, and these computer instructions can be one or more computer programs (including program codes).

[0377] According to another aspect of the present application, the present application also provides a computer program product or a computer program, the computer program product or the computer program includes a computer instruction, and the computer instruction is stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer executes the filtering method provided in the various optional modes mentioned above.

[0378] It should be understood that the computer device involved in the present application can be any device or apparatus capable of performing data processing, for example, including but not limited to: a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. In addition, the computer instructions involved in the present application can be stored in a computer-readable storage medium, or can be transmitted between one computer-readable storage medium and another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0379] According to another aspect of the present application, the present application further provides a code stream, which may be a code stream decoded using the filtering method provided by the present application or a code stream generated using the filtering method provided by the present application.

[0380] Those of ordinary skill in the art will appreciate that the units and process steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0381] Finally, it should be noted that the above content is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A filtering method, characterized in that: include: Acquire first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by encoding or decoding; When filtering the first unit to be filtered in the first information to be filtered using a neural network filter, at least one reference image in a reference image set is determined; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs; Based on the first information to be filtered and the at least one reference image, the first unit to be filtered is filtered using the neural network filter to obtain a filtered unit of the first unit to be filtered.

2. The method according to claim 1, characterized in that When the at least one reference image constitutes a first reference image combination, the first reference image combination is a reference image combination with the smallest filtering cost among multiple reference image combinations, and the multiple reference image combinations include reference image combinations obtained by combining reference images in the reference image set; the filtering cost is determined based on at least one of the following indicators: peak signal-to-noise ratio PSNR, multi-scale structural similarity MS-SSIM, and rate-distortion optimization RDO cost.

3. The method according to claim 1 or 2, characterized in that: The determining of at least one reference image comprises: Obtaining at least one first flag by decoding the code stream; The at least one first flag is used to indicate the at least one reference image in the reference image set; When the first unit to be filtered is an image to be filtered, the i-th first flag among the at least one first flag is used to indicate the i-th reference image of the image to be filtered in the reference image set; or, when the first unit to be filtered is an image block to be filtered, the i-th first flag among the at least one first flag is used to indicate the reference image to which the i-th reference image block of the image block to be filtered belongs in the reference image set.

4. The method according to claim 3, characterized in that The obtaining of at least one first mark comprises: By decoding the code stream, a second flag is obtained; If the second flag indicates that the first unit to be filtered is filtered using a reference image, the at least one first flag is obtained by decoding the code stream.

5. The method according to claim 1, characterized in that When the number of the at least one reference image is greater than an integer of 1, the reference images in the at least one reference image are different from each other.

6. The method according to claim 1 or 2, characterized in that: The filtering of the first unit to be filtered using the neural network filter based on the first information to be filtered and the at least one reference image, and obtaining a filtered unit of the first unit to be filtered, comprises: When the first unit to be filtered is an image to be filtered, taking the first information to be filtered and the at least one reference image as input, filtering the first unit to be filtered using the neural network filter to obtain a filtered unit of the first unit to be filtered; or When the first unit to be filtered is an image block to be filtered, the first information to be filtered and at least one reference image block in the at least one reference image that matches the image block to be filtered are taken as input, and the first unit to be filtered is filtered using the neural network filter to obtain a filtered unit of the first unit to be filtered.

7. The method according to claim 6, characterized in that The reference image block that matches the image block to be filtered in any one of the at least one reference images includes at least one of the following: an image block in any one of the reference images whose position is the same as that of the first unit to be filtered, an image block in any one of the reference images whose position is adjacent to the first unit to be filtered, and an image block in any one of the reference images whose distance from the first unit to be filtered is a preset distance.

8. The method according to claim 1 or 2, characterized in that: The first information to be filtered also includes at least one of the following: The prediction unit corresponding to the first unit to be filtered, the quantization parameter of the image to be filtered to which the first unit to be filtered belongs, and the frame type of the image to be filtered to which the first unit to be filtered belongs.

9. The method according to claim 1 or 2, characterized in that: The filtering of the first unit to be filtered using the neural network filter based on the first information to be filtered and the at least one reference image, and obtaining a filtered unit of the first unit to be filtered, comprises: Based on the first information to be filtered and the at least one reference image, the neural network filter is used to filter at least one of the following items of the first unit to be filtered, and a filtered unit of the first unit to be filtered is obtained: The luminance component of the first unit to be filtered and the chrominance component of the first unit to be filtered.

10. The method according to claim 1, characterized in that When the first information to be filtered includes the first unit to be filtered obtained by encoding, the determining at least one reference image in the reference image set includes: Acquire a plurality of reference image combinations; the plurality of reference image combinations include reference image combinations obtained by combining reference images in the reference image set; determining a filtering cost using each of the plurality of reference image combinations; Based on the filtering cost of each reference image combination, a reference image included in the reference image combination with the smallest filtering cost is determined as the at least one reference image.

11. A filtering device, characterized in that: include: An acquisition unit, configured to acquire first information to be filtered; the first information to be filtered includes a first unit to be filtered obtained by encoding or decoding; A determination unit, configured to determine at least one reference image in a reference image set when filtering a first unit to be filtered in the first information to be filtered using a neural network filter; the reference image set includes at least one of the following: a reference image in a reference image list of the first unit to be filtered, and an image to be filtered to which the first unit to be filtered belongs; A filtering unit is used to filter the first unit to be filtered using the neural network filter based on the first information to be filtered and the at least one reference image, and obtain a filtered unit of the first unit to be filtered.

12. An electronic device, characterized in that: include: a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein the computer program, when executed by the processor, implements: the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein the computer program causes a computer to execute: the method according to any one of claims 1 to 10.