Image filtering method and apparatus, and device
By aligning the block division scheme in the actual use process with the training process, the neural network-based filters achieve enhanced filtering performance and improved image quality.
Patent Information
- Application Number
- JP2024522042
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-05-18
- Filing Date
- 2023-03-01
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2043-03-01
AI Technical Summary
Neural network-based filters used for image filtering in video compression exhibit suboptimal performance due to mismatch between block division schemes in the training and actual use processes.
The image filtering method involves dividing the image into blocks using the same block division scheme as employed during the neural network-based filter's training process, ensuring consistent block division in both stages to enhance filtering performance.
This approach optimizes the filtering effect of neural network-based filters by aligning the block division in the actual use process with the training process, thereby improving image quality.
Smart Images

Figure 0007765159000001 
Figure 0007765159000002 
Figure 0007765159000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to a Chinese patent application filed on May 18, 2022, bearing application number 202210551039.7 and entitled "Image filtering method, apparatus, and device therefor," the entire contents of which are incorporated herein by reference.
[0002] The present embodiment relates to the field of image processing technology, and more particularly to an image filtering method and apparatus, and a device therefor. [Background technology]
[0003] Digital video technology can be incorporated into various video devices, such as digital televisions, smartphones, computers, e-book readers, or video players. With the development of video technology, the amount of data contained in video data increases. To facilitate the transmission of video data, video devices implement video compression techniques to transmit or store video data more efficiently.
[0004] In video compression technology, image loss occurs, and to reduce the loss, the reconstructed image is filtered. With the rapid development of neural network technology, in some scenarios, neural network-based filters are used to filter the reconstructed image. However, the filtering effect of some neural network-based filters is poor. Summary of the Invention
[0005] This application provides an image filtering method and its device and equipment, and its technical solutions are as follows:
[0006] In one aspect, the present application provides an image filtering method, the method comprising: obtaining an image to be filtered; determining a neural network based filter; Dividing the image to be filtered according to a block division scheme corresponding to the neural network-based filter to obtain N image blocks to be filtered (N is a positive integer), where the block division scheme is a block division scheme of a training image used by the neural network-based filter in a training process; and filtering each of the N image blocks to be filtered using the neural network-based filter to obtain a filtered image.
[0007] In another aspect, the present application provides an image filtering apparatus, the apparatus comprising: an acquisition unit configured to acquire a to-be-filtered image; a segmentation unit configured to determine a neural network-based filter, and segment the image to be filtered according to a block segmentation scheme corresponding to the neural network-based filter to obtain N image blocks to be filtered, where N is a positive integer, and the block segmentation scheme is a block segmentation scheme of a training image used by the neural network-based filter in a training process; a filtering unit configured to filter each of the N image blocks to be filtered using the neural network-based filter to obtain a filtered image.
[0008] In another aspect, there is provided an electronic device comprising: a memory configured to store a computer program; and a processor configured to execute a method according to an embodiment of each aspect of the present application by calling and executing the computer program stored in the memory.
[0009] In another aspect, a chip configured to implement the method of each aspect of the present application is provided, wherein the chip includes a processor configured to call and execute a computer program from a memory to cause a device incorporating the chip to perform the method according to the embodiment of each aspect of the present application.
[0010] In another aspect, a computer-readable storage medium is provided that is configured to store a computer program for causing a computer to perform a method according to an embodiment of each aspect of the present application.
[0011] In another aspect, a computer program product is provided that includes computer program instructions for causing a computer to perform a method according to an embodiment of each aspect of the present application.
[0012] In another aspect, there is provided a computer program which, when executed on a computer, causes the computer to carry out a method according to an embodiment of each aspect of the present application. [Effects of the Invention]
[0013] In summary, the present application obtains an image to be filtered, determines a neural network-based filter, divides the image to be filtered according to a block division scheme corresponding to the neural network-based filter to obtain N (N is a positive integer) image blocks to be filtered, and filters each of the N image blocks to be filtered using the neural network-based filter to obtain a filtered image. The above block division scheme is the block division scheme of the training image used by the neural network-based filter in the training process. That is, in the present embodiment, the block division scheme in the actual use process of the neural network-based filter is matched with the block division scheme used during training, so that the neural network filter can achieve optimal filtering performance and thereby improve the image filtering effect. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a schematic diagram of an application scenario according to an embodiment of the present invention. [Figure 2] 1 is an exemplary block diagram of a video codec system according to an embodiment of the present invention; [Figure 3] FIG. 1 is a schematic diagram of an encoding framework according to an embodiment of the present invention; [Figure 4] FIG. 1 is a schematic diagram of a decoding framework according to an embodiment of the present invention; [Figure 5] FIG. 1 is a schematic diagram of image types. [Figure 6] FIG. 1 is a schematic diagram of image types. [Figure 7] FIG. 1 is a schematic diagram of image types. [Figure 8] FIG. 1 is a schematic diagram of filtering according to an embodiment of the present invention. [Figure 9] This shows the block division method of an image. [Figure 10] This shows the block division method of an image. [Figure 11] This shows the block division method of an image. [Figure 12] This shows the block division method of an image. [Figure 13] 1 is a flowchart of an image filtering method according to one embodiment of the present application; [Figure 14] FIG. 2 is a schematic diagram of block division of an image according to an embodiment of the present invention. [Figure 15] FIG. 2 is a schematic diagram of block division of an image according to an embodiment of the present invention. [Figure 16] FIG. 2 is a schematic diagram of block division of an image according to an embodiment of the present invention. [Figure 17] FIG. 2 is a schematic diagram of block division of an image according to an embodiment of the present invention. [Figure 18] FIG. 10 is a schematic diagram of another image block division according to an embodiment of the present invention. [Figure 19] FIG. 10 is a schematic diagram of another image block division according to an embodiment of the present invention. [Figure 20] FIG. 10 is a schematic diagram of another image block division according to an embodiment of the present invention. [Figure 21] FIG. 10 is a schematic diagram of another image block division according to an embodiment of the present invention. [Figure 22] FIG. 10 is a schematic diagram of another image block division according to an embodiment of the present invention. [Figure 23] FIG. 10 is a schematic diagram of another image block division according to an embodiment of the present invention. [Figure 24] FIG. 10 is a schematic diagram of another image block division according to an embodiment of the present invention. [Figure 25] FIG. 10 is a schematic diagram of another image block division according to an embodiment of the present invention. [Figure 26] FIG. 2 is a schematic diagram of image block expansion according to an embodiment of the present invention; [Figure 27] FIG. 2 is a schematic diagram of image block expansion according to an embodiment of the present invention; [Figure 28] FIG. 2 is a schematic diagram of image block expansion according to an embodiment of the present invention; [Figure 29] FIG. 2 is a schematic diagram of image block expansion according to an embodiment of the present invention; [Figure 30] FIG. 10 is a schematic diagram of another image block expansion according to an embodiment of the present invention. [Figure 31] FIG. 10 is a schematic diagram of another image block expansion according to an embodiment of the present invention. [Figure 32] FIG. 10 is a schematic diagram of another image block expansion according to an embodiment of the present invention. [Figure 33] FIG. 10 is a schematic diagram of another image block expansion according to an embodiment of the present invention. [Figure 34] FIG. 2 is a schematic diagram of a spatial domain reference image block according to an embodiment of the present invention; [Figure 35] 2 is a schematic diagram of a spatial domain reference image block according to an embodiment of the present invention; [Figure 36] FIG. 10 is a schematic diagram of another spatial domain reference image block according to an embodiment of the present invention; [Figure 37] FIG. 10 is a schematic diagram of another spatial domain reference image block according to an embodiment of the present invention; [Figure 38] FIG. 10 is a schematic diagram of another spatial domain reference image block according to an embodiment of the present invention; [Figure 39] FIG. 10 is a schematic diagram of another spatial domain reference image block according to an embodiment of the present invention; [Figure 40] FIG. 10 is a schematic diagram of another spatial domain reference image block according to an embodiment of the present invention; [Figure 41] FIG. 10 is a schematic diagram of another spatial domain reference image block according to an embodiment of the present invention; [Figure 42] FIG. 2 is a schematic diagram of a time-domain reference image block according to an embodiment of the present invention; [Figure 43] FIG. 2 is a schematic diagram of a time-domain reference image block according to an embodiment of the present invention; [Figure 44] FIG. 10 is a schematic diagram of another time-domain reference image block according to an embodiment of the present invention; [Figure 45] FIG. 10 is a schematic diagram of another time-domain reference image block according to an embodiment of the present invention; [Figure 46] 1 is an exemplary flowchart of an image filtering method according to one embodiment of the present application; [Figure 47] 1 is an exemplary block diagram of an image filtering device according to one embodiment of the present application; [Figure 48] FIG. 1 is an exemplary block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0015] 1 is a schematic diagram of an application scenario according to an embodiment of the present invention, which includes an electronic device 100, which is equipped with a neural network-based filter 200. After acquiring an image to be filtered, the electronic device 100 inputs the image to be filtered into the neural network-based filter 200 for filtering.
[0016] In some embodiments, the electronic device 100 includes a display device, and thus the electronic device 100 can display the filtered image via the display device.
[0017] The present embodiment is not limited to a specific type of electronic device 100, and may be any device having a data processing function.
[0018] In some embodiments, electronic device 100 may be a terminal device, including, for example, a smartphone, a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a vehicle computer, etc.
[0019] In some embodiments, the electronic device 100 may be a server. The server may consist of one or more servers. When the server consists of multiple servers, there may be at least two servers for providing different services and / or at least two servers for providing the same service, for example, the same service may be provided in a load balancing manner, but the present embodiment is not limited thereto.
[0020] Optionally, the server may be an independent physical server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides fundamental cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data and artificial intelligence platforms, etc. The server may be a node of a blockchain.
[0021] In some embodiments, the server is a cloud server with powerful computing resources, and has highly virtualized and highly distributed characteristics.
[0022] In some embodiments, the image to be filtered is collected by an image collection device. For example, the image collection device transmits the collected image to be filtered to the electronic device 100, which then filters the captured image to be filtered using a neural network-based filter. In another example, the electronic device 100 has an image collection function, which allows the electronic device 100 to collect images and input the collected image to be filtered to a neural network-based filter for filtering.
[0023] In some embodiments, the electronic device 100 may be an encoding device, and the image to be filtered can be understood as a reconstructed image, which encodes the current image and then reconstructs it to obtain a reconstructed image, and then inputs the reconstructed image into a neural network-based filter for filtering.
[0024] In some embodiments, the electronic device 100 may be a decoding device, which performs image reconstruction after decoding a codestream to obtain a reconstructed image, and then inputs the reconstructed image into a neural network-based filter for filtering.
[0025] The present embodiments are applicable to any scenario where an image needs to be filtered.
[0026] In some embodiments, embodiments of the present disclosure may be applicable to various scenarios, including, but not limited to, cloud technology (e.g., cloud gaming), artificial intelligence, intelligent transportation, assisted driving, and the like.
[0027] In some embodiments, the present application can be applied to an image codec field, a video codec field, a hardware video codec field, a dedicated circuit video codec field, a real-time video codec field, etc. For example, the solution of the present application can be combined with an Audio Video coding Standard (AVS), such as the H.264 / Audio Video Coding (AVC) standard, the H.265 / High Efficiency Video Coding (HEVC) standard, and the H.266 / Versatile Video Coding (VVC) standard. Alternatively, the present solution can be combined with other proprietary or industry standards, including ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, and ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC, which includes the extensions Scalable Video Coding (SVC) and Multiview Video Coding (MVC)). It should be understood that the present technology is not limited to any particular codec standard or technology.
[0028] For ease of understanding, first, a video codec system according to an embodiment of the present invention will be introduced with reference to FIG.
[0029] FIG. 2 is an exemplary block diagram of a video codec system according to an embodiment of the present application. Note that FIG. 2 is merely an example, and video codec systems according to the embodiment of the present application include, but are not limited to, those shown in FIG. 2. As shown in FIG. 2, the video codec system includes an encoding device 110 and a decoding device 120. Here, the encoding device is configured to encode (or compress) video data to generate a codestream and transmit the codestream to a decoding device. The decoding device decodes the codestream generated by the encoding device to obtain decoded video data.
[0030] The encoding device 110 of the present embodiment may be understood as a device having a video encoding function, and the decoding device 120 may be understood as a device having a video decoding function, i.e., the encoding device 110 and the decoding device 120 of the present embodiment may include a wider range of devices, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, vehicle computers, etc.
[0031] In some embodiments, encoding device 110 may transmit encoded video data (e.g., a codestream, etc.) to decoding device 120 over channel 130. Channel 130 may include one or more media and / or devices that may transmit encoded video data from encoding device 110 to decoding device 120.
[0032] In one example, channel 130 includes one or more communication media that enable encoding device 110 to transmit encoded video data directly to decoding device 120 in real time. In this example, encoding device 110 may modulate the encoded video data in accordance with a communication standard and transmit the modulated video data to decoding device 120. Here, the communication media includes a communication medium such as a radio frequency spectrum, and optionally, the communication media may include a wired communication medium such as one or more physical transmission lines.
[0033] In another example, channel 130 includes a storage medium that can store the video data encoded by encoding device 110. The storage medium can include various locally accessible data storage media, such as optical disks, DVDs, flash memory, etc. In this example, decoding device 120 can obtain the encoded video data from the storage medium.
[0034] In another example, channel 130 may include a storage server, which may store video data encoded by encoding device 110. In this example, decoding device 120 may download the stored encoded video data from the storage server. Optionally, the storage server may store the encoded video data and transmit the encoded video data to decoding device 120 to a web server (e.g., for a website), a File Transfer Protocol (FTP) server, or the like.
[0035] In some embodiments, encoding device 110 comprises a video encoder 112 and an output interface 113, where output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0036] In some embodiments, encoding device 110 may further comprise a video source 111 in addition to video encoder 112 and input interface 113 .
[0037] The video source 111 may comprise at least one of a video collection device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system, where the video input interface is configured to receive video data from a video content provider and the computer graphics system is configured to generate the video data.
[0038] The video encoder 112 encodes video data from the video source 111 to generate a codestream. The video data may include one or more pictures or a sequence of pictures. The codestream contains coding information for the pictures or sequences of pictures in the form of a bitstream. The coding information may include coded image data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. An SPS may contain parameters that apply to one or more sequences. A PPS may contain parameters that apply to one or more pictures. A syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the codestream.
[0039] Video encoder 112 transmits the encoded video data directly to decoding device 120 via output interface 113. The encoded video data may also be stored on a storage medium or storage server so that it can be subsequently read by decoding device 120.
[0040] In some embodiments, the decoding device 120 comprises an input interface 121 and a video decoder 122 .
[0041] In some embodiments, decoding device 120 further comprises a display device 123 in addition to input interface 121 and video decoder 122 .
[0042] Here, the input interface 121 includes a receiver and / or a modem, and can receive encoded video data via a channel 130.
[0043] The video decoder 122 is configured to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to a display device 123 .
[0044] Display device 123 displays the decoded video data and may be integrated with decoding device 120 or located external to decoding device 120. Display device 123 may include a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0045] Furthermore, FIG. 2 is just an example, and the technical solution of the embodiment of the present application is not limited to FIG. 2, for example, the technology of the present application may be applied to one-sided video encoding or one-sided video decoding.
[0046] In the following, a video coding framework according to an embodiment of the present invention is introduced.
[0047] FIG. 3 is a schematic diagram of an encoding framework according to an embodiment of the present invention.
[0048] The coding framework may be used to perform lossy compression on images, or it may be used to perform lossless compression on images, which may be visually lossless compression or mathematically lossless compression.
[0049] The coding framework is applicable to image data in luminance-chrominance (YCbCr, YUV) format.
[0050] For example, the encoding framework reads video data and, for each frame of the video data, divides the image into several coding tree units (CTUs), which may also be referred to as "tree blocks," "largest coding units" (LCUs), or "coding tree blocks" (CTBs). Each CTU may be associated with a block of pixels of the same size in the image. Each pixel may correspond to one luminance (or luma) sample and two chrominance (or chroma) samples. Thus, each CTU may be associated with one luminance sample block and two chroma sample blocks. For example, the size of a CTU may be 128x128, 64x64, 32x32, etc. A CTU may be further divided into several coding units (CUs) for encoding, which may be rectangular or square blocks. A CU can be further divided into prediction units (PUs) and transform units (TUs), which allows for separation of coding, prediction, and transformation for more flexible processing. In one example, a CTU is divided into CUs in a quadtree fashion, and a CU is divided into TUs and PUs in a quadtree fashion.
[0051] Video encoders and decoders can support various PU sizes. Assuming that a particular CU has a size of 2Nx2N, the video encoder and decoder can support PU sizes of 2Nx2N or NxN for intra prediction, and symmetric PUs of 2Nx2N, 2NxN, Nx2N, NxN, or similar sizes for inter prediction. The video encoder and decoder can also support asymmetric PUs of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.
[0052] 3, the coding framework includes a prediction unit 11, a residual generation unit 12, a transform unit 13, a quantization unit 14, an inverse quantization unit 15, an inverse transform unit 16, a reconstruction unit 17, a filtering unit 18, and an entropy coding unit 19. The prediction unit 11 includes an inter prediction unit 101 and an intra prediction unit 102. The inter prediction unit 101 includes a motion estimation unit 1011 and a motion compensation unit 1012. It should be noted that the coding framework may include more, fewer, or different functional components.
[0053] Optionally, in this application, the current block is also referred to as a current coding unit (CU) or a current prediction unit (PU), etc. The prediction block is also referred to as a predicted image block or an image prediction block, and the reconstructed image block is also referred to as a reconstructed block or an image reconstructed image block.
[0054] Here, after receiving a video, the encoding side divides each frame of the video into a plurality of image blocks to be coded. For a current image block to be coded, the prediction unit 11 first predicts the current image block to be coded based on a reference reconstructed image block to obtain prediction information for the current image block to be coded. Here, the encoding side can use inter-prediction or intra-prediction technology to obtain the prediction information.
[0055] For example, the motion estimation unit 1011 in the inter prediction unit 101 may search a reference image in a list of reference images to find a reference block for the image block to be coded. The motion estimation unit 1011 may generate an index indicating the reference block and a motion vector indicating a spatial displacement between the image block to be coded and the reference block. The motion estimation unit 1011 may output the index of the reference block and the motion vector as motion information for the image block to be coded. The motion compensation unit 1012 may obtain prediction information for the image block to be coded based on the motion information for the image block to be coded.
[0056] The intra prediction unit 102 may employ an intra prediction mode to generate prediction information for the current image block to be coded. Currently, there are 15 intra prediction modes, including a planar mode, a DC mode, and 13 angle prediction modes. The intra prediction unit 102 may also employ an intra block copy (IBC) technique, an intra string copy (ISC) technique, etc.
[0057] HEVC uses 35 intra prediction modes in total, including Planar, DC, and 33 angle modes. VVC uses 67 intra prediction modes in total, including Planar, DC, and 65 angle modes. AVS3 uses 66 intra prediction modes in total, including DC, Planar, Bilinear, and 63 angle modes.
[0058] The residual generation unit 12 is configured to subtract the prediction information from the original signal of the current image block to be coded to obtain a residual signal, where after prediction, the amplitude of the residual signal is much smaller than that of the original signal.
[0059] The transform unit 13 and the quantization unit 14 are configured to perform transform and quantization operations on the residual signal. After transform and quantization, transform and quantization coefficients are obtained.
[0060] The entropy encoding unit 19 is configured to encode the quantized coefficients and other coding indication information using an entropy encoding technique to obtain a codestream.
[0061] Furthermore, the encoding side needs to reconstruct the current image block to be coded so as to provide reference pixels for encoding of the subsequent image block to be coded. Illustratively, after obtaining the transform quantization coefficients of the current image block to be coded, the inverse quantization unit 15 and the inverse transform unit 16 perform inverse quantization and inverse transform on the transform quantization coefficients of the current image block to be coded to obtain a reconstructed residual signal, and the reconstruction unit 17 adds prediction information corresponding to the current image block to the reconstructed residual signal to obtain a reconstructed signal of the current image block to be coded, and obtains a reconstructed image block based on the reconstructed signal.
[0062] Furthermore, the filtering unit 18 may filter the reconstructed image block, where deblocking filtering (DBF), sample adaptive offset (SAO), or adaptive loop filter (ALF), etc., may be employed, where the reconstructed image block can be used to predict a subsequent image block to be coded.
[0063] In some embodiments, the reconstructed image blocks may be stored in a decoded image cache, and the inter prediction unit 101 may perform inter prediction on PUs of other images using reference images containing the reconstructed pixel blocks. Furthermore, the intra prediction unit 102 may perform intra prediction on other PUs in the same image as the CU using the reconstructed image blocks in the decoded image cache.
[0064] FIG. 4 is a schematic diagram of a decoding framework according to an embodiment of the present invention.
[0065] 4, the decoding framework includes an entropy decoding unit 21, a prediction unit 22, an inverse quantization unit 23, an inverse transform unit 24, a reconstruction unit 25, and a filtering unit 26. The prediction unit 22 includes a motion compensation unit 221 and an intra prediction unit 222.
[0066] For example, after the decoding side obtains a codestream, the entropy decoding unit 21 first performs entropy decoding on the codestream to obtain the transform quantization coefficients of the current image block awaiting reconstruction. Then, the inverse quantization unit 23 and the inverse transform unit 24 perform inverse quantization and inverse transform on the transform quantization coefficients to obtain a reconstructed residual signal of the current image block awaiting reconstruction. The prediction unit 22 predicts the current image block awaiting reconstruction to obtain prediction information for the current image block awaiting reconstruction. If the prediction unit 22 employs inter-prediction, the motion compensation unit 221 may construct a first reference image list (list 0) and a second reference image list (list 1) based on syntax elements parsed from the codestream. Furthermore, the entropy decoding unit 21 may analyze motion information of the image block awaiting reconstruction. The motion compensation unit 221 may determine one or more reference blocks for the image block awaiting reconstruction based on the motion information. The motion compensation unit 221 may generate prediction information for the image block awaiting reconstruction based on the one or more reference blocks. If the prediction unit 22 employs intra prediction, the entropy decoding unit 21 can analyze the index of the intra prediction mode used, and the intra prediction unit 222 can employ the intra prediction mode according to the index to perform intra prediction and obtain prediction information for the image block to be reconstructed. The intra prediction unit 222 can also employ IBC or ISC techniques, etc.
[0067] Furthermore, the reconstruction unit 25 is configured to add the prediction information and the above-mentioned reconstruction residual signal to obtain a reconstruction signal of the current image block to be reconstructed, and then obtain a current reconstructed image block corresponding to the current image block to be reconstructed based on the reconstruction signal, where the current reconstructed image block can be used to predict other subsequent image blocks to be reconstructed. Similar to the above-mentioned encoding side, optionally, the filtering unit 26 on the decoding side can filter the current reconstructed image block.
[0068] If necessary, the codestream can include block division information determined by the encoding side, and mode or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering. The decoding side analyzes the codestream and determines the same block division information, mode or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering as the encoding side by analyzing the codestream based on existing information, thereby ensuring that the decoded image obtained by the encoding side is the same as the decoded image obtained by the decoding side.
[0069] The above description is the basic process of a video codec in a block-based hybrid coding framework, and as technology develops, some modules or steps of the framework or process may be optimized, and the present application is suitable for, but is not limited to, the basic process of a video codec in the block-based hybrid coding framework.
[0070] As can be seen from the prediction methods of the video codecs described above, coded images are divided into fully intra-coded images and inter-coded images, and images include frames, slices, and tiles as shown in Figures 5 to 7. The dashed blocks in Figure 5 represent boundaries of the largest coding unit (CTU), the solid black lines in Figure 6 represent slices, and the solid black lines in Figure 7 represent tile boundaries.
[0071] Here, all the reference information for the prediction of a fully intra-coded picture is obtained from the spatial domain information of the current picture, while the prediction process of an inter-codeable picture can refer to the temporal domain reference information of other reference frames.
[0072] As can be seen from the above knowledge about video codecs, traditional loop filters, including DBF, SAO, and ALF, mainly filter the reconstructed image to reduce block artifacts, linking artifacts, etc., thereby improving the quality of the reconstructed image. Ideally, the filter would restore the reconstructed image to the original image. However, many of the filtering coefficients in traditional filters are artificially designed, leaving much room for optimization. Given the excellent performance of deep learning tools in image processing, deep learning-based filters are often applied to the loop filter module.
[0073] The main technology in the present embodiment is a neural network-based filter such as a neural network loop filter (NNLF). As shown in Figure 8, the image to be filtered is input to the trained neural network loop filter (NNLF) for filtering to obtain a filtered image.
[0074] In some embodiments, the training process of a neural network-based filter requires inputting an image and specifying a target image as an optimization target to train the filter parameters. During training, the input image and the target image are spatially aligned. Typically, the input image is selected from the reconstructed distorted image, and the target image is selected from the original image as the optimal optimization target.
[0075] A loss function is used in the model training process. The loss function measures the difference between the predicted value and the actual value. The larger the loss value, the greater the difference. The goal of training is to reduce the loss. For deep learning-based coding tools, commonly used loss functions include the L1 norm loss function, the L2 norm loss function, and the smooth L1 loss function.
[0076] When using a neural network-based filter in practice, the image is usually divided into sub-images, and the sub-images are input to the neural network-based filter in stages for filtering, rather than directly inputting the entire frame image for filtering.
[0077] Therefore, in the training process of the neural network-based filter, the input image and the target image are divided to obtain matching pairs consisting of input image blocks and target image blocks, and then the neural network-based filter is trained based on the matching pairs.
[0078] In some embodiments, when dividing the input image and the target image, a random cropping method is adopted from one frame of image to select matching pairs consisting of input image blocks and target image blocks, and a random cropping method is adopted to randomly crop the input image and the target image to obtain image blocks numbered 1 to 5, as shown in Figures 9 and 10. In the actual filtering process, filtering is usually performed using a CTU as the basic unit, for example, five CTUs A to E are used to perform actual filtering, as shown in Figures 11 and 12.
[0079] As can be seen from the above description, the division method of image blocks in the training process of the neural network-based filter does not match the division method of image blocks in the actual use process, which reduces the filtering effect of the neural network-based filter.
[0080] To solve the above technical problems, in the embodiment of the present application, in the actual filtering process, the image to be filtered is divided according to the same block division method of the training image used by the neural network-based filter in the training process to obtain N image blocks to be filtered, and for each of the N image blocks to be filtered, the neural network-based filter is used to filter the image block to be filtered to obtain a final filtered image. That is, in the embodiment of the present application, the block division method in the actual use process of the neural network-based filter is consistent with the block division method used during training, so that the neural network filter can achieve optimal filtering performance, thereby improving the image filtering effect.
[0081] The following describes the technical solutions of the embodiments of the present application in detail through several embodiments, and the following several embodiments can be combined with each other, and the same or similar concepts or processes may not be described repeatedly in some embodiments.
[0082] FIG. 13 is a flowchart of an image filtering method according to one embodiment of the present application. As shown in FIG. 13, the embodiment of the present application includes the following steps:
[0083] In step S801, an image to be filtered is acquired.
[0084] In the present embodiment, the method of obtaining the image to be filtered includes, but is not limited to, the following several situations:
[0085] In Situation 1, in the case of a non-video codec scenario, in one example, the image to be filtered may be collected by an image acquisition device, e.g., photographed by a camera, or in another example, the image to be filtered may be generated by an image generation device, e.g., rendered by an image rendering device.
[0086] In the case of the video codec scenario, the method of obtaining the image to be filtered includes at least the following methods:
[0087] In Method 1, on the encoding side, the pre-filtering image may be an image before encoding, that is, on the encoding side, before encoding the pre-filtering image, the pre-filtering image is first filtered and then the filtered image is encoded. For example, before encoding the pre-filtering image 1 on the encoding side, the pre-filtering image 1 is first input to a neural network-based filter for filtering to obtain a filtered image 2. Next, the image 2 is divided into blocks to obtain a plurality of coding blocks, and each coding block is predicted using a prediction method such as inter prediction or intra prediction to obtain a predicted block of the coding block, and the difference between the coding block and the predicted block is obtained to obtain a residual block, and the residual block is transformed and quantized to obtain quantized coefficients, and finally the quantized coefficients are encoded.
[0088] In Method 2, the encoding side may use a reconstructed image as the image to be filtered. That is, the current image is reconstructed to obtain a reconstructed image of the current image, and the reconstructed image is determined as the image to be filtered. For example, the encoding side divides the current image into blocks to obtain multiple coding blocks, predicts each coding block using a prediction method such as inter-prediction or intra-prediction to obtain a predicted block of the coding block, calculates the difference between the coding block and the predicted block to obtain a residual block, transforms and quantizes the residual block to obtain quantized coefficients, and finally encodes the quantized coefficients. Furthermore, the encoding side performs inverse quantization on the quantization parameter to obtain transform coefficients of the coding block, performs inverse transform on the transform coefficients to obtain a residual block, and adds the residual block to the predicted block to obtain a reconstructed block of the coding block. The reconstructed blocks of all coding blocks in the current image are combined to obtain a reconstructed image of the current image. The reconstructed image is then input as the image to be filtered into a neural network-based filter for filtering to obtain a filtered image.
[0089] In Method 3, on the decoding side, the image to be filtered may be a reconstructed image, i.e., the current image is reconstructed to obtain a reconstructed image of the current image, and the reconstructed image is determined as the image to be filtered. For example, the decoding side decodes the received code stream to obtain quantized coefficients of a current block in the current image, then performs inverse quantization on the quantized coefficients to obtain transform coefficients of the current block, and performs inverse transform on the transform coefficients to obtain a residual block. Furthermore, the current block is predicted using a prediction method such as inter prediction or intra prediction to obtain a predicted block of the current block, and the residual block is added to the predicted block to obtain a reconstructed block of the current block. The reconstructed blocks of all blocks in the current image are combined to obtain a reconstructed image of the current image. The reconstructed image is then input as the image to be filtered into a neural network-based filter for filtering to obtain a filtered image.
[0090] In addition, in the present embodiment, the method for determining the image to be filtered includes, but is not limited to, the above-mentioned method. The present application can also obtain the image to be filtered by other methods, and the present embodiment is not limited thereto.
[0091] In step S802, a neural network-based filter is determined, and the image to be filtered is divided according to a block division method corresponding to the neural network-based filter to obtain N image blocks to be filtered.
[0092] Here, the above block division method is the block division method of the training image used by the neural network-based filter in the training process, that is, the image to be filtered is divided according to the same method as the block division method of the training image used by the neural network-based filter in the training process to obtain N image blocks to be filtered (N is a positive integer).
[0093] In the present embodiment, before using a neural network-based filter to filter the image to be filtered, the filter must first be determined.
[0094] In some embodiments, the neural network-based filter is a pre-configured or default filter, and thus, the pre-configured or default neural network-based filter can be directly used for filtering.
[0095] In some embodiments, the present invention includes a plurality of neural network-based candidate filters, hereinafter abbreviated as candidate filters, from which one candidate filter can be determined as the neural network-based filter of the present invention.
[0096] In one example, the network structures of at least two of the plurality of candidate filters are not identical.
[0097] In another example, at least two of the multiple candidate filters use different block division schemes for the training image during the training process. For example, during training of candidate filter 1, one CTU in the input image is determined as one input block, and one CTU at the same position in the target image is determined as one target block. Candidate filter 1 is then trained using the input block as an input and the target block as a target. In another example, during training of candidate filter 2, two CTUs in the input image are determined as one input block, and two CTUs at the same position in the target image are determined as one target block. Candidate filter 2 is then trained using the input block as an input and the target block as a target.
[0098] That is, the information such as the network structure, training parameters, and training methods of the plurality of candidate filters in the present embodiment are not exactly the same.
[0099] In the above, the manner of determining the neural network-based filter from a plurality of candidate filters includes, but is not limited to, the following manners.
[0100] In Method 1, any one of the candidate filters is determined as the neural network-based filter in this step.
[0101] In Method 2, among the plurality of candidate filters, the candidate filter with the highest filtering effect is selected as the neural network-based filter of the present application. Illustratively, each candidate filter of the plurality of candidate filters is used to filter the image to be filtered, to obtain a filtered image corresponding to each candidate filter, and these filtered images are compared to determine the filtered image with the best effect. Furthermore, the candidate filter corresponding to the filtered image with the best effect is determined as the neural network-based filter in this step.
[0102] In the present embodiment, there is no limitation on the method for determining the image effect of the filtered image, and the image effect of the filtered image can be determined by determining image indices such as image resolution, sharpness, artifacts, etc.
[0103] In Method 3, among the plurality of candidate filters, the candidate filter with the smallest distortion is determined as the neural network-based filter of the present application. Illustratively, each candidate filter of the plurality of candidate filters is used to filter the image to be filtered to obtain a filtered image corresponding to each candidate filter, and the filtered image corresponding to each candidate filter is compared with the image to be filtered to determine the distortion corresponding to each candidate filter, and the candidate filter with the smallest distortion is determined as the neural network-based filter in this step.
[0104] The embodiment of the present application is not limited to a method for determining the distortion corresponding to the candidate filter. For example, the difference between the filtered image corresponding to the candidate filter and the image to be filtered is determined as the distortion corresponding to the candidate filter.
[0105] In the present embodiment, the method for determining the neural network-based filter includes, but is not limited to, the above methods.
[0106] According to the above method, after determining the neural network-based filter, obtain the block division scheme of the training image that the neural network-based filter uses in the training process. For example, the file of the neural network-based filter includes the block division scheme of the training image that the neural network-based filter uses in the training process, and thus the block division scheme of the training image that the neural network-based filter uses in the training process can be directly read from the file of the neural network-based filter.
[0107] In the present embodiment, in order to improve the filtering performance of the neural network-based filter, the block division method used in the actual use of the neural network-based filter is made to match the block division method used during training.
[0108] Based on this, in the actual filtering process, in the embodiment of the present application, the block division method of the training image used by the neural network-based filter in the training process is used to divide the image to be filtered into N image blocks to be filtered. That is, in the embodiment of the present application, during actual filtering, the block division method of the training image used by the neural network-based filter in the training process is used to divide the image to be filtered into N image blocks to be filtered. For example, in the training process of the neural network-based filter, each CTU in the training image is determined as a training image block to perform model training, and in this way, during actual filtering, each CTU in the image to be filtered is also determined as an image block to be filtered for filtering. This makes the block division method used in the actual use of the neural network-based filter consistent with the block division method used in the training process, allowing the neural network-based filter to achieve optimal filtering performance and improve the filtering effect.
[0109] The present embodiment is not limited to a specific type of block division scheme for the training images that the neural network-based filter uses in the training process, and any block division scheme may be used.
[0110] In some embodiments, if the block division scheme of the training image used by the above neural network-based filter in training includes determining M (M is a positive integer) CTUs in the training image as one training image block, in the above step S802, dividing the image to be filtered according to the block division scheme corresponding to the neural network-based filter to obtain N image blocks to be filtered includes the following step S802-A.
[0111] In step S802-A, M CTUs in the image to be filtered are determined as one image block to be filtered, thereby obtaining N image blocks to be filtered.
[0112] That is, N image blocks to be filtered are determined from the image to be filtered, with M CTUs being treated as one image block to be filtered.
[0113] The present embodiment is not limited to the specific value of M above.
[0114] In Example 1, M=1, and as shown in FIG. 14, in the training process of the neural network-based filter, one CTU in the training image is determined as one training image block to train the neural network-based filter, where the area of one smallest dashed block in FIG. 14 is one CTU.
[0115] In one possible implementation, when the neural network-based filter is obtained by training using a supervised training method, the training image includes an input image and a target image, where the target image can be understood as a teacher image. In the training process, as shown in FIG. 15, one CTU in the input image is determined as one input image block, and as shown in FIG. 16, one CTU in the target image is determined as one target image block. The input image block and the target image block form a matching pair, that is, the position of the input image block in the input image in one matching pair matches the position of the target image block in the target image. Next, the input image block is input to the neural network-based filter and filtered to obtain a filtered image block corresponding to the input image block. The filtered image block of the input image block is compared with the target image block to calculate a loss, and the parameters of the neural network-based filter are adjusted based on the loss. Next, referring to the above method, the neural network-based filter is trained using the input image block in the next matching pair as input and the target image block in the next matching pair as target to obtain a trained neural network-based filter.
[0116] As can be seen from the above description, in Example 1, in the training process of the above neural network-based filter, one CTU in the training image is determined as one training image block, and correspondingly, in the actual filtering process, one CTU in the image to be filtered is determined as one image block to be filtered, as shown in Figure 17, and in this way, block division is performed on the image to be filtered to obtain N image blocks to be filtered.
[0117] In Example 2, M=4, and as shown in FIG. 18, in the training process of the neural network-based filter, four CTUs of the training image are determined as one training image block to train the neural network-based filter.
[0118] In one possible implementation, when the neural network-based filter is obtained by training using a supervised training method, the training image includes an input image and a target image, where the target image can be understood as a teacher image. In a specific training process, as shown in FIG. 19, four CTUs in the input image are determined as one input image block, and as shown in FIG. 20, four CTUs in the target image are determined as one target image block. The input image block and the target image block form a matching pair, that is, the position of the input image block in one matching pair in the input image matches the position of the target image block in the target image. Next, the input image block is input to the neural network-based filter and filtered to obtain a filtered image block corresponding to the input image block. The filtered image block of the input image block is compared with the target image block to calculate a loss, and the parameters of the neural network-based filter are adjusted based on the loss. Next, referring to the above method, the neural network-based filter is continuously trained using the input image block in the next matching pair as input and the target image block in the next matching pair as target to obtain a trained neural network-based filter.
[0119] As can be seen from the above description, in Example 2, in the training process of the above neural network-based filter, four CTUs in the training image are determined as one training image block, and correspondingly, in the actual filtering process, four CTUs in the image to be filtered are determined as one image block to be filtered, as shown in Figure 21, and in this way, N image blocks to be filtered are obtained.
[0120] Although the above description is given as an example in which M is 1 or 4, in some embodiments, M may be any positive integer such as 2, 3, 5, etc., and the present embodiment is not limited thereto.
[0121] In some embodiments, if the block division method of the training image used by the above neural network-based filter in training includes determining P (P is a positive integer) incomplete CTUs in the training image as one training image block, in the above step S802, dividing the image to be filtered according to the same method as the block division method of the training image used by the neural network-based filter in the training process to obtain N image blocks to be filtered includes the following step S802-B.
[0122] In S802-B, P incomplete CTUs in the image to be filtered are determined as one image block to be filtered, thereby obtaining N image blocks to be filtered.
[0123] That is, P incomplete CTUs are regarded as one image block to be filtered, and N image blocks to be filtered are determined from the image to be filtered.
[0124] The present embodiment is not limited to the specific value of P above.
[0125] For example, assuming P=4, as shown in FIG. 22, in the training process of the neural network-based filter, four incomplete CTUs in the training image are determined as one training image block to train the neural network-based filter.
[0126] In one possible implementation, when the neural network-based filter is obtained by training using a supervised training method, the training image includes an input image and a target image, where the target image can be understood as a teacher image. In a specific training process, as shown in FIG. 23, four incomplete CTUs in the input image are determined as one input image block, and as shown in FIG. 24, four incomplete CTUs in the target image are determined as one target image block, and the input image block and the target image block form one matching pair. Next, the input image block is input to the neural network-based filter and filtered to obtain a filtered image block corresponding to the input image block. The filtered image block of the input image block is compared with the target image block to calculate a loss, and the parameters of the neural network-based filter are adjusted based on the loss. Next, referring to the above method, the neural network-based filter is trained using the input image block in the next matching pair as input and the target image block in the next matching pair as target to obtain a trained neural network-based filter.
[0127] As can be seen from the above description, in this example, in the training process of the above neural network-based filter, four incomplete CTUs in the training image are determined as one training image block, and correspondingly, in the actual filtering process, four incomplete CTUs in the image to be filtered are determined as one image block to be filtered, as shown in Figure 25, and in this way, N image blocks to be filtered are obtained.
[0128] Although the above description uses the example of P being 4, in some embodiments, P may be any positive integer such as 1, 2, 3, 5, etc., and the present embodiment is not limited thereto.
[0129] In the above example, the four incomplete CTUs are adjacent, but in some embodiments, when the above P is greater than 1, the P residual CTUs may not be completely adjacent, that is, all of the P residual CTUs may not be adjacent to each other, or some of the P residual CTUs may be adjacent to each other and some may not be adjacent to each other.
[0130] In the above embodiment, the block division method of the training image includes determining M CTUs in the training image as one training image block, or determining P incomplete CTUs in the training image as one training image block. Note that the block division method of the training image according to the embodiment of the present application includes, but is not limited to, the above examples, and the embodiment of the present application is not limited thereto.
[0131] After the above steps, the image to be filtered is divided using the block division method of the training image used by the neural network-based filter in the training process to obtain N image blocks to be filtered, and then the following step S803 is performed.
[0132] In step S803, a neural network-based filter is used to filter the N image blocks to be filtered respectively to obtain a filtered image.
[0133] In this embodiment, the image to be filtered is divided using the same block division method as that of the training image to obtain N image blocks to be filtered. For each of the N image blocks to be filtered, the image block to be filtered is input to a neural network-based filter to be filtered to obtain a filtered image block for the image block to be filtered. According to the above method, a filtered image block for each of the N image blocks to be filtered can be determined, and the filtered image blocks for these N image blocks to be filtered constitute a filtered image.
[0134] In the filtering process, filtering based on adjacent content outside the boundary of the image block can improve the filtering effect of the boundary region of the image block. In view of this, in order to further improve the filtering effect of the neural network-based filter, in some embodiments, the neural network-based filter is obtained by training using an extended image block of the training image block. That is, in the filter training process, in addition to dividing the training image according to the block division scheme of the training image to obtain the training image block, the training image block can also be extended outward to obtain an extended image block, and the extended image block can be used to train the neural network-based filter. In this case, the above step S803 includes the following steps S803-A1 to S803-A3.
[0135] In step S803-A1, for each of the N image blocks to be filtered, the image block to be filtered is extended according to the extension method of the training image block to obtain an extended image block to be filtered.
[0136] In step S803-A2, the dilated to-be-filtered image block is filtered using a neural network-based filter to obtain a filtered dilated image block.
[0137] In step S803-A3, an image area corresponding to the image block to be filtered in the filtered extended image block is determined as the filtered image block of the image block to be filtered.
[0138] In this embodiment, to further improve the filtering effect of the image to be filtered, the image blocks to be filtered are extended using the same extension method as the training image blocks. Specifically, in the actual filtering process, the image to be filtered is divided using the same block division method as the training image to obtain N image blocks to be filtered. For each of the N image blocks to be filtered, the image block to be filtered is extended outward using the same extension method as the training image blocks to obtain an extended image block to be filtered. Next, the extended image block to be filtered is input to a neural network-based filter for filtering to obtain a filtered extended image block of the extended image block to be filtered. Since the filtered extended image block does not match the size of the image block to be filtered, the filtered extended image block is cropped. Specifically, the image region in the filtered extended image block corresponding to the image block to be filtered is determined as the filtered image block of the image block to be filtered. According to the above steps, a filtered image block for each of the N image blocks to be filtered can be determined, and the filtered image blocks of these N image blocks to be filtered are stitched together to obtain a final filtered image.
[0139] The present embodiment is not limited to a specific extension method for training image blocks.
[0140] In some embodiments, the extension method for the training image block includes extending at least one boundary region of the training image block outward, and in this case, in step S803-A1, extending the image block to be filtered according to the extension method for the training image block to obtain an extended image block to be filtered includes extending at least one boundary region of the image block to be filtered outward to obtain an extended image block to be filtered.
[0141] In some examples, as shown in FIG. 26, the training process of the neural network-based filter involves extending the perimeter of the training image block outward and using the extended training image block to train the neural network-based filter.
[0142] In one possible implementation, when the neural network-based filter is obtained by training using a supervised training method, the training image includes an input image and a target image, where the target image can be understood as a teacher image. In a specific training process, as shown in FIG. 27, one CTU of the input image is determined as one input image block, and the periphery of the input image block is expanded outward to obtain an expanded input image block. As shown in FIG. 28, one CTU of the target image is determined as one target image block, and the periphery of the target image block is expanded outward to obtain an expanded target image block. Next, the expanded input image block is input to the neural network-based filter and filtered to obtain a filtered image block corresponding to the expanded input image block. The filtered image block of the expanded input image block is compared with the expanded target image block to calculate a loss, and the parameters of the neural network-based filter are adjusted based on the loss. Next, referring to the above method, the neural network-based filter is continuously trained using the dilated input image block in the next matching pair as input and the dilated target image block in the next matching pair as target to obtain a trained neural network-based filter.
[0143] As can be seen from the above description, in this example, in the training process of the above neural network-based filter, one CTU in the training image is determined as one training image block, and the periphery of the training image block is extended outward to obtain an extended training image block. Correspondingly, in the actual filtering process, as shown in Figure 29, one CTU in the image to be filtered is determined as one image block to be filtered, and then the periphery of the image block to be filtered is extended outward using the above training image block extension method to obtain an extended image block to be filtered, and then the extended image block to be filtered is input to the neural network-based filter for filtering to obtain a filtered extended image block, and further, an image region of the filtered extended image block corresponding to the image block to be filtered is determined as the filtered image block of the image block to be filtered.
[0144] In some examples, as shown in FIG. 30, the training process of the neural network-based filter involves extending the left and top boundaries of the training image block outward, and using the extended training image block to train the neural network-based filter.
[0145] In one possible implementation, when the neural network-based filter is obtained by training using a supervised training method, the training image includes an input image and a target image, where the target image can be understood as a teacher image. In a specific training process, as shown in FIG. 31, one CTU of the input image is determined as an input image block, and the left and upper boundaries of the input image block are extended outward to obtain an extended input image block. As shown in FIG. 32, one CTU of the target image is determined as a target image block, and the left and upper boundaries of the target image block are extended outward to obtain an extended target image block. Next, the extended input image block is input to the neural network-based filter for filtering to obtain a filtered image block corresponding to the extended input image block. The filtered image block of the extended input image block is compared with the extended target image block to calculate a loss, and the parameters of the neural network-based filter are adjusted based on the loss. Next, referring to the above method, the neural network-based filter is trained using the dilated input image block in the next matching pair as input and the dilated target image block in the next matching pair as target to obtain a trained neural network-based filter.
[0146] As can be seen from the above description, in this example, in the training process of the above neural network-based filter, one CTU in the training image is determined as one training image block, and the left and upper boundaries of the training image block are extended outward to obtain an extended training image block. Correspondingly, in the actual filtering process, as shown in Figure 33, one CTU in the image to be filtered is determined as one image block to be filtered, and then the left and upper boundaries of the image block to be filtered are extended outward using the above training image block extension method to obtain an extended image block to be filtered, and then the extended image block to be filtered is input to the neural network-based filter for filtering to obtain a filtered extended image block, and then the image area of the filtered extended image block corresponding to the image block to be filtered is determined as the filtered image block of the image block to be filtered.
[0147] In some examples, the training image block extension scheme further includes extending other boundaries of the training image block outward, although embodiments of the present application are not limited thereto.
[0148] In some embodiments, to further improve the filtering effect of the neural network-based filter, in addition to inputting an input image block, a reference image block for the input image block is also input during the training process. To maintain consistency between the actual filtering process and the training process, the above step S803 includes the following steps S803-B1 to S803-B2:
[0149] In step S803-B1, for each of the N image blocks to be filtered, a reference image block for the image block to be filtered is determined.
[0150] In step S803-B2, the image block to be filtered and the reference image block of the image block to be filtered are input to a neural network-based filter for filtering, and a filtered image block of the image block to be filtered is obtained.
[0151] In the present embodiment, in the training process of the neural network-based filter, when the input information includes an input image block and a reference image block of the input image block, in the actual filtering process, in addition to the image block to be filtered, the input information further includes a reference image block of the image block to be filtered.
[0152] The present embodiment is not limited to the method for determining the reference image block of the image block to be filtered.
[0153] In some embodiments, the manner in which the reference image blocks of the image blocks to be filtered are determined is different from the manner in which the reference image blocks of the training image blocks are determined.
[0154] In some embodiments, the method for determining the reference image block of the image block to be filtered is the same as the method for determining the reference image block of the input image block. In this case, in the above step S803-B1, determining the reference image block of the image block to be filtered includes the following steps S803-B11 to S803-B12:
[0155] In step S803-B11, a determination method of a reference image block of an input image block is obtained, and the determination method is used to determine a corresponding reference image block based on at least one of spatial domain information and temporal domain information of the input image block.
[0156] In step S803-B12, a reference image block for the image block to be filtered is determined according to the method for determining a reference image block for the input image block.
[0157] In some embodiments, the method for determining the reference image block of an input image block can be read from the file of the neural network-based filter. In some embodiments, if the training device for the neural network-based filter is the same as the actual filtering device, the device stores the method for determining the reference image block of an input image block.
[0158] After obtaining the method for determining the reference image block of the input image block, the method for determining the reference image block of the input image block is used to determine the reference image block of the image block to be filtered, that is, in this embodiment, the method for determining the reference image block of the image block to be filtered is the same as the method for determining the reference image block of the input image block.
[0159] The present embodiment is not limited to a specific type of reference image block.
[0160] In some embodiments, when the reference image block of the input image block includes at least one of the time domain reference image block and the spatial domain reference image block of the input image block, the reference image block of the image block to be filtered includes at least one of the time domain reference image block and the spatial domain reference image block of the image block to be filtered.
[0161] Here, the spatial domain reference image block may select an image area at a fixed relative position of the current input image block, i.e., the spatial domain reference image block of the input image block is in the same frame as the input image block, i.e., both are in the input image.
[0162] The difference between the time domain reference image block and the spatial domain reference image block is that the time domain reference image block and the current input image block are in different frames, and the reference position of the time domain reference image block may select a reference block at the same spatial position as the current input image.
[0163] In this embodiment, the type of the reference image block of the image block to be filtered is the same as the type of the reference image block of the input image block.
[0164] In Example 1, if the reference image block of the input image block includes a spatial domain reference image block, the reference image block of the image block to be filtered also includes a spatial domain reference image block. In this case, the above step S803-B12 includes the following steps:
[0165] In step S803-B12-A, the spatial domain reference image block of the image block to be filtered is determined according to the method for determining the spatial domain reference image block of the input image block.
[0166] In this example, if the reference image block of the input image block includes a spatial domain reference image block, the spatial domain reference image block of the image block to be filtered is determined according to the method for determining the spatial domain reference image block of the input image block, thereby enabling the spatial domain reference image block of the image block to be filtered to be accurately determined.
[0167] In one possible implementation, a method for determining a spatial domain reference image block of an input image block is used to determine a spatial domain reference image block of an image block to be filtered, thereby ensuring that the input information in the training process of the neural network-based filter is consistent with the input information in the actual filtering process, and improving the filtering performance of the neural network-based filter.
[0168] The present embodiment is not limited to a specific type of spatial domain reference image block.
[0169] In some embodiments, when the spatial domain reference image block of the input image block includes at least one of the image block located to the upper left, the image block located to the left, and the image block located at the top of the input image block in the input image, the spatial domain reference image block of the image block to be filtered includes at least one of the image block located to the upper left, the image block located to the left, and the image block located at the top of the image block to be filtered in the image to be filtered. In this case, step S803-B12-A includes determining at least one of the image block located to the upper left, the image block located to the left, and the image block located at the top of the image block to be filtered in the image to be filtered as the spatial domain reference image block of the image block to be filtered.
[0170] For example, as shown in Figure 34, if the spatial domain reference image block of the input image block includes an image block located to the upper left of the input image block in the input image, as shown in Figure 35, the image block located to the upper left of the image block to be filtered in the image to be filtered is determined as the spatial domain reference image block of the image to be filtered.
[0171] In another example, as shown in Figure 36, if the spatial domain reference image block of the input image block includes an image block located to the left of the input image block in the input image, as shown in Figure 37, an image block located to the left of the image block to be filtered in the image to be filtered is determined as the spatial domain reference image block of the image to be filtered.
[0172] In another example, as shown in Figure 38, if the spatial domain reference image block of the input image block includes an image block located above the input image block in the input image, as shown in Figure 39, the image block located above the image block to be filtered in the image to be filtered is determined as the spatial domain reference image block of the image to be filtered.
[0173] In another example, as shown in Figure 40, if the spatial domain reference image blocks of the input image block include an image block located to the upper left, an image block located to the left, and an image block located at the top of the input image block in the input image, as shown in Figure 41, the image block located to the upper left, an image block located to the left, and an image block located at the top of the image block to be filtered in the image to be filtered are determined as the spatial domain reference image blocks of the image to be filtered.
[0174] In this embodiment, in order to further improve the filtering effect, in addition to inputting an input image block, a spatial domain reference image block of the input image block is also input during the training process, thereby improving the filtering effect of the neural network-based filter. In this way, in the actual filtering process, in order to maintain consistency between the input information in the actual filtering process and the input information in the training process, the same determination method as the determination method of the spatial domain reference image block of the input image block is adopted to determine the spatial domain reference image block of the image block to be filtered, and then the image block to be filtered and the spatial domain reference image block of the image block to be filtered are input into the neural network-based filter to achieve filtering of the image to be filtered.
[0175] In Example 2, if the reference image block of the input image block includes a time-domain reference image block, the reference image block of the image block to be filtered also includes a time-domain reference image block. In this case, the above step S803-B12 includes the following steps:
[0176] In step S803-B12-B, the time-domain reference image block of the image block to be filtered is determined according to the manner of determining the time-domain reference image block of the input image block.
[0177] In this example, if the reference image block of the input image block includes a time-domain reference image block, the time-domain reference image block of the image block to be filtered is determined according to the method for determining the time-domain reference image block of the input image block, thereby enabling the time-domain reference image block of the image block to be filtered to be accurately determined.
[0178] In one possible implementation, a method for determining a time-domain reference image block of an input image block is used to determine a time-domain reference image block of an image block to be filtered, thereby ensuring that the input information in the training process of the neural network-based filter is consistent with the input information in the actual filtering process, and improving the filtering performance of the neural network-based filter.
[0179] In some embodiments, step S803-B12-B includes determining a reference image of the image to be filtered, and determining an image block at a position corresponding to the image block to be filtered in the reference image of the image to be filtered as a time-domain reference image block of the image block to be filtered.
[0180] For example, as shown in Figures 42 and 43, the time-domain reference image block of an input image block is an image block at a position corresponding to the input image block in the reference image of the input image. That is, the position of the time-domain reference image block of the input image block in the reference image of the input image is the same as the position of the input image block in the input image. In this case, as shown in Figures 44 and 45, the process of determining the time-domain reference image block of the image block to be filtered is to first determine the reference image of the filtered image, and then determine the image block at a position corresponding to the image block to be filtered in the reference image of the filtered image as the time-domain reference image block of the image block to be filtered.
[0181] The present embodiment is not limited to the type of reference image; for example, when the method of the present embodiment is applied to the encoding side, the reference image of the above-mentioned image to be filtered may be any encoded image, and when the method of the present embodiment is applied to the decoding side, the reference image of the above-mentioned image to be filtered may be any decoded image.
[0182] In this embodiment, in order to further improve the filtering effect, in addition to inputting an input image block, a time-domain reference image block of the input image block is also input during the training process, thereby improving the filtering effect of the neural network-based filter. In this way, in the actual filtering process, in order to maintain consistency between the input information in the actual filtering process and the input information in the training process, the same determination method as that for the time-domain reference image block of the input image block is adopted to determine the time-domain reference image block of the image block to be filtered, and then the image block to be filtered and the time-domain reference image block of the image block to be filtered are input into the neural network-based filter, thereby realizing filtering of the image to be filtered.
[0183] In some embodiments, when the reference image block of the input image block includes a spatial domain reference image block and a time domain reference image block, the reference image block of the image block to be filtered also includes a spatial domain reference image block and a time domain reference image block, where the process of determining the spatial domain reference image block and the time domain reference image block of the image block to be filtered can refer to the above process of determining the spatial domain reference image block and the time domain reference image block, and will not be repeated here.
[0184] According to the above method, the neural network-based filter is used to filter each of the N image blocks to be filtered to obtain a filtered image.
[0185] In some embodiments, the filtering method of the present embodiment may be applied to a loop filtering module, in which case the method of the present embodiment further includes generating a reference image for prediction based on the filtered image and storing the generated reference image in a decoding cache for use as a reference image for a subsequent decoded image. Here, the manner of generating a reference image based on the filtered image may be to directly use the filtered image as a reference image, or to reprocess the filtered image, for example, by filtering it in another manner, and use the reprocessed image as a reference image. Optionally, the filtered image may be displayed by a display device.
[0186] In some embodiments, the method of the present application can also be applied to video post-processing, i.e., generating a display image based on the filtered image, inputting the generated display image to a display device for display, and skipping storing the filtered image or a reprocessed image of the filtered image in the decoding cache. That is, in the present application, a display image generated based on the filtered image is input to a display device for display, but is not stored in the decoding cache as a reference image. For example, after determining a reconstructed image of the current image by decoding video, the reconstructed image is stored in the decoding cache as a reference image, or the reconstructed image is filtered using at least one filter, such as DBF, SAO, and ALF, using a conventional loop filtering method, and the filtered image is stored in the decoding cache as a reference image. Next, the reconstructed image is used as an image to be filtered, and the reconstructed image is filtered using a neural network-based filter according to the method of the present application to obtain a filtered image. A display image is generated based on the filtered image, and the display image is input to a display device for display.
[0187] In some embodiments, the filtered image may be further filtered by at least one of the following filters: DBF, SAO, and ALF.
[0188] In some embodiments, the image to be filtered in the present embodiment may be an image that has been filtered by at least one filter such as DBF, SAO, and ALF, and then the method of the present embodiment further filters the filtered image using a neural network-based filter.
[0189] In the image filtering method according to the embodiment of the present application, an image to be filtered is obtained, a neural network-based filter is determined, the image to be filtered is divided according to the same block division method of the training image used by the neural network-based filter in the training process to obtain N (N is a positive integer) image blocks to be filtered, and the neural network-based filter is used to filter each of the N image blocks to be filtered to obtain a filtered image. That is, in the embodiment of the present application, the block division method in the actual use process of the neural network-based filter is matched with the block division method used during training, so that the neural network filter can achieve optimal filtering performance, thereby improving the image filtering effect.
[0190] 46 is an exemplary flowchart of an image filtering method according to one embodiment of the present application, which may be understood as one specific example of the filtering method shown in FIG.
[0191] As shown in FIG. 46, the image filtering method according to the present embodiment includes the following steps:
[0192] In step S901, an image to be filtered is acquired.
[0193] In the case of a non-video codec scenario, the image to be filtered may be collected by an image acquisition device or rendered by an image rendering device.
[0194] In the case of a video codec scenario, the image to be filtered may be a reconstructed image.
[0195] For the specific implementation of the above step S901, please refer to the description of the above step S801, and the description will not be repeated here.
[0196] In step S902, a neural network-based filter is determined, and the image to be filtered is divided according to a block division method corresponding to the neural network-based filter to obtain N image blocks to be filtered.
[0197] Here, the block division method is the block division method of the training image used by the neural network-based filter in the training process, that is, the image to be filtered is divided according to the same method as the block division method of the training image used by the neural network-based filter in the training process to obtain N image blocks to be filtered (N is a positive integer).
[0198] For example, the block division method of the training image used by the neural network-based filter in the training process is to determine one CTU in the training image as one training image block, and then, during actual filtering, one CTU in the image to be filtered is determined as one image block to be filtered, thereby obtaining N image blocks to be filtered.
[0199] For the specific implementation of the above step S902, please refer to the description of the above step S802, and the description will not be repeated here.
[0200] In step S903, for each of the N image blocks to be filtered, a reference image block of the image block to be filtered is determined according to the method for determining a reference image block of the input image block.
[0201] In this embodiment, the method of determining the reference image block of the image block to be filtered is the same as the method of determining the reference image block of the input image block.
[0202] In some embodiments, when the reference image block of the input image block includes at least one of the time domain reference image block and the spatial domain reference image block of the input image block, the reference image block of the image block to be filtered includes at least one of the time domain reference image block and the spatial domain reference image block of the image block to be filtered.
[0203] For the specific implementation of the above step S903, please refer to the description of the above step S803-B12, and the description will not be repeated here.
[0204] In step S904, the image block to be filtered and the reference image block of the image block to be filtered are input to a neural network-based filter for filtering to obtain a filtered image.
[0205] In this embodiment, in order to further improve the filtering effect, in addition to inputting an input image block, a reference image block of the input image block is also input during the training process, thereby improving the filtering effect of the neural network-based filter. In this way, in the actual filtering process, in order to maintain consistency between the input information in the actual filtering process and the input information in the training process, the same determination method as the determination method of the reference image block of the input image block is adopted to determine the reference image block of the image block to be filtered, and then the image block to be filtered and the reference image block of the image block to be filtered are input into the neural network-based filter to achieve filtering of the image to be filtered.
[0206] According to the image filtering method of the present embodiment, the image to be filtered is divided into N image blocks to be filtered using the same block division method as the training image used by the neural network-based filter in the training process, and for each of the N image blocks to be filtered, a reference image block of the image block to be filtered is determined according to the method of determining the reference image block of the input image block, and the image block to be filtered and the reference image block of the image block to be filtered are input into the neural network-based filter for filtering to obtain a filtered image. That is, in the present embodiment, the block division method in the actual use process of the neural network-based filter is consistent with the block division method used during training, and the method of determining the reference image block of the input image block is consistent with the method of determining the reference image block of the image block to be filtered, thereby further improving the filtering effect of the neural network-based filter.
[0207] It should be understood that Figures 13-46 are merely examples of the present application and should not be construed as limiting the present application.
[0208] An embodiment of the method of the present application has been described in detail above with reference to FIGS. 13 to 46, and an embodiment of the apparatus of the present application will now be described with reference to FIGS.
[0209] 47 is an exemplary block diagram of an image filtering device 10 according to one embodiment of the present application. The device 10 may be an electronic device or a part of an electronic device. As shown in FIG. 47, the image filtering device 10 includes: an acquisition unit 11 configured to acquire a ready-to-filter image; a division unit 12 configured to determine a neural network-based filter, and divide the image to be filtered according to a block division scheme corresponding to the neural network-based filter to obtain N (N is a positive integer) image blocks to be filtered, where the block division scheme is a block division scheme of a training image used by the neural network-based filter in a training process; and a filtering unit 13 configured to filter each of the N image blocks to be filtered using a neural network based filter to obtain a filtered image.
[0210] In some embodiments, the block division scheme of the training image includes determining M CTUs in the training image (M is a positive integer) as one training image block; The division unit 12 is configured to determine M CTUs in the image to be filtered as one image block to be filtered, thereby obtaining N image blocks to be filtered.
[0211] In some embodiments, the block division scheme of the training image includes determining P incomplete CTUs in the training image (P is a positive integer) as one training image block; The division unit 12 is configured to determine the P incomplete CTUs in the image to be filtered as one image block to be filtered, thereby obtaining N image blocks to be filtered.
[0212] In some embodiments, the neural network based filter is obtained by training it using an extended image block of the training image block; The filtering unit 13 is configured to, for each image block to be filtered of the N image blocks to be filtered, extend the image block to be filtered according to the extension scheme of the training image block to obtain an extended image block to be filtered, filter the extended image block to be filtered using a neural network-based filter to obtain a filtered extended image block, and determine an image region in the filtered extended image block corresponding to the image block to be filtered as the filtered image block corresponding to the image block to be filtered.
[0213] In some embodiments, the training image block expansion scheme comprises expanding at least one boundary region of the training image block outward; The filtering unit 13 is configured to extend outward at least one boundary region of the image block to be filtered to obtain an extended image block to be filtered.
[0214] In some embodiments, the training image comprises an input image, and input data for training the neural network-based filter comprises an input image block and a reference image block for the input image block, the input image block being obtained by performing image segmentation on the input image using a block segmentation scheme; The filtering unit 13 is configured to determine, for each of the N image blocks to be filtered, a reference image block of the image block to be filtered, and input the image block to be filtered and the reference image block of the image block to be filtered into a neural network-based filter for filtering to obtain a filtered image block of the image block to be filtered.
[0215] In some embodiments, the filtering unit 13 is configured to obtain a determination scheme of a reference image block of an input image block, the determination scheme being used to determine a corresponding reference image block based on at least one of spatial domain information and time domain information of the input image block, and to determine a reference image block of an image block to be filtered according to the determination scheme of the reference image block of the input image block.
[0216] In some embodiments, the reference image block of the input image block includes at least one of a time domain reference image block and a spatial domain reference image block of the input image block, and the reference image block of the image block to be filtered includes at least one of a time domain reference image block and a spatial domain reference image block of the image block to be filtered.
[0217] In some embodiments, the reference image blocks of the input image block also include spatial domain reference image blocks of the input image block; The filtering unit 13 is configured to determine the spatial domain reference image block of the image block to be filtered according to a method for determining the spatial domain reference image block of the input image block.
[0218] In some embodiments, the spatial domain reference image block of the input image block includes at least one of an image block located to the upper left of the input image block, an image block located to the left of the input image block, and an image block located above the input image block; The filtering unit 13 is configured to determine at least one of the image block located to the upper left of the image block to be filtered, the image block located to the left, and the image block located at the top in the image to be filtered as a spatial domain reference image block of the image block to be filtered.
[0219] In some embodiments, the reference image blocks of the input image block also include time-domain reference image blocks of the input image block; The filtering unit 13 is configured to determine the time-domain reference image block of the image block to be filtered according to a manner of determining the time-domain reference image block of the input image block.
[0220] In some embodiments, the time-domain reference image block of the input image block comprises an image block in a reference image of the input image at a position corresponding to the input image block; The filtering unit 13 is specifically configured to determine a reference image of the image to be filtered, and determine an image block in the reference image of the image to be filtered that is at a position corresponding to the image block to be filtered as a time-domain reference image block of the image block to be filtered.
[0221] In some embodiments, the training images include an input image and a target image corresponding to the input image, and the neural network-based filter is obtained by training using the input image blocks as input data and the target image blocks as targets, the input image blocks are obtained by performing image segmentation on the input image using the block segmentation method of the training images, and the target image blocks are obtained by performing image segmentation on the target image using the block segmentation method of the training images.
[0222] In some embodiments, the acquisition unit 11 is configured to reconstruct a current image to obtain a reconstructed image of the current image, and determine the reconstructed image as the image to be filtered.
[0223] In some embodiments, the filtering unit 13 is further configured to generate a reference image for prediction based on the filtered image and store it in the decoding cache.
[0224] In some embodiments, the filtering unit 13 is further configured to generate a display image based on the filtered image, input the display image to a display device for display, and skip storing the filtered image or a reprocessed image of the filtered image in the decoding cache.
[0225] 48 is an exemplary block diagram of an electronic device 30 configured to perform the above method embodiment, according to an embodiment of the present disclosure. As shown in FIG. 48, the electronic device 30 includes: The device may include a memory 31 and a processor 32, and the memory 31 is configured to store a computer program 33 and transmit the computer program 33 to the processor 32. In other words, the processor 32 can implement the method in the embodiment of the present application by calling and executing the computer program 33 from the memory 31.
[0226] For example, the processor 32 may be configured to carry out the steps of the above-described method according to instructions in the computer program 33 .
[0227] In some embodiments of the present application, the processor 32: This may include, but is not limited to, a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, or the like.
[0228] In some embodiments of the present application, the memory 31 includes: This may include, but is not limited to, volatile and / or nonvolatile memory. Nonvolatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), used as an external cache. By way of illustrative, but non-limiting example, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM).
[0229] In some embodiments of the present application, the computer program 33 can be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method of the present application. The one or more modules may be a series of computer program instruction segments that can complete a specific function, and the instruction segments are used to explain the execution process of the computer program 33 in the electronic device.
[0230] As shown in FIG. 48, the electronic device 30 further A transceiver 34 may be provided which may be connected to the processor 32 or the memory 31 .
[0231] The processor 32 controls the transceiver 34 to communicate with other devices, for example, to transmit information or data to other devices or to receive information or data transmitted by other devices. The transceiver 34 may include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas may be one or more.
[0232] It should be understood that the components of the electronic device 30 are connected by a bus system, which may include a power bus, a control bus, and a status signal bus in addition to a data bus.
[0233] According to one aspect of the present application, there is provided a computer storage medium having a computer program stored thereon, the computer program causing a computer to perform the method in the above method embodiment.
[0234] An embodiment of the present application further provides a computer program product comprising instructions for causing a computer to perform the method in the above method embodiment.
[0235] According to another aspect of the present application, there is provided a computer program product or computer program comprising computer instructions stored on a computer-readable storage medium, wherein a processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions to cause the electronic device to perform the method in the method embodiment described above.
Claims
1. 1. A method of image filtering performed by an electronic device, comprising: obtaining an image to be filtered; determining a neural network based filter; Dividing the image to be filtered according to a block division scheme corresponding to the neural network-based filter to obtain N image blocks to be filtered, where N is a positive integer, and the block division scheme is a block division scheme of a training image used by the neural network-based filter in a training process; filtering each of the N image blocks to be filtered using the neural network-based filter to obtain a filtered image; Including, The block division method of the training image includes determining P incomplete coding tree units in the training image as one training image block, where P is a positive integer; Dividing the image to be filtered according to a block division scheme corresponding to the neural network-based filter to obtain N image blocks to be filtered, The image filtering method includes determining the N image blocks to be filtered from the image to be filtered in a manner that P incomplete coding tree units are one image block to be filtered.
2. the neural network-based filter is obtained by training it using an extended image block of a training image block; The step of filtering each of the N image blocks to be filtered using the neural network-based filter to obtain a filtered image includes: For each of the N image blocks to be filtered, extend the image block to be filtered according to the extension method of the training image block to obtain an extended image block to be filtered; filtering the dilated to-be-filtered image block using the neural network-based filter to obtain a filtered dilated image block; determining an image region in the filtered extended image block corresponding to the image block to be filtered as a filtered image block corresponding to the image block to be filtered; The image filtering method of claim 1 .
3. The training image block expansion method includes expanding at least one boundary region of the training image block outward; The step of expanding the image block to be filtered according to the expansion manner of the training image block to obtain an expanded image block to be filtered includes: dilating at least one boundary region of the image block to be filtered outward to obtain the dilated image block to be filtered; The image filtering method according to claim 2 .
4. the training image includes an input image, and input data for training the neural network-based filter includes an input image block and a reference image block for the input image block, the input image block being obtained by performing image segmentation on the input image using the block segmentation scheme; The step of filtering each of the N image blocks to be filtered using the neural network-based filter to obtain a filtered image includes: For each of the N image blocks to be filtered, determining a reference image block for the image block to be filtered; inputting the image block to be filtered and a reference image block of the image block to be filtered into the neural network-based filter for filtering to obtain a filtered image block of the image block to be filtered; The image filtering method of claim 1 .
5. The step of determining a reference image block for the image block to be filtered includes: obtaining a determination scheme of a reference image block for the input image block, the determination scheme being used to determine a corresponding reference image block based on at least one of spatial domain information and temporal domain information of the input image block; determining a reference image block of the image block to be filtered according to a method for determining a reference image block of the input image block; 5. The image filtering method of claim 4.
6. the reference image blocks of the input image block include at least one of a time domain reference image block and a spatial domain reference image block of the input image block; the reference image block of the image block to be filtered includes at least one of a time domain reference image block and a spatial domain reference image block of the image block to be filtered; 6. The image filtering method according to claim 5.
7. The step of determining a reference image block of the image block to be filtered according to a method of determining a reference image block of the input image block includes: determining a spatial domain reference image block of the image block to be filtered according to a spatial domain reference image block determination manner of the input image block; 7. The image filtering method of claim 6.
8. the spatial domain reference image block of the input image block includes at least one of an image block located to the upper left of the input image block, an image block located to the left of the input image block, and an image block located above the input image block; The step of determining a spatial domain reference image block of the image block to be filtered according to a spatial domain reference image block determination manner of the input image block includes: determining at least one of an image block located to the upper left of the image block to be filtered, an image block located to the left of the image block to be filtered, and an image block located above the image block to be filtered as a spatial domain reference image block of the image block to be filtered; The image filtering method according to claim 7.
9. The step of determining a reference image block of the image block to be filtered according to a method of determining a reference image block of the input image block includes: determining a time-domain reference image block of the image block to be filtered according to a time-domain reference image block determination manner of the input image block; 6. The image filtering method according to claim 5.
10. the time-domain reference image block of the input image block includes an image block in a reference image of the input image at a position corresponding to the input image block; The step of determining a time-domain reference image block of the image block to be filtered according to a manner of determining a time-domain reference image block of the input image block includes: determining a reference image for the image to be filtered; determining an image block in a reference image of the image to be filtered, the image block being located at a position corresponding to the image block to be filtered, as a time-domain reference image block of the image block to be filtered; 10. The image filtering method of claim 9.
11. the training images include an input image and a target image corresponding to the input image, the neural network-based filter is obtained by training using input image blocks as input data and target image blocks as targets, the input image blocks are obtained by performing image segmentation on the input image using a block segmentation scheme of the training images, and the target image blocks are obtained by performing image segmentation on the target image using the block segmentation scheme of the training images. The image filtering method of claim 1 .
12. The step of acquiring the image to be filtered includes: reconstructing an original image to obtain a reconstructed image of the original image; determining the reconstructed image as the image to be filtered; The image filtering method of claim 1 .
13. An image filtering device, comprising: an acquisition unit configured to acquire a to-be-filtered image; a segmentation unit configured to determine a neural network-based filter, and segment the image to be filtered according to a block segmentation scheme corresponding to the neural network-based filter to obtain N image blocks to be filtered, where N is a positive integer, and the block segmentation scheme is a block segmentation scheme of a training image used by the neural network-based filter in a training process; a filtering unit configured to filter each of the N image blocks to be filtered using the neural network-based filter to obtain a filtered image; Equipped with The block division method of the training image includes determining P incomplete coding tree units in the training image as one training image block, where P is a positive integer; The image filtering device, wherein the division unit is configured to determine the N image blocks to be filtered from the image to be filtered in a manner that P incomplete coding tree units are one image block to be filtered.
14. Electronic equipment configured to perform the image filtering method according to any one of claims 1 to 12.
15. A computer program product for causing a computer to carry out the image filtering method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Image restoration filter and learning apparatus
JP2019087778A
Encoding program, decoding program, encoding device, decoding device, encoding method, and decoding method
JP2020198463A
Signal processing device and signal processing method, system, learning method, and program
JP2021090135A
Convolutional-neutral-network based filter for video coding
US20210329286A1
Image filter device, image decoding device, and image coding device
WO2019031410A1