Decoding device, encoding device, decoding method, and encoding method

The decoding and encoding devices apply multiple loop filtering processes in parallel, addressing delays and improving efficiency and image quality by using both non-neural and neural network filtering techniques independently, thereby optimizing video coding for increased digital data handling.

WO2026150720A1PCT designated stage Publication Date: 2026-07-16PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2025-12-09
Publication Date
2026-07-16

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently handling the increasing volume of digital video data, particularly in terms of encoding efficiency, image quality, processing load, and circuit size, without causing delays in loop filtering processes.

Method used

A decoding and encoding device that applies multiple loop filtering processes, including both non-neural network and neural network processes, where non-neural network processes are performed independently of neural network outputs, allowing parallel execution and reducing delays.

Benefits of technology

The solution enables improved encoding efficiency, image quality, reduced processing load, and faster processing speeds by allowing parallel execution of loop filtering processes without waiting for neural network outputs, thus enhancing overall video coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025042877_16072026_PF_FP_ABST
    Figure JP2025042877_16072026_PF_FP_ABST
Patent Text Reader

Abstract

A decoding device (200) comprises a circuit (b1) and a memory (b2) connected to the circuit (b1). In an operation, the circuit (b1) applies a plurality of loop filter processes to a decoding-target current block. The plurality of loop filter processes include one or more non-neural network loop filter processes, each of which is a loop filter process that does not use a neural network, and one or more neural network loop filter processes, each of which is a loop filter process that uses a neural network. Said one or more non-neural network loop filter processes are applied without using output data obtained through said one or more neural network loop filter processes.
Need to check novelty before this filing date? Find Prior Art

Description

Decoding device, encoding device, decoding method, and encoding method

[0001] This disclosure relates to a decoding device, etc.

[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Coding). With this advancement, there is a constant need to provide improvements and optimizations to video coding technology to handle the ever-increasing volume of digital video data in various applications. This disclosure relates to further advancements, improvements, and optimizations in video coding.

[0003] Non-patent document 1 relates to an example of a conventional standard concerning the video coding technology described above.

[0004] H. 266 (ISO / IEC 23090-3) / VVC (Versatile Video Coding)

[0005] Regarding the encoding methods described above, there is a need for proposals for new methods to improve encoding efficiency, image quality, processing load, circuit size, or to appropriately select elements or actions such as filters, blocks, size, motion vectors, reference pictures, or reference blocks.

[0006] This disclosure provides a configuration or method that can contribute to one or more of the following: improved encoding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, improved processing applied to the decoded image, or provision of new processing. This disclosure may also include configurations or methods that can contribute to other benefits not mentioned above.

[0007] For example, a decoding device according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit applies a plurality of loop filtering processes to the current block to be decoded during operation, and the plurality of loop filtering processes include one or more non-neural network loop filtering processes, each of which does not use a neural network, and one or more neural network loop filtering processes, each of which uses a neural network, and the one or more non-neural network loop filtering processes are applied without using the output data obtained by the one or more neural network loop filtering processes.

[0008] Each embodiment, or any part thereof, of the present disclosure enables at least one of the following: improved encoding efficiency, improved image quality, reduced encoding / decoding processing load, reduced circuit size, or improved encoding / decoding processing speed. Alternatively, each embodiment, or any part thereof, of the present disclosure enables appropriate selection of components / operations such as filters, blocks, sizes, motion vectors, reference pictures, and reference blocks in encoding and decoding. The present disclosure also includes disclosures of configurations or methods that may provide benefits other than those mentioned above, such as configurations or methods that improve encoding efficiency while suppressing an increase in processing load.

[0009] Further advantages and effects of one aspect of this disclosure will be made apparent from the specification and drawings. Such advantages and / or effects may be obtained by several embodiments and features described in the specification and drawings, but not all of them are necessarily provided to obtain one or more advantages and / or effects.

[0010] These general or specific embodiments may be implemented as a system, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of system, method, integrated circuit, computer program, and recording medium.

[0011] A configuration or method relating to one aspect of this disclosure may contribute to one or more of the following: improved encoding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. A configuration or method relating to one aspect of this disclosure may also contribute to other benefits not mentioned above.

[0012] This is a schematic diagram showing an example of the configuration of a transmission system according to an embodiment. This is a block diagram showing an example of the implementation of an encoding device according to an embodiment. This is a block diagram showing an example of the configuration of an encoding device according to an embodiment. This is a flowchart showing an example of the overall encoding process by the encoding device according to an embodiment. This is a flowchart showing an example of the processing for each block by the encoding device according to an embodiment. This is a block diagram showing an example of the implementation of a decoding device according to an embodiment. This is a block diagram showing an example of the configuration of a decoding device according to an embodiment. This is a flowchart showing an example of the overall decoding process by the decoding device according to an embodiment. This is a flowchart showing an example of the processing for each block by the decoding device according to an embodiment. This is a schematic diagram showing another example of the configuration of a transmission system according to an embodiment. This is a schematic diagram showing an example of the configuration of a transmitting device according to an embodiment. This is a block diagram showing an example of the implementation of a transmitting device according to an embodiment. This is a schematic diagram showing an example of the configuration of a receiving device according to an embodiment. This is a block diagram showing an example of the implementation of a receiving device according to an embodiment. This is a schematic diagram showing an example of the configuration of a bitstream generation device according to an embodiment. This is a block diagram showing an example of the implementation of a bitstream generation device according to an embodiment. This is a schematic diagram showing an example of the configuration of a storage medium and a computer according to an embodiment. This is a schematic diagram showing an example of the configuration of a computer according to an embodiment. This is a diagram showing an example of the hierarchical structure of data in a stream. This is a diagram showing an example of the configuration of a slice. This is a diagram showing an example of the configuration of a tile. This is a diagram showing an example of the configuration of a time-scalable stream. This is a diagram showing an example of the configuration of a stream using multilayer encoding. This is a block diagram showing an example of the configuration of a loop filter section. This is a diagram showing an example of the filter shape used in an ALF (adaptive loop filter). This is a diagram showing another example of the filter shape used in an ALF. This is a diagram showing another example of the filter shape used in an ALF. This is a diagram showing an example of a CCALF. This is a diagram showing the filter shape of a CCALF. This is a block diagram showing an example of the detailed configuration of a loop filter section that functions as a DBF. This is a flowchart showing an example of the processing in the DBF processing section. This is a flowchart showing another example of the processing in the DBF processing section.This figure shows an example of a DBF having filter characteristics symmetric with respect to block boundaries. This figure illustrates an example of a block boundary on which DBF processing is performed. This figure shows an example of a block boundary strength Bs value. This block diagram shows an example of the configuration of a loop filter section. This flowchart shows an example of processing performed in the prediction section of an encoding device. This flowchart shows another example of processing performed in the prediction section of an encoding device. This flowchart shows another example of processing performed in the prediction section of an encoding device. This conceptual diagram shows another example of a reference picture. This conceptual diagram shows another example of generating a generated reference picture. This conceptual diagram shows another example of generating a generated reference picture. This conceptual diagram shows an example of a generated picture. This conceptual diagram shows another example of generating a generated picture. This conceptual diagram shows another example of generating a generated picture. This conceptual diagram shows a pipeline processing. This block diagram shows an example of a configuration used in NN processing. This block diagram shows another example of a configuration used in NN processing. This conceptual diagram shows a patch used in NN processing. This conceptual diagram shows an extended patch containing an extended region. This conceptual diagram shows an image containing two patches with overlapping regions. This conceptual diagram shows patches defined in each picture. This conceptual diagram shows an example of a patch generation method. This conceptual diagram shows another example of a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a diagram showing a first configuration of the loop filter section of a decoding device according to a reference example. This is a diagram showing a second configuration of the loop filter section of a decoding device according to a reference example. This is a diagram showing an example of realizing data transfer between a CPU and a GPU using a data bus to which DMA technology is applied. This is a flowchart outlining the loop filtering process by the encoding device of this disclosure. This is a flowchart outlining the loop filtering process by the decoding device of this disclosure. This is a diagram for explaining an example of the application order of the loop filtering process performed by the decoding device of the first embodiment. This is a block diagram showing a simplified representation of the configuration of the loop filter section according to the first embodiment. This is a diagram showing a first configuration example of the loop filter section according to the first embodiment. This is a diagram showing a second configuration example of the loop filter section according to the first embodiment.This is another block diagram showing a simplified configuration of the loop filter unit according to the first embodiment. This is a diagram showing a third configuration example of the loop filter unit according to the first embodiment. This is a diagram showing a fourth configuration example of the loop filter unit according to the first embodiment. This is a diagram showing the correspondence between the order of image display and the order of image decoding. This is the first diagram for explaining an example of timing when the loop filter unit performs one or more neural network loop filtering operations. This is the second diagram for explaining an example of timing when the loop filter unit performs one or more neural network loop filtering operations. This is the third diagram for explaining an example of timing when the loop filter unit performs one or more neural network loop filtering operations. This is the fourth diagram for explaining an example of timing when the loop filter unit performs one or more neural network loop filtering operations. This is a diagram showing the correspondence between the order of image display and the order of image decoding. This is the first diagram for explaining an example of timing when the loop filter unit performs one or more neural network loop filtering operations. This is the second diagram for explaining an example of timing when the loop filter unit performs one or more neural network loop filtering operations. This is the third diagram for explaining an example of timing when the loop filter unit performs one or more neural network loop filtering operations. This is the fourth figure illustrating an example of timing when the loop filter unit performs one or more neural network loop filter processes. This figure shows a fifth configuration example of the loop filter unit according to the first embodiment. This figure illustrates an example of the application order of loop filter processing performed by the decoding device according to the second embodiment. This is a simplified block diagram showing the configuration of the loop filter unit according to the second embodiment. This figure shows a first configuration example of the loop filter unit according to the second embodiment. This figure shows a second configuration example of the loop filter unit according to the second embodiment. This is another simplified block diagram showing the configuration of the loop filter unit according to the second embodiment. This figure shows a third configuration example of the loop filter unit according to the second embodiment. This figure shows a fourth configuration example of the loop filter unit according to the second embodiment. This figure shows a fifth configuration example of the loop filter unit according to the second embodiment. This figure shows a sixth configuration example of the loop filter unit according to the second embodiment.This is a diagram showing a seventh configuration example of the loop filter unit according to the second embodiment. This is a diagram showing an eighth configuration example of the loop filter unit according to the second embodiment. This is a block diagram showing a simplified configuration of the loop filter unit according to a modified version of the second embodiment. This is a diagram showing a first configuration example of the loop filter unit according to a modified version of the second embodiment. This is a diagram showing a second configuration example of the loop filter unit according to a modified version of the second embodiment. This is another block diagram showing a simplified configuration of the loop filter unit according to a modified version of the second embodiment. This is a diagram showing a third configuration example of the loop filter unit according to a modified version of the second embodiment. This is a diagram showing a fourth configuration example of the loop filter unit according to a modified version of the second embodiment. This is a first block diagram showing a simplified configuration of the loop filter unit according to the third embodiment. This is a second block diagram showing a simplified configuration of the loop filter unit according to the third embodiment. This is a third block diagram showing a simplified configuration of the loop filter unit according to the third embodiment. This is a fourth block diagram showing a simplified configuration of the loop filter unit according to the third embodiment. This is a fifth block diagram showing a simplified configuration of the loop filter unit according to the third embodiment. This is a sixth block diagram showing a simplified configuration of the loop filter section according to the third embodiment. This is a seventh block diagram showing a simplified configuration of the loop filter section according to the third embodiment. This figure shows an example of the location in the bitstream where encoded NNLF processing granularity information is included. This figure shows another example of the location in the bitstream where encoded NNLF processing granularity information is included. This figure shows an example of NNLF processing granularity information encoded in the bitstream. This is a table showing an example of the granularity value of the NNLF processing indicated by the NNLF processing granularity information. This figure shows an example of the location in the bitstream where encoded computing power information is included. This figure shows another example of the location in the bitstream where encoded computing power information is included. This figure shows an example of computing power information to be encoded in the bitstream. This figure shows an example of the location in the bitstream where encoded flags are included. This figure shows another example of the location in the bitstream where encoded flags are included. This figure shows an example of flags to be encoded in the bitstream. This figure shows an example of the location in the bitstream where encoded flags are included.This figure shows another example of the location where encoded flags are included in a bitstream. This figure shows an example of flags to be encoded in a bitstream. This figure shows an example of the location where encoded NNLF processing delay information is included in a bitstream. This figure shows another example of the location where encoded NNLF processing delay information is included in a bitstream. This figure shows an example of NNLF processing delay information to be encoded in a bitstream. This figure shows an example of the overall configuration of a content supply system that realizes a content distribution service. This figure shows an example of the configuration of a content distribution system. This figure shows an example of a web page display screen. This figure shows an example of a web page display screen. This figure shows an example of a smartphone. This block diagram shows an example of a smartphone configuration.

[0013] [Introduction] This disclosure relates to applying a neural network-based loop filtering process to a current block to be decoded or encoded. More specifically, this disclosure relates to applying multiple loop filtering processes to a current block to be decoded or encoded. In this disclosure, multiple loop filtering processes include one or more non-neural network loop filtering processes, each of which does not use a neural network, and one or more neural network loop filtering processes, each of which uses a neural network. Furthermore, this disclosure provides an apparatus and method that can materialize the application of multiple loop filtering processes to a current block to be decoded or encoded.

[0014] The following sections explain some terms and describe techniques that can accommodate the application of multiple loop filtering processes in this disclosure.

[0015] For example, a proposed configuration for a decoding device involves applying some non-neural network loop filtering to the current block to be decoded, then performing some other non-neural network loop filtering and neural network loop filtering in parallel, and finally applying the remaining non-neural network loop filtering. Similarly, a proposed configuration for an encoding device involves applying some non-neural network loop filtering to the current block to be encoded, then performing some other non-neural network loop filtering and neural network loop filtering in parallel, and finally applying the remaining non-neural network loop filtering.

[0016] However, in the above proposed decoding device, some non-neural network loop filtering and neural network loop filtering are performed in parallel, and the remaining non-neural network loop filtering is performed using the output data obtained from the neural network loop filtering. As a result, data is exchanged in blocks. Therefore, when the above proposed decoding device applies multiple loop filtering to a single picture, a significant delay occurs until the multiple loop filtering is completed. Similarly, when the above proposed encoding device applies multiple loop filtering to a single picture, a significant delay occurs until the multiple loop filtering is completed.

[0017] Therefore, the decoding device of Example 1 comprises a circuit and a memory connected to the circuit, wherein the circuit applies a plurality of loop filtering processes to the current block to be decoded during operation, and the plurality of loop filtering processes include one or more non-neural network loop filtering processes, each of which does not use a neural network, and one or more neural network loop filtering processes, each of which uses a neural network, and the one or more non-neural network loop filtering processes are applied without using the output data obtained by the one or more neural network loop filtering processes, thus the decoding device.

[0018] As a result, the decoder can perform one or more non-neural network loop filtering processes without depending on the output data obtained from one or more neural network loop filtering processes. Therefore, the decoder can smoothly perform one or more non-neural network loop filtering processes and one or more neural network loop filtering processes, and can suppress the delays that occur until multiple loop filtering processes are completed.

[0019] Furthermore, the decoding device in Example 2 may be the decoding device in Example 1, wherein the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes is used as input data for one of the one or more neural network loop filtering processes.

[0020] As a result, the decoding device performs a single neural network loop filter using the output data obtained from the last non-neural network loop filter applied to the current block, thereby improving the quality of the reconstructed image and increasing the image compression ratio.

[0021] Furthermore, the decoding device of Example 3 may be the decoding device of Example 1, wherein the input data for one of the one or more non-neural network loop filtering processes is the same as the input data for one of the one or more neural network loop filtering processes, and the circuit is a decoding device that combines the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes with the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes.

[0022] This allows the decoder to perform one or more non-neural network loop filtering processes and one or more neural network loop filtering processes in parallel. Therefore, the decoder can further reduce the delay that occurs while multiple loop filtering processes are being completed.

[0023] Furthermore, the decoding device in Example 4 may be the decoding device in Example 3, wherein the coupling is one of the following: addition, weighted addition, and calculation of the absolute value of the difference.

[0024] This allows the decoder to combine the output data obtained from the last non-neural network loop filter applied to the current block with the output data obtained from the last neural network loop filter applied to the current block by addition, weighted addition, or calculation of the absolute value of the difference. Therefore, the decoder can improve the quality of the reconstructed image and improve the image compression ratio by utilizing the output data obtained from one or more neural network loop filters.

[0025] Furthermore, the decoding device in Example 5 may be any decoding device from Examples 1 to 4, wherein each of the one or more non-neural network loop filter processes is one of LMCS (Luma Mapping with Chroma Scaling) processing, DBF (Deblocking Filter) processing, SAO (Sample Adaptive Offset) processing, and ALF (Adaptive Loop Filter) processing.

[0026] As a result, the decoding device can perform at least one of the following processes: LMCS processing, DBF processing, SAO processing, and ALF processing, using one or more non-neural network processes.

[0027] Furthermore, the decoding device in Example 6 may be any decoding device from Examples 1 to 5, wherein the circuit uses the output data obtained by the non-neural network loop filtering process that is applied last to the current block among the one or more non-neural network loop filtering processes as the display image.

[0028] This allows the decoder to use the output data obtained from the last non-neural network loop filtering process applied to the current block as the display image, without waiting for the results of one or more neural network loop filtering processes. Therefore, the decoder can further reduce the delay in displaying the reconstructed image.

[0029] Furthermore, the decoding device in Example 7 may be any of the decoding devices in Examples 1 to 5, wherein the circuit uses the output data obtained by the neural network loop filtering process that is applied last to the current block among the one or more neural network loop filtering processes as the display image.

[0030] As a result, the decoding device uses the output data obtained from the neural network loop filtering process applied last to the current block as the display image, enabling it to display a reconstructed image with superior image quality.

[0031] Further, the decoding device of Example 8 is any one of the decoding devices of Examples 1 to 7, the memory includes a DPB (Decoded Picture Buffer), and the circuit stores, in the DPB, as a candidate for a reference image for a subsequent picture, output data obtained by the non-neural network loop filter process that is last applied to the current block among the one or more non-neural network loop filter processes. It may be a decoding device.

[0032] As a result, the decoding device can perform the decoding process of the subsequent picture by using, as a reference image, the output data obtained by the non-neural network loop filter process without waiting for the results of the one or more neural network loop filter processes. Therefore, the decoding device can further suppress the delay that occurs until a plurality of loop filter processes are completed.

[0033] Further, the decoding device of Example 9 is any one of the decoding devices of Examples 1 to 7, the memory includes a DPB (Decoded Picture Buffer), and the circuit stores, in the DPB, as a candidate for a reference image for a subsequent picture, output data obtained by the neural network loop filter process that is last applied to the current block among the one or more neural network loop filter processes. It may be a decoding device.

[0034] As a result, the decoding device can use, as a reference image, the output data obtained by the one or more neural network loop filter processes, and thus can use a reconstructed image with better image quality in the decoding process of the subsequent picture. Therefore, the decoding device can improve the quality of the reconstructed image and the image compression rate.

[0035] Further, the decoding device of Example 10 is any one of the decoding devices of Examples 1 to 9, and the circuit may be a decoding device that performs the one or more neural network loop filter processes on the data of the specific area when the data of the specific area is obtained as input data for the one or more neural network loop filter processes.

[0036] As a result, when the decoder obtains the data in the specific area, since the decoder performs one or more neural network loop filter processes on the data in the specific area, the one or more neural network loop filter processes can be started at an appropriate timing. Then, the decoder can perform one or more neural network loop filter processes in an appropriate data unit corresponding to the data in the specific area.

[0037] Further, the decoder of Example 11 may be the decoder of Example 10, where the specific area is any one of a processing block, a CTU (Coding Tree Unit), a slice, a tile, a subpicture, and a picture.

[0038] As a result, when the decoder obtains the data in any unit of a processing block, a CTU, a slice, a tile, a subpicture, and a picture, the decoder can perform one or more neural network loop filter processes.

[0039] Further, the decoder of Example 12 may be the decoder of Example 10 or 11, where the circuit decodes information indicating the specific area from the bitstream.

[0040] As a result, the decoder can specify the unit of data for which one or more neural network loop filter processes are to be performed based on the decoded information.

[0041] Further, the decoder of Example 13 may be any of the decoders of Examples 1 to 12, where the circuit decodes information from the bitstream indicating whether to use, as the display image, (i) the output data obtained by the non-neural network loop filter process that was last applied to the current block among the one or more non-neural network loop filter processes, or (ii) the output data obtained by the neural network loop filter process that was last applied to the current block among the one or more neural network loop filter processes when using the picture including the current block as the display image.

[0042] This allows the decoding device to identify the output data to be used as the display image based on the decoded information.

[0043] Furthermore, the decoding device of Example 14 may be any decoding device of Examples 1 to 13, wherein the circuit decodes from the bitstream information indicating whether, when using the picture including the current block as a reference image for a subsequent picture, (i) the output data obtained by the non-neural network loop filtering process applied last to the current block among the one or more non-neural network loop filtering processes is used as the reference image, or (ii) the output data obtained by the neural network loop filtering process applied last to the current block among the one or more neural network loop filtering processes is used as the reference image.

[0044] This allows the decoding device to identify the output data to be used as a reference image based on the decoded information.

[0045] Furthermore, the decoding device in Example 15 may be any decoding device from Examples 1 to 14, wherein the circuit decodes information from the bitstream indicating whether or not to apply the one or more neural network loop filter processes to the current block.

[0046] This allows the decoding device to decide whether or not to perform one or more neural network loop filtering processes based on the decoded information.

[0047] Furthermore, the decoding device of Example 16 may be the decoding device of Example 1 or 2, wherein the circuit includes a first processor which is a CPU (Central Processing Unit) and a second processor which is one of a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a TPU (Tensor Processing Unit), the first processor performs one or more non-neural network loop filtering processes, and the second processor performs one or more neural network loop filtering processes in parallel with the one or more non-neural network loop filtering processes performed by the first processor.

[0048] This allows the decoder to perform one or more non-neural network loop filtering processes and one or more neural network loop filtering processes in parallel. Therefore, the decoder can further reduce the delay that occurs while multiple loop filtering processes are being completed.

[0049] Furthermore, the decoding device of Example 17 may be any decoding device of Examples 1, 3, and 4, wherein the circuit includes a first processor which is a CPU (Central Processing Unit) and a second processor which is one of a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a TPU (Tensor Processing Unit), and the circuit performs one or more non-neural network loop filtering processes in a predetermined processing order, the first processor performs the Nth non-neural network loop filtering processes from the beginning in the processing order of the one or more non-neural network loop filtering processes, the second processor performs the remaining non-neural network loop filtering processes not performed by the first processor, and performs one or more neural network loop filtering processes in parallel with the remaining non-neural network loop filtering processes.

[0050] This allows the decoder to perform non-neural network loop filtering and one or more neural network loop filtering processes in parallel, which are performed by the second processor. Therefore, the decoder can further reduce the delays that occur while multiple loop filtering processes are being completed.

[0051] Furthermore, the encoding device of Example 18 comprises a circuit and a memory connected to the circuit, wherein the circuit applies a plurality of loop filter processes to the current block to be encoded during operation, and the plurality of loop filter processes include one or more non-neural network loop filter processes, each of which does not use a neural network, and one or more neural network loop filter processes, each of which uses a neural network, and the one or more non-neural network loop filter processes are applied without using the output data obtained by the one or more neural network loop filter processes, thus making it an encoding device.

[0052] As a result, the encoding device can perform one or more non-neural network loop filtering processes without depending on the output data obtained from one or more neural network loop filtering processes. Therefore, the encoding device can smoothly perform one or more non-neural network loop filtering processes and one or more neural network loop filtering processes, and can suppress the delays that occur until multiple loop filtering processes are completed.

[0053] Furthermore, the encoding device of Example 19 may be the encoding device of Example 18, wherein the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes is used as input data for one of the one or more neural network loop filtering processes.

[0054] As a result, the encoding device performs a single neural network loop filter using the output data obtained from the last non-neural network loop filter applied to the current block, thereby improving the quality of the reconstructed image and increasing the image compression ratio.

[0055] Furthermore, the encoding device of Example 20 may be the encoding device of Example 18, wherein the input data for one of the one or more non-neural network loop filtering processes is the same as the input data for one of the one or more neural network loop filtering processes, and the circuit is an encoding device that combines the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes with the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes.

[0056] This allows the encoding device to perform one or more non-neural network loop filtering processes and one or more neural network loop filtering processes in parallel. Therefore, the encoding device can further reduce the delay that occurs while multiple loop filtering processes are being completed.

[0057] Furthermore, the encoding device in Example 21 may be the encoding device in Example 20, wherein the coupling is one of the following: addition, weighted addition, and calculation of the absolute value of the difference.

[0058] This allows the encoding device to combine the output data obtained from the last non-neural network loop filter applied to the current block with the output data obtained from the last neural network loop filter applied to the current block by addition, weighted addition, or calculation of the absolute value of the difference. Therefore, the encoding device can improve the quality of the reconstructed image and improve the image compression ratio by utilizing the output data obtained from one or more neural network loop filters.

[0059] Furthermore, the encoding device of Example 22 may be any encoding device from Examples 18 to 21, wherein each of the one or more non-neural network loop filter processes is one of LMCS (Luma Mapping with Chroma Scaling) processing, DBF (Deblocking Filter) processing, SAO (Sample Adaptive Offset) processing, and ALF (Adaptive Loop Filter) processing.

[0060] As a result, the encoding device can perform at least one of the following processes: LMCS processing, DBF processing, SAO processing, and ALF processing, using one or more non-neural network processes.

[0061] Furthermore, the encoding device in Example 23 may be any of the encoding devices in Examples 18 to 22, wherein the circuit uses the output data obtained by the non-neural network loop filtering process that is applied last to the current block among the one or more non-neural network loop filtering processes as the display image.

[0062] This allows the encoding device to use the output data obtained from the last non-neural network loop filtering process applied to the current block as the display image without waiting for the results of one or more neural network loop filtering processes. Therefore, the encoding device can further reduce the delay in displaying the reconstructed image. Note that the encoding device using the output data as the display image may also involve encoding information such as the display timing when the output data is used as the display image.

[0063] Furthermore, the encoding device of Example 24 may be any of the encoding devices of Examples 18 to 22, wherein the circuit uses the output data obtained by the neural network loop filtering process that is applied last to the current block among the one or more neural network loop filtering processes as the display image.

[0064] As a result, the encoding device uses the output data obtained from the neural network loop filtering process applied last to the current block as the display image, enabling it to display a reconstructed image with superior image quality. Note that the encoding device's use of the output data as the display image may also involve encoding information such as the display timing when the output data is used as the display image.

[0065] Furthermore, the encoding device of Example 25 may be any encoding device of Examples 18 to 24, wherein the memory includes a DPB (Decoded Picture Buffer), and the circuit stores the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes in the DPB as a candidate reference image for the subsequent picture.

[0066] This allows the encoding device to perform encoding on subsequent pictures by using the output data obtained from non-neural network loop filtering as a reference image, without waiting for the results of one or more neural network loop filtering processes. Therefore, the encoding device can further reduce the delay that occurs while multiple loop filtering processes are being completed.

[0067] Furthermore, the encoding device of Example 26 may be any encoding device of Examples 18 to 24, wherein the memory includes a DPB (Decoded Picture Buffer), and the circuit stores the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes in the DPB as a candidate reference image for the subsequent picture.

[0068] As a result, the encoding device can use output data obtained from one or more neural network loop filter processes as a reference image, enabling it to use a reconstructed image with superior image quality in the encoding process of subsequent pictures. Therefore, the encoding device can improve the quality of the reconstructed image and improve the image compression ratio.

[0069] Furthermore, the encoding device of Example 27 may be any of the encoding devices of Examples 18 to 26, wherein the circuit is an encoding device that performs the one or more neural network loop filter processing on the data in the specific region when the data in the specific region is obtained as input data for the one or more neural network loop filter processing.

[0070] As a result, the encoding device performs one or more neural network loop filtering processes on the data in a specific region as soon as data in that region is obtained, allowing one or more neural network loop filtering processes to be started at the appropriate timing. Furthermore, the encoding device can perform one or more neural network loop filtering processes on appropriate data units corresponding to the data in the specific region.

[0071] Furthermore, the encoding device of Example 28 may be the encoding device of Example 27, wherein the specific region is one of a processing block, a CTU (Coding Tree Unit), a slice, a tile, a subpicture, and a picture.

[0072] This allows the encoding device to perform one or more neural network loop filtering operations as soon as it obtains data in any unit of processing block, CTU, slice, tile, subpicture, or picture.

[0073] Furthermore, the encoding device in Example 29 may be the encoding device in Example 27 or 28, wherein the circuit encodes information indicating the specific region into a bitstream.

[0074] This allows the encoding device to specify a unit of data for the decoding device to perform one or more neural network loop filtering processes.

[0075] Furthermore, the encoding device of Example 30 may be any encoding device of Examples 18 to 29, wherein the circuit encodes information into a bitstream indicating whether, when using a picture including the current block as a display image, (i) the output data obtained by the non-neural network loop filter processing applied last to the current block among the one or more non-neural network loop filter processing is used as the display image, or (ii) the output data obtained by the neural network loop filter processing applied last to the current block among the one or more neural network loop filter processing is used as the display image.

[0076] This allows the encoding device to specify the output data that the decoding device will use as the display image.

[0077] Furthermore, the encoding device of Example 31 may be any encoding device of Examples 18 to 30, wherein the circuit encodes information into a bitstream indicating whether, when using the picture including the current block as a reference image for a subsequent picture, (i) the output data obtained by the non-neural network loop filter processing applied last to the current block among the one or more non-neural network loop filter processings is used as the reference image, or (ii) the output data obtained by the neural network loop filter processing applied last to the current block among the one or more neural network loop filter processings is used as the reference image.

[0078] This allows the encoding device to specify the output data that the decoding device will use as a reference image.

[0079] Furthermore, the encoding device of Example 32 may be any encoding device from Examples 18 to 31, wherein the circuit encodes information into a bitstream indicating whether or not to apply the one or more neural network loop filter processes to the current block.

[0080] This allows the encoding device to specify whether or not the decoding device performs one or more neural network loop filtering operations.

[0081] Furthermore, the encoding device of Example 33 may be the encoding device of Example 18 or 19, wherein the circuit includes a first processor which is a CPU (Central Processing Unit) and a second processor which is one of a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a TPU (Tensor Processing Unit), wherein the first processor performs one or more non-neural network loop filtering processes, and the second processor performs one or more neural network loop filtering processes in parallel with the one or more non-neural network loop filtering processes performed by the first processor.

[0082] This allows the encoding device to perform one or more non-neural network loop filtering processes and one or more neural network loop filtering processes in parallel. Therefore, the encoding device can further reduce the delay that occurs while multiple loop filtering processes are being completed.

[0083] Furthermore, the encoding device of Example 34 may be any encoding device of Examples 18, 20, and 21, wherein the circuit includes a first processor which is a CPU (Central Processing Unit) and a second processor which is any of a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a TPU (Tensor Processing Unit), and the circuit performs one or more non-neural network loop filtering processes in a predetermined processing order, the first processor performs the Nth non-neural network loop filtering processes from the beginning in the processing order of the one or more non-neural network loop filtering processes, the second processor performs the remaining non-neural network loop filtering processes not performed by the first processor, and performs one or more neural network loop filtering processes in parallel with the remaining non-neural network loop filtering processes.

[0084] This allows the encoding device to perform non-neural network loop filtering and one or more neural network loop filtering processes in parallel, which are performed by the second processor. Therefore, the encoding device can further reduce the delay that occurs while multiple loop filtering processes are being completed.

[0085] Furthermore, the decoding method of Example 35 applies multiple loop filtering processes to the current block to be decoded, and the multiple loop filtering processes include one or more non-neural network loop filtering processes, each of which does not use a neural network, and one or more neural network loop filtering processes, each of which uses a neural network, and the one or more non-neural network loop filtering processes are applied without using the output data obtained by the one or more neural network loop filtering processes.

[0086] As a result, the decoding method can perform one or more non-neural network loop filtering processes without depending on the output data obtained from one or more neural network loop filtering processes. Therefore, the decoding method can smoothly perform one or more non-neural network loop filtering processes and one or more neural network loop filtering processes, and can suppress the delays that occur until multiple loop filtering processes are completed.

[0087] Furthermore, the encoding method of Example 36 applies multiple loop filter processes to the current block to be encoded, and the multiple loop filter processes include one or more non-neural network loop filter processes, each of which does not use a neural network, and one or more neural network loop filter processes, each of which uses a neural network, and the one or more non-neural network loop filter processes are applied without using the output data obtained by the one or more neural network loop filter processes.

[0088] As a result, the encoding method can perform one or more non-neural network loop filtering processes without depending on the output data obtained from one or more neural network loop filtering processes. Therefore, the encoding method can smoothly perform one or more non-neural network loop filtering processes and one or more neural network loop filtering processes, and can suppress the delays that occur until multiple loop filtering processes are completed.

[0089] Furthermore, these comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0090] [Definition of Terms] Each term may be defined as follows, for example:

[0091] (1) Image: A unit of data composed of a collection of pixels, consisting of pictures and smaller blocks, and includes both still images and videos.

[0092] (2) Chroma Chroma represents a sample sequence or single sample that represents one of two color difference signals. The two color difference signals may be represented by the symbols Cb and Cr. Chroma is an adjective, for example, represented by the symbols Cb and Cr, or U and V. The term chrominance can also be used instead of chroma.

[0093] (3) Luminance (luma) Luminance refers to a sample sequence or a single sample representing a monochrome signal associated with the primary colors. Luminance is an adjective represented by the symbols Y or L. The term luminance can also be used instead of lumina.

[0094] (4) Component A component represents a single array or sample. A component may be, for example, one of the three arrays (i.e., luminance and two color differences) that make up a color format picture, or a single sample, or one array or a single sample that makes up a monochrome format picture.

[0095] (5) A picture is an image processing unit composed of a set of picture samples, and is sometimes called a frame or field. A picture can be a set of luminance samples, a set of two color difference samples, or a set of both luminance samples and a set of two color difference samples. A set of samples can also be described as a matrix of samples.

[0096] (5-1) I-Picture An I-picture is a picture encoded using only intra-prediction. I-pictures can be decoded independently. I-pictures are also called intra-pictures, intra-frames, keypictures, or keyframes.

[0097] (5-2) P-Picture A P-Picture is a picture encoded using single prediction. In other words, only one picture is referenced to encode the P-Picture. Single prediction is also called unidirectional prediction.

[0098] (5-3) B-Picture A B-picture is a picture encoded using bidirectional prediction. Bidirectional prediction is also expressed as dual prediction and means a prediction that references multiple pictures. Bidirectional prediction may refer to multiple pictures in different directions from the picture being processed (for example, a forward picture and a backward picture in different temporal directions), or it may refer to pictures in the same direction from the picture being processed.

[0099] Note that P-pictures and B-pictures are also called inter-pictures or inter-frames.

[0100] (6) A block is a processing unit of a set containing a specific number of pixels, and its name is not restricted, as shown in the following examples. It is also not restricted in shape, and includes not only rectangles made up of M x N pixels and squares made up of M x M pixels, but also other shapes.

[0101] (Examples of blocks) Slice / Tile / Brick CTB (Coding Tree Block) / CTU (Coding Tree Unit) / Superblock / Segment / Basic division unit / CU / Processing block unit / Prediction block unit (PU) / Translation block unit (TU) / Unit / Subblock / VPDU / Hardware processing division unit

[0102] A CTB is a processing unit into which components are divided. A CTB contains samples. A CTB is also called a CTU, superblock, or basic division unit. A CTB may be, for example, an N x N square block.

[0103] A Coding Block (CB) is a processing unit obtained by dividing a Coding Block (CTB). For example, it may be an M x N rectangular block. A CB is also called a Coding Unit (CU).

[0104] (7) A pixel / sample is the smallest unit of a picture, and includes not only pixels at integer positions but also pixels at decimal positions that are generated based on pixels at integer positions. A pixel / sample is a fundamental element that makes up a picture.

[0105] (8) Pixel value / sample value: An intrinsic value of a pixel, which includes not only luminance value, chrominance value, and RGB gradation, but also depth value, or binary values ​​of 0 or 1.

[0106] (9) Flags A flag is a variable or a single-bit syntax element. For example, a flag can only take one of two values: 0 or 1.

[0107] (10) A symbol or code used to transmit signal information, which includes not only discretized digital signals but also analog signals that take continuous values.

[0108] (11) Stream / Bitstream: Refers to a sequence of digital data or a flow of digital data. A stream / bitstream may consist of a single stream, or it may be divided into multiple layers and composed of multiple streams. It also includes cases where data is transmitted via serial communication over a single transmission path, as well as cases where data is transmitted via packet communication over multiple transmission paths.

[0109] (12) In the case of difference / difference scalar quantities, it is sufficient that the difference operation is included in addition to the simple difference (x - y), and this includes the absolute value of the difference (|x - y|), the squared difference (x^2 - y^2), the square root of the difference (√(x - y)), the weighted difference (ax - by: a, b is a constant), and the offset difference (x - y + a: a is the offset). Note that the difference operation also includes cases where bit shifts are performed (x >> 1 - y >> 1).

[0110] (13) In the case of summation scalar quantities, it is sufficient that the sum operation is included in addition to the simple sum (x + y), and this includes the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), weighted sum (ax + by: a, b is a constant), and offset sum (x + y + a: a is the offset). Note that the sum operation also includes cases where bit shifts are performed (x >> 1 + y >> 1).

[0111] (14) Based on: This includes cases where factors other than the subject being based on are taken into consideration. It also includes cases where the result is obtained not only by obtaining a direct result, but also by obtaining an intermediate result.

[0112] (15) Using: This includes cases where elements other than the target of use are taken into consideration. It also includes cases where the result is obtained not only by obtaining the result directly, but also by obtaining the result via an intermediate result.

[0113] (16) Prohibit, forbid. This can be rephrased as not being allowed. Also, not prohibiting or allowing something does not necessarily mean it is an obligation.

[0114] (17) To restrict (limit, restriction / restrict / restricted) This can be rephrased as not being allowed. Also, not being prohibited or being permitted does not necessarily mean being obligated. Furthermore, it is sufficient if it is prohibited in part quantitatively or qualitatively, and it also includes cases where it is prohibited entirely.

[0115] (18) MV (motion vector) A two-dimensional vector used in interpretation, which represents the offset from the coordinates of the decoded image to the coordinates of the reference image.

[0116] (19) Data assigned to a specific column or row in a database, such as an index table or list. For example, a reference index is an index for a reference picture list, where one reference picture is assigned to the reference picture index.

[0117] [Explanation of Descriptions] In the drawings, the same reference number indicates the same or similar component. Also, the size and relative position of components in the drawings are not necessarily depicted to a constant scale.

[0118] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all general or specific examples. The numerical values, shapes, materials, components, arrangement and connection forms of components, steps, relationships and sequences of steps shown in the following embodiments are examples only and are not intended to limit the scope of the claims.

[0119] The following describes embodiments of encoding and decoding devices. These embodiments are examples of encoding and decoding devices to which the processes and / or configurations described in each aspect of this disclosure can be applied. The processes and / or configurations can also be implemented in encoding and decoding devices different from those in the embodiments. For example, with respect to the processes and / or configurations applicable to the embodiments, one of the following may be implemented:

[0120] (1) Any of the multiple components of the encoding or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0121] (2) In the encoding or decoding device of the embodiment, any changes such as addition, replacement, or deletion of functions or processes performed by some of the multiple components of the encoding or decoding device may be made. For example, any of the functions or processes may be replaced or combined with other functions or processes described in any of the embodiments of this disclosure.

[0122] (3) In the methods performed by the encoding or decoding apparatus of the embodiment, any modifications such as additions, replacements, and deletions may be made to some of the processes included in the method. For example, any of the processes in the method may be replaced with or combined with other processes described in any of the embodiments of this disclosure.

[0123] (4) Some of the multiple components constituting the encoding or decoding device of the embodiment may be combined with components described in any of the embodiments of this disclosure, or with components that have some of the functions described in any of the embodiments of this disclosure, or with components that perform some of the processing performed by the components described in any of the embodiments of this disclosure.

[0124] (5) Components that provide some of the functions of the encoding or decoding device of the embodiment, or components that perform some of the processing of the encoding or decoding device of the embodiment, may be combined with or replaced with components described in any of the embodiments of this disclosure, components that provide some of the functions described in any of the embodiments of this disclosure, or components that perform some of the processing described in any of the embodiments of this disclosure.

[0125] (6) In a method performed by an encoding or decoding device of an embodiment, any of the processes included in the method may be replaced or combined with a process described in any of the embodiments of the present disclosure, or any of the similar processes.

[0126] (7) Some of the processes included in the methods performed by the encoding or decoding device of the embodiment may be combined with the processes described in any of the embodiments of this disclosure.

[0127] (8) The methods of carrying out the processes and / or configurations described in each aspect of the present disclosure are not limited to the encoding or decoding devices of the embodiments. For example, the processes and / or configurations may be carried out in devices used for purposes other than the video encoding or video decoding disclosed in the embodiments.

[0128] [System Configuration] Figure 1 is a schematic diagram showing an example of the configuration of the transmission system according to this embodiment.

[0129] A transmission system Trs is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs includes, for example, an encoding device 100, a network Nw, and a decoding device 200, as shown in Figure 1.

[0130] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding process.

[0131] The original image input to the encoding device 100 before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image. Furthermore, the image is a higher-level concept than sequences, pictures, and blocks, and is not limited by spatial and temporal domains unless otherwise specified. The image consists of a sequence of pixels or pixel values, and the signal or pixel values ​​representing the image are also called samples.

[0132] Furthermore, the stream may also be called a bitstream, encoded bitstream, compressed bitstream, or encoded signal. In addition, the encoding device 100 may be called an image encoding device or a video encoding device, and the encoding method by the encoding device 100 may be called an encoding method, an image encoding method, or a video encoding method.

[0133] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting.

[0134] Furthermore, the network may be replaced by a storage medium that records streams, such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc®).

[0135] The decoding device 200 generates a decoded image, which is, for example, an uncompressed image, by decoding the stream transmitted by the network Nw. For example, the decoding device decodes the stream according to a decoding method that corresponds to the encoding method by the encoding device 100.

[0136] The decoding device 200 may also be called an image decoding device or a video decoding device, and the decoding method performed by the decoding device 200 may be called a decoding method, an image decoding method, or a video decoding method.

[0137] [Encoding device] Next, the encoding device 100 according to the embodiment will be described.

[0138] [Implementation Example of Encoding Device] Figure 2 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, the multiple components of the encoding device 100 shown in Figure 3, which will be described later, are implemented by the processor a1 and memory a2 shown.

[0139] Processor a1 is a circuit that performs information processing and is a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit for encoding images. Processor a1 may be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits. Furthermore, for example, processor a1 may play the role of multiple components of the encoding device 100 shown in Figure 3, which will be described later, excluding the component for storing information.

[0140] Memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to encode an image is stored. Memory a2 may be an electronic circuit and may be connected to the processor a1. Memory a2 may also be included in the processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.

[0141] For example, memory a2 may store the image to be encoded, or it may store a stream corresponding to the encoded image. Alternatively, memory a2 may store a program for processor a1 to encode the image.

[0142] Furthermore, for example, memory a2 may play the role of an information storage component among the multiple components of the encoding device 100 shown in Figure 3, which will be described later. Specifically, memory a2 may play the role of the block memory 118 and frame memory 122 shown in Figure 3, which will be described later. More specifically, memory a2 may store a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.).

[0143] The generated parameters disclosed in the embodiments may be stored in memory a2 and referenced in the processing of processor a1. The generated parameters may be stored in memory a2 and may or may not be encoded. Whether or not to encode them is determined appropriately based on the relationship between the increase in the amount of code due to encoding the parameters and the reduction in the processing load of the decoding device 200 due to receiving the parameters.

[0144] Furthermore, in the encoding device 100, not all of the multiple components shown in Figure 3 described later are to be implemented, nor are all of the multiple processes described above to be performed. Some of the multiple components shown in Figure 3 described later may be included in other devices, and some of the multiple processes described later may be performed by other devices.

[0145] The following is a general description of the encoding device 100, followed by a description of the components included in the encoding device 100.

[0146] [Example of Encoding Device Configuration] Figure 3 is a block diagram showing an example of the configuration of an encoding device 100 according to an embodiment. The encoding device 100 encodes images in block units.

[0147] As shown in Figure 3, the encoding device 100 is a device that encodes an image in block units and comprises a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, and a prediction unit. The prediction unit comprises an intra-prediction unit 124, an inter-prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130.

[0148] These components are implemented, for example, by the processor a1 and memory a2 of the encoding device 100 shown in Figure 2 above.

[0149] [Overall Encoding Process Flow] Figure 4A is a flowchart showing an example of the overall encoding process by the encoding device 100. Figure 4B is a flowchart showing an example of the block-by-block processing by the encoding device 100.

[0150] First, the division unit 102 of the encoding device 100 divides each picture contained in the original image into multiple blocks. For example, the division unit 102 first divides the picture into blocks of a fixed size (for example, 128 x 128 pixels) (step Sa_1a). These fixed-size blocks are sometimes called coding tree units (CTUs).

[0151] Then, the division unit 102 selects a division pattern for the fixed-size block and further divides the fixed-size block into multiple blocks to constitute the selected division pattern (step Sa_2a).

[0152] The further divided blocks are of a variable size, for example, 64 x 64 pixels or less. The vertical and horizontal pixel counts of the variable size can be any combination of 4, 8, 16, 32, or 64. These variable-sized blocks are sometimes called coding units (CUs), prediction units (PUs), or transformation units (TUs).

[0153] In various implementation examples, CU, PU, ​​and TU do not need to be distinguished, and some or all of the blocks within a picture may be processing units for CU, PU, ​​or TU.

[0154] Then, the encoding device 100 performs the processing shown in steps Sa_3 to Sa_11 in Figure 4B for each of the multiple blocks (step Sa_3a). Here, an example of a CTU size of 128 x 128 pixels is given, but other sizes are also possible. For example, the size of the CTU may be 256 x 256 pixels, or a larger size such as 512 x 512 pixels. In this case, the vertical and horizontal size of the divided blocks, i.e., the number of pixels, may be greater than 64 pixels, such as 128 pixels or 256 pixels.

[0155] Furthermore, the division unit 102 outputs parameters indicating the division pattern to the conversion unit 106, the inverse conversion unit 114, the intra-prediction unit 124, the inter-prediction unit 126, and the entropy coding unit 110. The conversion unit 106 may convert the prediction residuals based on these parameters, and the intra-prediction unit 124 and the inter-prediction unit 126 may generate a prediction image based on these parameters. The entropy coding unit 110 may also perform entropy coding on these parameters.

[0156] Then, the encoding device 100 processes each of the multiple blocks (step Sa_3a). Specifically, the encoding device 100 processes steps Sa_3 to Sa_11 as shown in Figure 4B. Then, the encoding device 100 determines whether or not the encoding of the entire picture is complete (step Sa_4a), and if it determines that it is not complete (No. in step Sa_4a), it repeats the processing from step Sa_2a.

[0157] Next, the processing for each block will be explained (Sa_3 to Sa_11 in Figure 4B). First, the prediction unit generates a predicted image of the current block (step Sa_3). Specifically, the prediction unit generates a predicted image of the current block by referring to a reconstructed image generated by encoding and then decoding other blocks. The predicted image is also called a predicted signal, predicted block, or predicted sample.

[0158] The reconstructed image may be, for example, the image of the reference picture, or it may be the image of an encoded block (i.e., the other block mentioned above) within the current picture, which is the picture containing the current block. An encoded block within the current picture is, for example, an adjacent block to the current block.

[0159] The predicted image is, for example, an intra-prediction image (intra-prediction signal) generated based on intra-prediction, or an inter-prediction image (inter-prediction signal) generated based on inter-prediction.

[0160] Intra prediction is a method of predicting the current block by referring to blocks within the current picture, and is also called in-screen prediction. Specifically, the intra prediction unit 124 performs intra prediction by referring to the pixel values ​​(e.g., luminance values ​​or chrominance values, etc.) of blocks adjacent to the current block that are included in the current picture stored in the block memory 118. As a result, the intra prediction unit 124 generates an intra prediction image and outputs the intra prediction image to the prediction control unit 128.

[0161] Inter-prediction is a method of predicting the current block by referring to a reference picture different from the current picture, and is also called inter-screen prediction. Specifically, the inter-prediction unit 126 generates an inter-predicted image by referring to a reference picture stored in the frame memory 122 and performing inter-prediction of the current block, and outputs the inter-predicted image to the prediction control unit 128.

[0162] Note that prediction processing using the reconstructed image may not be performed on some blocks, such as the first block of the first image to be encoded. In that case, the original image will be output in the next subtraction process without performing subtraction on the original image. The interpretation unit 126 may also generate a predicted image by performing prediction processing without using a reference image.

[0163] Next, the subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in block units that are input from the division unit 102 and divided by the division unit 102. In other words, the subtraction unit 104 generates the difference between the current block and the predicted image as the predicted residual (step Sa_4).

[0164] The prediction residual is also called the prediction error. The original image is the input signal to the encoding device 100, and is, for example, a signal representing the image of each picture that makes up the video (e.g., a luminance (luma) signal and two chroma (chroma) signals). If the prediction process is skipped, the prediction residual becomes the value of the original image. For example, the prediction process is skipped for the first block in the processing order.

[0165] Next, the conversion unit 106 applies a conversion process to the predicted residual to generate conversion coefficients and outputs the conversion coefficients to the quantization unit 108 (step Sa_5).

[0166] The transformation process performed by the transformation unit 106 is, for example, an orthogonal transformation such as a discrete cosine transform (DCT) or discrete sine transform (DST) that transforms the predicted residual in the spatial domain into transformation coefficients in the frequency domain. However, other transformation processes such as weblet transforms, non-orthogonal transforms, or transformation processes expressed by matrix operations such as NSST (non-separable secondary transform) as defined in VVC may also be used.

[0167] The conversion unit 106 may perform a conversion process selected from among a plurality of candidate conversion processes. The conversion unit 106 may output conversion coefficients generated by sequentially applying a plurality of conversion processes to the predicted residual, or it may output the predicted residual as is without performing any conversion processes.

[0168] Depending on the processing performed by the conversion unit 106, information indicating whether or not a conversion process is applied to the predicted residuals, and / or information indicating the conversion type, may be encoded in the bitstream. This information is, for example, signaled at the CU level, but is not limited to the CU level and may be encoded at other levels (e.g., sequence level, picture level, slice level, brick level, or CTU level).

[0169] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106 (step Sa_6). Specifically, the quantization unit 108 quantizes the conversion coefficients based on the quantization parameter (QP) corresponding to the conversion coefficients. The quantization unit 108 then outputs a plurality of quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy coding unit 110 and the inverse quantization unit 112. In other words, the quantization unit 108 generates quantization coefficients and outputs them to the entropy coding unit 110 and the inverse quantization unit 112.

[0170] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the error in the quantization coefficient (quantization error) increases.

[0171] Furthermore, the quantization unit 108 may quantize the conversion coefficients based on a quantization matrix. In other words, quantization parameters (QP) and / or a quantization matrix may be used for quantization. The quantization parameters (QP) and the quantization matrix may be encoded, for example, at the sequence level, picture level, slice level, brick level, or CTU level.

[0172] The quantization unit 108 performs quantization processing in a predetermined scanning order. This predetermined scanning order is the order for quantization / inverse quantization of the conversion coefficients. For example, the predetermined scanning order is defined as ascending order of frequency (from low frequency to high frequency) or descending order (from high frequency to low frequency).

[0173] Next, the entropy coding unit 110 generates a stream (step Sa_7) by coding (specifically, entropy coding) the plurality of quantization coefficients and the prediction parameters related to the generation of the predicted image. For this entropy coding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding), which is used in VVC, may be used. Alternatively, a coding method that is a modified version of CABAC used in VVC may be used.

[0174] Furthermore, the processing of the entropy encoding unit 110 does not necessarily have to be performed after the quantization process; for example, it may be performed outside the loop of processing for each block.

[0175] Next, the inverse quantization unit 112 and the inverse transformation unit 114 restore the predicted residuals by performing inverse quantization and inverse transformation on a plurality of quantization coefficients (steps Sa_8 and Sa_9).

[0176] Next, the summing unit 116 reconstructs the current block by adding the predicted image to the recovered predicted residual (step Sa_10). This generates a reconstructed image. The reconstructed image is also called a reconstructed block, and the reconstructed image generated by the encoding device 100 is also called a local decoded block or local decoded image.

[0177] Next, the loop filter unit 120 performs filtering on the reconstructed image as needed (step Sa_11). Specifically, the loop filter unit 120 applies loop filtering to the reconstructed image output from the adder unit 116 and outputs the filtered reconstructed image to the frame memory 122.

[0178] Loop filters are filters used to reduce block noise that occurs at block boundaries, and include, for example, adaptive loop filters (ALF), deblocking filters (DF or DBF), and sample adaptive offset (SAO). Here, loop filters are described as in-loop filters used within the coding loop, but loop filters may also be out-loop filters used outside the coding loop.

[0179] In the example described above, the encoding device 100 selects one division pattern for a fixed-size block and encodes each block according to that division pattern. However, it may also encode each block according to multiple division patterns. In this case, the encoding device 100 may evaluate the cost of each of the multiple division patterns and, for example, select the stream obtained by encoding according to the division pattern with the smallest cost as the final output stream.

[0180] Furthermore, the processes in steps Sa_1a to Sa_4a and Sa_3 to Sa_11 may be performed sequentially by the encoding device 100, some of these processes may be performed in parallel, and the order may be changed.

[0181] The encoding process performed by such an encoding device 100 is a hybrid encoding using predictive encoding and transformative encoding. Furthermore, predictive encoding is performed by an encoding loop consisting of a subtraction unit 104, a transformer unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transformer unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra-prediction unit 124, an inter-prediction unit 126, and a prediction control unit 128. In other words, the prediction processing unit consisting of the intra-prediction unit 124 and the inter-prediction unit 126 constitutes a part of the encoding loop.

[0182] [Decoding Device] Next, a decoding device 200 capable of decoding the stream output from the encoding device 100 will be described.

[0183] [Implementation Example of Decryption Device] Figure 5 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and memory b2. For example, the multiple components of the decoding device 200 shown in Figure 6, which will be described later, are implemented by the processor b1 and memory b2 shown in Figure 5.

[0184] Processor b1 is a circuit that performs information processing and is a circuit that can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit for decoding streams. Processor b1 may be a processor such as a CPU. Alternatively, processor b1 may be a collection of multiple electronic circuits. Furthermore, for example, processor b1 may play the role of multiple components of the decoding device 200 shown in Figure 6, etc., described later, excluding the component for storing information.

[0185] Memory b2 is a dedicated or general-purpose memory in which information for the processor b1 to decode the stream is stored. Memory b2 may be an electronic circuit and may be connected to the processor b1. Memory b2 may also be included in the processor b1. Memory b2 may also be a collection of multiple electronic circuits. Memory b2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory b2 may also be a non-volatile memory or a volatile memory.

[0186] For example, memory b2 may store an image or a stream. Alternatively, memory b2 may store a program for processor b1 to decode the stream.

[0187] Furthermore, for example, memory b2 may play the role of an information storage component among the multiple components of the decoding device 200 shown in Figure 6 below. Specifically, memory b2 may play the role of a block memory 210 and a frame memory 214 shown in Figure 6 below. More specifically, memory b2 may store a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.).

[0188] The generated parameters disclosed in the embodiments may be stored in memory b2 and referenced during processing by processor b1. The parameters may be generated by a decoding process or not.

[0189] Furthermore, the decoding device 200 does not need to implement all of the components shown in Figure 6, etc., described later, nor does it need to perform all of the processes described above. Some of the components shown in Figure 6, etc., described later, may be included in other devices, and some of the processes described above may be performed by other devices.

[0190] The following is a general description of the decoding device 200, followed by a description of the components included in the decoding device 200. Note that detailed descriptions of some of the components included in the decoding device 200 that perform the same processing as the components included in the encoding device 100 may be omitted.

[0191] [Example of Decoding Device Configuration] Figure 6 is a block diagram showing an example of the configuration of a decoding device 200 according to an embodiment. The decoding device 200 is a device that decodes a stream, which is an encoded image, in block units.

[0192] As shown in Figure 6, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, a prediction unit, a prediction control unit 220, a prediction parameter generation unit 222, and a division determination unit 224. The prediction unit includes an intra-prediction unit 216 and an inter-prediction unit 218.

[0193] These components are implemented, for example, by the processor b1 and memory b2 of the decoding device 200 shown in Figure 5 above.

[0194] Furthermore, components included in the decoding device 200 may perform the same processing as components included in the encoding device 100, and their explanation may be omitted. For example, the inverse quantization unit 204, inverse transform unit 206, adder unit 208, block memory 210, frame memory 214, intra prediction unit 216, inter prediction unit 218, prediction control unit 220, and loop filter unit 212 perform the same processing as the inverse quantization unit 112, inverse transform unit 114, adder unit 116, block memory 118, frame memory 122, intra prediction unit 124, inter prediction unit 126, prediction control unit 128, and loop filter unit 120, respectively.

[0195] [Overall Decryption Process Flow] Figure 7A is a flowchart showing an example of the overall decryption process by the decryption device 200. Figure 7B is a flowchart showing an example of the block-by-block processing by the decryption device 200.

[0196] First, the entropy decoding unit 202 of the decoding device 200 acquires the stream output from the encoding device 100. Then, the entropy decoding unit 202 decodes (specifically, entropy decodes) the encoded quantization coefficients and prediction parameters of the current block contained in the bitstream (step Sp_1a).

[0197] The division determination unit 224 determines the division pattern for each of the multiple fixed-size blocks (128 x 128 pixels) contained in the picture based on the parameters input from the entropy decoding unit 202 (step Sp_2a). This division pattern is the division pattern selected by the encoding device 100. The decoding device 200 then performs block-specific processing for each of the multiple blocks. Specifically, it performs the processing in steps Sp_1 to Sp_5 for each of the multiple blocks (step Sp_3a).

[0198] The decoding device 200 then determines whether or not the decoding of the entire picture is complete (step Sp_4a). If it determines that it is not complete (No. in step Sp_4a), it repeats the process from step Sp_2a.

[0199] Next, we will explain the processing for each block (Sp_1 to Sp_5 in Figure 7B).

[0200] First, the inverse quantization unit 204 inversely quantizes the quantization coefficients of the current block, which are input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inversely quantizes the quantization coefficient based on the quantization parameter corresponding to that quantization coefficient (step Sp_1).

[0201] For example, the inverse quantization unit 204 may acquire a parameter indicating whether or not to perform inverse quantization, and quantization parameters (such as difference quantization parameters and QP index). The inverse quantization unit 204 may then decide whether or not to perform inverse quantization based on the acquired parameters, and may perform the inverse quantization process if it is decided to perform inverse quantization.

[0202] The inverse quantization unit 204 then outputs the inversely quantized quantization coefficients (i.e., transformation coefficients) of the current block to the inverse transformation unit 206.

[0203] Next, the inverse transform unit 206 restores the predicted residual by inversely transforming the transformation coefficients, which are input from the inverse quantization unit 204 (step Sp_2). Here, inverse transform refers to the reverse transformation process of the transformation process described in the encoding device 100. In other words, the inverse transform unit 206 can also be said to be performing a transformation process.

[0204] Next, the prediction unit, consisting of the intra-prediction unit 216, the inter-prediction unit 218, and the prediction control unit 220, generates a predicted image of the current block (step Sp_3). Here, the intra-prediction unit 216 and the inter-prediction unit 218 may perform the same processing as described above on the encoding device 100. Note that the predicted image needs to be generated before the addition process because it is used in the subsequent addition process, but it does not necessarily have to be done after the prediction residual restoration process. For example, the prediction image generation process may be performed in parallel with the prediction residual restoration process.

[0205] Next, the summing unit 208 reconstructs the current block into a reconstructed image (also called a decoded image block) by adding the predicted image to the predicted residual (step Sp_4). In other words, the summing unit 208 generates a reconstructed image of the current block by adding the predicted image to the predicted residual. The summing unit 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.

[0206] The block memory 210 is a block referenced in intra prediction and is a storage unit for storing blocks within the current picture. Specifically, the block memory 210 stores the reconstructed image output from the adder 208.

[0207] The loop filter unit 212 performs filtering on the reconstructed image (step Sp_5). The filtered reconstructed image is output to the frame memory 214 and the display device, etc.

[0208] In Figures 6 and 7B, the loop filter unit 212 processes the reconstructed image within the loop; that is, it outputs the filtered reconstructed image to the frame memory 214. Some or all of the processing in the loop filter unit 212 may be performed outside the loop. That is, filtering may be performed before outputting to a display device or the like.

[0209] Furthermore, the processes in steps Sp_1a to Sp_4a and Sp_1 to Sp_5 may be performed sequentially by the decoding device 200, some of these processes may be performed in parallel, or the order of these processes may be changed.

[0210] [Another Example of System Configuration] Figure 8 is a schematic diagram showing another example of the configuration of the transmission system according to this embodiment. The transmission system Trs is a system that transmits generated streams and receives transmitted streams. Such a transmission system Trs includes, for example, a transmitting device 2000, a network Nw, and a receiving device 3000, as shown in Figure 8. Note that the transmission system Trs does not need to have all of these components and may consist of only some of the devices.

[0211] The network Nw may be the Internet, a wide-area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network transmitting broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Furthermore, the network Nw may be replaced by a storage medium that records streams, such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc®).

[0212] The transmitting device 2000 transmits the stream over the network Nw. The transmitting device 2000 may also take the encoded stream as input and output the acquired stream. Alternatively, the transmitting device 2000 may have the encoding processing configuration disclosed as the encoding device 100, take the original image before encoding as input, encode the acquired image to generate a stream, and transmit it.

[0213] When the transmitting device 2000 generates a stream, the original input image before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image. The stream includes, for example, an encoded image and control information for decoding that encoded image. This encoding compresses the image.

[0214] The transmitting device 2000 can reduce the load on stream transmission by reducing the amount of data in the stream, which includes the encoded image.

[0215] The receiving device 3000 receives an encoded stream from the network Nw. The receiving device 3000 may store the received stream in memory, transfer the received stream to another device, or perform any processing on the received stream. For example, the receiving device 3000 may have a decoding process configuration disclosed as the decoding device 200 and decode the received stream.

[0216] The receiving device 3000 can reduce the delay of the stream by reducing the amount of data in the received stream.

[0217] [Transmitting Device] Figure 9 is a schematic diagram showing an example configuration of the transmitting device 2000 according to this embodiment. The transmitting device 2000 is a device that transmits a stream.

[0218] [Implementation Example of Transmitting Device] Next, a transmitting device 2000 according to the embodiment will be described. Figure 10 is a block diagram showing an implementation example of the transmitting device 2000. The transmitting device 2000 includes a processor a1 and a transmitting unit a3 that outputs a stream. The transmitting device 2000 may also further include a memory a2 as shown in Figure 2, or a receiving unit that acquires a video or stream.

[0219] Processor a1 is a circuit that performs information processing. Processor a1 may also be a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that transmits a stream. Processor a1 may also be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits.

[0220] Memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to transmit a stream is stored. Memory a2 may be an electronic circuit and may be connected to the processor a1. Memory a2 may also be included in the processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.

[0221] For example, memory a2 may store the input image or the generated stream. Memory a2 may also store a program for processor a1 to send the stream.

[0222] The stream transmitted by the transmitting device 2000 may be the stream (also called a bitstream) described in this disclosure.

[0223] Furthermore, in the transmitting device 2000, the processor a1 may perform some or all of the roles of the transmitting unit a3. Alternatively, the transmitting device 2000 may not have a transmitting unit a3, and the processor a1 may output the stream.

[0224] [Receiving device] Figure 11 is a schematic diagram showing an example of the configuration of a receiving device 3000 according to this embodiment. The receiving device 3000 is a device that receives a stream.

[0225] [Implementation Example of Receiving Device] Next, a receiving device 3000 according to the embodiment will be described. Figure 12 is a block diagram showing an implementation example of the receiving device 3000. The receiving device 3000 includes a processor b1 and a receiving unit b4 that receives a stream. The receiving device 3000 may also further include a memory b2 as shown in Figure 5, or a transmitting unit that transmits a video or stream.

[0226] Processor b1 is a circuit that performs information processing. Processor b1 may also be a circuit that can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that receives a stream. Processor b1 may also be a processor such as a CPU. Alternatively, processor b1 may be a collection of multiple electronic circuits.

[0227] Memory b2 is a dedicated or general-purpose memory that stores information for processor b1 to receive streams. Memory b2 may be an electronic circuit and may be connected to processor b1. Memory b2 may also be included in processor b1. Memory b2 may also be a collection of multiple electronic circuits. Memory b2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory b2 may also be a non-volatile memory or a volatile memory.

[0228] For example, memory b2 may store the input stream or the decoded image. Alternatively, memory b2 may store a program for processor b1 to receive the stream.

[0229] The stream received by the receiving device 3000 may be the stream (also called a bitstream) described in this disclosure.

[0230] Furthermore, in the receiving device 3000, the processor b1 may perform some or all of the functions of the receiving unit b4. Alternatively, the receiving device 3000 may not have a receiving unit b4, and the processor b1 may receive the stream.

[0231] [Bitstream Generation Device] Figure 13 is a schematic diagram showing an example of the configuration of a bitstream generation device according to this embodiment. The bitstream generation device 1000 is a device that generates a bitstream and transmits the generated stream.

[0232] An image is input to the bitstream generator 1000. The bitstream generator 1000 generates a stream by encoding the input image and outputs the stream. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding process.

[0233] The original image input to the bitstream generator 1000 before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image.

[0234] Furthermore, an image is a broader concept than sequences, pictures, and blocks, and is not limited to spatial or temporal domains unless otherwise specified. An image consists of a sequence of pixels or pixel values, and the signal or pixel values ​​representing that image are also called samples. A stream may also be called a bitstream, encoded bitstream, compressed bitstream, or encoded signal.

[0235] Furthermore, the bitstream generation device may also be called the encoding device, image encoding device, or video encoding device, and the method for generating a bitstream by the bitstream generation device may also be called the encoding method, image encoding method, or video encoding method.

[0236] [Implementation Example of Bitstream Generation Device] Next, a bitstream generation device 1000 according to an embodiment will be described. Figure 14 is a block diagram showing an implementation example of the bitstream generation device 1000. The bitstream generation device 1000 includes a processor a1 and a memory a2. The bitstream generation device 1000 may also include an input unit for inputting moving images, and an output unit for outputting the generated bitstream.

[0237] Processor a1 is a circuit that performs information processing and is a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that generates a bitstream. Processor a1 may be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits.

[0238] Memory a2 is a dedicated or general-purpose memory in which information for processor a1 to generate a bitstream is stored. Memory a2 may be an electronic circuit and may be connected to processor a1. Memory a2 may also be included in processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.

[0239] For example, memory a2 may store the input image or the generated stream. Memory a2 may also store a program for processor a1 to generate the bitstream.

[0240] The bitstream generator 1000 may have the same configuration as the encoding device 100 in this disclosure. Specifically, the bitstream generator 1000 may implement the multiple components shown in Figure 3. However, not all of the multiple components shown in Figure 3 are to be implemented, nor are all of the multiple processes to be performed. Some of the multiple components shown in Figure 3 may be included in other devices, and some of the multiple processes described above may be performed by other devices.

[0241] Furthermore, for example, processor a1 may perform the role of multiple components of the encoding device 100 shown in Figure 3, excluding the component for storing information. Also, for example, memory a2 may perform the role of the component for storing information among the multiple components of the encoding device 100 shown in Figure 3.

[0242] Specifically, memory a2 may function as the block memory 118 and frame memory 122 shown in Figure 3. More specifically, memory a2 may store reconstructed images (specifically, reconstructed blocks or reconstructed pictures, etc.).

[0243] Furthermore, during operation, processor a1 uses memory a2 to generate a bitstream that includes parameters for causing the decoding device 200 to execute processing, and / or parameters for switching the processing of the decoding device 200 or other specific processing. For example, processor a1 generates these parameters, includes the generated parameters in the bitstream, and generates the bitstream.

[0244] Here, the "parameters" to be executed by the decoding device 200 may be parameters for executing any decoding process exemplified in this disclosure as a process to be executed. For example, the "parameters" to be executed by the decoding device 200 may be any parameter in the syntax described in this disclosure. Furthermore, the "process of the decoding device 200" or "other specific process" that can be switched may be any decoding process exemplified in this disclosure as a process performed by obtaining the syntax.

[0245] Furthermore, the generated parameters disclosed in the embodiments may be stored in memory a2 and referenced in the processing of processor a1. The generated parameters may be stored in memory a2 and may or may not be included in the bitstream. Whether or not to include them in the bitstream is appropriately determined based on the relationship between the increase in the amount of code due to encoding the parameters and the reduction in the processing load of the decoding device 200 due to receiving the parameters.

[0246] [Storage Medium] Figure 15A is a schematic diagram showing an example of the configuration of a storage medium and a computer according to this embodiment. The storage medium 4000 is a medium for storing bitstreams. The storage medium 4000 is connected to a processor b1 included in the computer 5000 and outputs a bitstream, which is an executable instruction for the computer 5000, to the computer 5000. For example, when the computer 5000 reads the bitstream from the storage medium 4000, the storage medium 4000 outputs the bitstream to the computer 5000.

[0247] The encoded stream stored in the storage medium 4000 is an instruction that the computer 5000 can execute. The processor b1 in the computer 5000 generates a reconstructed image by executing a bitstream decoding process based on the acquired bitstream. In this way, the computer 5000, having acquired a stream from the storage medium 4000, can decode the bitstream based on the parameters and other information contained in the encoded stream.

[0248] Furthermore, the storage medium 4000 may be the memory b2 described herein. For example, the memory b2 may store a program for the processor b1 to generate a bitstream. The storage medium 4000 may be a non-transitory computer-readable medium. The storage medium 4000 is also referred to as a recording medium.

[0249] The storage medium 4000 may include an input / output unit for inputting and outputting bitstreams. The storage medium 4000 may also include an input unit for receiving bitstreams and an output unit for outputting bitstreams.

[0250] Furthermore, the computer 5000 may also be the decoding device 200. In other words, the computer 5000 may be read as the decoding device 200.

[0251] Figure 15B is a schematic diagram showing an example of the configuration of a computer according to this embodiment. The schematic diagram shown in Figure 15B differs from the example in Figure 15A in that the storage medium 4000 is included in the computer 5000. The storage medium 4000 and the processor b1 may be the same as those described in Figure 15A.

[0252] Here, the bitstream includes parameters that indicate instructions for the computer 5000 to execute. The "parameters" may include parameters that cause the computer 5000 to perform processing, and / or parameters that switch the processing or other specific processing of the computer 5000.

[0253] The “parameter” may be any parameter that performs any decryption process exemplified in this disclosure as an example of a process to be performed. For example, it may be any parameter in the syntax described in this disclosure. The process to be performed and the process to be switched may be any decryption process exemplified in this disclosure as a process performed by obtaining the syntax.

[0254] [Examples of Parameters] In the above, for the bitstream generator 1000, "parameters that cause the decoding device 200 to execute processing, and / or parameters that switch the processing of the decoding device 200 or other specific processing" are described. Also, for the storage medium 4000, "parameters that cause the computer 5000 to execute processing, and / or parameters that switch the processing of the computer 5000 or other specific processing" are described. These parameters may be, for example, the following parameters.

[0255] The "parameters" may include, for example, parameters indicating the method of dividing the block. The computer 5000 may divide the block based on the parameters indicating the method of dividing the block and perform any of the processes illustrated in this disclosure on the divided blocks.

[0256] The "parameters" may include, for example, parameters indicating the prediction mode to be applied. The computer 5000 may determine the prediction mode to be applied based on the parameters indicating the prediction mode to be applied, and may perform any prediction process exemplified in this disclosure corresponding to the determined prediction mode.

[0257] The "parameters" may include, for example, a parameter indicating whether or not process y can be performed within range x. The computer 5000 may determine whether or not process y can be performed on a target block based on the parameter indicating whether or not process y can be performed within range x.

[0258] A flag (parameter) indicating whether or not process y can be performed within range x is shown, for example, as xxx_yyy_enabled_flag / xxx_yyy_disabled_flag. Here, yyy may represent process y, which is indicated by the flag as either or not. Process y may also be any process exemplified in this disclosure.

[0259] xxx may represent the level to which parameters are assigned. For example, xxx may correspond to sps, pps, ph, sh, cu (Coding Unit), or block, and this level may indicate the range x. More specifically, if it is sps, the range x is a sequence; if it is pps, the range x is a picture; if it is ph, the range x is a picture; and if it is sh, the range x is a slice. Note that the levels to which parameters are assigned are not limited to these.

[0260] If the parameter indicates that process y is not feasible, then process y will not be performed in the blocks included in range x. If the parameter indicates that process y is feasible, then process y is feasible in the blocks included in range x. In other words, if the parameter indicates that process y is feasible, then process y may or may not be performed in the blocks included in range x.

[0261] The flag indicating whether or not process y can be performed may be defined for multiple headers. For example, process y may be permitted for a sequence, but not for a particular picture within that sequence.

[0262] In the above case, the sps_yyy_enabled_flag (or sps_yyy_disabled_flag) for the sequence may have a value indicating that process y can be performed. Furthermore, the ph_yyy_enabled_flag (or ph_yyy_disabled_flag) for a certain picture included in the sequence may have a value indicating that process y cannot be performed.

[0263] As described above, the levels correspond to sps, pps, ph, sh, cu (Coding Unit), or blocks, etc. An example definition for each level is shown below. Note that the definitions for each level can be changed.

[0264] The description of `sps` (Sequence Parameter Set) corresponds to a syntax structure containing syntax elements that apply to zero or more CLVS (Coded Layer Video Sequences). Here, the number of zero or more CLVSs is determined by the content of the syntax elements contained in the PPS referenced by the syntax elements contained in each picture header. In other words, `sps` corresponds to the sequence level.

[0265] The description of pps (Picture Parameter Set) corresponds to a syntax structure containing syntax elements that apply to zero or more encoded pictures. Here, zero or more encoded pictures are determined by the syntax elements contained in each picture header. In other words, pps corresponds to the picture level.

[0266] ph (PH: Picture Header) corresponds to a syntax structure containing syntax elements that apply to all slices of an encoded picture. In other words, ph corresponds to the picture level.

[0267] The entry for sh (SH: Slice Header) corresponds to the portion of the encoded slice that contains all tiles or data elements related to CTU rows within a slice. In other words, sh corresponds to the slice level.

[0268] The notation `cu` (Coding Unit) corresponds to the coding block of the sample and the syntax structure used to encode that sample. In other words, `cu` corresponds to the coding unit level or the block level.

[0269] Here, the above coding blocks correspond to (i) the coding block for the luminance sample and the two corresponding coding blocks for the chrominance sample of a picture having three sample sequences in single-tree mode, (ii) the coding block for the luminance sample of a picture having three sample sequences in dual-tree mode, (iii) the two coding blocks for the chrominance sample of a picture having three sample sequences in dual-tree mode, or (iv) the coding block for the sample of a monochrome picture.

[0270] Although a flag indicating whether or not process y can be performed has been described above, a flag indicating whether or not to perform process y (xxx_yyy_flag) may be used in a similar manner. Furthermore, the flag indicating whether or not process y can be performed and the flag indicating whether or not to perform process y may be used in combination.

[0271] [Data Structure] Figure 16 shows an example of the hierarchical structure of data in a stream. A stream includes, for example, a video sequence. This video sequence includes, for example, a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), an SEI (Supplemental Enhancement Information), and multiple pictures, as shown in Figure 16(a).

[0272] VPS includes encoding parameters common to multiple layers in a video composed of multiple layers, and encoding parameters related to the multiple layers included in the video, or to individual layers.

[0273] The SPS includes parameters used for the sequence, i.e., encoding parameters that the decoding device 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of the picture. Multiple SPSs may exist.

[0274] The PPS includes parameters used for the picture, i.e., encoding parameters that the decoding device 200 references to decode each picture in the sequence. For example, the encoding parameters may include a reference value for the quantization width used to decode the picture and a flag indicating the application of weighted prediction. There may be multiple PPSs. Also, the SPS and PPS are sometimes simply referred to as parameter sets.

[0275] The picture may include a picture header and one or more slices, as shown in Figure 16(b). The picture header includes encoding parameters that the decoding device 200 refers to in order to decode the one or more slices.

[0276] A slice includes a slice header and one or more Coding Tree Units (CTUs), as shown in Figure 16(c). The slice header includes coding parameters that the decoding device 200 references to decode the one or more CTUs.

[0277] A picture may not contain slices, but instead contain tile groups. In this case, a tile group may contain one or more tiles, and each tile may contain one or more CTUs. Furthermore, a tile may contain one or more slices, and a slice may contain one or more tiles.

[0278] A CTU is also called a superblock or basic partitioning unit. Such a CTU includes a CTU header and one or more CUs (Coding Units), as shown in Figure 16(d). The CTU header contains coding parameters that the decoding device 200 refers to in order to decode one or more CUs.

[0279] A CU may be divided into multiple smaller CUs. Furthermore, as shown in Figure 16(e), a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating the predicted residual, which will be described later.

[0280] Note that a CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but in SBT, for example, as described later, it may include multiple TUs smaller than the CU. Also, a CU may be processed for each VPDU (Virtual Pipeline Decoding Unit) that constitutes that CU. A VPDU is a fixed unit that can be processed in one stage when performing pipeline processing in hardware, for example.

[0281] Note that a stream does not necessarily have to have some of the hierarchical levels shown in Figure 16. Also, the order of these hierarchical levels may be changed, and any hierarchical level may be replaced by another hierarchical level.

[0282] Furthermore, the picture that is currently being processed by a device such as the encoding device 100 or the decoding device 200 is called the current picture. If the processing is encoding, the current picture is synonymous with the picture to be encoded; if the processing is decoding, the current picture is synonymous with the picture to be decoded.

[0283] Furthermore, a block, such as CU or CU, that is currently being processed by a device such as an encoding device 100 or a decoding device 200 is called the current block. If the processing is encoding, the current block is synonymous with the block to be encoded; if the processing is decoding, the current block is synonymous with the block to be decoded.

[0284] [Picture Composition: Slices / Tiles] To decode pictures in parallel, pictures may be composed of slices or tiles.

[0285] A slice is the basic encoding unit that makes up a picture. A picture is composed of, for example, one or more slices. A slice consists of one or more consecutive CTUs.

[0286] Figure 17 shows an example of a slice configuration. For example, a picture contains 11 x 8 CTUs and is divided into four slices (slices 1-4). Slice 1 consists of, for example, 16 CTUs, slice 2 consists of, for example, 21 CTUs, slice 3 consists of, for example, 29 CTUs, and slice 4 consists of, for example, 22 CTUs. Here, each CTU in the picture belongs to one of the slices.

[0287] The shape of a slice is a horizontal division of the picture. The boundaries of a slice do not have to be at the edges of the screen, but can be anywhere among the CTU boundaries within the screen. The processing order (encoding order or decoding order) of the CTUs within a slice is, for example, the raster scan order. A slice also includes a slice header and encoded data. The slice header may describe the characteristics of the slice, such as the CTU address of the beginning of the slice and the slice type.

[0288] A tile is a rectangular area that makes up a picture. Each tile may be assigned a number called a TileId in the order of the raster scan.

[0289] Figure 18 shows an example of a tile configuration. For example, a picture contains 11 x 8 CTUs and is divided into four rectangular tile regions (tiles 1-4). When tiles are used, the processing order of the CTUs is changed compared to when tiles are not used.

[0290] If tiles are not used, multiple CTUs within a picture are processed, for example, in raster scan order. If tiles are used, at least one CTU in each of the multiple tiles is processed, for example, in raster scan order. For example, as shown in Figure 18, the processing order of the multiple CTUs contained in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then from the left end of the second column of tile 1 to the right end of the second column of tile 1.

[0291] Note that one tile may contain one or more slices, and one slice may contain one or more tiles.

[0292] A picture may be composed of tilesets. A tileset may contain one or more tile groups, or one or more tiles. A picture may consist of only one of a tileset, a tile group, or a tile. For example, the order in which multiple tiles are scanned in raster order for each tileset is defined as the basic coding order of the tiles. Within each tileset, a collection of one or more tiles whose basic coding order is consecutive is defined as a tile group. Such a picture may be composed of the division unit 102 (see Figure 3) described later.

[0293] [Temporal Scalable Encoding] Figure 19 shows an example of a temporally scalable stream configuration.

[0294] The encoding device 100 may generate a temporally scalable stream by encoding multiple pictures in multiple layers, as shown in Figure 19. For example, the encoding device 100 achieves scalability by encoding each picture in its own layer, with enhancement layers existing above the base layer. This type of encoding of each picture is called temporally scalable encoding.

[0295] This allows the decoding device 200 to switch the frame rate of the image displayed by decoding the stream. In other words, the decoding device 200 decides which layers to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth.

[0296] As a result, the decoding device 200 can freely switch between decoding the same content in low-frame-rate and high-frame-rate formats. For example, a user of the stream can watch part of the video using their smartphone while on the go, and then watch the rest of the video using an internet TV or other device after returning home.

[0297] Each of the aforementioned smartphones and devices incorporates a decoding device 200, which may or may not have the same performance. In this case, if the device decodes up to the upper layers of the stream, the user can view high-frame-rate video after returning home. This eliminates the need for the encoding device 100 to generate multiple streams with the same content but different frame rates, thereby reducing the processing load.

[0298] In temporally scalable coding, layers are sometimes referred to as temporal layers or temporal sublayers.

[0299] [Multilayer Encoding] Figure 20 shows an example of a stream configuration using multilayer encoding. Each of AU0 to AU3 corresponds to an AU (Access Unit). Each of Layer0 to Layer3 corresponds to a layer in multilayer encoding. Each PU (Picture Unit) corresponds to a picture. Each of POC_1 to POC_3 corresponds to a POC (Picture Of Count) and indicates the display order.

[0300] As shown in Figure 20, the encoding device 100 may generate a stream in which spatial resolution, image quality, or multiplexed content can be scalably changed by encoding multiple pictures into multiple layers. For example, the encoding device 100 achieves scalability by encoding pictures layer by layer, with enhancement layers existing above the base layer. This type of encoding of each picture is called multi-layer encoding.

[0301] This allows the decoding device 200 to switch the spatial resolution, image quality, or multiplexed content of the image displayed by decoding the stream. In other words, the decoding device 200 decides which layers to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth.

[0302] As a result, the decoding device 200 can freely switch between decoding low-resolution and high-resolution content, low-quality and high-quality content, and basic and customized content for the same content. For example, a user of the stream can watch part of the video on their smartphone while on the go, and then watch the rest of the video on an internet TV or other device after returning home.

[0303] Furthermore, each of the aforementioned smartphones and devices incorporates a decoding device 200 with identical or different performance characteristics. In this case, if the device decodes up to the upper layers of the stream, the user can view high-definition video after returning home. This eliminates the need for the encoding device 100 to generate multiple streams with the same content but different image quality, thereby reducing the processing load.

[0304] [Metadata] Each layer of the stream may include metadata based on statistical information of the image. The decoding device 200 may generate high-resolution video by super-resolution the pictures of each layer based on the metadata. Super-resolution may be either an improvement in the signal-to-noise ratio (SN) at the same resolution, or an increase in resolution. The metadata may include information for identifying linear or nonlinear filter coefficients used in the super-resolution process, or information for identifying parameter values ​​in the filtering process, machine learning, or least-squares operation used in the super-resolution process.

[0305] Alternatively, the picture may be divided into tiles or the like, depending on the meaning of each object within it. In this case, the decoding device 200 may decode only a portion of the picture by selecting the tiles to be decoded.

[0306] Furthermore, object attributes (such as person, car, or ball) and their position within the picture (such as their coordinate position within the same picture) may be stored as metadata. In this case, the decoding device 200 can identify the position of a desired object based on the metadata and determine the tile containing that object. For example, the metadata is stored using a data storage structure different from pixel data, such as SEI in HEVC. This metadata may indicate, for example, the position, size, or color of the main object.

[0307] Furthermore, metadata may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. This allows the decoding device 200 to obtain information such as the time when a specific person appears in the video, and by using that time and the information in the picture units, it can identify the picture in which the object exists and the position of the object within that picture.

[0308] [Loop Filtering Section in Encoding Process] The loop filter section 120 in the encoding device 100 applies a filter to the reconstructed image output from the adder 116 and outputs the filtered reconstructed image to the frame memory 122. Here, the filter performed by the loop filter section 120 is a filter used within the encoding loop (also called a loop filter or in-loop filter).

[0309] Filtering is a technique that reduces encoding noise generated by quantization processing within the encoding loop. Filtering can directly reduce visually noticeable block distortion or ringing distortion, improving subjective and / or objective performance. Furthermore, the loop filter unit 120 prevents image quality degradation from propagating between frames by performing filtering within the loop.

[0310] Filters include, for example, adaptive loop filters (ALF), deblocking filters (DF or DBF), sample adaptive offset (SAO), luminance mapping with chroma scaling (LMCS), or any combination thereof.

[0311] Figure 21 is a block diagram showing an example of the configuration of the loop filter unit 120. The loop filter unit 120 includes, for example, an LMCS processing unit 120d, a DBF processing unit 120a, a SAO processing unit 120b, and an ALF processing unit 120c, as shown in Figure 21.

[0312] The LMCS processing unit 120d performs luminance mapping and color difference scaling on the reconstructed image. The DBF processing unit 120a performs DBF processing on the reconstructed image after LMCS processing. The SAO processing unit 120b performs SAO processing on the reconstructed image after DBF processing. In addition, the ALF processing unit 120c applies ALF processing to the reconstructed image after SAO processing. Details of ALF and DBF will be described later.

[0313] SAO processing is a process that improves image quality by reducing ringing (a phenomenon in which pixel values ​​are distorted in a wave-like manner around edges) and correcting pixel value misalignment. Examples of SAO processing include edge offset processing and band offset processing.

[0314] LMCS processing is a process that reassigns codewords assigned to rarely occurring pixel values ​​to codewords assigned to frequently occurring pixel values ​​when there is a large bias in the pixel value distribution of the luminance signal of the original image. Examples of LMCS processing include mapping the luminance signal using a piecewise linear model based on the pixel value distribution of the original image, and scaling the residual color difference signal according to the luminance pixel values.

[0315] Furthermore, the loop filter unit 120 does not necessarily have to include all of the processing units disclosed in Figure 21, and may include only some of them. Also, the loop filter unit 120 may perform the above-mentioned processing in an order different from the processing order disclosed in Figure 21.

[0316] [Loop Filter Section > Adaptive Loop Filter] In ALF, a least-squares error filter is applied to remove encoding distortion. For example, for each 2x2 pixel subblock within the current block, one filter selected from several filters is applied based on the direction and activity of the local gradient.

[0317] Specifically, first, subblocks (for example, 2x2 pixel subblocks) are classified into multiple classes (for example, 15 or 25 classes). The classification of subblocks is done, for example, based on the direction of the gradient and the activity level. In a specific example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the subblocks are classified into multiple classes.

[0318] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). The gradient activation value A is derived, for example, by adding the gradients in multiple directions and quantizing the sum.

[0319] Based on the results of this classification, a filter for the subblock is determined from among multiple filters.

[0320] For example, circularly symmetrical shapes are used for filters in ALF. Figures 22A to 22C show several examples of filter shapes used in ALF. Figure 22A shows a 5x5 diamond-shaped filter, Figure 22B shows a 7x7 diamond-shaped filter, and Figure 22C shows a 9x9 diamond-shaped filter.

[0321] Information indicating the shape of the filter is typically signaled at the picture level. However, the signaling of information indicating the shape of the filter is not limited to the picture level; it may be at other levels (e.g., sequence level, slice level, brick level, CTU level, or CU level).

[0322] The on / off status of the ALF may be determined, for example, at the picture level or the CU level. For example, the decision to apply the ALF may be made at the CU level for luminance, and at the picture level for color difference. Information indicating whether the ALF is on or off is usually signaled at the picture level or the CU level.

[0323] Note that the signaling of ALF on / off information is not limited to picture level or CU level, but may be at other levels (e.g., sequence level, slice level, brick level, or CTU level). If the ALF on / off information read from the stream indicates that ALF is on, one filter is selected from several filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.

[0324] Furthermore, as described above, one filter is selected from among several filters and ALF processing is applied to the subblock. For each of these filters (for example, up to 15 or 25 filters), the set of coefficients used in that filter is usually signaled at the picture level. However, the signaling of the coefficient set is not limited to the picture level and may be at other levels (for example, sequence level, slice level, brick level, CTU level, CU level, or subblock level).

[0325] [Loop Filter Section > CCALF] CC-ALF (Cross Component Adaptive Loop Filter) is a type of restoration filter that uses correlations between color components. The filtering results of eight luminance signals corresponding to the positions of the color difference signal being processed are added to the color difference signal. To enable parallel processing with ALF, the luminance signal before the application of ALF is used for filtering.

[0326] Figure 22D shows a configuration diagram of an example of a CCALF in this disclosure. Figure 22E shows an example of a filter shape of a CCALF in this disclosure.

[0327] One example of CC-ALF operates by applying a linear diamond-shaped filter (Figures 22D and 22E) to the luminance channel of each chromatic difference component. For example, the filter coefficients are transmitted via APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled by a variable block size and indicated by a context-encoded flag received for each block of samples. The block size and CC-ALF enable flag are received at the slice level for each chromatic difference component.

[0328] The syntax and semantics of CC-ALF are provided in the encoding standard. Block sizes of 16x16, 32x32, 64x64, and 128x128 may also be supported for color difference samples.

[0329] [Loop Filter Section > DBF] In DBF processing, the loop filter section 120 reduces distortion occurring at block boundaries by applying a filter to the block boundaries of the reconstructed image.

[0330] Figure 23A is a block diagram showing an example of the detailed configuration of the DBF processing unit 120a. The DBF processing unit 120a includes, for example, a boundary determination unit 1201, a filter determination unit 1203, a filter processing unit 1205, a processing determination unit 1208, a filter characteristic determination unit 1207, and switches 1202, 1204, and 1206.

[0331] Figure 23B is a flowchart showing an example of the processing performed by the DBF processing unit 120a.

[0332] First, the DBF processing unit 120a determines the boundary to be processed in the boundary determination unit 1201 (step S_0b).

[0333] Next, the DBF processing unit 120a determines whether or not to perform DBF (step S_1b). If it determines to perform DBF, it proceeds to the next determination. If the DBF processing unit 120a determines not to perform DBF, it terminates the process without performing DBF (step S_2b).

[0334] Specifically, for example, the DBF processing unit 120a may determine in the boundary determination unit 1201 whether or not the target pixel for DBF processing is located near a block boundary, and if the target pixel is not located near a block boundary, it may decide not to perform DBF. Alternatively, the DBF processing unit 120a may calculate a Bs value and decide not to perform DBF according to the Bs value. For example, if the Bs value is equal to 0, it indicates that DBF will not be performed, and if it is 1 or greater, it indicates that there is a possibility of performing DBF.

[0335] Next, the DBF processing unit 120a determines whether or not to perform DBF on the target pixel in the filter determination unit 1203 (step S_3b). In other words, the DBF processing unit 120a determines the type of filter to be applied. Specifically, for example, the filter determination unit 1203 determines whether or not to perform DBF processing on the target pixel based on the pixel values ​​of at least one surrounding pixel located around the target pixel. The image before filtering is an image consisting of the target pixel and at least one surrounding pixel located around that target pixel.

[0336] The filter determination unit 1203 may compare a value based on the pixel value of at least one peripheral pixel, and / or the difference between multiple pixels, etc., with a threshold value, and decide whether to execute the DBF to be determined if the value is greater than the threshold value, and whether to execute the DBF to be determined if the value is less than or equal to the threshold value. The threshold value is, for example, β or tC determined based on the quantization parameter QP. In other words, in the DBF process, for example, the DBF to be performed is selected based on quantization.

[0337] In one example of the process for determining whether or not to perform the DBF to be judged, it is first determined whether or not to perform DBF(a). If it is determined that DBF(a) should be performed, DBF(a) is performed (step S_4b). If it is determined that DBF(a) should not be performed, the next determination is made. In the next determination, it is determined whether or not to perform DBF(b). If it is determined that DBF(b) should be performed, DBF(b) is performed (step S_5b). If it is determined that DBF(b) should not be performed, DBF is not performed (step S_2b).

[0338] Furthermore, if it is decided not to implement DBF(b), there may be multiple DBF candidates, such as deciding whether or not to implement DBF(c). Here, DBF(a), which has been decided to implement, may be further subdivided, and it may be decided whether or not to implement DBF(a_1) or DBF(a_2).

[0339] Here, DBF(a), DBF(b), DBF(a_1), and DBF(a_2) may be short filter, long filter, weak filter, and strong filter, respectively.

[0340] Figure 23C is a flowchart showing another example of the processing performed by the DBF processing unit 120a.

[0341] First, the DBF processing unit 120a determines the boundary to be processed in the boundary determination unit 1201 (step S_0c). Next, the DBF processing unit 120a determines whether or not to perform DBF (step S_1c), and if it determines to perform DBF, it selects the DBF to perform (step S_2c). Then, the DBF processing unit 120a performs the selected DBF (step S_3c). If the DBF processing unit 120a determines in S_1c not to perform DBF, it terminates the process without performing DBF (step S_4c).

[0342] Note that the processing of S_1c may be the same as or different from the processing of S_1b. Also, the processing of S_2c may be the same as or different from the processing of S_3b.

[0343] Figure 24 shows an example of a DBF (Digital Block Filter) with symmetrical filter characteristics with respect to block boundaries. DBF is a process that reduces distortion at block boundaries by correcting the sample values ​​on both sides of the block boundary. The number of pixels to be filtered and the number of pixels used for filtering vary depending on the type of filter. The higher the filter strength, the greater the number of pixels to be filtered and the greater the number of pixels used for filtering.

[0344] When DBF processing is performed on the block boundary between block P and block Q adjacent to block P, the pixels to be processed are pn (n=0, ..., n) in block P and qm (m=0, ..., m) in block Q. Pixels with a value of n or m of 0 are adjacent to the boundary, while pixels with larger values ​​are further from the boundary.

[0345] For example, in a strong filter, as shown in Figure 24, processing is performed on pixels p0 to p2 in block P and pixels q0 to q2 in block Q. The respective pixel values ​​of pixels q0 to q2 are changed to pixel values ​​q'0 to q'2 by performing the calculation shown in the following equation.

[0346] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8 q'1=(p0+q0+q1+q2+2) / 4 q'2=(p0+q0+q1+3×q2+2×q3+4) / 8

[0347] In the above equations, p0 to p2 and q0 to q2 are the pixel values ​​of pixels p0 to p2 and pixels q0 to q2, respectively. Also, q3 is the pixel value of pixel q3, which is adjacent to pixel q2 on the opposite side of the block boundary. Furthermore, the coefficient multiplied by the pixel value of each pixel used in the deblocking filter process on the right-hand side of each of the above equations is the filter coefficient.

[0348] Furthermore, in DBF processing, clipping may be performed to ensure that the pixel value after calculation does not change beyond a threshold. In this clipping process, the pixel value after calculation using the above formula is clipped to "pre-calculation pixel value ± a × threshold (where a is an integer)" using a threshold determined from the quantization parameters. In other words, the amount of change in the pixel value before and after calculation is clipped so that it is within the range of ± a × threshold (where a is an integer). This prevents excessive smoothing. Specifically, for example, the pixel values ​​q'0 to q'2 are expressed by the following formula.

[0349] q'0=Clip3 (q0-3×tC, q0+3×tC, (p1+2×p0+2×q0+2×q1+q2+4)>>3) q'1=Clip3 (q1-2×tC, q1+2×tC, (p0+q0+q1+q2+2)>>2) q'2=Clip3 (q2-1×tC, q2+1×tC, (p0+q0+q1+3×q2+2×q3+4)>>3)

[0350] Furthermore, Clip3 may be defined as follows:

[0351]

[0352] Figure 25 is a diagram illustrating an example of a block boundary where DBF processing is performed. Figure 26 is a diagram illustrating an example of a block boundary strength Bs value.

[0353] The block boundaries on which DBF processing is performed are, for example, the CU, PU, ​​or TU boundaries of an 8x8 pixel block as shown in Figure 25. DBF processing is performed, for example, in units of 4 rows or 4 columns. First, for blocks P and Q shown in Figure 25, the Boundary Strength (Bs) value is determined as shown in Figure 26. The Bs value is determined independently for each color component.

[0354] In Figure 26, conditions listed higher up have higher priority, and evaluation is performed in order from the highest priority condition. If a condition is met, evaluation of lower priority conditions is not performed. The magnitude and range of block distortion differ depending on the boundary, and excessive smoothing can cause problems such as blurring. Therefore, in this example, the Bs value is determined according to the prediction modes of blocks P and Q, the presence or absence of motion vectors and transformation coefficients.

[0355] Furthermore, based on the determined Bs value, it may be decided whether or not to perform DBF processing of different strengths even for block boundaries belonging to the same image. DBF processing for chrominance signals is performed when the Bs value is 2. DBF processing for luminance signals is performed when the Bs value is 1 or greater and predetermined conditions are met. Note that the criteria for determining the Bs value are not limited to the example shown in Figure 26, and may be determined based on other parameters.

[0356] The above method for determining the Bs value and the method for performing DBF processing based on the Bs value are merely examples, and the method for determining the Bs value and the method for performing DBF processing based on the Bs value are not limited to the above example.

[0357] In DBF processing, it is determined whether or not to apply a filter to each of the luminance sample and chrominance sample, and the type of filter to be applied. Furthermore, in DBF processing, it is determined whether or not to apply a filter to each of the vertical and horizontal boundaries contained in each of the luminance sample and chrominance sample, and the type of filter to be applied.

[0358] Some or all of the processing may be performed on either the luminance sample or the chrominance sample, or on both. The luminance sample and the chrominance sample may undergo the same processing, or they may undergo different processing.

[0359] [Loop Filtering Unit in Decoding Process] The loop filtering unit 212 in the decoding device 200 applies a loop filter to the reconstructed image generated by the summing unit 208, and outputs the filtered reconstructed image to the frame memory 214 and the display device, etc.

[0360] Figure 27 is a block diagram showing an example of the configuration of the loop filter unit 212. The loop filter unit 212 has a configuration similar to that of the loop filter unit 120 of the encoding device 100. The loop filter unit 212 includes, for example, an LMCS processing unit 212d, a DBF processing unit 212a, a SAO processing unit 212b, and an ALF processing unit 212c, as shown in Figure 27.

[0361] The LMCS processing unit 212d performs luminance mapping and color difference scaling on the reconstructed image. The DBF processing unit 212a performs DBF processing on the reconstructed image after LMCS processing. The SAO processing unit 212b performs SAO processing on the reconstructed image after DBF processing. In addition, the ALF processing unit 212c applies ALF processing to the reconstructed image after SAO processing.

[0362] Note that the loop filter unit 212 does not necessarily have to include all of the processing units disclosed in Figure 27, and may include only some of them. Also, the loop filter unit 212 may perform the above-mentioned processing in an order different from the processing order disclosed in Figure 27.

[0363] [Prediction Unit (Intra Prediction Unit, Inter Prediction Unit, Prediction Control Unit)] Figure 28 is a flowchart showing an example of processing performed in the prediction unit of the encoding device 100. For example, the prediction unit consists of all or some of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra prediction unit 124 and the inter prediction unit 126.

[0364] The prediction unit generates a predicted image of the current block (step Sb_1). The predicted image may be, for example, an intra-prediction image (intra-prediction signal) or an inter-prediction image (inter-prediction signal). Specifically, the prediction unit generates a predicted image of the current block using a reconstructed image already obtained by generating predicted images for other blocks, generating prediction residuals, generating quantization coefficients, restoring prediction residuals, and adding the predicted images.

[0365] The reconstructed image may be, for example, the image of the reference picture, or it may be the image of an encoded block (i.e., the other block mentioned above) within the current picture, which is the picture containing the current block. An encoded block within the current picture is, for example, an adjacent block to the current block.

[0366] Figure 29 is a flowchart showing another example of the processing performed in the prediction unit of the encoding device 100.

[0367] The prediction unit generates a predicted image using a first method (step Sc_1a), a second method (step Sc_1b), and a third method (step Sc_1c). The first, second, and third methods are different methods for generating predicted images, and may be, for example, an interpretation method, an intraprediction method, and other prediction methods. These prediction methods may use the reconstructed images described above.

[0368] Next, the prediction unit evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the prediction unit calculates a cost C for each of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c, and evaluates the predicted images by comparing the costs C of those predicted images.

[0369] The cost C is calculated using the R-D optimization model formula, for example, C = D + λ × R. In this formula, D is the coding distortion of the predicted image, which can be expressed as, for example, the sum of the absolute differences between the pixel values ​​of the current block and the pixel values ​​of the predicted image. R is the bitrate of the stream, and λ is, for example, the Lagrange multiplier.

[0370] Next, the prediction unit selects one of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the prediction unit selects a method or mode for obtaining the final predicted image. For example, the prediction unit selects the predicted image with the smallest cost C based on the cost C calculated for those predicted images. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 may be based on parameters used in the coding process.

[0371] The encoding device 100 may signal information to identify the selected prediction image, scheme, or mode into a stream. This information may be, for example, a flag. Based on this information, the decoding device 200 can generate a prediction image according to the scheme or mode selected by the encoding device 100.

[0372] In the example shown in Figure 29, the prediction unit generates predicted images using each method and then selects one of the predicted images. However, the prediction unit may also select a method or mode based on the parameters used in the encoding process described above before generating those predicted images, and then generate the predicted images according to that method or mode.

[0373] For example, the first method and the second method are intra-prediction and inter-prediction, respectively, and the prediction unit may select the final predicted image for the current block from the predicted images generated according to these prediction methods.

[0374] Figure 30 is a flowchart showing another example of the processing performed in the prediction unit of the encoding device 100.

[0375] First, the prediction unit generates a predicted image by intra-prediction (step Sd_1a) and then generates a predicted image by inter-prediction (step Sd_1b). The predicted image generated by intra-prediction is also called an intra-prediction image, and the predicted image generated by inter-prediction is also called an inter-prediction image.

[0376] Next, the prediction unit evaluates both the intra-predicted image and the inter-predicted image (step Sd_2). The cost C mentioned above may be used for this evaluation. The prediction unit may then select the prediction image with the smallest cost C from the intra-predicted image and the inter-predicted image as the final prediction image for the current block (step Sd_3). In other words, a prediction method or mode for generating the prediction image for the current block is selected.

[0377] [Prediction Control Unit] The prediction control unit 128 selects either an intra-prediction image (an image or signal output from the intra-prediction unit 124) or an inter-prediction image (an image or signal output from the inter-prediction unit 126), and outputs the selected prediction image to the subtraction unit 104 and the addition unit 116.

[0378] [Prediction Parameter Generation Unit] The prediction parameter generation unit 130 may output information regarding intra-prediction, inter-prediction, and the selection of a predicted image in the prediction control unit 128 as prediction parameters to the entropy coding unit 110. The entropy coding unit 110 may generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantization coefficients input from the quantization unit 108. The prediction parameters may be used by the decoding device 200.

[0379] The decoding device 200 may receive and decode the stream and perform the same processing as the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0380] The prediction parameters may include a selected prediction signal (e.g., MV, prediction type, or prediction mode used in the intra-prediction unit 124 or the inter-prediction unit 126), or any index, flag, or value that is based on or indicates the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128.

[0381] [Generated Reference Picture] Figure 31A is a conceptual diagram showing another example of a reference picture. As illustrated in Figure 31A, a new image (PicA') generated based on an image (PicA) may be used as a reference picture. A reference picture generated in this way is defined as a generated reference picture. A generated reference picture may be used as a reference picture in the prediction process.

[0382] Additionally, the generated reference picture may be added to the same reference picture list as the reference picture list to which other reference pictures are added. Alternatively, a separate reference picture list may be generated, and the generated reference picture may be added to that separate reference picture list.

[0383] Figure 31B is a conceptual diagram illustrating another example of generating a generated reference picture. In the example in Figure 31B, a new image (PicAB') is generated by inputting two images (PicA and PicB) into an NPU (Neural Processing Unit) or a GPU (Graphics Processing Unit). The generated new image (PicAB') is used as a generated reference picture for interpretation.

[0384] Interpretation using this generated reference picture may also be called neural network intercoding (NN inter). Note that the NPU or GPU may be input to more than two images, not just two.

[0385] Figure 31C is a conceptual diagram illustrating another example of generating a generated reference picture. As illustrated in Figure 31C, the size of the input image and the size of the generated image may differ. In this example, the size of the generated image (PicA') is smaller than the size of the input image (PicA). The size of the generated image is not limited to this example and may be larger than the size of the input image. Image processing or the size of the generated image per frame may be based on the performance of the NPU or GPU. This size information may also be included in the bitstream.

[0386] In the example above, a new reference picture (also called a generated image or processed image) is generated from one or more encoded or decoded reference pictures using a neural network. The method for generating the new reference picture is not limited to the example above. For example, a new reference picture may be generated using a process such as image conversion, and the generated new reference picture may be defined as the generated reference picture.

[0387] Adding these generated reference pictures to the reference picture list and making them accessible may improve encoding efficiency.

[0388] In the example above, images are generated on a picture-by-picture basis, but the generation unit is not limited to this example. Reference data may be generated in predetermined units such as blocks, slices, or tiles contained within a picture.

[0389] [Generated Pictures] Generated pictures are images obtained through a generation process, unlike regular reference images (in other words, images decoded by intra / inter prediction). Here, as an example, we show how to generate a new reference picture by inputting one or more pictures into a neural network, but the method of generating a new reference picture is not limited to this.

[0390] The generation process may support the application of NN loop filter (NN LF) processing, which includes NN (Neural Network) processing, or it may support transformation processing (scaling, rotation, mirroring, shifting, etc.). Furthermore, the generation process may be a combination of multiple processes (for example, a process that generates a predicted image using an NN such as NN LF, a transformation process, and other processes).

[0391] The new reference picture generated by the generation process is also referred to as a generated reference picture, a generated picture, or a processed picture. Furthermore, such a new reference picture or its image may also be referred to as a generated reference image, a generated image, or a processed image. Adding the generated picture to the reference picture list and making it accessible may improve encoding efficiency.

[0392] First, we will explain the generated reference picture that is generated using the NN. The encoding device 100 and the decoding device 200 input one or more images to the NN and generate a generated picture.

[0393] Figure 32A is a conceptual diagram illustrating an example of a generated picture. As illustrated in Figure 32A, a new image (PicA') generated by inputting an image (PicA) into the NN may be used as a reference picture. For example, a reference picture generated by NN LF is an example of this.

[0394] Figure 32B is a conceptual diagram showing another example of generating a generated picture. As illustrated in Figure 32B, a new image (PicAB') generated by inputting two images (PicA and PicB) into the NN may be used as the reference picture. Note that there are not limited to two images input into the NN; more than two images can be input.

[0395] Figure 32C is a conceptual diagram showing another example of generating a generated picture. As illustrated in Figure 32C, if NN processing such as NN LF is not fast enough, a new image (PicA') smaller in size than the input image (PicA) may be used as a reference picture. In other words, the size of the input image and the size of the generated image may be different. Also, information about the size of the generated image may be included in the bitstream.

[0396] While this example shows generation at the picture level, the generation unit is not limited to this. The generated reference picture may be generated in predetermined units such as blocks, slices, or tiles contained within the picture.

[0397] Furthermore, the NN may be implemented using an NPU (Neural Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit).

[0398] Furthermore, as mentioned above, the newly generated image (PicA' or PicAB') may be used as a generated picture for inter prediction. Inter prediction using this generated picture may be called neural network intercoding (NN inter).

[0399] Furthermore, multiple neural networks (NNs) may be combined in the generation of the generated picture. In other words, for example, PicA and PicB may be input to NN1 to generate PicAB', and PicAB' may be input to another NN2 to generate PicAB''.

[0400] In typical image encoding on a CPU, a conventional pipeline is used, as shown in Figure 33A. Specifically, the entire encoding process is divided into several stages (0 to n), and processing is performed in block units within each stage. Furthermore, multiple stages are processed in parallel with each other.

[0401] In contrast, for generated pictures produced using NN processing, as can be seen from the configuration examples in Figures 33B and 33C, the image is first processed by the CPU, and then the data is passed to the NPU / GPU / TPU, etc. Then, NN processing is executed on the NPU / GPU / TPU side using a different pipeline than the CPU. In addition, processing may be executed on the NPU / GPU / TPU in units different from those of the CPU.

[0402] Due to overhead caused by data transfer and other processes, it may take some time for the generated picture produced by the NN processing to become available as a reference image for interpretation.

[0403] Patches may be defined to generate the generated picture by NN processing. A patch refers to a predetermined pixel region containing a predetermined number of pixels, as shown in Figure 34A. The shape of the patch may be a square, a non-square rectangle, or any other shape, as shown in Figure 34A. For example, a patch may have a non-rectangular shape such as a triangle or a rhombus.

[0404] Figure 34B is a conceptual diagram showing a patch containing an extended region. A patch containing an extended region includes an extended region in the vertical and / or horizontal directions relative to the overall size of the patch.

[0405] In the example shown in Figure 34B, the extension area has equal size in the vertical and horizontal directions, resulting in a square patch. The sizes of the vertical and horizontal extension areas are not limited to the example in Figure 34B and may be different. Also, the extension areas do not have to be equal, for example, at the top and bottom, or left and right. Furthermore, information regarding the shape or size of the patch may be added to the header area of ​​the bitstream, or it may be added to the bitstream as metadata (also called meta information).

[0406] Figure 34C is a conceptual diagram showing an image containing patches 1A and 1B that have overlapping regions. Patches 1A and 1B overlap each other across multiple pixels (=set of samples), as shown by the dotted lines in Figure 34C. Information indicating the shape (width and height) or size (number of pixels included in the overlapping region) of this overlapping region may be stored in the bitstream.

[0407] Furthermore, in the example in Figure 34C, each patch contains a 256x256 block. Therefore, each patch can also be described as a patch with an extended area relative to the 256x256 block in Figure 34B. In the example in Figure 34C, the 256x256 block in patch 1A and the 256x256 block in patch 1B are defined so as not to overlap.

[0408] Furthermore, while Figure 34C shows an example of two patches defined within a single picture, in an NN interface, for example, two patches may be defined in each of two pictures. In that case, the positions of the patches defined in the first and second pictures (for example, patch 1A in the first picture and patch 2A in the second picture, which is not shown) may be the same or different.

[0409] Figure 34D is a conceptual diagram showing the patches defined in each picture. Specifically, it shows multiple patches defined in an example such as an NN interface, including the first and third patches in the first picture, the second and fourth patches in the second picture, and the fifth and sixth patches in the third picture. The juxtaposed points of the third, fourth, and sixth patches are indicated by an "x" mark.

[0410] Figure 34E is a conceptual diagram showing an example of a method for generating the sixth patch. The sixth patch is obtained by cropping the seventh patch, which is generated by inputting the third and fourth patches into an NN generator. Similarly, the fifth patch may be obtained by cropping an image (the eighth patch, not shown) generated by inputting the first and second patches into an NN generator.

[0411] Furthermore, the first and second pictures may be ordinary reference pictures. In contrast, the third picture, which includes the fifth and sixth patches, may be a generated picture. In one embodiment, the encoding order and display order may be first picture → second picture → third picture. In another embodiment, the display order may be first picture → third picture → second picture.

[0412] Next, the generated picture produced using the transformation process will be described. The encoding device 100 and the decoding device 200 generate a generated picture by performing a transformation process on a single image. The single image input to the transformation process may be a normal reference image (in other words, an image decoded by intra / inter prediction) or an image generated by an NN. Furthermore, the image generated by the transformation process may be used as input to the NN.

[0413] Figure 35A is a conceptual diagram illustrating another example of a generated picture. As illustrated in Figure 35A, a new image (PicA') generated by scaling (enlarging or reducing) an image (PicA) may be used as the reference picture.

[0414] In the case of upscaling (enlarging), a new image PicA' is generated by enlarging the image and then cropping the resulting enlarged image. In the case of downscaling (reducing), a new image PicA' is generated by reducing the image and then padding the resulting reduced image. Note that enlargement can also be expressed as expansion.

[0415] Figure 35B is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 35B, a new image (PicA') generated by rotating an image (PicA) may be used as the reference picture. For example, a new image PicA' is generated by rotating the image by an angle θ in a specified direction and then cropping and padding the rotated image. The rotation direction may always be clockwise or always counterclockwise, and the rotation direction may also be specified as a parameter included in the bitstream along with the rotation angle.

[0416] Figure 35C is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 35C, a new image (PicA') generated by mirroring an image (PicA) may be used as a reference picture. In Figure 35C, the left-hand sample of the input image (PicA) is mirrored horizontally.

[0417] Mirroring is not limited to the example above; the sample on the right may be mirrored to the left, or vertical mirroring may be performed (mirroring the upper sample to the lower side, or mirroring the lower sample to the upper side). In this example, when mirroring, half of the width x of the input image (PicA) (or half of the height y in the case of vertical mirroring) is always mirrored, and the width / height of the mirroring is uniformly defined, but the width / height of the mirroring may also be specified by parameters.

[0418] Figure 35D is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 35D, a new image (PicA') generated by shifting an image (PicA) may be used as the reference picture. In Figure 35D, the input image (PicA) is shifted to the right. In other words, the left side of the input image is padded and the right side is cropped. The shift is not limited to this example; the right side of the sample may be shifted to the left, or a vertical shift (shifted upwards or downwards) may be performed.

[0419] In the above, shifting the image to the right corresponds to shifting the image sample, i.e., the image content, to the right, and to shifting the processing range of the image to the left. Similarly, shifting the image to the left, up, and down corresponds to shifting the image sample, i.e., the image content, to the left, up, and down, respectively, and to shifting the processing range of the image to the right, down, and up, respectively.

[0420] Furthermore, multiple transformations may be combined. For example, an image (PicA) may be scaled to produce an image (Scaled PicA), and then shifted to produce an image (Shifted PicA).

[0421] [NN Loop Filter] The NN Loop Filter (NN LF: Neural Network Loop Filter) corresponds to a filtering process performed on an input image (such as a PicA) using a neural network (NN). The image generated by the NN LF may be used as a reference image for interpretation, or it may be used as the display image as is.

[0422] [NN Inter Prediction] NN Inter Prediction (Neural Network Inter Prediction) corresponds to inter prediction performed using a neural network (NN). Specifically, in NN inter prediction, one or more images (PicA and / or PicB, etc.) are used as input to the NN, and an image is generated as a generated image (PicAB', etc.). The generated image is then used as a reference image for inter prediction. Since the image is generated using an NN, the aforementioned NN processing may also be used.

[0423] The generated reference image used in NN interpretation is not an encoded or decoded image, and therefore may not contain motion information, reference image index, encoding information, or any combination thereof. Furthermore, since the generated prediction image is generated by predicting the current picture, it may have the same POC (Proof of Concept) as the current picture (CurrentPic).

[0424] Note that while this example shows the use of newly generated images using a neural network as an example of NN interpretation, the NN interpretation process is not limited to this example. For example, images to which NN LF processing has been applied may also be used.

[0425] [Translation Inter Prediction] Translation Inter Prediction corresponds to an inter prediction performed using a transformation process. Specifically, in translation inter prediction, a transformation process is applied to one or more images (PicA and / or PicB, etc.) to generate an image (PicAB', etc.). This generated image is then used as the reference image for inter prediction. The transformation process described above may also be used as the transformation process.

[0426] The generated reference image referenced in the conversion interface prediction is not an encoded or decoded image, and therefore may not contain motion information, reference image index, encoding information, or any combination thereof. Furthermore, since the generated prediction image is generated by predicting the current picture, it may have the same POC (Proof of Concept) as the current picture (CurrentPic).

[0427] [Prior Art Regarding Loop Filter Processing Unit Including NNLF Processing] Several configurations have been proposed for a loop filter unit including NNLF (Neural Network Loop Filter) processing in order to improve the decoding efficiency of a decoding device. Figure 36 is a diagram showing a first configuration of the loop filter unit of a decoding device according to a reference example. Figure 37 is a diagram showing a second configuration of the loop filter unit of a decoding device according to a reference example.

[0428] In the configuration shown in Figure 36, the loop filter unit performs LMCS processing on the current block to be decoded in the reconstructed image, then performs DBF processing and NNLF processing in parallel, and then sequentially applies SAO processing and ALF processing to the data obtained by combining the output data obtained by DBF processing and the output data obtained by NNLF processing. Note that LMCS processing, DBF processing, SAO processing, and ALF processing each correspond to non-neural network loop filter processing.

[0429] Furthermore, in the configuration shown in Figure 37, the loop filter unit sequentially applies LMCS processing, DBF processing, and SAO processing to the current block to be decoded included in the reconstructed image, then performs processing using a fixed filter within the ALF processing and NNLF processing in parallel, and then applies processing using Online within the ALF processing to the data obtained by combining the output data obtained by processing using a fixed filter and the output data obtained by NNLF processing.

[0430] In the configuration shown in Figures 36 and 37, the decoding device is configured such that the first processor, the CPU, performs other decoding processes as well as LMCS processing, DBF processing, SAO processing, and ALF processing (i.e., one or more non-neural network loop filter processes). The CPU is mainly configured to perform the above five processes in block units. On the other hand, in the configuration shown in Figures 36 and 37, in order to speed up the NNLF processing, the decoding device is configured so that one of the second processors, such as the NPU, GPU, and TPU, performs the NNLF processing. Therefore, in the above configuration, in order to perform NNLF processing, it is necessary to transfer data from the CPU to the GPU, etc. Furthermore, in the above configuration, some non-neural network loop filter processes use the output data obtained by the NNLF processing, so it is also necessary to transfer data from the GPU, etc., to the CPU. In other words, the decoding device in the above configuration performs some non-neural network loop filtering, then performs some other non-neural network loop filtering and neural network loop filtering in parallel, and then performs the remaining non-neural network loop filtering. Therefore, in the above configuration example, data is exchanged in blocks between the CPU and the GPU (NPU or TPU). As a result, if the decoding device in the above configuration applies multiple loop filtering to a single picture, a significant delay occurs until the multiple loop filtering processes are completed. Also, if the loop filtering section of the encoding device has the same configuration as the loop filtering section of the above configuration, a significant delay occurs until the multiple loop filtering processes are completed.

[0431] This type of problem also occurs in other neural network-related algorithms.

[0432] [Technical Challenges] As described above, when configuring a loop filter unit including NNLF processing, in order to avoid compromising the generality of multiple loop filter processing, one or more non-neural network loop filter processing is configured to be performed by the CPU, and the NNLF processing is configured to be performed by a GPU or the like. However, in such specifications, if the prior art shown in the reference examples of Figures 36 and 37 is applied as is, it is not possible to achieve the same processing speed as the conventional system. Therefore, this disclosure proposes several configurations for the loop filter unit 212 that can achieve the same processing speed as the conventional system in the above specifications. Furthermore, this disclosure realizes data transfer between the CPU and the GPU or the like using a data bus that supports technologies such as PCIe (Peripheral Component Interconnect Express) or DMA (Direct Memory Access).

[0433] Figure 38 shows an example of data transfer between the CPU and GPU using a data bus corresponding to DMA technology. In this specific example, both the output data obtained by ALF processing and the output data obtained by NNLF processing are stored in the DPB (Decoded Picture Buffer) 400 at different timings and used as candidate reference images for subsequent pictures.

[0434] As shown in Figure 38, the data bus may transfer the same data to multiple hardware devices multiple times using DMA technology. In the example in Figure 38, the data bus uses DMA technology to transfer output data obtained by ALF processing to both the DPB400 and the GPU. Specifically, first, the output data obtained by ALF processing is stored in the DPB400 from the CPU via DMA (in other words, the data bus uses DMA technology). Next, the same data stored in the DPB400 is transferred from the DPB400 to the GPU via DMA, and the data transferred to the GPU is used as input data for NNLF. Then, the output data obtained by NNLF processing is stored in the DPB400 from the GPU via DMA. Finally, either the output data obtained by ALF processing or the output data obtained by NNLF processing is transferred from the DPB400 to the display 500 via DMA, and the display 500 displays the output data as a display image.

[0435] In the specific example shown in Figure 38, output data transferred from the DPB 400 to the display 500 via DMA is shown as a display image, but this is not the only example. For example, output data transferred from the CPU to the display 500 via DMA without being stored in the DPB 400 (specifically, output data obtained by ALF processing) may be displayed as a display image. Alternatively, for example, output data transferred from the GPU to the display 500 via DMA without being stored in the DPB 400 (specifically, output data obtained by NNLF processing) may be displayed as a display image.

[0436] [Overview of Loop Filtering by Encoding Device] Figure 39 is a flowchart showing an overview of loop filtering by the encoding device 100 of this disclosure.

[0437] The symbolization device 100 encodes the first sample group into a bitstream (S1001). Note that the first sample group is image data (specifically, a picture including the current block, etc.). For example, the first sample group includes samples filtered by ALF processing. In such a case, in step S1001, the symbolization device 100 repeatedly executes a plurality of loop filter processes in block units to encode the first sample group in picture units into a bitstream.

[0438] The symbolization device 100 processes the second sample group by performing NNLF processing on the first sample group (S1002). Note that the processing procedure of the second sample group includes at least a filtering step. Also, the second sample group is image data (specifically, a picture including the current block, etc.). In addition, the NNLF processing can include at least one of a convolutional layer, a recurrent layer, an attention layer, an embedding layer, a spiking layer, a memory layer, and a special layer.

[0439] The symbolization device 100 encodes the third sample group into a bitstream using the second sample group as a reference image (S1003). Note that when the symbolization device 100 encodes the third sample group, the first sample group may be used as a reference image instead of the second sample group. Also, the third sample group is image data (specifically, a picture including the current block, etc.). Similar to step S1001, in step S1003, the symbolization device 100 repeatedly executes a plurality of loop filter processes in block units to encode the third sample group in picture units into a bitstream.

[0440] Also, the first sample group, the second sample group, and the third sample group can each be any of a processing block, a CTU, a slice, a tile, a subpicture, and a picture. Also, the first sample group, the second sample group, and the third sample group can each be at other granularity levels other than the above. Also, the first sample group, the second sample group, and the third sample group correspond to output data.

[0441] Also, the first sample group, the second sample group, and the third sample group may each have the same particle size level. Also, the first sample group and the third sample group may have a particle size level different from that of the second sample group.

[0442] The encoding device 100 can use the output data of any one of the first sample group, the second sample group, and the third sample group as a display image. In the present disclosure, the encoding device 100 using the output data as a display image may be to encode information such as the display timing when the output data is used as a display image.

[0443] Also, the plurality of sample groups described in FIG. 39 may be combined as input data for the NNLf process in step S1002. For example, the encoding device 100 may use the first sample group and the third sample group as input data for the NNLf process.

[0444] Also, in step S1001, the encoding device 100 stores the encoded first sample group in the picture buffer. Also, in step S1002, the encoding device 100 reads out and uses the first sample group from the picture buffer.

[0445] Also, in step S1002, the encoding device 100 stores the second sample group obtained by the NNLf process in the picture buffer. Also, in step S1003, the encoding device 100 reads out and uses the second sample group from the picture buffer.

[0446] [Outline of Loop Filter Processing by Decoding Device] FIG. 40 is a flowchart showing the outline of loop filter processing by the decoding device 200 of the present disclosure.

[0447] The decoding device 200 decodes a first sample group from the bitstream (S2001). The first sample group is image data (specifically, a picture including the current block). For example, the first sample group includes samples filtered by ALF processing. In such a case, in step S2001, the decoding device 200 decodes a first sample group in picture units from the bitstream by repeatedly performing multiple loop filtering processes in block units.

[0448] The decoding device 200 processes the second sample group by performing NNLF processing on the first sample group (S2002). The processing procedure for the second sample group includes at least a filtering step. The second sample group is image data (specifically, a picture including the current block, etc.). The NNLF processing may include at least one layer from among convolutional layers, recursive layers, attention layers, embedding layers, spiking layers, memory layers, and special layers.

[0449] The decoding device 200 uses the second sample group as a reference image to decode the third sample group from the bitstream (S2003). Note that when the decoding device 200 decodes the third sample group, the first sample group may be used as a reference image instead of the second sample group. The third sample group is image data (specifically, a picture including the current block, etc.). In step S2003, similar to step S2001, the decoding device 200 repeatedly performs multiple loop filtering processes in block units to decode the third sample group in picture units from the bitstream.

[0450] Furthermore, the first, second, and third sample groups may each be a processing block, CTU, slice, tile, subpicture, and picture, respectively. Also, the first, second, and third sample groups may each be at other granularity levels not mentioned above. Furthermore, the first, second, and third sample groups correspond to the output data.

[0451] Furthermore, the first, second, and third sample groups may each have the same particle size level. Also, the first and third sample groups may have different particle size levels than the two sample groups.

[0452] The decoding device 200 can use the output data from any of the first, second, or third sample groups as a display image.

[0453] Furthermore, the multiple sample groups described in Figure 40 may be combined as input data for the NNLF processing in step S2002. For example, the decoding device 200 may use the first sample group and the third sample group as input data for the NNLF processing.

[0454] Furthermore, in step S2001, the decoding device 200 stores the decoded first sample group in the DPB 400. Then, in step S2002, the decoding device 200 reads the first sample group from the DPB 400 and uses it.

[0455] Furthermore, in step S2002, the decoding device 200 stores the second sample group that has undergone NNLF processing in the DPB 400. Then, in step S2003, the decoding device 200 reads the second sample group from the DPB 400 and uses it.

[0456] (First Embodiment) [Application Order of Loop Filtering by Decoder] Figure 41 is a diagram illustrating an example of the application order of loop filtering performed by the decoder 200 in the first embodiment.

[0457] At time t=0, the decoding device 200 decodes the first sample group from the bitstream using one or more non-neural network loop filter processes (specifically, LMCS processing / DBF processing / SAO processing / ALF processing) and other decoding processes (corresponding to the process in step S2001 of Figure 40). For example, the decoding device 200 decodes the first sample group in picture units from the bitstream by repeatedly executing multiple loop filter processes in block units. One or more non-neural network loop filters are performed sequentially by the CPU (first processor).

[0458] The first sample group may be in block units.

[0459] Next, the first sample group, which is the output data obtained by the above process, is used as input data for one or more neural network loop filtering (NNLF) processes (corresponding to the process in step S2002 of Figure 40). The NNLF is performed by the GPU (second processor). Therefore, the first sample group is transferred from the CPU to the GPU via a data bus compatible with technologies such as PCIe or DMA, and used as input data for the NNLF process. The second sample group, which is the output data obtained by the NNLF process, is stored in the DPB400. The second sample group is a picture containing the current block to be decoded. Therefore, the decoding device 200 performs NNLF processing on a picture-by-picture basis and outputs on a picture-by-picture basis.

[0460] The decoding device 200 may perform NNLF processing on a block-by-block basis and output on a picture-by-picture basis.

[0461] Then, at time t=1, when decoding the third sample group from the bitstream, the decoding device 200 uses the second sample group stored in the DPB 400 at time t=0 as a reference image (corresponding to the process in step S2003 of Figure 40). For example, the decoding device 200 decodes the third sample group in picture units from the bitstream by repeatedly executing multiple loop filtering processes in block units. The third sample group may also be in block units.

[0462] Figure 42 shows a simplified block diagram illustrating the application order of the loop filter processing shown in Figure 41. Figure 42 is a simplified block diagram illustrating the configuration of the loop filter unit 212 according to the first embodiment.

[0463] As shown in Figure 42, the loop filter unit 212 applies one or more non-neural network loop filtering processes (specifically, LMCS processing / DBF processing / SAO processing / ALF processing) to the first sample group, and then applies one or more neural network loop filtering processes (NNLF processing) to the second sample group. Furthermore, the loop filter unit 212 shown in Figure 42 performs the one or more non-neural network loop filtering processes without using the output data obtained from the one or more neural network loop filtering processes. The first sample group also includes samples that are output data obtained from the ALF processing.

[0464] The following provides an example of a specific configuration of the loop filter unit 212 according to the first embodiment proposed in this disclosure.

[0465] [First Configuration Example of Loop Filter Unit According to the First Embodiment] Figure 43 is a diagram showing a first configuration example of the loop filter unit 212 according to the first embodiment. Figure 43 is a conceptual diagram showing a first data transfer example performed between the CPU and the GPU. In the first configuration example shown in Figure 43, output data from one or more non-neural network loop filter processes, specifically, a first sample group or a third sample group which are output data obtained by ALF processing, are used as the display image.

[0466] Other decoding processes and one or more non-neural network loop filtering processes (LMCS processing / DBF processing / SAO processing / ALF processing) are applied to the bitstream (S101). The other decoding processes and one or more non-neural network loop filtering processes are performed sequentially by the CPU.

[0467] The output data obtained by ALF processing (for example, the first sample group) is transferred from the CPU to the data bus 300 corresponding to DMA technology (S102). Subsequently, the output data obtained by ALF processing is further transferred from the data bus 300 to the GPU (S103) and used as input data for NNLF processing.

[0468] Meanwhile, the output data obtained by ALF processing is also transmitted to the display 500 for use as a display image (S104).

[0469] Furthermore, after the loop filter unit 212 applies NNLF processing to the input data, the output data obtained by the NNLF processing (for example, the second sample group) is transferred from the GPU to the data bus 300 (S105). Subsequently, the output data obtained by the NNLF processing is further transferred from the data bus 300 to the DPB 400 (S106) and used as a reference image for subsequent pictures (for example, the third sample group). Alternatively, the output data obtained by the NNLF processing may be stored in an RPL (Reference Picture List) via the data bus 300 and used as a reference image for subsequent pictures.

[0470] [Second Configuration Example of Loop Filter Unit According to the First Embodiment] Figure 44 is a diagram showing a second configuration example of the loop filter unit 212 according to the first embodiment. Figure 44 is a conceptual diagram showing a second data transfer example performed between the CPU and the GPU. In the second configuration example shown in Figure 44, the output data of one or more neural network loop filter processes, specifically the second sample group which is output data obtained by NNLF processing, is used as the display image. In other words, the difference between the first configuration example of the loop filter unit 212 and the second configuration example of the loop filter unit 212 lies in the difference in the output data used as the display image.

[0471] Other decoding processes and one or more non - neural network loop filter processes (LMCS process / DBF process / SAO process / ALF process) are applied to the bitstream (S111). Note that the other decoding processes and one or more non - neural network loop filter processes are sequentially performed by the CPU.

[0472] The output data (e.g., the first sample group) obtained by the ALF process is transferred from the CPU to the data bus 300 corresponding to the DMA technology (S112). Then, the output data obtained by the ALF process is further transferred from the data bus 300 to the GPU (S113) and used as input data for the NNLF process.

[0473] After the loop filter unit 212 applies the NNLF process to the input data, the output data (e.g., the second sample group) obtained by the NNLF process is transferred from the GPU to the data bus 300 (S114). Then, the output data obtained by the NNLF process is transferred from the data bus 300 to the DPB 400 (S115) and used as a reference image for a subsequent picture (e.g., the third sample group). Note that the output data obtained by the NNLF process may be stored in the RPL via the data bus 300 and used as a reference image for a subsequent picture.

[0474] Also, the output data (e.g., the second sample group) obtained by the NNLF process is transmitted to the display 500 for use as a display image (S116).

[0475] [Advantages of the first and second configuration examples of the loop filter unit according to the first aspect] As described above, by adding one or more neural network loop filter processes (NNLF processes) to the loop filter unit 212, the decoding device 200 can improve the quality of the reconstructed image and the image compression rate. Also, the decoding device 200 can perform sequential processing and suppress the complication of control.

[0476] Furthermore, the first and second configuration examples of the loop filter unit 212 according to the first embodiment are configured to apply one or more neural network loop filter processes to the current block to be decoded after applying all non-neural network loop filter processes to the current block to be decoded. Therefore, in the loop filter unit 212, all non-neural network loop filter processes are executed on the CPU before one or more neural network loop filter processes are executed on the GPU. After that, in the loop filter unit 212, the output data (specifically, the picture including the current block to be decoded) is transferred from the CPU to the GPU only once, and NNLF processing is applied to the output data on the GPU. Specifically, the loop filter unit 212 according to the first embodiment does not perform some non-neural network loop filter processes and neural network loop filter processes in parallel. In other words, in the loop filter unit 212 according to the first embodiment, one or more non-neural network loop filter processes are performed without using the output data obtained by one or more neural network loop filter processes. In contrast, in the configurations shown in the reference examples of Figures 36 and 37, data is transferred multiple times in block units between the CPU and the GPU. This is because the output data obtained from one or more non-neural network loop filter processes (in the example of Figure 36, the output data obtained from the LMCS process) needs to be transferred from the CPU to the GPU to be used as input data for the NNLF process, and the output data obtained from the NNLF process needs to be transferred from the GPU to the CPU to be used as input data for the remaining non-neural network loop filter processes (in the example of Figure 36, the SAO process).

[0477] Therefore, the loop filter unit 212 according to the first embodiment can reduce the number of data exchanges between the CPU and the GPU, etc., when decoding a picture including the current block to be decoded, and can reduce the overhead during data transfer between the CPU and the GPU.

[0478] Furthermore, the first and second configuration examples of the loop filter unit 212 according to the first embodiment are configured to apply all non-neural network loop filter processing to the current block to be decoded, and then apply one or more neural network loop filter processing to the current block to be decoded. Therefore, the loop filter unit 212 can free up the CPU operation while the NNLF processing is being performed on the GPU. Specifically, in the loop filter unit 212 according to the first embodiment, while the GPU (second processor) is performing NNLF processing, the CPU (first processor) can perform other processing such as intra / inter prediction and decoding of subsequent pictures. In contrast, in the configurations according to the reference examples in Figures 36 and 37, one or more neural network loop filter processing is provided in parallel with some non-neural network loop filter processing that is in the middle of one or more non-neural network loop filter processing. Therefore, in the example configuration, the remaining non-neural network loop filtering processes (SAO and ALF in the example of Figure 36) that are performed after the partial non-neural network loop filtering process cannot be carried out until one or more neural network loop filtering processes are completed.

[0479] Thus, the loop filter unit 212 according to the first embodiment is configured such that there is no non-neural network loop filter process that uses output data obtained from one or more neural network loop filter processes. Therefore, the loop filter unit 212 according to the first embodiment can independently perform one or more non-neural network loop filter processes without depending on the output data obtained from one or more neural network loop filter processes. In other words, the loop filter unit 212 according to the first embodiment can proceed with CPU processing without waiting for the completion of one or more neural network loop filter processes, and can perform CPU processing and GPU processing in parallel.

[0480] Furthermore, in the first configuration example of the loop filter unit 212 according to the first embodiment (Figure 43), the output data obtained by ALF processing is displayed on the display 500. Therefore, in the first configuration example, the reconstructed image that has undergone one or more non-neural network loop filter processing performed by the CPU can be displayed immediately without waiting for the results of one or more neural network loop filter processing performed by the GPU. In addition, the first configuration example can suppress the delay in displaying the reconstructed image even when image processing by one or more neural network loop filter processing takes a long time.

[0481] Furthermore, in the second configuration example of the loop filter unit 212 according to the first embodiment (Figure 44), the output data obtained by NNLF processing is displayed on the display 500. Therefore, in the second configuration example, when the time required for one or more neural network loop filter processes performed by the GPU is relatively short, compared to the case where the output data obtained by one or more non-neural network loop filter processes is displayed on the display 500 from the start of bitstream decoding, the output data obtained by one or more neural network loop filters can be displayed on the display 500 at almost the same timing, and a reconstructed image with better image quality can be displayed on the display 500. The case where the time required for one or more neural network loop filter processes is relatively short includes, for example, when one or more neural network loop filter processes are low-complexity neural network loop filter processes, or when a high-speed GPU is used for one or more neural network loop filter processes.

[0482] While Figures 43 and 44 illustrate an example configuration of the loop filter unit 212 when the CPU performs one or more non-neural network loop filter processes and the GPU performs one or more neural network loop filter processes, the configuration is not limited to this. Specifically, the configuration of the loop filter unit 212 may be such that the GPU performs some of the non-neural network loop filter processes among the one or more non-neural network loop filter processes. More specifically, the configuration of the loop filter unit 212 may be such that the GPU performs non-neural network loop filter processes that are provided consecutively immediately before the NNLF process. For example, the configuration of the loop filter unit 212 may be such that the GPU performs only the ALF process among the one or more non-neural network loop filter processes. Alternatively, for example, the configuration of the loop filter unit 212 may be such that the GPU performs the SAO process and the ALF process among the one or more non-neural network loop filter processes. For example, the configuration of the loop filter section 212 may be such that a GPU or the like performs DBF processing, SAO processing, and ALF processing among the non-neural network loop filter processing of one or more.

[0483] Furthermore, while an example in which only output data obtained from one or more neural network loop filtering processes is stored in the DPB 400 has been explained using Figures 43 and 44, data other than output data obtained from one or more neural network loop filtering processes may also be stored in the DPB 400. In another configuration example of the loop filter unit 212 according to the first embodiment, in addition to output data obtained from one or more neural network loop filtering processes, output data obtained from one or more non-neural network loop filtering processes is stored in the DPB 400. Figure 45 is another block diagram that simplifies the configuration of the loop filter unit 212 according to the first embodiment.

[0484] The loop filter unit 212 shown in Figure 45 differs from the loop filter unit 212 shown in Figure 42 in that, in addition to the output data obtained from one or more neural network loop filter processes, it stores the output data from one or more non-neural network loop filter processes in the DPB 400. Specifically, the loop filter unit 212 shown in Figure 45 stores the output data obtained from ALF processing and the output data obtained from NNLF processing in the DPB 400. As a result, the loop filter unit 212 shown in Figure 45 can use the output data obtained from ALF processing and the output data obtained from NNLF processing as candidate reference images for subsequent pictures. Below, variations of the specific configuration example of the loop filter unit 212 shown in Figure 45 will be illustrated.

[0485] [Third Configuration Example of Loop Filter Unit According to the First Embodiment] Figure 46 is a diagram showing a third configuration example of the loop filter unit 212 according to the first embodiment. Figure 46 is a conceptual diagram showing a third data transfer example performed between the CPU and the GPU.

[0486] Other decoding processes and one or more non-neural network loop filtering processes (LMCS processing / DBF processing / SAO processing / ALF processing) are applied to the bitstream (S201). The other decoding processes and one or more non-neural network loop filtering processes are performed sequentially by the CPU.

[0487] The output data obtained by the ALF processing (e.g., the first sample group) is transferred from the CPU to the data bus 300 corresponding to the DMA technology (S202). Subsequently, the output data obtained by the ALF processing is transferred from the data bus 300 to the DPB 400 (S203) and used as a candidate reference image for the subsequent picture (e.g., the third sample group). During this data transfer, the CPU can perform other processing, such as decoding the subsequent picture.

[0488] Simultaneously, the output data transferred to the DPB400 (output data obtained by ALF processing) is transferred from the DPB400 to the data bus 300 (S204), and then transferred from the data bus 300 to the GPU (S205). The output data transferred to the GPU is then used as input data for NNLF processing.

[0489] After the loop filter unit 212 performs NNLF processing on the input data, the output data obtained by the NNLF processing is transferred from the GPU to the data bus 300 (S206), and then transferred from the data bus 300 to the DPB 400 (S207). The output data transferred to the DPB 400 is then used as a reference image for subsequent pictures.

[0490] Furthermore, the output data obtained by NNLF processing (for example, the second sample group) is also transmitted to the display 500 for use as a display image (S208).

[0491] [Fourth Configuration Example of the Loop Filter Unit According to the First Embodiment] Figure 47 is a diagram showing a fourth configuration example of the loop filter unit 212 according to the first embodiment. Figure 47 is a conceptual diagram showing a fourth data transfer example performed between the CPU and the GPU.

[0492] Other decoding processes and one or more non-neural network loop filters (LMCS processing / DBF processing / SAO processing / ALF processing) are applied to the bitstream (S211). The other decoding processes and one or more non-neural network loop filter processes are performed sequentially by the CPU.

[0493] The output data obtained by ALF processing (for example, the first sample group) is transferred from the CPU to the data bus 300 corresponding to DMA technology (S212). Subsequently, the output data transferred to the data bus 300 is transferred from the data bus 300 to the DPB 400 (S213) and used as a candidate reference image for the subsequent picture. During this data transfer, the CPU performs other processing, such as decoding the subsequent picture.

[0494] Furthermore, the output data obtained by ALF processing (for example, the second sample group) is transmitted to the display 500 for use as a display image (S214).

[0495] At the same time, the output data transferred to the DPB400 (output data obtained by ALF processing) is transferred to the data bus 300 (S215), and then transferred from the data bus 300 to the GPU (S216). The output data transferred to the GPU is then used as input data for NNLF processing.

[0496] After the loop filter unit 212 performs NNLF processing on the input data, the output data obtained by the NNLF processing is transferred from the GPU to the data bus 300 (S217), and then transferred from the data bus 300 to the DPB 400 (S218). The output data transferred to the DPB 400 is then used as a reference image for subsequent pictures.

[0497] [Timing of One or More Neural Network Loop Filter Processing] As explained above, the third and fourth configuration examples of the loop filter unit 212 differ from the first and second configuration examples of the loop filter unit 212 in that, in addition to the output data of one or more neural network loop filter processing, the output data of one or more non-neural network loop filter processing is stored in the DPB 400. Therefore, in the third and fourth configuration examples, by using the output data of one or more non-neural network loop filter processing as a reference image for the subsequent picture, the decoding process of the subsequent picture can be advanced before the output data of one or more neural network loop filter processing is obtained. The timing of one or more neural network loop filter processing performed by the third and fourth configuration examples will be explained below with specific examples.

[0498] [Example 1 of Timing for One or More Neural Network Loop Filter Processing] Figures 48A to 48E are diagrams illustrating examples of timing when the loop filter unit 212 performs one or more neural network loop filter processing in a decoding order of RA (Random Access). Figure 48A is a diagram showing the correspondence between the order in which images are displayed and the order in which images are decoded. Figure 48B is the first diagram illustrating an example of timing when the loop filter unit 212 performs one or more neural network loop filter processing. Figure 48C is the second diagram illustrating an example of timing when the loop filter unit 212 performs one or more neural network loop filter processing. Figure 48D is the third diagram illustrating an example of timing when the loop filter unit 212 performs one or more neural network loop filter processing. Figure 48E is the fourth diagram illustrating an example of timing when the loop filter unit 212 performs one or more neural network loop filter processing.

[0499] As shown in Figure 48A, the decoding device 200 displays the images in the order of POC0, POC1, POC2, POC3, and POC4. Furthermore, the loop filter unit 212, which is described in Figures 48A to 48E, decodes the images in the decoding order of the RA configuration, so after decoding POC0, it decodes the images in the order of POC4, POC2, POC1, and POC3.

[0500] The timing of when the loop filter unit 212 performs one or more neural network loop filtering operations will be explained in detail below using Figures 48B to 48E. In the following explanation, it will be assumed that the time it takes for the loop filter unit 212 to complete one or more neural network loop filtering operations is equivalent to the time it takes for the loop filter unit 212 to decode one picture obtained by applying one or more non-neural network loop filtering operations. Specifically, the maximum processing time (n) required for the loop filter unit 212 to complete one or more neural network loop filtering operations will be assumed to be equal to or less than (n ≤ p) the decoding time (p) of the image obtained by applying one or more non-neural network loop filtering operations.

[0501] Figure 48B is a diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter operations on the POC0. Figure 48B(a) is a diagram illustrating the target images to which the loop filter unit 212 applies one or more non-neural network loop filter operations and the target images to which it applies one or more neural network loop filter operations at time t1. Figure 48B(b) is a diagram illustrating the timing when the loop filter unit 212 performs one or more neural network loop filter operations. Figure 48B(c) is a diagram showing the images stored in the DPB400 at time t1. Figure 48B(d) is a diagram illustrating the figures shown in Figures 48B(a) to (c).

[0502] First, we will use Figure 48B(d) to explain the figures shown in Figures 48B(a) to (c). As shown in Figure 48B(d), four rectangular figures are shown. In the explanation of Figure 48B, the differences between the rectangular figures are represented by the presence or absence and type of hatching. For example, the leftmost rectangular figure in Figure 48B(d) (the figure without hatching) represents a picture that has not been decoded by CPU processing (one or more non-neural network loop filter processing). The second rectangular figure from the left (the figure shown with grid hatching) represents a picture that has been decoded by CPU processing. The third rectangular figure from the left (the figure shown with vertical line hatching) represents a picture that has not been decoded by GPU processing (one or more neural network loop filter processing). The fourth rectangular figure from the left (the figure shown with horizontal line hatching) represents a picture that has been decoded by GPU processing.

[0503] As shown in Figures 48B(a) and (c), at time t1, POC0 (for example, output data obtained by ALF processing) is stored in DPB400. In this case, the loop filter unit 212 uses POC0 (output data obtained by ALF processing) as input data for one or more neural network loop filter processes and starts generating POC0p as output data for one or more neural network loop filter processes. Furthermore, when the loop filter unit 212 generates POC4 by performing one or more non-neural network loop filter processes, it uses only POC0 (output data obtained by ALF processing) as a reference image. The reason the loop filter unit 212 uses only POC0 as a reference image at this time is that POC0p is not stored in DPB400, and only POC0 is stored in DPB400.

[0504] Furthermore, as shown in Figure 48B(b), the processes for generating POC0p and POC4 are performed simultaneously from time t1. In other words, the loop filter unit 212 performs the GPU processing for generating POC0p and the CPU processing for generating POC4 in parallel. Note that the example in Figure 48B(b) shows the case where the time required for GPU processing (n) and the time required for CPU processing (p) are the same. Also, the processes for generating POC0p and POC4 do not necessarily have to be performed simultaneously. The process for generating POC0p may start earlier than the process for generating POC4, or the process for generating POC0p may start later than the process for generating POC4.

[0505] Figure 48C is a diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes on the POC 4. Figure 48C(a) is a diagram illustrating the target images to which the loop filter unit 212 applies one or more non-neural network loop filter processes and the target images to which it applies one or more neural network loop filter processes at time t2. Figure 48C(b) is a diagram illustrating the timing when the loop filter unit 212 performs one or more neural network loop filter processes. Figure 48C(c) is a diagram showing the images stored in the DPB 400 at time t2. Figure 48C(d) is a diagram illustrating the figures shown in Figures 48C(a) to (c). Note that the types of figures shown in Figure 48C(d) are the same as those explained in Figure 48B(d), so the explanation is omitted.

[0506] As shown in Figures 48C(a) and (c), at time t2, the loop filter unit 212 completes one or more neural network loop filter processes on POC0 and generates POC0p (for example, output data obtained by NNLF processing). The generated POC0p is then stored in the DPB 400. In the example in Figure 48C(c), POC0 is overwritten with POC0p, but both images of POC0 and POC0p may be stored in the DPB 400.

[0507] Furthermore, at time t2, the loop filter unit 212 completes one or more non-neural network loop filter processes on the bitstream and generates POC4 (for example, output data obtained by ALF processing). The generated POC4 is then stored in the DPB400.

[0508] Next, at time t2, the loop filter unit 212 starts generating POC4p as output data for one or more neural network loop filter processes, using POC4 (for example, output data obtained by ALF processing) as input data for one or more neural network loop filter processes.

[0509] Furthermore, at time t2, POC0p and POC4 are stored in DPB400. In this case, when the loop filter unit 212 generates POC2 by performing one or more non-neural network loop filter processes, it uses POC0p (for example, output data obtained by NNLF processing) and POC4 (for example, output data obtained by ALF processing) as reference images. However, at the time of generating POC2, POC4p is not stored in DPB400, so the loop filter unit 212 cannot use POC4p (output data obtained by NNLF processing) as a reference image.

[0510] Furthermore, as shown in Figure 48C(b), the processes for generating POC4p and POC2 are performed simultaneously from time t2. In other words, the loop filter unit 212 performs the GPU processing for generating POC4p and the CPU processing for generating POC2 in parallel. Note that the example in Figure 48C(b) shows the case where the time required for GPU processing (n) and the time required for CPU processing (p) are the same. Also, the processes for generating POC4p and POC2 do not necessarily have to be performed simultaneously. The process for generating POC4p may start earlier than the process for generating POC2, or the process for generating POC4p may start later than the process for generating POC2.

[0511] Figure 48D is a diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes on the POC2. Figure 48D(a) is a diagram illustrating the target images to which the loop filter unit 212 applies one or more non-neural network loop filter processes and the target images to which it applies one or more neural network loop filter processes at time t3. Figure 48D(b) is a diagram illustrating the timing when the loop filter unit 212 performs one or more neural network loop filter processes. Figure 48D(c) is a diagram showing the images stored in the DPB400 at time t3. Figure 48D(d) is a diagram illustrating the figures shown in Figures 48D(a) to (c). Note that the types of figures shown in Figure 48D(d) are the same as those explained in Figure 48B(d), so the explanation is omitted.

[0512] As shown in Figures 48D(a) and (c), at time t3, the loop filter unit 212 completes one or more neural network loop filter processes on POC4 and generates POC4p (for example, output data obtained by NNLF processing). The generated POC4p is then stored in the DPB400. In the example in Figure 48D(c), POC4 is overwritten with POC4p, but both images of POC4 and POC4p may be stored in the DPB400.

[0513] Furthermore, at time t3, the loop filter unit 212 completes one or more non-neural network loop filter processes on the bitstream and generates POC2 (for example, output data obtained by ALF processing). The generated POC2 is then stored in the DPB400.

[0514] Next, at time t3, the loop filter unit 212 starts generating POC2p as output data for one or more neural network loop filter processes, using POC2 (output data obtained by ALF) as input data for one or more neural network loop filter processes.

[0515] Furthermore, at time t3, POC0p, POC4p, and POC2 are stored in DPB400. In this case, when the loop filter unit 212 generates POC1 by performing one or more non-neural network loop filter processes, it uses POC0p (for example, output data obtained by NNLF processing), POC4p (for example, output data obtained by NNLF processing), and POC2 (for example, output data obtained by ALF processing) as reference images. However, at the time of generating POC1, POC2p is not stored in DPB400, so the loop filter unit 212 cannot use POC2p (for example, output data obtained by NNLF processing) as a reference image.

[0516] Furthermore, as shown in Figure 48D(b), the processes for generating POC2p and POC1 are performed simultaneously from time t3. In other words, the loop filter unit 212 performs the GPU processing for generating POC2p and the CPU processing for generating POC1 in parallel. Note that the example in Figure 48D(b) shows the case where the time required for GPU processing (n) and the time required for CPU processing (p) are the same. Also, the processes for generating POC2p and POC1 do not necessarily have to be performed simultaneously. The process for generating POC2p may start earlier than the process for generating POC1, or the process for generating POC2p may start later than the process for generating POC1.

[0517] Figure 48E is a diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes on the POC1. Figure 48E(a) is a diagram illustrating the target images to which the loop filter unit 212 applies one or more non-neural network loop filter processes and the target images to which it applies one or more neural network loop filter processes at time t4. Figure 48E(b) is a diagram illustrating the timing when the loop filter unit 212 performs one or more neural network loop filter processes. Figure 48E(c) is a diagram showing the images stored in the DPB400 at time t4. Figure 48E(d) is a diagram illustrating the figures shown in Figures 48E(a) to (c). Note that the types of figures shown in Figure 48E(d) are the same as those explained in Figure 48B(d), so the explanation is omitted.

[0518] As shown in Figures 48E(a) and (c), at time t4, the loop filter unit 212 completes one or more neural network loop filter processes on POC2 and generates POC2p (for example, output data obtained by NNLF processing). The generated POC2p is then stored in the DPB400. In the example in Figure 48E(c), POC2 is overwritten with POC2p, but both images of POC2 and POC2p may be stored in the DPB400.

[0519] Furthermore, at time t4, the loop filter unit 212 completes one or more non-neural network loop filter processes on the bitstream and generates POC1 (for example, output data obtained by ALF processing). The generated POC1 is then stored in the DPB400.

[0520] Next, at time t4, the loop filter unit 212 starts generating POC1p as output data for one or more neural network loop filter processes, using POC1 (for example, output data obtained by ALF processing) as input data for one or more neural network loop filter processes.

[0521] Furthermore, at time t4, POC0p, POC4p, POC2p, and POC1 are stored in DPB400. In this case, when the loop filter unit 212 generates POC3 by performing one or more non-neural network loop filter processes, it uses POC0p (e.g., output data obtained by NNLF processing), POC4p (e.g., output data obtained by NNLF processing), POC2p (e.g., output data obtained by NNLF processing), and POC1 (e.g., output data obtained by ALF processing) as reference images. However, at the time of generating POC3, POC1p is not stored in DPB400, so the loop filter unit 212 cannot use POC1p (output data obtained by NNLF processing) as a reference image.

[0522] Furthermore, as shown in Figure 48E(b), the processes for generating POC1p and POC3 are performed simultaneously from time t4. In other words, the loop filter unit 212 performs the GPU processing for generating POC1p and the CPU processing for generating POC3 in parallel. Note that the example in Figure 48E(b) shows the case where the time required for GPU processing (n) and the time required for CPU processing (p) are the same. Also, the processes for generating POC1p and POC3 do not necessarily have to be performed simultaneously. The process for generating POC1p may start earlier than the process for generating POC3, or the process for generating POC1p may start later than the process for generating POC3.

[0523] [Example 2 of the timing when one or more neural network loop filter processes are performed] Figures 49A to 49E are diagrams illustrating examples of the timing when the loop filter unit 212 performs one or more neural network loop filter processes in an LD (Low-Delay) decoding order. Figure 49A is a diagram illustrating the correspondence between the order in which images are displayed and the order in which images are decoded. Figure 49B is the first diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes. Figure 49C is the second diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes. Figure 49D is the third diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes. Figure 49E is the fourth diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes.

[0524] As shown in Figure 49A, the decoding device 200 displays the images in the order of POC0, POC1, POC2, POC3, and POC4. Furthermore, the loop filter unit 212, which is described in Figures 49A to 49E, decodes the images in the decoding order of the LD configuration, so after decoding POC0, it decodes the images in the order of POC1, POC2, POC3, and POC4.

[0525] The timing of when the loop filter unit 212 performs one or more neural network loop filtering operations will be explained in detail below using Figures 49B to 49E. In the following explanation, it will be assumed that the time it takes for the loop filter unit 212 to complete one or more neural network loop filtering operations is equivalent to the time it takes for the loop filter unit 212 to decode one picture obtained by applying one or more non-neural network loop filtering operations. Specifically, the maximum processing time (n) required for the loop filter unit 212 to complete one or more neural network loop filtering operations will be explained as being equal to or less than (n ≤ p) the decoding time (p) of the image obtained by applying one or more non-neural network loop filtering operations.

[0526] Figure 49B is a diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes on the POC0. Figure 49B(a) is a diagram illustrating the target images to which the loop filter unit 212 applies one or more non-neural network loop filter processes and the target images to which one or more neural network loop filter processes are applied at time t1. Figure 49B(b) is a diagram showing the images stored in the DPB400 at time t1. Figure 49B(c) is a diagram illustrating the figures shown in Figures 49B(a) to (b).

[0527] First, we will use Figure 49B(c) to explain the shapes shown in Figures 49B(a) and (b). As shown in Figure 49B(c), four rectangular shapes are shown. In the explanation of Figure 49B, the differences between the rectangular shapes are represented by the presence or absence and type of hatching. For example, the leftmost rectangular shape in Figure 49B(c) (the shape without hatching) represents a picture that has not been decoded by CPU processing (one or more non-neural network loop filter processing). The second rectangular shape from the left (the shape shown with grid hatching) represents a picture that has been decoded by CPU processing. The third rectangular shape from the left (the shape shown with vertical line hatching) represents a picture that has not been decoded by GPU processing (one or more neural network loop filter processing). The fourth rectangular shape from the left (the shape shown with horizontal line hatching) represents a picture that has been decoded by GPU processing.

[0528] As shown in Figures 49B(a) and (b), at time t1, POC0 (for example, output data obtained by ALF processing) is stored in DPB400. In this case, the loop filter unit 212 uses POC0 (for example, output data obtained by ALF processing) as input data for one or more neural network loop filter processes to generate POC0p as output data for one or more neural network loop filter processes.

[0529] Furthermore, when the loop filter unit 212 performs one or more non-neural network loop filter processes to generate POC1, it uses only POC0 (output data obtained by ALF) as the reference image. The reason the loop filter unit 212 uses only POC0 as the reference image at this time is that POC0p is not stored in the DPB400, and only POC0 is stored in the DPB400.

[0530] Although not shown in Figure 49B, the processes for generating POC0p and POC1 are performed simultaneously from time t1. In other words, the loop filter unit 212 performs the GPU processing for generating POC0p and the CPU processing for generating POC1 in parallel. Figure 49B illustrates the case where the time required for GPU processing and the time required for CPU processing are the same. Furthermore, the processes for generating POC0p and POC1 do not necessarily have to be performed simultaneously. The process for generating POC0p may start earlier than the process for generating POC1, or the process for generating POC0p may start later than the process for generating POC1.

[0531] Figure 49C is a diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes on the POC1. Figure 49C(a) is a diagram illustrating the target images to which the loop filter unit 212 applies one or more non-neural network loop filter processes and the target images to which it applies one or more neural network loop filter processes at time t2. Figure 49C(b) is a diagram showing the images stored in the DPB400 at time t2. Figure 49C(c) is a diagram illustrating the figures shown in Figures 49C(a) to (b). Note that the types of figures shown in Figure 49C(c) are the same as those explained in Figure 49B(c), so the explanation is omitted.

[0532] As shown in Figures 49C(a) and (b), at time t2, the loop filter unit 212 completes one or more neural network loop filtering processes on POC0 and generates POC0p (output data obtained by NNLF). The generated POC0p is then stored in the DPB 400. In the example in Figure 49C(b), POC0 is overwritten with POC0p, but both images of POC0 and POC0p may be stored in the DPB 400.

[0533] Furthermore, at time t2, the loop filter unit 212 completes one or more non-neural network loop filter processes on the bitstream and generates POC1 (for example, output data obtained by ALF processing). The generated POC1 is then stored in the DPB 400.

[0534] Next, at time t2, the loop filter unit 212 starts generating POC1p as output data for one or more neural network loop filter processes, using POC1 (for example, output data obtained by ALF processing) as input data for one or more neural network loop filter processes.

[0535] Furthermore, at time t2, POC0p and POC1 are stored in DPB400. In this case, when the loop filter unit 212 generates POC2 by performing one or more non-neural network loop filter processes, it uses POC0p (for example, output data obtained by NNLF processing) and POC1 (for example, output data obtained by ALF processing) as reference images. However, at the time of generating POC2, POC1p is not stored in DPB400, so the loop filter unit 212 cannot use POC1p (for example, output data obtained by NNLF processing) as a reference image.

[0536] Although not shown in Figure 49C, the processes for generating POC1p and POC2 are performed simultaneously from time t2. In other words, the loop filter unit 212 performs the GPU processing for generating POC1p and the CPU processing for generating POC2 in parallel. Figure 49C illustrates the case where the time required for GPU processing and the time required for CPU processing are the same. Furthermore, the processes for generating POC1p and POC2 do not necessarily have to be performed simultaneously. The process for generating POC1p may start earlier than the process for generating POC2, or the process for generating POC1p may start later than the process for generating POC2.

[0537] Figure 49D is a diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes on the POC2. Figure 49D(a) is a diagram illustrating the target images to which the loop filter unit 212 applies one or more non-neural network loop filter processes and the target images to which it applies one or more neural network loop filter processes at time t3. Figure 49D(b) is a diagram showing the images stored in the DPB400 at time t3. Figure 49D(c) is a diagram illustrating the figures shown in Figures 49D(a) to (b). Note that the types of figures shown in Figure 49D(c) are the same as those explained in Figure 49B(c), so the explanation is omitted.

[0538] As shown in Figures 49D(a) and (b), at time t3, the loop filter unit 212 completes one or more neural network loop filtering processes on POC1 and generates POC1p (output data obtained from NNLF). The generated POC1p is then stored in the DPB400. In the example in Figure 49D(b), POC1 is overwritten with POC1p, but both images of POC1 and POC1p may be stored in the DPB400.

[0539] Furthermore, at time t3, the loop filter unit 212 completes one or more non-neural network loop filter processes on the bitstream and generates POC2 (for example, output data obtained by ALF processing). The generated POC2 is then stored in the DPB400.

[0540] Next, at time t3, the loop filter unit 212 starts generating POC2p as output data for one or more neural network loop filter processes, using POC2 (for example, output data obtained by ALF processing) as input data for one or more neural network loop filter processes.

[0541] Furthermore, at time t3, POC0p, POC1p, and POC2 are stored in DPB400. In this case, when the loop filter unit 212 generates POC3 by performing one or more non-neural network loop filter processes, it uses POC0p (for example, output data obtained by NNLF processing), POC1p (for example, output data obtained by NNLF processing), and POC2 (for example, output data obtained by ALF processing) as reference images. However, at the time of generating POC3, POC2p is not stored in DPB400, so the loop filter unit 212 cannot use POC2p (for example, output data obtained by NNLF processing) as a reference image.

[0542] Although not shown in Figure 49D, the processes for generating POC2p and POC3 are performed simultaneously from time t3. In other words, the loop filter unit 212 performs the GPU processing for generating POC2p and the CPU processing for generating POC3 in parallel. Figure 49D illustrates the case where the time required for GPU processing and the time required for CPU processing are the same. Furthermore, the processes for generating POC2p and POC3 do not necessarily have to be performed simultaneously. The process for generating POC2p may start earlier than the process for generating POC3, or the process for generating POC2p may start later than the process for generating POC3.

[0543] Figure 49E is a diagram illustrating an example of the timing when the loop filter unit 212 performs one or more neural network loop filter processes on the POC 3. Figure 49E(a) is a diagram illustrating the target images to which the loop filter unit 212 applies one or more non-neural network loop filter processes and the target images to which it applies one or more neural network loop filter processes at time t4. Figure 49E(b) is a diagram showing the images stored in the DPB 400 at time t4. Figure 49E(c) is a diagram illustrating the figures shown in Figures 49E(a) to (b). Note that the types of figures shown in Figure 49E(c) are the same as those explained in Figure 49B(c), so the explanation is omitted.

[0544] As shown in Figures 49E(a) and (b), at time t4, the loop filter unit 212 completes one or more neural network loop filtering processes on POC2 and generates POC2p (output data obtained from NNLF). The generated POC2p is then stored in the DPB400. In the example in Figure 49E(b), POC2 is overwritten with POC2p, but both images of POC2 and POC2p may be stored in the DPB400.

[0545] Furthermore, at time t4, the loop filter unit 212 completes one or more non-neural network loop filter processes on the bitstream and generates POC3 (for example, output data obtained by ALF processing). The generated POC3 is then stored in the DPB400.

[0546] Next, at time t4, the loop filter unit 212 starts generating POC3p as output data for one or more neural network loop filter processes, using POC3 (for example, output data obtained by ALF processing) as input data for one or more neural network loop filter processes.

[0547] Furthermore, at time t4, POC0p, POC1p, POC2p, and POC3 are stored in DPB400. In this case, when the loop filter unit 212 generates POC4 by performing one or more non-neural network loop filter processes, it uses POC0p (e.g., output data obtained by NNLF processing), POC1p (e.g., output data obtained by NNLF processing), POC2p (e.g., output data obtained by NNLF processing), and POC3 (e.g., output data obtained by ALF processing) as reference images. However, at the time of generating POC4, POC3p is not stored in DPB400, so the loop filter unit 212 cannot use POC3p (e.g., output data obtained by NNLF processing) as a reference image.

[0548] [Advantages of the Third and Fourth Configuration Examples of the Loop Filter Unit According to the First Embodiment] In the third and fourth configuration examples of the loop filter unit 212 according to the first embodiment, one or more non-neural network loop filter processes are performed without using output data obtained by one or more neural network loop filter processes. As a result, in the loop filter unit 212, while the GPU (second processor) is performing one or more neural network loop filter processes on the picture containing the current block to be decoded, the CPU (first processor) can perform one or more non-neural network loop filter processes on subsequent pictures. Therefore, the loop filter unit 212 can suppress the delay that occurs until multiple loop filter processes are completed.

[0549] Furthermore, in the third and fourth configuration examples of the loop filter unit 212 according to the first embodiment, the decoding process for subsequent pictures can be performed by using the output data obtained by one or more non-neural network loop filter processes as a reference image, without waiting for the results of one or more neural network loop filter processes. As a result, the loop filter unit 212 can suppress the delay that occurs until multiple loop filter processes are completed.

[0550] Furthermore, in the third and fourth configuration examples of the loop filter unit 212 according to the first embodiment, output data obtained by one or more non-neural network loop filter processes can be rewritten with output data obtained by one or more neural network loop filter processes and stored in the DPB 400. As a result, the loop filter unit 212 can use reconstructed images with better image quality for intra-prediction or inter-prediction in the decoding process of subsequent pictures. Therefore, the loop filter unit 212 can improve the quality of the reconstructed image and improve the image compression ratio. In addition, the decoding device 200 can perform sequential processing and suppress the complexity of control.

[0551] [Fifth Configuration Example of Loop Filter Unit According to the First Embodiment] The first to fourth configuration examples of the loop filter unit 212 according to the first embodiment described above describe the case in which all non-neural network loop filter processing is performed serially, but it is not limited to this. For example, the loop filter unit 212 may be configured to perform some of the processing of one or more non-neural network loop filter processing (LMCS processing / DBF processing / SAO processing / ALF processing) in parallel, and the remaining non-neural network loop filter processing is performed serially (for example, Figure 50). In such a loop filter unit 212, the GPU performs one or more neural network loop filter processing (NNLF processing) after the CPU has performed all non-neural network loop filter processing. Figure 50 is a diagram showing a fifth configuration example of the loop filter unit 212 according to the first embodiment.

[0552] As shown in Figure 50, the loop filter unit 212 is configured to perform ALF processing and DBF processing in parallel. The input data used for DBF processing and ALF processing is the output data obtained from LMCS processing. The loop filter unit 212 also combines the output data obtained from ALF processing and the output data obtained from DBF processing. The data obtained from the combination is used as input data for SAO processing.

[0553] Subsequently, the output data obtained by the SAO processing is transferred from the CPU to the GPU via a data bus 300 corresponding to DMA technology and used as input data for one or more neural network loop filtering (NNLF) processes.

[0554] Furthermore, the output data obtained through NNLF processing is stored in the DPB400 and used as a reference image for subsequent pictures.

[0555] Furthermore, the combinations of non-neural network loop filter processes that are performed in parallel among the one or more non-neural network loop filter processes are not limited to the combinations described in Figure 50. For example, the loop filter unit 212 may be configured to perform LMCS processing and SAO processing in parallel.

[0556] Furthermore, although the fifth configuration example shown in Figure 50 describes an example in which only the output data of one or more neural network loop filter processes is stored in the DPB 400, data other than the output data of one or more neural network loop filter processes may also be stored in the DPB 400. For example, the loop filter unit 212 may be configured to store in the DPB 400 the output data obtained by the non-neural network loop filter process that is last applied to the current block among one or more non-neural network loop filter processes, and the output data obtained by the neural network loop filter process that is last applied to the current block among one or more neural network loop filter processes. In this way, the loop filter unit 212 can use the output data obtained by the non-neural network loop filter process and the output data obtained by the neural network loop filter process as candidate reference images for subsequent pictures.

[0557] Furthermore, in the fifth configuration example shown in Figure 50, the output data obtained by the non-neural network loop filter processing applied last to the current block may be used as the display image, or the output data obtained by the neural network loop filter processing applied last to the current block may be used as the display image. For example, the loop filter unit 212 may use the output data obtained by the non-neural network loop filter processing applied last to the current block (SAO processing in the example of Figure 50) from among one or more non-neural network loop filter processing as the display image. Alternatively, for example, the loop filter unit 212 may use the output data obtained by the neural network loop filter processing applied last to the current block (NNLF processing in the example of Figure 50) from among one or more neural network loop filter processing as the display image.

[0558] [Advantages of the fifth configuration example of the loop filter unit according to the first embodiment] In the fifth configuration example of the loop filter unit 212 according to the first embodiment, some of the non-neural network loop filter processing (LMCS processing / DBF processing / SAO processing / ALF processing) among one or more non-neural network loop filter processing are performed in parallel. As a result, the loop filter unit 212 can shorten the time required to decode all images.

[0559] In the above description, the configuration example of the loop filter unit 212 according to the first embodiment (i.e., the loop filter unit 212 of the decoding device 200) has been mainly explained, but the loop filter unit 120 of the encoding device 100 can also be configured in the same way as the loop filter unit 212 according to the first embodiment. By doing so, the loop filter unit 120 of the encoding device 100 has the same advantages as the loop filter unit 212 according to the first embodiment, such as being able to suppress the delay that occurs until multiple loop filtering processes are completed.

[0560] Furthermore, in the explanation using Figures 38 to 50 above, the example given for the application order of one or more non-neural network loop filter processes is LMCS processing, DBF processing, SAO processing, and ALF processing, but this is not limited to this. The application order of one or more non-neural network loop filter processes is not limited to this and may be changed as appropriate, or multiple non-neural network loop filter processes may be performed in parallel.

[0561] Furthermore, in the explanation using Figures 38 to 50 above, examples of one or more non-neural network loop filtering processes were given, such as LMCS processing, DBF processing, SAO processing, and ALF processing, but the explanation is not limited to these. For example, the loop filter unit 212 may use at least one of the LMCS processing, DBF processing, SAO processing, and ALF processing as one or more non-neural network loop filtering processes. Alternatively, the loop filter unit 212 may use one or more loop filtering processes that do not use neural networks other than LMCS processing, DBF processing, SAO processing, and ALF processing as one or more non-neural network loop filtering processes.

[0562] Furthermore, while the explanation using Figures 38 to 50 above illustrates the case where one NNLF process is used as one or more neural network loop filter processes, it is not limited to this. For example, the loop filter unit 212 may use multiple NNLF processes that have undergone the same training as one or more neural network loop filter processes, or it may use multiple NNLF processes that have undergone different training as one or more neural network loop filter processes. When using multiple neural network loop filter processes, the loop filter unit 212 may be configured to apply the multiple neural network loop filter processes sequentially. Also, when using multiple neural network loop filter processes, the loop filter unit 212 may be configured so that some of the multiple neural network loop filter processes are performed in parallel.

[0563] (Second Embodiment) The loop filter unit 212 according to the second embodiment differs from the loop filter unit 212 according to the first embodiment in that it performs one or more neural network loop filter processes in parallel with one or more non-neural network loop filter processes. Figure 51 is a diagram illustrating an example of the application order of loop filter processes performed by the decoding device 200 according to the second embodiment.

[0564] At time t=0, the decoding device 200 decodes the first sample group from the bitstream using other decoding processes such as the entropy decoding unit 202, the inverse quantization unit 204, and the inverse transform unit 206 (corresponding to step S2001 in Figure 40). For example, the decoding device 200 decodes the first sample group in picture units from the bitstream by repeatedly executing the other decoding processes in block units. The CPU (first processor) also performs other decoding processes.

[0565] Next, the first sample group, which is the output data obtained by the above process, is used as input data for one or more neural network loop filter processes (NNLF processes) (corresponding to step S2002 in Figure 40). The NNLF process is performed by the GPU (second processor). Therefore, the first sample group is transferred from the CPU to the GPU via a data bus 300 that supports technologies such as DMA, and is used as input data for the NNLF process. The NNLF process is performed by the GPU, which is a separate circuit from the CPU, so that it is performed in parallel with one or more non-neural network loop filter processes.

[0566] The GPU (second processor) performs one or more neural network loop filtering processes simultaneously, while the CPU (first processor) sequentially performs one or more non-neural network loop filtering processes (LMCS processing / DBF processing / SAO processing / ALF processing) on ​​the first sample group.

[0567] After one or more non-neural network loop filtering processes are completed, the decoding device 200 transfers the output data obtained from the non-neural network loop filtering process that is last applied to the first sample group (in the example in Figure 21, the output data obtained from the ALF process) from the CPU to the GPU via the data bus. The decoding device 200 then combines the output data obtained from the NNLF process and the output data obtained from the ALF process to generate a second sample group. The generated second sample group is then stored in the DPB 400. The second sample group is a picture containing the current block to be decoded. Therefore, the decoding device 200 performs NNLF processing on a picture-by-picture basis and outputs on a picture-by-picture basis.

[0568] Then, at time t=1, when decoding the third sample group from the bitstream, the decoding device 200 uses the second sample group stored in the DPB 400 at time t=0 as a reference image (corresponding to step S2003 in Figure 40). For example, the decoding device 200 decodes the third sample group in picture units from the bitstream by repeatedly executing other decoding processes in block units. The other decoding processes are executed sequentially by the CPU (first processor).

[0569] Figure 52 shows a simplified block diagram illustrating the application order of the loop filter processing shown in Figure 51. Figure 52 is a simplified block diagram illustrating the configuration of the loop filter unit 212 according to the second embodiment.

[0570] As shown in Figure 52, the loop filter unit 212 performs one or more non-neural network loop filter processes (specifically, LMCS process / DBF process / SAO process / ALF process) and one or more neural network loop filter processes (specifically, NNLF process) in parallel. At this time, the input data for the LMCS process, which is one of the one or more non-neural network loop filter processes, is the same as the input data for the NNLF process, which is one of the one or more neural network loop filter processes. The loop filter unit 212 also combines the output data obtained by the ALF process, which is the last process applied to the first sample group among the one or more non-neural network loop filter processes, with the output data obtained by the NNLF process, which is the last process applied to the first sample group among the one or more neural network loop filter processes, to generate a second sample group, and stores the generated second sample group in the DPB 400.

[0571] The following provides an example of a specific configuration of the loop filter unit 212 according to the second embodiment proposed in this disclosure.

[0572] [First Configuration Example of Loop Filter Unit According to the Second Embodiment] Figure 53 is a diagram showing a first configuration example of the loop filter unit 212 according to the second embodiment. Figure 53 is a conceptual diagram showing a first data transfer example performed between the CPU and the GPU.

[0573] The decoding device 200 transfers the output data obtained from other decoding processes (such as the inverse quantization unit 204 and the inverse transform unit 206) from the CPU to the data bus 300 corresponding to DMA technology (S301). Subsequently, the output data is transferred from the data bus 300 to the GPU (S302) and used as input data for NNLF processing.

[0574] At the same time that data obtained from other decoding processes is being transferred from the CPU to the GPU, the loop filter unit 212 acquires the output data obtained from the other decoding processes as input data for one or more non-neural network loop filter processes (LMCS process / DBF process / SAO process / ALF process). Then, one or more non-neural network loop filter processes are performed by the CPU (S303).

[0575] The output data obtained by the ALF processing is transferred from the CPU to the data bus 300 (S304). Subsequently, the output data obtained by the ALF processing is further transferred from the data bus 300 to the arithmetic unit (S305). In the example shown in Figure 53, the functions of the arithmetic unit can be implemented by a GPU.

[0576] Subsequently, the loop filter unit 212 inputs the output data obtained by the NNLF processing to the calculation unit (S306), and the output data obtained by the ALF processing and the output data obtained by the NNLF processing are combined.

[0577] The output data obtained by the arithmetic unit is transferred to the data bus 300 (S307). Subsequently, the output data obtained by the arithmetic unit is transferred from the data bus 300 to the DPB 400 (S308) and used as a reference image for the subsequent picture.

[0578] Furthermore, the output data obtained by the ALF processing is also transmitted to the display 500 for use as a display image (S309).

[0579] [Second Configuration Example of Loop Filter Unit According to the Second Embodiment] Figure 54 is a diagram showing a second configuration example of the loop filter unit 212 according to the second embodiment. Figure 54 is a conceptual diagram showing a second data transfer example performed between the CPU and the GPU.

[0580] The decoding device 200 transfers the output data obtained from other decoding processes (such as the inverse quantization unit 204 and the inverse transform unit 206) from the CPU to the data bus 300 corresponding to DMA technology (S311). Subsequently, the output data is transferred from the data bus 300 to the GPU (S312) and used as input data for NNLF processing.

[0581] At the same time that data obtained from other decoding processes is being transferred from the CPU to the GPU, the loop filter unit 212 acquires the output data obtained from the other decoding processes as input data for one or more non-neural network loop filter processes (LMCS process / DBF process / SAO process / ALF process). Then, one or more non-neural network loop filter processes are performed by the CPU (S313).

[0582] The output data obtained by the ALF processing is transferred from the CPU to the data bus 300 (S314). Subsequently, the output data obtained by the ALF processing is further transferred from the data bus 300 to the arithmetic unit (S315). In the example of Figure 54, the functions of the arithmetic unit can be implemented by a GPU.

[0583] Subsequently, the loop filter unit 212 inputs the output data obtained by the NNLF processing to the calculation unit (S316), and the output data obtained by the ALF processing and the output data obtained by the NNLF processing are combined.

[0584] The output data obtained by the arithmetic unit is transferred to the data bus 300 (S317). Subsequently, the output data obtained by the arithmetic unit is transferred from the data bus 300 to the DPB 400 (S318) and used as a reference image for the subsequent picture.

[0585] Furthermore, the output data obtained by the calculation unit is also transmitted to the display 500 for use as a display image (S319).

[0586] In the examples of Figures 53 and 54, the arithmetic unit's functions are described as being implemented by a GPU, but this is not the only example. The arithmetic unit's functions may be implemented by, for example, a DMA. Alternatively, the arithmetic unit's functions may be implemented by a CPU. When the arithmetic unit's functions are implemented by a CPU, the output data obtained by the NNLF processing must be in picture units. This is because, when the output data obtained by the NNLF processing is in picture units, the number of times data is transferred from the GPU to the CPU can be kept to a minimum compared to the configurations in the reference examples of Figures 36 and 37, allowing multiple loop filter processes to be completed without causing significant delays.

[0587] [Advantages of the first and second configuration examples of the loop filter section according to the second embodiment] As described above, by adding one or more neural network loop filter processes (NNLF processes) to the loop filter section 212, the decoding device 200 can improve the quality of the reconstructed image and improve the image compression ratio.

[0588] Furthermore, in the first and second configuration examples of the loop filter unit 212 according to the second embodiment, one or more neural network loop filter processes are added in parallel with one or more non-neural network loop filter processes. The loop filter unit 212 also combines the output data obtained by the neural network loop filter process that is applied last to the first sample group among the one or more neural network loop filter processes with the output data obtained by the non-neural network loop filter process that is applied last to the first sample group among the one or more non-neural network loop filter processes. Therefore, the GPU (second processor) can perform one or more neural network loop filter processes in parallel with all the neural network loop filter processes performed by the CPU (first processor). On the other hand, in the configuration according to the reference example in Figures 36 and 37, after performing some non-neural network loop filter processes, another part of non-neural network loop filter processes and neural network loop filter processes are performed in parallel, and then the remaining non-neural network loop filter processes are performed. Therefore, in the example configuration, non-neural network loop filtering, which is performed after neural network loop filtering, cannot be performed until the neural network loop filtering is complete.

[0589] Thus, the first and second configuration examples of the loop filter unit 212 according to the second embodiment can perform multiple loop filter processing in the same amount of time as when performing only one or more non-neural network loop filter processing. Furthermore, the first and second configuration examples of the loop filter unit 212 according to the second embodiment can improve the overall quality of the resulting reconstructed image.

[0590] Furthermore, in the first configuration example of the loop filter unit 212 according to the second embodiment (Figure 53), the output data obtained by ALF processing is displayed on the display 500. Therefore, in the first configuration example, the reconstructed image that has undergone one or more non-neural network loop filter processing performed by the CPU can be displayed immediately without waiting for the results of one or more neural network loop filter processing performed by the GPU. In addition, the first configuration example can suppress the delay in displaying the reconstructed image even when image processing by one or more neural network loop filter processing takes a long time.

[0591] Furthermore, in the second configuration example of the loop filter unit 212 according to the second embodiment (Figure 54), the output data obtained by the coupling by the calculation unit is displayed on the display 500. Therefore, in the second configuration example, when the time required for one or more neural network loop filter processes performed by the GPU is relatively short, the output data obtained by one or more neural network loop filter processes can be displayed on the display 500 at almost the same timing as when the output data obtained by one or more non-neural network loop filter processes is displayed on the display 500, and a reconstructed image with better image quality can be displayed on the display 500. The case where the time required for one or more neural network loop filter processes is relatively short includes, for example, when one or more neural network loop filter processes are low-complexity neural network loop filter processes, or when a high-speed GPU is used for one or more neural network loop filter processes.

[0592] Furthermore, while Figures 53 and 54 illustrate an example in which only output data obtained from one or more neural network loop filtering processes is stored in the DPB 400, data other than the output data obtained from one or more neural network loop filtering processes may also be stored in the DPB 400. In another configuration example of the loop filter unit 212 according to the second embodiment, in addition to the output data from one or more neural network loop filtering processes, the output data from one or more non-neural network loop filtering processes is stored in the DPB 400. Figure 55 is another block diagram that simplifies the configuration of the loop filter unit 212 according to the second embodiment.

[0593] The loop filter unit 212 shown in Figure 55 differs from the loop filter unit 212 shown in Figure 52 in that, in addition to the output data of one or more neural network loop filter processes, it stores the output data of one or more non-neural network loop filter processes in the DPB 400. Specifically, the loop filter unit 212 shown in Figure 55 stores the output data obtained by ALF processing and the output data obtained by NNLF processing in the DPB 400. As a result, the loop filter unit 212 shown in Figure 55 can use the output data obtained by ALF processing and the output data obtained by NNLF processing as candidate reference images for subsequent pictures. Below, a specific configuration example of the loop filter unit 212 shown in Figure 55 will be illustrated.

[0594] [Third Configuration Example of Loop Filter Unit According to the Second Embodiment] Figure 56 is a diagram showing a third configuration example of the loop filter unit 212 according to the second embodiment. Figure 56 is a conceptual diagram showing a third data transfer example performed between the CPU and the GPU.

[0595] The decoding device 200 transfers the output data obtained from other decoding processes (such as the inverse quantization unit 204 and the inverse transform unit 206) from the CPU to the data bus 300 corresponding to DMA technology (S401). Subsequently, the output data is transferred from the data bus 300 to the GPU (S402) and used as input data for NNLF processing.

[0596] At the same time that data obtained from other decoding processes is being transferred from the CPU to the GPU, the loop filter unit 212 acquires the output data obtained from the other decoding processes as input data for one or more non-neural network loop filter processes (LMCS process / DBF process / SAO process / ALF process). Then, one or more non-neural network loop filter processes are performed sequentially by the CPU (S403).

[0597] The output data obtained by the ALF processing is transferred to the data bus 300 (S404). Subsequently, the output data obtained by the ALF processing is further transferred from the data bus 300 to the DPB 400 (S405). The output data transferred to the DPB 400 is then used as a reference image for the subsequent picture. During this data transfer, the CPU can perform other processing, such as decoding the subsequent picture.

[0598] Furthermore, the output data transferred to the DPB400 (output data obtained by ALF processing) is transferred from the DPB400 to the data bus 300 (S406), and then the output data obtained by ALF processing is further transferred from the data bus 300 to the arithmetic unit (S407). In the example of Figure 56, the functions of the arithmetic unit can be realized by a GPU.

[0599] Subsequently, the loop filter unit 212 inputs the output data obtained by the NNLF processing to the calculation unit (S408), and the output data obtained by the ALF processing and the output data obtained by the NNLF processing are combined.

[0600] The output data obtained by the arithmetic unit is transferred to the data bus 300 (S409). Subsequently, the output data obtained by the arithmetic unit is transferred from the data bus 300 to the DPB 400 (S410) and used as a reference image for the subsequent picture.

[0601] Furthermore, the output data obtained by the calculation unit is also transmitted to the display 500 for use as a display image (S411).

[0602] [Fourth Configuration Example of Loop Filter Unit According to the Second Embodiment] Figure 57 is a diagram showing a fourth configuration example of the loop filter unit 212 according to the second embodiment. Figure 57 is a conceptual diagram showing a fourth data transfer example performed between the CPU and the GPU.

[0603] The decoding device 200 transfers the output data obtained from other decoding processes (such as the inverse quantization unit 204 and the inverse transformation unit 206) from the CPU to the data bus 300 corresponding to DMA technology (S421). Subsequently, the output data is transferred from the data bus 300 to the GPU (S422) and used as input data for NNLF processing.

[0604] At the same time that data obtained from other decoding processes is being transferred from the CPU to the GPU, the loop filter unit 212 acquires the output obtained from the other decoding processes as input data for one or more non-neural network loop filter processes (LMCS process / DBF process / SAO process / ALF process). Then, one or more neural network loop filter processes are performed sequentially by the CPU (S423).

[0605] The output data obtained by the ALF processing is transferred from the CPU to the data bus 300 (S424). Subsequently, the output data obtained by the ALF processing is further transferred from the data bus 300 to the DPB 400 (S425). The output data transferred to the DPB 400 is then used as a reference image for the subsequent picture. During this data transfer, the CPU can perform other processing, such as decoding the subsequent picture.

[0606] Furthermore, the output data transferred to the DPB400 (output data obtained by ALF processing) is transferred from the DPB400 to the data bus 300 (S426), and then the output data obtained by ALF processing is further transferred from the data bus 300 to the arithmetic unit (S427). In the example of Figure 57, the functions of the arithmetic unit can be realized by a GPU.

[0607] Subsequently, the loop filter unit 212 inputs the output data obtained by the NNLF processing to the calculation unit (S428), and the output data obtained by the ALF processing and the output data obtained by the NNLF processing are combined.

[0608] The output data obtained by the arithmetic unit is transferred to the data bus 300 (S429). Subsequently, the output data obtained by the arithmetic unit is transferred from the data bus 300 to the DPB 400 (S430) and used as a reference image for the subsequent picture.

[0609] Furthermore, the output data obtained by the ALF processing is also transmitted to the display 500 for use as a display image (S431).

[0610] In the examples of Figures 56 and 57, the arithmetic unit's functions are described in an example where they are implemented by a GPU, but this is not the only example. The arithmetic unit's functions may be implemented by, for example, a DMA. Alternatively, the arithmetic unit's functions may be implemented by a CPU. When the arithmetic unit's functions are implemented by a CPU, the output data obtained by the NNLF processing must be in picture units. This is because, when the output data obtained by the NNLF processing is in picture units, the number of times data is transferred from the GPU to the CPU can be kept to a minimum compared to the configurations in the reference examples of Figures 36 and 37, allowing multiple loop filter processes to be completed without causing significant delays.

[0611] [Advantages of the Third and Fourth Configuration Examples of the Loop Filter Unit According to the Second Embodiment] In the third and fourth configuration examples of the loop filter unit 212 according to the second embodiment, one or more neural network loop filter processes are added in parallel with one or more non-neural network loop filter processes. For example, while the GPU (second processor) is decoding the picture including the current block, the CPU (first processor) can decode subsequent pictures. Therefore, the loop filter unit 212 is a suitable configuration when the time required for one or more neural network loop filter processes is longer than the time required for one or more non-neural network loop filter processes.

[0612] Furthermore, in the third and fourth configuration examples of the loop filter unit 212 according to the second embodiment, the decoding process for subsequent pictures can be performed by using the output data obtained by one or more non-neural network loop filter processes as a reference image, without waiting for the results of one or more neural network loop filter processes. As a result, the loop filter unit 212 can suppress the delay that occurs until multiple loop filter processes are completed.

[0613] Furthermore, in the third and fourth configuration examples of the loop filter unit 212 according to the second embodiment, output data obtained by one or more non-neural network loop filter processes can be rewritten with output data obtained by one or more neural network loop filter processes and stored in the DPB 400. As a result, the loop filter unit 212 can use reconstructed images with superior image quality through intra-prediction or inter-prediction in the subsequent picture decoding process. Therefore, the loop filter unit 212 can contribute to improving image quality.

[0614] The loop filter unit 212, as described using Figures 52 and 55, performs one or more neural network loop filter processes (NNLF processes) and one or more non-neural network loop filter processes (LMCS processes / DBF processes / SAO processes / ALF processes) in parallel. The reason for proposing the loop filter unit 212 shown in Figures 52 and 55 in this disclosure is the time required for NNLF processing. For example, if the time required for NNLF processing is long, increasing the number of non-neural network loop filter processes performed in parallel with NNLF processing can reduce the impact of image display delays caused by the long time required for NNLF processing.

[0615] However, if the time required for NNLF processing is relatively short, it is not a problem to reduce the number of non-neural network loop filter processes performed in parallel with NNLF processing. Below, examples of the configuration of the loop filter unit 212 in which the number of one or more neural network loop filter processes and one or more non-neural network loop filter processes performed in parallel are changed will be explained with reference to Figures 58 to 61.

[0616] Figure 58 shows a fifth configuration example of the loop filter unit 212 according to the second embodiment. As shown in Figure 58, the loop filter unit 212 performs three non-neural network loop filter processes, DBF processing, SAO processing, and ALF processing, and NNLF processing in parallel. At this time, the input data for the DBF processing, which is one or more non-neural network loop filter processes, is the same as the input data for the NNLF processing. Specifically, the output data obtained by LMCS processing is used as the input data for the DBF processing and the NNLF processing, respectively. Note that the other processing performed by the loop filter unit 212 is substantially the same as the configuration examples shown in Figures 52 and 55, so its explanation is omitted.

[0617] Figure 59 shows a sixth configuration example of the loop filter unit 212 according to the second embodiment. As shown in Figure 59, the loop filter unit 212 performs two non-neural network loop filter processes, SAO processing and ALF processing, and NNLF processing in parallel. At this time, the input data for the SAO processing, which is one of the one or more non-neural network loop filter processes, is the same as the input data for the NNLF processing. Specifically, the output data obtained by DBF processing is used as the input data for the SAO processing and NNLF processing, respectively. Note that the other processing performed by the loop filter unit 212 is substantially the same as the configuration examples shown in Figures 52 and 55, so a description is omitted.

[0618] Figure 60 shows a seventh configuration example of the loop filter unit 212 according to the second embodiment. As shown in Figure 60, the loop filter unit 212 performs one non-neural network loop filter process of ALF processing and an NNLF process in parallel. At this time, the input data for the ALF process, which is one of the one or more non-neural network loop filter processes, is the same as the input data for the NNLF process. Specifically, the output data obtained by the SAO process is used as the input data for the ALF process and the NNLF process, respectively. Note that the other processes performed by the loop filter unit 212 are substantially the same as the configuration examples shown in Figures 52 and 55, so their explanation is omitted.

[0619] Figure 61 shows an eighth configuration example of the loop filter unit 212 according to the second embodiment. As shown in Figure 61, the loop filter unit 212 performs the online processing within the ALF processing and the NNLF processing in parallel. At this time, the input data for the online processing within the ALF processing, which is one of the one or more non-neural network loop filter processing, is the same as the input data for the NNLF processing. Specifically, the output data obtained by the processing using a fixed filter within the ALF processing is used as the input data for the online processing within the ALF processing and the NNLF processing, respectively. The filters incorporated into the ALF processing include fixed filters, Gaussian filters, and residual filters. Furthermore, the other processing performed by the loop filter unit 212 is almost the same as the configuration examples shown in Figures 52 and 55, so its explanation is omitted.

[0620] Furthermore, in another variation of the configuration example of the loop filter unit 212 according to the second embodiment, the multiple configuration examples shown in Figures 52, 55, and 58 to 61 may be referenced. For example, the output data of multiple non-neural network loop filter processes may be combined to be used as input data for one or more neural networks.

[0621] [Granularity of sample groups subject to multiple loop filtering processes] The granularity of sample groups subject to multiple loop filtering processes performed by the loop filtering unit 212 according to the second embodiment described above (in other words, data in a specific region) is not necessarily limited to picture units. That is, the first sample group, the second sample group, and the third sample group may be data at CPU, block, slice, tile, picture, subpicture, or other granularity levels. For example, the loop filtering unit 212 according to the second embodiment may apply one or more neural network loop filtering processes to the first sample group (corresponding to data in a specific region) at the CTU level.

[0622] Here, by comparing with the configuration example shown in Figure 58, we will explain the advantages of applying one or more neural network loop filtering processes to the first sample group in units of CTUs.

[0623] In the configuration example shown in Figure 58, when the loop filter unit 212 performs one or more non-neural network loop filter processes, it transfers only the output data obtained from the non-neural network loop filter process applied last to the current block from the CPU to memory such as RMA. On the other hand, in the configuration example shown in Figure 58, the loop filter unit 212 transfers the output data obtained from the LMCS process to the GPU or the like via the data bus 300, and then sequentially performs one or more neural network loop filter processes. In such a case, the loop filter unit 212 needs to transfer the output data obtained from the LMCS process in order to perform one or more neural network loop filter processes, which causes a delay due to data transfer. Furthermore, when the loop filter unit 212 applies one or more neural network loop filter processes to a first sample group on a picture-by-picture basis, it is necessary to combine the output data obtained from one or more non-neural network loop filter processes with the output data obtained from one or more neural network loop filter processes on a picture-by-picture basis. Therefore, the CPU (first processor) has to wait until the GPU (second processor) has completed one or more neural network loop filtering processes, resulting in a delay of up to one picture.

[0624] On the other hand, if the loop filter unit 212 applies one or more neural network loop filter processes to a first sample group in units of CTUs, the output data in units of CTUs can be output to the arithmetic unit, so the GPU can reduce the CPU waiting time to a maximum delay of 1 CTU. Thus, if the loop filter unit 212 applies one or more neural network loop filter processes to a first sample group in units of CTUs, there is the advantage that the CPU waiting time can be further reduced. In other words, the loop filter unit 212 can perform one or more neural network loop filter processes in appropriate data units corresponding to data in a specific region.

[0625] Note that the first, second, and third sample groups may each have the same particle size level. Also, the first and third sample groups may have different particle size levels than the second sample group.

[0626] (Modification of the second embodiment) The loop filter unit 212 according to the modification of the second embodiment is the same as the loop filter unit 212 according to the second embodiment in that it performs one or more neural network loop filter processes in parallel with one or more non-neural network loop filter processes. However, the loop filter unit 212 according to the modification of the second embodiment differs from the loop filter unit 212 according to the second embodiment in that it performs some of the non-neural network loop filter processes, which are performed in parallel with one or more neural network loop filter processes, on a GPU.

[0627] Figure 62 is a simplified block diagram showing the configuration of the loop filter unit 212 according to a modified example of the second embodiment. As shown in Figure 62, the loop filter unit 212 performs the online processing within the ALF processing and the NNLF processing in parallel, similar to the loop filter unit 212 shown in Figure 61.

[0628] In the loop filter unit 212 according to the modified second embodiment shown in Figure 62, the GPU performs NNLF processing as well as processing using Online within ALF processing. Such a loop filter unit 212 transfers the output data obtained by the fixed filter incorporated into the ALF processing from the CPU to the data bus 300, and then transfers it from the data bus 300 to the GPU only once. Therefore, the loop filter unit 212 shown in Figure 62 can reduce the overhead of multiple data transfers between the CPU and the GPU.

[0629] Furthermore, the data transferred from the data bus 300 to the GPU is duplicated and used as input data for both the Online and NNLF processes within the ALF process. In addition, the output data obtained from the Online process within the ALF process and the output data obtained from the NNLF process are combined within the GPU.

[0630] The following provides an example of a specific configuration of the loop filter section 212 shown in Figure 62.

[0631] [First Configuration Example of Loop Filter Unit According to Modification of the Second Embodiment] Figure 63 is a diagram showing a first configuration example of the loop filter unit 212 according to modification of the second embodiment. Figure 63 is a diagram conceptually showing a first data transfer example performed between the CPU and the GPU.

[0632] Other decoding processes and one or more non-neural network loop filter processes are applied to the bitstream (S501). The output data obtained by processing with fixed filters in the ALF process is transferred from the CPU to the data bus 300 corresponding to DMA technology (S502). Subsequently, the output data obtained by processing with fixed filters in the ALF process is transferred from the data bus 300 to the GPU (S503 and S504) and used as input data for both the Online and NNLF processes in the ALF process. The loop filter unit 212 inputs the output data obtained by processing with Online in the ALF process to the arithmetic unit (S505), and inputs the output data obtained by NNLF processing to the arithmetic unit (S506). The output data obtained by processing with Online in the ALF process and the output data obtained by NNLF processing are combined.

[0633] The output data obtained by the arithmetic unit is transferred from the GPU to the data bus 300 (S507). Subsequently, the output data obtained by the arithmetic unit is transferred from the data bus 300 to the DPB 400 (S508) and used as a reference image for the subsequent picture.

[0634] Furthermore, the output data obtained by the calculation unit is also transmitted to the display 500 for use as a display image (S509).

[0635] [Second Configuration Example of Loop Filter Unit According to Modification of the Second Embodiment] Figure 64 is a diagram showing a second configuration example of the loop filter unit 212 according to modification of the second embodiment. Figure 64 is a diagram conceptually showing a second data transfer example performed between the CPU and the GPU.

[0636] The second configuration example of the loop filter unit 212 shown in Figure 64 differs from the first configuration example of the loop filter unit 212 shown in Figure 63 in that the output data obtained by processing using Online within the ALF processing is transmitted to the display 500 for use as a display image (S519). Note that the processing performed in S511 to S518 shown in Figure 64 is the same as the processing performed in S501 to S508 shown in Figure 63, so the explanation is omitted.

[0637] [Advantages of the First and Second Configuration Examples of the Loop Filter Unit According to the Modified Example of the Second Embodiment] In the first and second configuration examples of the loop filter unit 212 according to the modified example of the second embodiment, non-neural network loop filtering and one or more neural network loop filtering processes can be performed in parallel by the second processor. Therefore, in the first and second configuration examples, the delay that occurs until multiple loop filtering processes are completed can be further reduced.

[0638] In the first configuration example of the loop filter unit 212 according to a modification of the second embodiment, the output data obtained by processing using Online in the ALF processing and the output data obtained by NNLF processing are combined and displayed on the display 500. Therefore, in this first configuration example, a reconstructed image with better image quality can be displayed on the display 500, which has the advantage of enhancing the user experience.

[0639] Furthermore, in the second configuration example of the loop filter unit 212 according to a modification of the second embodiment, if the time required for processing using Online within the ALF processing is shorter than the time required for NNLF processing, the output data obtained by processing using Online within the ALF processing is displayed on the display 500 in advance so that the user can confirm the output. Therefore, in this second configuration example, the reconstructed image can be displayed even while NNLF processing is being performed on the output data transferred from the data bus 300, which has the advantage of reducing delays in image display and minimizing impact on the user experience.

[0640] Furthermore, although Figures 62 to 64 illustrate an example in which only the output data obtained by the loop filtering process performed by the GPU is stored in the DPB 400, data other than the output data obtained by the loop filtering process performed by the GPU may also be stored in the DPB 400. Figure 65 is another block diagram that simplifies the configuration of the loop filtering unit 212 according to a modified example of the second embodiment.

[0641] The loop filter unit 212 shown in Figure 65 differs from the loop filter unit 212 shown in Figure 62 in that, in addition to the output data obtained by the loop filter processing performed by the GPU, it also stores the output data obtained by the loop filter processing performed by the CPU in the DPB 400. Specifically, the loop filter unit 212 shown in Figure 65 stores in the DPB 400 the combined data obtained by the processing using Online within the ALF processing and the output data obtained by the NNLF processing, as well as the output data obtained by the processing using a fixed filter within the ALF processing. As a result, the loop filter unit 212 shown in Figure 65 can use the combined data and the output data obtained by the processing using a fixed filter within the ALF processing as candidate reference images for subsequent pictures. Below, variations of the specific configuration example of the loop filter unit 212 shown in Figure 65 will be illustrated.

[0642] [Third Configuration Example of Loop Filter Unit According to Modification of the Second Embodiment] Figure 66 is a diagram showing a third configuration example of the loop filter unit 212 according to modification of the second embodiment. Figure 66 is a diagram conceptually showing a third data transfer example performed between the CPU and the GPU.

[0643] Other decoding processes and one or more non-neural network loop filter processes are applied to the bitstream (S601). The output data obtained by processing with fixed filters in the ALF process is then transferred from the CPU to the data bus 300 corresponding to DMA technology (S602). Subsequently, the output data obtained by processing with fixed filters in the ALF process is further transferred from the data bus 300 to the DPB 400 (S603). The output data transferred to the DPB 400 is then used as a candidate reference image for the subsequent picture. During this data transfer, the CPU can perform other processing, such as decoding the subsequent picture.

[0644] Furthermore, the output data transferred to the DPB400 (output data obtained by processing using a fixed filter within the ALF processing) is transferred from the DPB400 to the GPU via the data bus 300 for use as input data for processing using Online and NNLF processing within the ALF processing (S604-S606).

[0645] Subsequently, the loop filter unit 212 inputs the output data obtained by the Online processing within the ALF processing and the output data obtained by the NNLF processing to the calculation unit (S607-S608). The output data obtained by the Online processing within the ALF processing and the output data obtained by the NNLF processing are then combined by the calculation unit, and the output data obtained by the combination by the calculation unit is transferred from the GPU to the DPB 400 via the data bus 300 (S609-S610). The output data transferred to the DPB 400 is then used as a candidate reference image for the subsequent picture.

[0646] Furthermore, the output data obtained by the calculation unit is also transmitted to the display 500 for use as a display image (S611).

[0647] [Fourth Configuration Example of Loop Filter Unit According to Modification of the Second Embodiment] Figure 67 is a diagram showing a fourth configuration example of the loop filter unit 212 according to modification of the second embodiment. Figure 67 is a diagram conceptually showing a fourth data transfer example performed between the CPU and the GPU.

[0648] The fourth configuration example of the loop filter unit 212 shown in Figure 67 differs from the first configuration example of the loop filter unit 212 shown in Figure 66 in that the output data obtained by processing using Online within the ALF processing is transmitted to the display 500 for use as a display image (S631). Note that the processing performed in S621 to S630 shown in Figure 67 is the same as the processing performed in S601 to S610 shown in Figure 66, so its explanation is omitted.

[0649] [Advantages of the third and fourth configuration examples of the loop filter unit according to the modification of the second embodiment] In the third and fourth configuration examples of the loop filter unit 212 according to the modification of the second embodiment, non-neural network loop filtering and one or more neural network loop filtering processes performed by the second processor can be performed in parallel. Therefore, in the third and fourth configuration examples, the delay that occurs until multiple loop filtering processes are completed can be further reduced.

[0650] Furthermore, in the third and fourth configuration examples of the loop filter unit 212 according to the modification of the second embodiment, the subsequent picture decoding process can be performed by using the output data obtained by the loop filter processing performed by the CPU as a reference image, without waiting for the result of the loop filter processing performed by the GPU. As a result, the loop filter unit 212 can suppress the delay that occurs until multiple loop filter processing is completed.

[0651] Furthermore, in the third and fourth configuration examples of the loop filter unit 212 according to the modification of the second embodiment, the output data obtained by the loop filter processing performed by the CPU can be rewritten with the output data obtained by the loop filter processing performed by the GPU and stored in the DPB 400. As a result, the loop filter unit 212 can use a reconstructed image with superior image quality through intra-prediction or inter-prediction in the subsequent picture decoding process. Therefore, the loop filter unit 212 can contribute to improving image quality.

[0652] Furthermore, in the third configuration example of the loop filter unit 212 according to a modification of the second embodiment, the output data obtained by processing using Online in the ALF processing and the output data obtained by NNLF processing are combined and displayed on the display 500. Therefore, this third configuration example has the advantage that a reconstructed image with better image quality can be displayed on the display 500, thereby enhancing the user experience.

[0653] Furthermore, in the fourth configuration example of the loop filter unit 212 according to a modification of the second embodiment, if the time required for processing using Online within the ALF processing is shorter than the time required for NNLF processing, the output data obtained by processing using Online within the ALF processing is displayed on the display 500 in advance so that the user can confirm the output. Therefore, in this fourth configuration example, the reconstructed image can be displayed even while NNLF processing is being performed on the output data transferred from the data bus 300, which has the advantage of reducing delays in image display and minimizing impact on the user experience.

[0654] In the above description, the configuration example of the loop filter unit 212 according to the second embodiment (i.e., the loop filter unit 212 of the decoding device 200) has been mainly explained, but the loop filter unit 120 of the encoding device 100 can also be configured in the same way as the loop filter unit 212 according to the second embodiment. By doing so, the loop filter unit 120 of the encoding device 100 has the same advantages as the loop filter unit 212 according to the second embodiment, such as being able to suppress the delay that occurs until multiple loop filtering processes are completed.

[0655] Furthermore, in the explanation using Figures 51 to 67 above, the example given for the application order of one or more non-neural network loop filter processes is LMCS processing, DBF processing, SAO processing, and ALF processing, but this is not limited to this. The application order of one or more non-neural network loop filter processes is not limited to this and may be changed as appropriate, or multiple non-neural network loop filter processes may be performed in parallel.

[0656] Furthermore, in the explanation using Figures 51 to 67 above, examples of one or more non-neural network loop filtering processes were given, such as LMCS processing, DBF processing, SAO processing, and ALF processing, but the explanation is not limited to these. For example, the loop filter unit 212 may use at least one of the LMCS processing, DBF processing, SAO processing, and ALF processing as one or more non-neural network loop filtering processes. Alternatively, the loop filter unit 212 may use one or more loop filtering processes that do not use neural networks other than LMCS processing, DBF processing, SAO processing, and ALF processing as one or more non-neural network loop filtering processes.

[0657] Furthermore, while the explanation using Figures 51 to 67 above illustrates the case where one NNLF process is used as one or more neural network loop filter processes, it is not limited to this. For example, the loop filter unit 212 may use multiple NNLF processes that have undergone the same training as one or more neural network loop filter processes, or it may use multiple NNLF processes that have undergone different training as one or more neural network loop filter processes. When multiple neural network loop filter processes are used, the loop filter unit 212 may be configured to apply the multiple neural network loop filter processes sequentially. Also, when multiple neural network loop filter processes are used, the loop filter unit 212 may be configured so that some of the multiple neural network loop filter processes are performed in parallel.

[0658] (Third Embodiment) The loop filter unit 212 according to the third embodiment differs from the loop filter unit 212 according to the second embodiment in that it performs one or more non-neural network loop filter processes using output data obtained from one or more neural network loop filter processes. For example, the loop filter unit 212 according to the third embodiment can be configured to use the output data obtained from one or more neural network loop filter processes as input data for online processing within the ALF process. Figure 68 is a first block diagram that shows a simplified configuration of the loop filter unit 212 according to the third embodiment.

[0659] As shown in Figure 68, the loop filter unit 212 performs LMCS processing, DBF processing, SAO processing, and processing using fixed filters within ALF processing in parallel with NNLF processing. At this time, the input data for LMCS processing is the same as the input data for NNLF processing. Specifically, the input data for LMCS processing and NNLF processing uses output data obtained from other decoding processes. Furthermore, the loop filter unit 212 combines the output data obtained from processing using fixed filters within ALF processing with the output data obtained from NNLF processing, and uses the combined data as input data for processing using Online within ALF processing.

[0660] In this configuration of the loop filter unit 212, the output data obtained by NNLF processing is combined with the output data obtained by processing using a fixed filter within ALF processing. This allows residual artifacts to be reduced by using the output data obtained by NNLF processing before applying learning-based optimization in the Online processing within ALF processing. By suppressing error propagation to the loop filter processing performed after NNLF processing, the Online processing within ALF processing can focus on scene-specific improvements. Therefore, the configuration of the loop filter unit 212 shown in Figure 68 can achieve both improved reconstructed image quality and improved image compression ratio.

[0661] In the configuration of the loop filter unit 212 shown in Figure 68, the GPU performs processing using Online within the ALF processing. In this case, the output data obtained by processing using the fixed filter within the ALF processing is transferred from the CPU to the GPU via the data bus 300 and combined with the output data obtained by the NNLF processing. The combined output data is then input to the processing using Online within the ALF processing.

[0662] Furthermore, the CPU may perform processing using Online within the ALF processing. In such a case, the output data obtained by NNLF processing is transferred from the GPU to the CPU via the data bus 300 and combined with the output data obtained by processing using a fixed filter within the ALF processing. The combined output data is then input to the processing using Online within the ALF processing.

[0663] Furthermore, the following will illustrate and explain another example of the configuration of the loop filter unit 212 according to the third embodiment.

[0664] Figure 69 is a second block diagram that shows a simplified configuration of the loop filter section 212 according to the third embodiment.

[0665] As shown in Figure 69, the loop filter unit 212 performs DBF processing, SAO processing, processing using fixed filters within ALF processing, and NNLF processing in parallel. At this time, the input data for DBF processing is the same as the input data for NNLF processing. Specifically, the output data obtained by LMCS processing is used as the input data for DBF processing and NNLF processing. Furthermore, the loop filter unit 212 combines the output data obtained by processing using fixed filters within ALF processing with the output data obtained by NNLF processing, and uses the combined data as the input data for processing using Online within ALF processing.

[0666] Figure 70 is a third block diagram that shows a simplified configuration of the loop filter section 212 according to the third embodiment.

[0667] As shown in Figure 70, the loop filter unit 212 performs SAO processing, processing using a fixed filter within ALF processing, and NNLF processing in parallel. At this time, the input data for SAO processing is the same as the input data for NNLF processing. Specifically, the output data obtained by DBF processing is used as the input data for SAO processing and NNLF processing. Furthermore, the loop filter unit 212 combines the output data obtained by processing using a fixed filter within ALF processing and the output data obtained by NNLF processing, and uses the combined data as the input data for processing using Online within ALF processing.

[0668] Figure 71 is a fourth block diagram that shows a simplified configuration of the loop filter section 212 according to the third embodiment.

[0669] As shown in Figure 71, the loop filter unit 212 performs processing using a fixed filter within the ALF process and NNLF processing in parallel. At this time, the input data for processing using a fixed filter within the ALF process is the same as the input data for NNLF processing. Specifically, the output data obtained by SAO processing is used as the input data for processing using a fixed filter within the ALF process and for NNLF processing. Furthermore, the loop filter unit 212 combines the output data obtained by processing using a fixed filter within the ALF process and the output data obtained by NNLF processing, and uses the combined data as the input data for processing using Online within the ALF process.

[0670] Furthermore, the loop filter unit 212 according to the third embodiment may be configured to use, for example, output data obtained by one or more neural network loop filter processes as output data obtained by ALF processing. Figure 72 is a fifth block diagram that simplifies the configuration of the loop filter unit 212 according to the third embodiment.

[0671] As shown in Figure 72, the loop filter unit 212 performs LMCS processing, DBF processing, SAO processing, and NNLF processing in parallel. At this time, the input data for LMCS processing is the same as the input data for NNLF processing. Specifically, the output data obtained from other decoding processes is used as the input data for LMCS processing and NNLF processing, respectively. Furthermore, the loop filter unit 212 combines the output data obtained from SAO processing and the output data obtained from NNLF processing, and uses the combined data as the input data for ALF processing.

[0672] In this configuration of the loop filter unit 212, the output data obtained by NNLF processing is combined with the output data obtained by SAO processing, which is performed immediately before ALF processing. This allows NNLF processing to function as a powerful pre-filter that reduces artifacts before ALF processing is applied. As a result, the residuals of the entire frame are reduced, allowing ALF processing to further reduce distortion by focusing on fine-grained adjustments at the block level based on the content of the frame. Consequently, in the configuration of the loop filter unit 212 shown in Figure 72, detail preservation is improved, resulting in improved reconstruction quality and improved image compression ratio.

[0673] In the configuration of the loop filter unit 212 shown in Figure 72, the GPU performs ALF processing. In this case, the output data obtained by SAO processing is transferred from the CPU to the GPU via the data bus 300 and combined with the output data obtained by NNLF processing. The combined output data is then input to the ALF processing.

[0674] Alternatively, the CPU may perform ALF processing. In this case, the output data obtained by NNLF processing is transferred from the GPU to the CPU via the data bus 300 and combined with the output data obtained by SAO processing. In this case, the number of times data is transferred from the GPU to the CPU is kept to a minimum compared to the configurations in the reference examples of Figures 36 and 37. Therefore, in order for multiple loop filter processing to be completed without causing significant delays, the data from the GPU to the CPU must be in picture units. The combined output data is then input to the ALF processing.

[0675] Figure 73 is a sixth block diagram that shows a simplified configuration of the loop filter section 212 according to the third embodiment.

[0676] As shown in Figure 73, the loop filter unit 212 performs DBF processing, SAO processing, and NNLF processing in parallel. At this time, the input data for DBF processing is the same as the input data for NNLF processing. Specifically, the output data obtained by LMCS processing is used as the input data for DBF processing and NNLF processing. Furthermore, the loop filter unit 212 combines the output data obtained by SAO processing and the output data obtained by NNLF processing, and uses the combined data as the input data for ALF processing.

[0677] Figure 74 is a seventh block diagram that shows a simplified configuration of the loop filter section 212 according to the third embodiment.

[0678] As shown in Figure 74, the loop filter unit 212 performs SAO processing and NNLF processing in parallel. At this time, the input data for SAO processing is the same as the input data for NNLF processing. Specifically, the output data obtained by DBF processing is used as the input data for both SAO processing and NNLF processing. Furthermore, the loop filter unit 212 combines the output data obtained by SAO processing and the output data obtained by NNLF processing, and uses the combined data as the input data for ALF processing.

[0679] In the above description, the configuration example of the loop filter unit 212 according to the third embodiment (i.e., the loop filter unit 212 of the decoding device 200) has been mainly explained, but the loop filter unit 120 of the encoding device 100 can also be configured in the same way as the loop filter unit 212 according to the third embodiment. By doing so, the loop filter unit 120 of the encoding device 100 has the same advantages as the loop filter unit 212 according to the third embodiment, such as being able to suppress the delay that occurs until multiple loop filtering processes are completed.

[0680] Furthermore, in the explanation using Figures 68 to 74 above, the example given for the application order of one or more non-neural network loop filter processes is LMCS processing, DBF processing, SAO processing, and ALF processing, but this is not limited to this. The application order of one or more non-neural network loop filter processes is not limited to this and may be changed as appropriate, or multiple non-neural network loop filter processes may be performed in parallel.

[0681] Furthermore, in the explanation using Figures 68 to 74 above, examples of one or more non-neural network loop filtering processes were given, such as LMCS processing, DBF processing, SAO processing, and ALF processing, but the explanation is not limited to these. For example, the loop filter unit 212 may use at least one of the LMCS processing, DBF processing, SAO processing, and ALF processing as one or more non-neural network loop filtering processes. Alternatively, the loop filter unit 212 may use one or more loop filtering processes that do not use neural networks other than LMCS processing, DBF processing, SAO processing, and ALF processing as one or more non-neural network loop filtering processes.

[0682] Furthermore, while the explanation using Figures 68 to 74 above illustrates the case where o...

Claims

1. A decoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, applies a plurality of loop filtering processes to the current block to be decoded, the plurality of loop filtering processes include one or more non-neural network loop filtering processes, each of which does not use a neural network, and one or more neural network loop filtering processes, each of which uses a neural network, and the one or more non-neural network loop filtering processes are applied without using the output data obtained by the one or more neural network loop filtering processes.

2. The decoding device according to claim 1, wherein the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes is used as input data for one of the one or more neural network loop filtering processes.

3. The decoding device according to claim 1, wherein the input data for one of the one or more non-neural network loop filtering processes is the same as the input data for one of the one or more neural network loop filtering processes, and the circuit combines the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes with the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes.

4. The decoding device according to claim 3, wherein the combination is one of addition, weighted addition, and calculation of the absolute value of the difference.

5. The decoding device according to any one of claims 1 to 4, wherein each of the one or more non-neural network loop filtering processes is one of LMCS (Luma Mapping with Chroma Scaling), DBF (Deblocking Filter), SAO (Sample Adaptive Offset), and ALF (Adaptive Loop Filter).

6. The decoding device according to any one of claims 1 to 4, wherein the circuit uses the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes as a display image.

7. The decoding device according to any one of claims 1 to 4, wherein the circuit uses the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes as a display image.

8. The decoding device according to any one of claims 1 to 4, wherein the memory includes a DPB (Decoded Picture Buffer), and the circuit stores in the DPB the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes as a candidate reference image for the subsequent picture.

9. The decoding device according to any one of claims 1 to 4, wherein the memory includes a DPB (Decoded Picture Buffer), and the circuit stores in the DPB the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes as a candidate reference image for the subsequent picture.

10. The decoding device according to any one of claims 1 to 4, wherein the circuit performs the one or more neural network loop filter processing on the data in the specific region when data in the specific region is obtained as input data for the one or more neural network loop filter processing.

11. The decoding apparatus according to claim 10, wherein the specific region is any of a processing block, a CTU (Coding Tree Unit), a slice, a tile, a subpicture, and a picture.

12. The decoding device according to claim 10, wherein the circuit decodes information indicating the specific region from the bitstream.

13. The decoding device according to any one of claims 1 to 4, wherein the circuit decodes from a bitstream information indicating whether, when using a picture including the current block as a display image, (i) the output data obtained by the non-neural network loop filtering process applied last to the current block among the one or more non-neural network loop filtering processes is used as the display image, or (ii) the output data obtained by the neural network loop filtering process applied last to the current block among the one or more neural network loop filtering processes is used as the display image.

14. The decoding device according to any one of claims 1 to 4, wherein the circuit decodes from a bitstream information indicating whether, when using the picture including the current block as a reference image for a subsequent picture, (i) the output data obtained by the non-neural network loop filtering process applied last to the current block among the one or more non-neural network loop filtering processes is used as the reference image, or (ii) the output data obtained by the neural network loop filtering process applied last to the current block among the one or more neural network loop filtering processes is used as the reference image.

15. The decoding device according to any one of claims 1 to 4, wherein the circuit decodes from the bitstream information indicating whether or not to apply the one or more neural network loop filtering processes to the current block.

16. The decoding device according to claim 1 or 2, wherein the circuit includes a first processor which is a CPU (Central Processing Unit) and a second processor which is one of a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a TPU (Tensor Processing Unit), the first processor performs one or more non-neural network loop filtering processes, and the second processor performs one or more neural network loop filtering processes in parallel with the one or more non-neural network loop filtering processes performed by the first processor.

17. The decoding device according to any one of claims 1, 3, and 4, wherein the circuit includes a first processor which is a CPU (Central Processing Unit), and a second processor which is one of a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a TPU (Tensor Processing Unit), the circuit performs one or more non-neural network loop filtering processes in a predetermined processing order, the first processor performs the Nth non-neural network loop filtering processes from the beginning in the processing order of the one or more non-neural network loop filtering processes, the second processor performs the remaining non-neural network loop filtering processes not performed by the first processor, and performs one or more neural network loop filtering processes in parallel with the remaining non-neural network loop filtering processes.

18. An encoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, applies a plurality of loop filter processes to the current block to be encoded, the plurality of loop filter processes include one or more non-neural network loop filter processes, each of which does not use a neural network, and one or more neural network loop filter processes, each of which uses a neural network, and the one or more non-neural network loop filter processes are applied without using the output data obtained by the one or more neural network loop filter processes.

19. The encoding device according to claim 18, wherein the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes is used as input data for one of the one or more neural network loop filtering processes.

20. The encoding device according to claim 18, wherein the input data for one of the one or more non-neural network loop filtering processes is the same as the input data for one of the one or more neural network loop filtering processes, and the circuit combines the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes with the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes.

21. The encoding apparatus according to claim 20, wherein the combination is one of addition, weighted addition, and calculation of the absolute value of the difference.

22. The encoding apparatus according to any one of claims 18 to 21, wherein each of the one or more non-neural network loop filtering processes is one of LMCS (Luma Mapping with Chroma Scaling), DBF (Deblocking Filter), SAO (Sample Adaptive Offset), and ALF (Adaptive Loop Filter).

23. The encoding device according to any one of claims 18 to 21, wherein the circuit uses the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes as a display image.

24. The encoding device according to any one of claims 18 to 21, wherein the circuit uses the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes as a display image.

25. The encoding device according to any one of claims 18 to 21, wherein the memory includes a DPB (Decoded Picture Buffer), and the circuit stores in the DPB the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes as a candidate reference image for the subsequent picture.

26. The encoding device according to any one of claims 18 to 21, wherein the memory includes a DPB (Decoded Picture Buffer), and the circuit stores in the DPB the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes as a candidate reference image for the subsequent picture.

27. The encoding device according to any one of claims 18 to 21, wherein the circuit performs the one or more neural network loop filter processing on the data in the specific region when data in the specific region is obtained as input data for the one or more neural network loop filter processing.

28. The encoding apparatus according to claim 27, wherein the specific region is any of a processing block, a CTU (Coding Tree Unit), a slice, a tile, a subpicture, and a picture.

29. The encoding device according to claim 27, wherein the circuit encodes information indicating the specific region into a bitstream.

30. The encoding device according to any one of claims 18 to 21, wherein the circuit encodes information into a bitstream indicating whether (i) the output data obtained by the non-neural network loop filtering process that is last applied to the current block among the one or more non-neural network loop filtering processes is used as the display image, or (ii) the output data obtained by the neural network loop filtering process that is last applied to the current block among the one or more neural network loop filtering processes is used as the display image.

31. The encoding device according to any one of claims 18 to 21, wherein the circuit encodes information into a bitstream indicating whether, when using the picture including the current block as a reference image for a subsequent picture, (i) the output data obtained by the non-neural network loop filtering process applied last to the current block among the one or more non-neural network loop filtering processes is used as the reference image, or (ii) the output data obtained by the neural network loop filtering process applied last to the current block among the one or more neural network loop filtering processes is used as the reference image.

32. The encoding device according to any one of claims 18 to 21, wherein the circuit encodes information in a bitstream indicating whether or not to apply the one or more neural network loop filtering processes to the current block.

33. The encoding apparatus according to claim 18 or 19, wherein the circuit includes a first processor which is a CPU (Central Processing Unit) and a second processor which is one of a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a TPU (Tensor Processing Unit), the first processor performs one or more non-neural network loop filtering processes, and the second processor performs one or more neural network loop filtering processes in parallel with the one or more non-neural network loop filtering processes performed by the first processor.

34. The encoding device according to any one of claims 18, 20, and 21, wherein the circuit includes a first processor which is a CPU (Central Processing Unit), and a second processor which is one of a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a TPU (Tensor Processing Unit), the circuit performs one or more non-neural network loop filtering processes in a predetermined processing order, the first processor performs the Nth non-neural network loop filtering processes from the beginning in the processing order of the one or more non-neural network loop filtering processes, the second processor performs the remaining non-neural network loop filtering processes not performed by the first processor, and performs one or more neural network loop filtering processes in parallel with the remaining non-neural network loop filtering processes.

35. A decoding method comprising applying multiple loop filtering processes to the current block to be decoded, wherein the multiple loop filtering processes include one or more non-neural network loop filtering processes, each of which does not use a neural network, and one or more neural network loop filtering processes, each of which uses a neural network, and the one or more non-neural network loop filtering processes are applied without using the output data obtained by the one or more neural network loop filtering processes.

36. An encoding method comprising applying multiple loop filtering processes to the current block to be encoded, wherein each of the multiple loop filtering processes includes one or more non-neural network loop filtering processes, each of which does not use a neural network, and one or more neural network loop filtering processes, each of which uses a neural network, and the one or more non-neural network loop filtering processes are applied without using the output data obtained by the one or more neural network loop filtering processes.