Parallel Processing of Image Regions Using Neural Networks - Decoding, Post-Filtering, and RDOQ
By using a neural network with multiple sub-networks to process input tensors in a tile-based manner, the method addresses the challenges of balancing memory and computational resources in video coding, achieving efficient encoding and decoding with improved picture quality.
Patent Information
- Application Number
- JP2024569033
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-07-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-01
AI Technical Summary
Existing video coding technologies face challenges in efficiently encoding and decoding neural network-based bitstreams, particularly in balancing memory resources and computational complexity, while maintaining high compression ratios and picture quality.
The proposed solution involves a method for encoding and decoding picture data using a neural network with at least two sub-networks. This method processes the input tensor by dividing it into tiles and applying different sub-networks to each tile, allowing for varying tile sizes to optimize processing efficiency based on hardware limitations and picture content.
This approach enables efficient encoding and decoding of picture data, reducing memory footprint and processing frequency while maintaining high compression ratios and improving picture quality by reducing artifacts at tile boundaries.
Smart Images

Figure 2025516914000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to the field of encoding and decoding pictures or videos, and more particularly, to encoding and decoding neural network-based bitstreams.
Background Art
[0002] Video coding (video encoding and decoding) is used in a wide range of digital video applications such as, for example, broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat and video conferencing, DVDs and Blu-ray (registered trademark) discs, video content acquisition and editing systems, and camcorders for security applications.
[0003] The amount of video data required to depict even a relatively short video can be quite large, which can pose difficulties when the data is to be streamed or otherwise communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. Also, since memory resources can be limited, the size of the video can also be a problem when the video is stored on a storage device. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Since network resources are limited and the demand for higher video quality is increasing, improved compression and decompression techniques that improve the compression ratio without sacrificing much or any picture quality are desirable.
[0004] Neural network (NN) and deep learning (DL) technologies that utilize artificial neural networks have been used for some time in the technical fields of encoding and decoding of videos, images (e.g., still images), etc.
[0005] It is desirable to further improve the efficiency of such picture coding (video picture coding or still image coding) based on a trained network (e.g., neural network NN) that takes into account the limitations of the available memory and / or processing speed of the decoder and / or encoder. SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION
[0006] Some embodiments of the present disclosure provide methods and apparatuses for encoding and / or decoding pictures in an efficient manner, thereby reducing the memory footprint and the required operating frequency of the processing unit. In particular, the present disclosure enables a trade-off between memory resources and computational complexity within an NN-based video / picture encoding-decoding framework applicable to moving images and still images. MEANS FOR SOLVING THE PROBLEM
[0007] The above and other objects are achieved by the subject matter of the independent claims. Further implementations are apparent from the dependent claims, the description, and the drawings.
[0008] According to an aspect of the present disclosure, a method for encoding an input tensor representing picture data is provided, the method including processing the input tensor by a neural network including at least a first sub-network and a second sub-network, the processing including applying the first sub-network to a first tensor in a spatial dimension by dividing the first tensor into a first plurality of tiles and processing the first plurality of tiles by the first sub-network; and applying the second sub-network to a second tensor in the spatial dimension after applying the first sub-network by dividing the second tensor into a second plurality of tiles and processing the second plurality of tiles by the second sub-network, wherein at least two respective co-located tiles of the first plurality of tiles and the second plurality of tiles are of different sizes. As a result, by varying the tile size for different sub-networks, the input tensor representing picture data can be encoded more efficiently. Further, hardware limitations and requirements can be taken into account.
[0009] In some exemplary implementations, tiles of the first plurality of tiles adjacent in at least one dimension of the spatial dimension overlap partially, and / or tiles of the second plurality of tiles adjacent in at least one dimension of the spatial dimension overlap partially. Thus, the quality of the reconstructed picture can be improved, especially along the tile boundaries. Thus, picture artifacts can be reduced.
[0010] In a further implementation, tiles of the first plurality of tiles are processed independently by the first sub-network, and / or tiles of the second plurality of tiles are processed independently by the second sub-network. For example, at least two tiles of the first plurality of tiles are processed in parallel by the first sub-network, and / or at least two tiles of the second plurality of tiles are processed in parallel by the second sub-network. As a result, the encoding of the picture data can be performed faster.
[0011] According to one implementation, splitting the first tensor includes determining the tile sizes within a first plurality of tiles based on a first predetermined condition, and / or splitting the second tensor includes determining the tile sizes within a second plurality of tiles based on a second predetermined condition. For example, the first predetermined condition and / or the second predetermined condition are based on available decoder hardware resources and / or motion present in the picture data. Thus, the tile sizes can be adapted and optimized according to the available encoder and / or decoder resources and / or according to the picture content.
[0012] In one example, the first subnetwork performs processing by one or more layers including at least one convolutional layer and at least one pooling layer, and / or the second subnetwork performs processing by one or more layers including at least one convolutional layer and at least one pooling layer. Thus, since the convolutional network is particularly suitable for processing data in the spatial dimension, the input tensor data can be processed efficiently.
[0013] In a further example, the first subnetwork and the second subnetwork perform respective processes that are part of picture or video compression. For example, the first subnetwork and / or the second subnetwork perform one of picture encoding by a convolutional subnetwork; rate distortion optimized quantization (RDOQ); and picture filtering. Thus, the encoding of picture data can include subnetwork processing in multiple related stages, improving the encoding efficiency.
[0014] According to an implementation example, the input tensor is a picture or a sequence of pictures that includes one or more components, at least one of which is a color component. This can enable the encoding of color components. In one example, the input tensor has at least two components, namely a first component and a second component; the first subnetwork divides the first component into a third plurality of tiles, divides the second component into a fourth plurality of tiles, and at least two co-located tiles of the third plurality of tiles and the fourth plurality of tiles have different sizes; and / or the second subnetwork divides the first component into a fifth plurality of tiles, divides the second component into a sixth plurality of tiles, and at least two co-located tiles of the fifth plurality of tiles and the sixth plurality of tiles have different sizes. As a result, multiple components can undergo encoding processing for each tile with different tile sizes for the components, which can lead to further improvement in encoding efficiency and / or hardware implementation.
[0015] A further implementation example includes generating a bitstream by including the output of the processing by the neural network in the bitstream. The implementation further includes including in the bitstream an indication of the size of the tiles in the first plurality of tiles and / or an indication of the size of the tiles in the second plurality of tiles. Thus, by providing the indication, the encoder and decoder can set the tile size in a corresponding adaptive manner.
[0016] According to an aspect of the present disclosure, a method for decoding a tensor representing picture data is provided. The method includes processing an input tensor representing picture data by a neural network including at least a first sub-network and a second sub-network. The processing includes applying the first sub-network to a first tensor in the spatial dimension by dividing the first tensor into a first plurality of tiles and processing the first plurality of tiles by the first sub-network; and applying the second sub-network to a second tensor in the spatial dimension after applying the first sub-network by dividing the second tensor into a second plurality of tiles and processing the second plurality of tiles by the second sub-network, wherein at least two respective co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes. As a result, by varying the tile size for different sub-networks, the input tensor representing picture data can be decoded more efficiently. Further, hardware limitations and requirements can be taken into account.
[0017] In some exemplary implementations, tiles of the first plurality of tiles adjacent in at least one dimension of the spatial dimension overlap partially, and / or tiles of the second plurality of tiles adjacent in at least one dimension of the spatial dimension overlap partially. Thus, the quality of the reconstructed picture can be improved, especially along the tile boundaries. Thus, picture artifacts can be reduced.
[0018] In a further implementation, tiles of the first plurality of tiles are processed independently by the first sub-network, and / or tiles of the second plurality of tiles are processed independently by the second sub-network. For example, at least two tiles of the first plurality of tiles are processed in parallel by the first sub-network, and / or at least two tiles of the second plurality of tiles are processed in parallel by the second sub-network. As a result, the encoding of picture data can be performed faster.
[0019] According to one implementation, splitting the first tensor includes determining the tile sizes within a plurality of first tiles based on a first predetermined condition, and / or splitting the second tensor includes determining the tile sizes within a plurality of second tiles based on a second predetermined condition. For example, the first predetermined condition and / or the second predetermined condition are based on available decoder hardware resources and / or motion present in the picture data. Thus, the tile sizes can be adapted and optimized according to the available encoder and / or decoder resources and / or according to the picture content.
[0020] In one example, the first subnetwork performs processing by one or more layers including at least one convolutional layer and at least one pooling layer, and / or the second subnetwork performs processing by one or more layers including at least one convolutional layer and at least one pooling layer. Thus, since the convolutional network is particularly suitable for processing data in the spatial dimension, the input tensor data can be processed efficiently.
[0021] In a further example, the first subnetwork and the second subnetwork perform respective processes that are part of decompressing a picture or video. For example, the first subnetwork and / or the second subnetwork perform one of picture decoding by a convolutional subnetwork and picture filtering. Thus, the decoding of picture data includes subnetwork processing in a plurality of related stages and can improve the coding efficiency.
[0022] According to one implementation, the input tensor is a picture or a sequence of pictures that includes one or more components where at least one is a color component. This can enable the decoding of the color components. In one example, the input tensor has at least two components, namely a first component and a second component, the first subnetwork divides the first component into a third plurality of tiles, divides the second component into a fourth plurality of tiles, and at least two respective co-located tiles of the third plurality of tiles and the fourth plurality of tiles have different sizes; and / or the second subnetwork divides the first component into a fifth plurality of tiles, divides the second component into a sixth plurality of tiles, and at least two respective co-located tiles of the fifth plurality of tiles and the sixth plurality of tiles have different sizes. As a result, the plurality of components can undergo decoding processing tile by tile using different tile sizes for the components, which can lead to further improvement in encoding efficiency and / or hardware implementation.
[0023] A further example implementation includes extracting an input tensor from a bitstream for processing by a neural network. Thereby, the input tensor can be extracted at high speed.
[0024] According to one implementation, the second subnetwork performs picture post-filtering, and for at least two tiles of the second plurality of tiles, one or more parameters of the post-filtering are different and are extracted from the bitstream. Thus, the decoding of the picture data includes subnetwork processing in a plurality of related stages and can improve the encoding efficiency. Further, the post-filtering is performed using filter parameters adapted to the tile size, improving the quality of the reconstructed picture data.
[0025] In one example, it further includes parsing from the bitstream an indication of the tile size in the first plurality of tiles and / or an indication of the tile size in the second plurality of tiles. Thus, by providing the indication, the encoder and decoder can set the tile size in a corresponding adaptive manner.
[0026] According to an aspect of the present disclosure, there is provided a computer program stored in a non-transitory medium that includes code for performing any of the steps of the foregoing aspects of the present disclosure when executed on one or more processors.
[0027] According to an aspect of the present disclosure, there is provided a processing device for encoding an input tensor representing picture data. The processing device includes a processing circuit configured to process the input tensor by a neural network including at least a first sub-network and a second sub-network. The processing includes applying the first sub-network to a first tensor in the spatial dimension by dividing the first tensor into a first plurality of tiles and processing the first plurality of tiles by the first sub-network; and after applying the first sub-network, applying the second sub-network to a second tensor in the spatial dimension by dividing the second tensor into a second plurality of tiles and processing the second plurality of tiles by the second sub-network, wherein at least two co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes.
[0028] According to an aspect of the present disclosure, there is provided a processing device for encoding an input tensor representing picture data. The processing device includes one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors. The programming configures the encoder to perform a method related to encoding an input tensor representing picture data when executed by the one or more processors.
[0029] According to one aspect of the present disclosure, a processing device for decoding a tensor representing picture data is provided. The processing device includes a processing circuit configured to process an input tensor representing picture data by a neural network including at least a first sub-network and a second sub-network. The processing includes applying the first sub-network to a first tensor in the spatial dimension by dividing the first tensor into a first plurality of tiles and processing the first plurality of tiles by the first sub-network; and after applying the first sub-network, applying the second sub-network to a second tensor in the spatial dimension by dividing the second tensor into a second plurality of tiles and processing the second plurality of tiles by the second sub-network, wherein at least two co-located tiles of the first plurality of tiles and the second plurality of tiles are of different sizes.
[0030] According to one aspect of the present disclosure, a processing device for decoding a tensor representing picture data is provided. The processing device includes one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors. The programming configures the decoder to execute a method related to decoding a tensor representing picture data when executed by the one or more processors.
[0031] The present disclosure is applicable to both end-to-end AI codecs and hybrid AI codecs. In a hybrid AI codec, for example, filtering operations (filtering of reconstructed pictures) can be performed by a neural network (NN). The present disclosure is applied to such NN-based processing modules. Generally, the present disclosure can be applied to all or part of the video compression and decompression processes when at least part of the processing includes an NN and such an NN includes convolutional or transposed convolutional operations. For example, the present disclosure is applicable to individual processing tasks that are performed as part of the processing executed by an encoder and / or a decoder, including in-loop filtering and / or post-filtering and / or pre-filtering.
[0032] Note that the present disclosure is not limited to a specific framework. Further, the present disclosure is not limited to image or video compression and can also be applied to object detection, image generation, and recognition systems.
[0033] The present invention can be implemented in hardware (HW) and / or software (SW). Further, the HW-based implementation may be combined with the SW-based implementation.
[0034] For clarity, any one of the foregoing embodiments can be combined with one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0035] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, objects, and advantages will become apparent from this specification, the drawings, and the claims.
Brief Description of the Drawings
[0036] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings.
FIG. 1A
FIG. 1B
FIG. 2
FIG. 3
FIG. 4
FIG. 5
FIG. 6A
FIG. 6B
FIG. 7
FIG. 8
FIG. 9
FIG. 9A
FIG. 9B
FIG. 10
FIG. 11
FIG. 12
FIG. 13
FIG. 14
FIG. 15
FIG. 16
FIG. 17A
FIG. 17B
FIG. 18
FIG. 19
FIG. 20
FIG. 21
FIG. 22
FIG. 23
FIG. 24
FIG. 25
FIG. 26
FIG. 27
[0037] In the following, some embodiments of the present disclosure will be described with reference to the figures. FIGS. 1 - 3 refer to a video coding system and method that may be used with more specific embodiments of the invention described in further figures. Specifically, the embodiments described with respect to FIGS. 1 - 3 may be used with encoding / decoding techniques further described below that utilize neural networks to encode and / or decode a bitstream.
[0038] In the following description, reference is made to the accompanying drawings that form a part hereof and illustrate specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that the embodiments may be used in other aspects and may include structural or logical changes not shown in the figures. Accordingly, the following detailed description should not be construed in a limiting sense, and the scope of the present disclosure is defined by the appended claims.
[0039] For example, it is understood that the disclosure related to the described method also applies to the corresponding device or system configured to execute the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units for executing the described one or more method steps, such as functional units (e.g., one unit for executing the one or more steps, or a plurality of units each executing one or more of the plurality of steps), even if such one or more units are not explicitly described or shown in the drawings. On the other hand, for example, if a specific device is described based on one or more units, such as a certain functional unit, the corresponding method may include one step of executing the functions of the one or more units (e.g., one step of executing the functions of the one or more units, or a plurality of steps each executing the functions of one or more of the plurality of units), even if such one or more steps are not explicitly described or shown in the drawings. Furthermore, it is understood that the various exemplary embodiments and / or aspect features described herein can be combined with each other unless otherwise specified.
[0040] Video coding generally refers to the processing of a sequence of pictures that form a video or video sequence. Instead of the term "picture", the terms "frame" or "image" may be used as synonyms in the field of video coding. Video coding (or generally coding) includes two parts, video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture for more efficient storage and / or transmission. Video decoding is performed on the destination side and typically involves the reverse process compared to the encoder to reconstruct the video picture. Embodiments referring to the "coding" of a video picture (or generally a picture) should be understood to relate to the "encoding" or "decoding" of the video picture or each video sequence. The combination of the encoding part and the decoding part is also called a codec (coding and decoding).
[0041] In the case of lossless video coding, the original video picture can be reconstructed, i.e., the reconstructed video picture has the same quality as the original video picture (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy video coding, further compression, e.g., by quantization, is performed to reduce the amount of data representing the video picture, and this cannot be fully reconstructed at the decoder. That is, the quality of the reconstructed video picture is lower or worse compared to the quality of the original video picture.
[0042] Some video coding standards belong to the group of "lossy hybrid video coders" (i.e., combining spatial and temporal prediction in the sample domain and 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is typically divided into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, the video is typically processed, i.e., encoded, at the block (video block) level, which, for example, uses spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to generate a prediction block, subtracts the prediction block from the current block (the currently processed / to-be-processed block) to obtain a residual block, transforms the residual block, quantizes the residual block in the transform domain to reduce the amount of data to be transmitted (compression). On the other hand, in the decoder, the reverse process compared to the encoder is applied to the encoded or compressed block to reconstruct the current block for presentation. Further, the encoder duplicates the decoder processing loop so that both generate the same prediction (e.g., intra prediction and inter prediction) and / or reconstruction for processing, i.e., coding, subsequent blocks. Recently, part or all of the encoding and decoding chain has been implemented by using a neural network, or generally, any machine learning or deep learning framework.
[0043] In the following embodiments of the video coding system 10, the video encoder 20 and the video decoder 30 are described with reference to FIG. 1.
[0044] Figure 1A is a schematic block diagram showing an exemplary coding system 10 that can utilize the techniques of the present application, for example, a video coding system 10 (or simply coding system 10). The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) of the video coding system 10 represent examples of devices that can be configured to perform the techniques according to various examples described in the present application.
[0045] As shown in Figure 1A, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to a destination device 14, for example, to decode the encoded picture data 13.
[0046] The source device 12 includes an encoder 20 and further optionally includes a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22. Some embodiments of the present disclosure (for example, those related to initial rescaling or rescaling between two progressive layers) may be implemented by the encoder 20. Some embodiments (for example, those related to initial rescaling) may be implemented by the picture preprocessor 18.
[0047] The picture source 16 may have, or may be, any kind of picture capture device such as a camera that captures pictures of the real world, and / or any kind of picture generation device such as a computer graphics processor for generating computer animation pictures, or any kind of other device that acquires and / or provides pictures of the real world, computer-generated pictures (such as screen content, virtual reality (VR) pictures) and / or any combination thereof (such as augmented reality (AR) pictures). The picture source may be any kind of memory or storage device that stores any of the aforementioned pictures.
[0048] Distinguished from the processing executed by the pre-processor 18 and the pre-processing unit 18, the picture or picture data 17 may also be called raw picture or raw picture data 17.
[0049] The pre-processor 18 is configured to receive the (raw) picture data 17 and perform pre-processing on the picture data 17 to obtain a pre-processed picture 19 or pre-processed picture data 19. The pre-processing executed by the pre-processor 18 can include, for example, trimming, color format conversion (such as from RGB to YCbCr or generally from RGB to YUV), color correction, or noise removal. It can be understood that the pre-processing unit 18 may be an optional component. Hereinafter, color space components (for example, R, G, B for the RGB space and Y, U, V for the YUV space) are also referred to as color channels. Further, in the YCbCr color space, Y represents luminance (or luma), and U, V, Cb, Cr represent chrominance (or chroma) channels (components).
[0050] The video encoder 20 is configured to receive the pre - processed picture data 19 and provide the encoded picture data 21 (further details are described below, for example, based on FIG. 4). The encoder 20 can be implemented via the processing circuit 46 to embody various modules discussed with respect to the encoder 20 of FIG. 4 and / or any other encoder system or subsystem described herein.
[0051] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) through the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.
[0052] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and additionally, i.e., optionally, includes a communication interface or communication unit 28, a post - processor 32 (or post - processing unit 32), and a display device 34.
[0053] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof) from, for example, directly from the source device 12 or any other source, such as a storage device, e.g., a storage device of the encoded picture data, and provide the encoded picture data 21 to the decoder 30.
[0054] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 13 between source device 12 and destination device 14 via a direct communication link, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network or any combination thereof, or via any type of private and public network, or via any combination of any of these types.
[0055] Communication interface 22 may be configured to, for example, package encoded picture data 21 into an appropriate format, such as a packet, and / or process the encoded picture data using any type of transmission encoding or processing for transmission through a communication link or communication network.
[0056] Communication interface 28, which forms the counterpart of communication interface 22, may be configured to, for example, receive the transmitted data and process the transmitted data using any type of corresponding transmission decoding or processing and / or unpacking to obtain the encoded picture data 21.
[0057] Both communication interface 22 and communication interface 28 may be configured as a unidirectional communication interface, as indicated by the arrow pointing from source device 12 to destination device 14 for communication channel 13 in FIG. 1A, or as a bidirectional communication interface, for example, to send and receive messages, such as to set up a connection, receive and confirm any other information related to the communication link and / or data transmission, such as the transmission of encoded picture data, and exchange it.
[0058] Decoder 30 is configured to receive the encoded picture data 21 and provide the decoded picture data 31 or the decoded picture 31 (further details are described below, for example, based on FIGS. 3 and 5). Decoder 30 may be implemented via processing circuit 46 to embody various modules discussed with respect to decoder 30 of FIG. 5 and / or any other decoder system or subsystem described herein.
[0059] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as the reconstructed picture data), for example, the decoded picture 31, to obtain the post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing for preparing the decoded picture data 31 for display by, for example, the display device 34.
[0060] Some embodiments of the present disclosure may be implemented by the decoder 30 or by the post-processor 32.
[0061] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33, for example, to display a picture to a user or viewer. The display device 34 may be, for example, any type of display for displaying a reconstructed picture, such as an integrated or external display or monitor, or may have one. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0062] FIG. 1A shows the source device 12 and the destination device 14 as separate devices, but embodiments of the device may include both or both functions, i.e., the source device 12 or corresponding function and the destination device 14 or corresponding function. In such embodiments, the source device 12 or corresponding function and the destination device 14 or corresponding function may be implemented using the same hardware and / or software, or by separate hardware and / or software, or by any combination thereof.
[0063] As will be apparent to those skilled in the art based on this document, the functions or the presence and (exact) partitioning of the various units within the source device 12 and / or destination device 14 as shown in FIG. 1A may vary depending on the actual device and application.
[0064] Encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30) or both encoder 20 and decoder 30 may be implemented via processing circuitry (such as shown in FIG. 1B) including one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. Encoder 20 may be implemented via processing circuitry 46 to embody various modules and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented via processing circuitry 46 to embody various modules and / or any other decoder system or subsystem described herein. The processing circuitry may be configured to perform various operations as described hereinafter. As shown in FIG. 3, if the techniques are implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Either the video encoder 20 or the video decoder 30 may also be integrated as part of a combined encoder / decoder (codec) in a single device, for example, as shown in FIG. 1B.
[0065] The source device 12 and the destination device 14 may be any kind of handheld or fixed device, such as a notebook or laptop computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video game console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiver device, a broadcast transmitter device, etc., and may be equipped with any of a wide range of devices, may not use an operating system, or may use any kind of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 can be wireless communication devices.
[0066] In some cases, the video coding system 10 shown in FIG. 1A is merely an example, and the techniques of the present application may be applied to video coding settings (such as video encoding or video decoding) that do not necessarily include data communication between an encoding device and a decoding device. In other examples, the data is retrieved from local memory, streamed over a network, or the like. The video encoding device may encode the data and store it in memory, and / or the video decoding device may retrieve the data from memory and decode it. In some examples, encoding and decoding are performed by devices that simply encode data into memory and / or retrieve and decode data from memory without communicating with each other.
[0067] For the sake of convenience in explanation, in this document, some embodiments are described, for example, by referring to the reference software of High Efficiency Video Coding (HEVC), or Versatile Video Coding (VVC), which is the next-generation video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC.
[0068] Figure 2 is a schematic diagram of a video coding device 200 according to an embodiment of the present invention. The video coding device 200 is suitable for implementing the disclosed embodiments as described herein. In one embodiment, the video coding device 200 can be a decoder such as the video decoder 30 in FIG. 1A, or an encoder such as the video encoder 20 in FIG. 1A.
[0069] The video coding device 200 includes an inlet port 210 (or input port 210) and a receiver unit (Rx) 220 for receiving data; a processor, logic unit, or central processing unit (CPU) 230 for processing data; a transmitter unit (Tx) 240 and an outlet port 250 (or output port 250) for transmitting data; and a memory 260 for storing data. The video coding device 200 may also include opto-electrical (OE) components and electro-optical (EO) components coupled to the inlet port 210, the receiver unit 220, the transmitter unit 240, and the outlet port 250 for sending or receiving optical or electrical signals.
[0070] Processor 230 is implemented by hardware and software. Processor 230 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 230 communicates with an input port 210, a receiver unit 220, a transmitter unit 240, an output port 250, and a memory 260. Processor 230 includes a coding module 270. The coding module 270 implements the disclosed embodiments described above. For example, the coding module 270 performs, processes, prepares, or provides various coding operations. Thus, including the coding module 270 provides a substantial improvement to the functionality of the video coding device 200 and results in the conversion of the video coding device 200 to different states. Alternatively, the coding module 270 may be implemented as instructions stored in the memory 260 and executed by the processor 230.
[0071] Memory 260 may include one or more disks, tape drives, and solid state drives and may be used as an overflow data storage device to store such programs when a program is selected for execution and to store instructions and data read during program execution. Memory 260 may be, for example, volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0072] FIG. 3 is a simplified block diagram of an apparatus 300 that may be used as either or both of the source device 12 and the destination device 14 from FIG. 1, according to an exemplary embodiment.
[0073] The processor 302 within the device 300 can be a central processing unit. Alternatively, the processor 302 can be any other type of device or devices capable of manipulating or processing information, whether currently existing or developed in the future. The disclosed embodiments can be implemented using a single processor, such as processor 302 as illustrated, but the advantages in terms of speed and efficiency can be achieved using more than one processor.
[0074] The memory 304 within the device 300 can be, in one implementation, a read-only memory (ROM) device or a random access memory (RAM) device in some implementations. Any other suitable type of storage device can be used as the memory 304. The memory 304 can include code and data 306 that are accessed by the processor 302 using the bus 312. The memory 304 can further include an operating system 308 and an application program 310, and the application program 310 includes at least one program that enables the processor 302 to execute the methods described herein. For example, the application program 310 can include applications 1 to N, which further include video coding applications that execute the methods described herein.
[0075] The device 300 can also include one or more output devices, such as a display 318. The display 318 can be, in one example, a touch-sensitive display that combines a touch-sensitive element operable to sense touch input with a display. The display 318 can be coupled to the processor 302 via the bus 312.
[0076] Although shown here as a single bus, the bus 312 of the apparatus 300 can be composed of a plurality of buses. Further, the secondary storage 314 may be directly coupled to other components of the apparatus 300 or accessed via a network, and may include a single integrated unit such as a memory card or a plurality of units such as a plurality of memory cards. Thus, the apparatus 300 can be implemented in a wide variety of configurations.
[0077] FIG. 4 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the techniques of the present application. In the example of FIG. 4, the video encoder 20 includes an input 401 (or input interface 401), a residual calculation unit 404, a conversion processing unit 406, a quantization unit 408, an inverse quantization unit 410, an inverse conversion processing unit 412, a reconstruction unit 414, a loop filter unit 420, a decoded picture buffer (DPB) 430, a mode selection unit 460, an entropy encoding unit 470, and an output 472 (or output interface 472). The mode selection unit 460 may include an inter prediction unit 444, an intra prediction unit 454, and a partition division unit 462. The inter prediction unit 444 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 as shown in FIG. 4 may also be referred to as a hybrid video encoder or a video encoder by a hybrid video codec.
[0078] Encoder 20 may be configured to receive a picture 17 (or picture data 17), for example, a video or a sequence of pictures forming a video sequence, via, for example, input 401. The received picture or picture data may be pre-processed picture 19 (or pre-processed picture data 19). For simplicity, the following description refers to picture 17. Picture 17 is also referred to as the current picture or the picture to be coded (in particular, in video coding, to distinguish the current picture from other pictures, for example, pictures that have been encoded and / or decoded previously in the same video sequence, i.e., the video sequence that also includes the current picture).
[0079] (Digital) pictures are, or can be regarded as, two-dimensional arrays or matrices of samples having intensity values. Samples in the array are sometimes also called pixels (short for picture elements). The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are used, i.e., the picture may be represented by, or may include, three sample arrays. In the RGB format or color space, the picture includes corresponding red, green, and blue sample arrays. However, in video coding, each pixel typically includes a luminance component represented by Y (sometimes L is also used instead) and two chrominance components represented by Cb and Cr, in a luminance and chrominance format or color space, such as represented by YCbCr. The luminance (or simply luma) component Y represents brightness or gray-level intensity (such as in a grayscale picture), and the two chrominance (or simply chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format may be converted or transformed to the YCbCr format, and vice versa, and this process is also known as color conversion or transformation. If the picture is monochrome, the picture may include only a luminance sample array. Thus, a picture can be, for example, an array of luma samples in a monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0080] An embodiment of the video encoder 20 may have a picture partitioning unit (not shown in FIG. 2) configured to partition picture 17 into a plurality of (typically non-overlapping) picture blocks 403. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB), or coding tree units (CTU) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size and the corresponding grid defining the block size for all pictures of the video sequence, or to vary the block size between pictures or subsets or groups of pictures and partition each picture into corresponding blocks. The abbreviation AVC represents Advanced Video Coding.
[0081] In a further embodiment, the video encoder may be configured to directly receive the blocks 403 of picture 17, for example, one, some, or all of the blocks forming picture 17. The picture block 403 may also be referred to as the current picture block or the picture block to be coded.
[0082] Similar to Picture 17, Picture Block 403 is a two - dimensional array or matrix of samples having intensity values (sample values) that is, or can be regarded as, of a smaller dimension than Picture 17. In other words, Block 403 can include, for example, one sample array (e.g., a luma array in the case of a monochrome Picture 17, or a luma or chroma array in the case of a color picture), or three sample arrays (e.g., a luma and two chroma arrays in the case of a color Picture 17), or any other number and / or type of arrays depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of Block 403 defines the size of Block 403. Thus, the block can be, for example, an M×N (M columns and N rows) array of samples, or an M×N array of transform coefficients.
[0083] An embodiment of video encoder 20 as shown in FIG. 4 may be configured to encode Picture 17 block - by - block, for example, encoding and prediction are performed for each Block 403.
[0084] An embodiment of video encoder 20 as shown in FIG. 4 may be further configured to divide and / or encode a picture by using slices (also called video slices), the picture may be divided into one or more slices (typically non - overlapping), or encoded using them, and each slice may include one or more blocks (e.g., CTUs).
[0085] An embodiment of the video encoder 20 as shown in FIG. 4 may be further configured to divide and / or encode a picture by using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles), where the picture may be divided into one or more tile groups (typically non-overlapping), or encoded using them, and each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles, and each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), e.g., complete or partial blocks.
[0086] FIG. 5 shows an example of a video decoder 30 configured to implement the techniques of the present application. The video decoder 30 is configured to receive encoded picture data 21 (e.g., an encoded bitstream 21), encoded, for example, by the encoder 20, to obtain a decoded picture 531. The encoded picture data or bitstream includes data representing picture blocks of the encoded video slice (and / or tile group or tile) for decoding the encoded picture data, e.g., and associated syntax elements.
[0087] Entropy decoding unit 504 parses the bitstream 21 (or generally the encoded picture data 21), and for example, performs entropy decoding on the encoded picture data 21 to obtain, for example, the quantized coefficients 309 and / or the decoded coding parameters (not shown in FIG. 3), for example, the inter prediction parameters (for example, the reference picture index and the motion vector), the intra prediction parameters (for example, the intra prediction mode or index), the transform parameters, the quantization parameters, the loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 504 may be configured to apply a decoding algorithm or method corresponding to the encoding method described for the entropy encoding unit 470 of the encoder 20. The entropy decoding unit 504 may be further configured to provide the inter prediction parameters, the intra prediction parameters, and / or other syntax elements to the mode application unit 360, and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level. In addition to or instead of the slice and its respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used.
[0088] The reconstruction unit 514 (for example, the adder or summer 514) may be configured to add the reconstructed residual block 513 to the prediction block 565 by adding, for example, the sample values of the reconstructed residual block 513 and the sample values of the prediction block 565 to obtain the reconstructed block 515 in the sample area.
[0089] An embodiment of the video decoder 30 as shown in FIG. 5 may be configured to divide and / or decode a picture by using slices (also called video slices), where the picture may be divided into one or more slices (typically non-overlapping) or decoded using them, and each slice may include one or more blocks (e.g., CTUs).
[0090] An embodiment of the video decoder 30 shown in FIG. 5 may be configured to divide and / or decode a picture by using tile groups (also called video tile groups) and / or tiles (also called video tiles), where the picture may be divided into one or more tile groups (typically non-overlapping) or decoded using them, and each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles, and each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), e.g., complete or partial blocks.
[0091] Other variations of the video decoder 30 may be used to decode the encoded picture data 21. For example, the decoder 30 can generate an output video stream without the loop filtering unit 520. For example, a non-conversion-based decoder 30 can directly inverse quantize the residual signal without the inverse transform processing unit 512 for certain blocks or frames. In another implementation, the video decoder 30 can have an inverse quantization unit 510 and an inverse transform processing unit 512 combined in a single unit.
[0092] In the encoder 20 and decoder 30, it should be understood that the processing result of the current step can be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0093] Specific, non-limiting, exemplary embodiments of the present invention are described below. Prior to that, some explanations and definitions are provided to assist in the understanding of the present disclosure.
[0094] Picture size Picture size refers to the width w or height h of the picture, or the width-height pair. The width and height of an image are typically measured by the number of luma samples.
[0095] Downsampling Downsampling is a process in which the sampling rate (sampling interval) of a discrete input signal is reduced. For example, if the input signal is an image having sizes h and w and the output of downsampling has sizes h2 and w2, at least one of the following holds. h2 < h w2 < w In one exemplary implementation, downsampling can be performed by retaining only every mth sample and discarding the rest of the input signal (e.g., an image).
[0096] Upsampling: Upsampling is a process in which the sampling rate (sampling interval) of a discrete input signal is increased. For example, if the input image has sizes h and w and the output of downsampling has sizes h2 and w2, at least one of the following holds. h2 > h w2 > w
[0097] Resampling: Both downsampling and upsampling processes are examples of resampling. Resampling is the process by which the sampling rate (sampling interval) of an input signal is changed. Resampling is a technique for changing the size (or rescaling) of an input signal.
[0098] During an upsampling or downsampling process, filtering may be applied to improve the accuracy of the resampled signal and reduce the aliasing effect. Interpolation filtering typically involves a weighted combination of sample values at sample positions around the resampling location. This can be implemented as follows. [Number] Here, f() refers to the resampled signal. (x r , y r ) are the resampling coordinates (coordinates of the resampled sample), C(k) is the interpolation filter coefficient, and s(x, y) is the input signal. The coordinates x, y are the coordinates of the samples of the input image. The summation operation is performed for (x r , y r ) in the neighborhood Ω r . In other words, the new sample f(x r , y r ) is obtained as a weighted sum of the input picture samples s(x, y). The weighting is performed by the coefficient C(k), where k indicates the position (index) of the filter coefficient within the filter mask. For example, in the case of a 1D filter, k takes values from 1 to the order of the filter. In the case of a 2D filter applicable to 2D images, k can be an index indicating one of all possible (non-zero) filter coefficients. The index is conventionally associated with a specific position of the coefficient within the filter mask (filter kernel).
[0099] Cropping: Trim (cut) the outer edges of the digital image. Cropping can be used to make the image smaller (with respect to the number of samples) and / or to change the aspect ratio (length to width) of the image. This can be understood as removing samples from the signal, typically samples at the boundaries of the signal.
[0100] Padding: Padding refers to increasing the size of the input (i.e., the input image) by generating new samples (e.g., at the boundaries of the image) by using predefined sample values or by using sample values at existing positions within the input image (e.g., copying or combining). The generated samples are approximations of actual sample values that do not exist.
[0101] Resizing: Resizing is a general term for changing the size of the input image. Resizing may be performed using one of the methods of padding or cropping. Alternatively, resizing may also be done by resampling.
[0102] Integer division: Integer division is a division in which the fractional part (remainder) is discarded.
[0103] Convolution: Convolution can be defined in one dimension for an input signal f() and a filter g() as follows.
Equation
[0104] Artificial neural network An artificial neural network (ANN) or connectionist system is a computing system vaguely inspired by the biological neural networks that make up animal brains. Such systems "learn" to perform tasks by considering examples, generally without being programmed with task-specific rules. For example, in image recognition, an artificial neural network can learn to identify images containing cats by analyzing images of examples manually labeled as "cat" or "not cat" and using the results to identify cats in other images. This is done without prior knowledge of cats, such as having fur, a tail, whiskers, and a cat-like face. Instead, the artificial neural network automatically generates the features to identify from the examples it processes.
[0105] An ANN is based on a collection of connected units or nodes called artificial neurons that roughly model neurons in a biological brain. Each connection can transmit a signal to other neurons, like a synapse in a biological brain. An artificial neuron that receives a signal can then process the signal and transmit a signal to the neurons connected to it. In the implementation of an ANN, the "signal" in a connection is a real number, and the output of each neuron can be calculated by some non - linear function of the sum of its inputs. These connections are called edges. Neurons and edges typically have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal in the connection. A neuron may have a threshold such that a signal is sent only if the aggregated signal exceeds that threshold. Typically, neurons are aggregated into layers. Different layers can perform different transformations on their inputs. The signal progresses from the first layer (input layer) to the last layer (output layer), possibly crossing multiple layers several times.
[0106] The original goal of the ANN approach was to solve problems in the same way as the human brain. Over time, the interest has shifted to performing specific tasks, leading to a divergence from biology. ANNs have been used in a variety of tasks, including computer vision, speech recognition, machine translation, social network filtering, playing board and video games, medical diagnosis, and even activities that have traditionally been considered human specialties such as painting.
[0107] Downsampling layer: A layer of a neural network that results in a reduction of at least one of the dimensions of the input. In general, the input can have three or more dimensions, which can include the number of channels, width, and height. The downsampling layer typically refers to a reduction in the width and / or height dimensions. This can be accomplished using operations such as convolution (possibly with a stride), averaging, max - pooling, etc.
[0108] Upsampling layer: A layer of a neural network that brings about an increase in at least one of the dimensions of the input. Generally, the input can have three or more dimensions, and the dimensions can include the number of channels, width, and height. An upsampling layer typically refers to an increase in the width and / or height dimensions. This can be implemented using operations such as deconvolution, replication, etc.
[0109] Feature map: A feature map is generated by applying a filter (kernel) or feature detector to an input image or the feature map output of previous layers. Feature map visualization provides insights into the internal representation for a specific input for each convolutional layer within the model. Generally, a feature map is the output of a neural network layer. A feature map typically contains one or more feature elements.
[0110] Convolutional neural network The name "Convolutional Neural Network" (CNN) indicates that the network uses a mathematical operation called convolution. Convolution is a special type of linear operation. A convolutional network is simply a neural network that uses convolution instead of general matrix multiplication in at least one layer. A convolutional neural network consists of an input layer, an output layer, and multiple hidden layers. The input layer is the layer where the input is provided for processing.
[0111] For example, the neural network of FIG. 6A is a CNN. The hidden layer of a CNN typically consists of a series of convolutional layers (e.g., conv layers 601-612 in FIG. 6A) that are convolved with multiplication or other dot products. The result of a layer is one or more feature maps, which may also be called channels. There may be subsampling that participates in some or all of the layers. As a result, the feature maps become smaller. The activation function in a CNN may be, as already exemplified above, a RELU (Rectified Linear Unit) layer or a GDN layer, and then additional convolutions such as a pooling layer, a fully connected layer, and a normalization layer may follow. These layers are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. These layers are colloquially called convolutions, but this is just by convention. Mathematically, it is technically a sliding dot product or cross-correlation. This is important for the indices in the matrix in terms of how the weights are determined at a particular index point.
[0112] When programming a CNN to process a picture or an image, the input is a tensor having a shape of (number of images) × (image width) × (image height) × (image depth) (e.g., an input tensor such as tensor x 614 in FIG. 6A). Then, after passing through the convolutional layer, the image is abstracted into a feature map (feature tensor) having a shape of (number of images) × (feature map width) × (feature map height) × (feature map channels). In FIG. 6A, such a feature map is, for example, y. The convolutional layer in a neural network should have the following attributes. A convolutional kernel defined by the width and height (hyperparameters). The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) should be equal to the number of channels (depth) of the input feature map. For example, conv Nx5x5 in FIG. 6A refers to a kernel of size 5x5 and N channels, where N is an integer greater than or equal to 1.
[0113] In the past, traditional multi-layer perceptron (MLP) models have been used for image recognition. However, due to the full connectivity between nodes, MLP models have suffered from high dimensionality and have not scaled well with higher resolution images. A 1000×1000 pixel image with RGB color channels has 3 million weights, which is too high to be processed efficiently at a scale with full connectivity. Also, such network architectures do not consider the spatial structure of the data and treat distant input pixels the same as neighboring pixels. This ignores the locality of reference in image data, both computationally and semantically. Thus, the full connectivity of neurons is wasteful for purposes such as image recognition that are dominated by spatially local input patterns.
[0114] Convolutional neural networks are biologically inspired variants of multi-layer perceptrons specifically designed to emulate the behavior of the visual cortex. CNN models alleviate the problems posed by MLP architectures by exploiting the strong spatially local correlations present in natural images. The convolutional layer is the core component of a CNN. The parameters of the layer consist of a set of learnable filters (the kernels described above), which have small receptive fields but extend over the full depth of the input volume. During the forward pass, each filter is convolved over the width and height of the input volume, computing the dot product between the entries of the filter and the input, and generating a 2D activation map for that filter. As a result, the network learns filters that activate when it detects some particular type of feature at some spatial location in the input.
[0115] By stacking the activation maps for all filters along the depth dimension, the complete output volume of the convolutional layer is formed. Thus, all entries within the output volume can also be interpreted as looking at small regions within the input and being the output of neurons that share neurons and parameters within the same activation map. The feature map, i.e., the activation map, is the output activation for a given filter. Feature map and activation have the same meaning. In some papers, it is called the activation map because it is a mapping corresponding to the activation of different parts of the image, and it is also called the feature map because it is a mapping of where certain features are found within the image. High activation means that a certain feature has been found.
[0116] Another important concept in CNN is pooling. Pooling is a form of non-linear downsampling. There are several non-linear functions for performing pooling, among which max pooling is the most common. It divides the input image into a set of non-overlapping rectangles and outputs the maximum value for each such sub-region.
[0117] Intuitively, the exact location of a feature is not as important as its approximate location relative to other features. This is the idea behind the use of pooling in convolutional neural networks. The pooling layer gradually reduces the spatial size of the representation, reduces the number of parameters, memory footprint, and computational amount in the network, and thus also serves to control overfitting. In a CNN architecture, it is common to periodically insert a pooling layer between successive convolutional layers. The pooling operation provides another form of translational invariance.
[0118] The pooling layer acts independently on all depth slices of the input and spatially resizes it. The most common form is a pooling layer with a 2×2 filter applied with a stride of 2, which down-samples by a factor of 2 along both width and height for each depth slice in the input, discarding 75% of the activations. In this case, each max operation is over four numbers. The depth dimension remains unchanged.
[0119] In addition to max pooling, the pooling unit can use other functions such as average pooling or L2-norm pooling. Average pooling has been historically often used but has become less preferred recently compared to max pooling which generally performs better in practice. There is a recent trend towards actively reducing the size of the representation, either by using smaller filters or by eliminating the pooling layer altogether. "Region of interest" pooling (also known as ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important component of convolutional neural networks for object detection based on the fast R-CNN architecture.
[0120] The above ReLU is short for rectified linear unit and applies a non-saturating activation function. It effectively removes negative values from the activation map by setting them to zero. This increases the decision function and the non-linear characteristics of the network as a whole without affecting the receptive field of the convolutional layer. Other functions such as the saturated hyperbolic tangent and sigmoid functions are also used to increase non-linearity. ReLU is often preferred over other functions because it can train neural networks several times faster without a significant penalty to generalization accuracy.
[0121] After several convolutional layers and max pooling layers, high-level inference in a neural network is performed through fully connected layers. Neurons in fully connected layers have connections to all activations in the previous layer, as seen in a normal (non-convolutional) artificial neural network. Thus, the activation can be computed as an affine transformation with matrix multiplication followed by a bias offset (vector addition of learned or fixed bias terms).
[0122] The "loss layer" specifies how to penalize the deviation between the predicted (output) label and the true label during training and is usually the final layer of the neural network. Different loss functions suitable for different tasks can be used. Softmax loss is used to predict a single class out of K mutually exclusive classes. Sigmoid cross-entropy loss is used to predict K independent probability values in [0,1]. Euclidean loss is used for regression to real-valued labels.
[0123] Subnetwork A neural network can contain multiple sub-networks. A sub-network consists of one or more layers. Different sub-networks have different input / output sizes, resulting in different memory requirements and / or computational complexities.
[0124] Pipeline A series of sub-networks that process specific components of an image. For example, the image component can be any of the R, G, B components. The component can also be one of the luma Y or chroma components U or V. An example can be a system with two pipelines, where the first pipeline processes only the luma component and the second pipeline processes the chroma component(s). One pipeline processes only one component, and the second component (e.g., the luma component) can be used as auxiliary information to assist in its processing (e.g., the processing of the chroma component). For example, a pipeline with a chroma component as output can have both potential representations of the luma component and the chroma component as input (conditional coding of the chroma component).
[0125] Rate-Distortion Optimized Quantization (RDOQ) RDOQ is an encoder-only technique, i.e., it is applied to the processing performed by the encoder rather than the decoder. Before writing to the bitstream, the parameters are quantized (descaling, rounding, etc.) to a specified standard fixed precision. Among the multiple methods of rounding, for example, the minimum RD cost variant as used in HEVC or VVC for transform coefficient coding can often be selected.
[0126] Conditional Color Separation (CCS) In an NN architecture for image / video coding / processing, CCS refers to the independent coding / processing of the primary color components (e.g., the luma component), while the secondary color components (e.g., chroma UV) are conditionally coded / processed using the primary component as auxiliary input.
[0127] Autoencoders and Unsupervised Learning An autoencoder is a type of artificial neural network used to learn efficient data coding without a teacher. A schematic diagram thereof is shown in FIG. 7. This may be considered a simplified representation of the CNN-based VAE (variational autoencoder) structure of FIGS. 6A or 6B. The purpose of the autoencoder is to learn a representation (encode) for a set of data, typically for dimensionality reduction, by training the network to ignore the signal "noise". Along with the reduction side, a reconstruction side is learned, where the autoencoder attempts to generate a representation as close as possible to its original input from the reduced encoding. This is where its name comes from.
[0128] In the simplest case, when given a single hidden layer, the encoder stage of the autoencoder takes the input x and maps it to h: h = σ(Wx + b) This image h is typically called the code, latent variable, or latent representation. Here, σ is an element-wise activation function such as the sigmoid function or rectified linear unit. W is the weight matrix. b is the bias vector. The weights and biases are typically initialized randomly and then sequentially updated iteratively through backpropagation during training. Subsequently, the decoder stage of the autoencoder maps h to a reconstruction x' of the same shape as x. x' = σ'(W'h' + b') Here, σ', W', b' for the decoder may be independent of the corresponding σ, W, b for the encoder.
[0129] Variational autoencoder models make strong assumptions about the distribution of the latent variables. They use a variational approach for learning the latent representation, resulting in an additional loss component and a specific estimator for a training algorithm called the stochastic gradient variational Bayes (SGVB) estimator. The data is generated by the directed graphical model p θ (x|h), and the encoder is the posterior distribution pθ Approximate q with respect to (h|x) φ Suppose that (h|x) is being learned. Here, φ and θ represent the parameters of the encoder (recognition model) and decoder (generation model), respectively. The probability distribution of the latent vector of the VAE typically matches the probability distribution of the training data much better than that of a standard autoencoder. The objective function of the VAE has the following form:
Number
Number
[0130] Recent advances in the field of artificial neural networks, particularly convolutional neural networks, have enabled researchers' interest in applying neural network-based techniques to the tasks of image and video compression. For example, end-to-end optimized image compression using a network based on the variational autoencoder (AVE) has been proposed.
[0131] Therefore, data compression is considered a fundamental and well-studied problem in engineering and is usually formulated with the goal of designing codes for a given discrete data ensemble with minimum entropy. This solution strongly depends on the knowledge of the probabilistic structure of the data, and thus this problem is closely related to probabilistic source modeling. However, since all practical codes must have a finite entropy, continuous-valued data (such as vectors of image pixel intensities) must be quantized into a finite set of discrete values, which introduces errors.
[0132] In this context, known as the lossy compression problem, one must trade off two competing costs, namely, the entropy (rate) of the discretized representation and the error (distortion) resulting from quantization. Different compression applications, such as data storage or transmission through a channel with limited capacity, require different rate-distortion tradeoffs. The simultaneous optimization of rate and distortion is difficult. Without additional constraints, the general problem of optimal quantization in high-dimensional spaces cannot be addressed.
[0133] For this reason, most existing image compression methods operate by linearly transforming the data vector into a suitable continuous-valued representation, quantizing its elements independently, and then encoding the resulting discrete representation using a lossless entropy code. This approach is called transform coding because of the central role of the transformation.
[0134] For example, JPEG uses the discrete cosine transform on blocks of pixels, while JPEG2000 uses multiscale orthogonal wavelet decomposition. Typically, the three components of a transform coding method - the transform, the quantizer, and the entropy coder - are optimized separately (often through manual parameter tuning). Modern video compression standards such as HEVC, VVC, and EVC also use the transformed representation to encode the residual signal after prediction. Several transforms such as the discrete cosine and sine transforms (DCT, DST), as well as a manually optimized low-frequency non-separable transform (LFNST), are used for that purpose.
[0135] Latent space: The latent space refers to the feature map generated in the bottleneck layer of the NN. This is shown in the examples shown in FIGS. 7 and 8. In the case of an NN topology where the purpose of the network is to reduce the dimension of the input signal (such as an autoencoder topology), the bottleneck layer typically refers to the layer where the dimension of the input signal is reduced to a minimum. The purpose of reducing the number of dimensions is typically to achieve a more compact representation of the input. Therefore, the bottleneck layer is a layer suitable for compression, and thus, in the case of a video coding application, the bitstream is generated based on the bottleneck layer.
[0136] An autoencoder topology typically consists of an encoder and a decoder connected to each other in the bottleneck layer. The purpose of the encoder is to reduce the dimension of the input and make it more compact (or more intuitive). The purpose of the decoder is to reverse the operation of the encoder and thus reconstruct the input as best as possible based on the bottleneck layer.
[0137] Variational autoencoder (VAE) framework The VAE framework can be considered a non-linear transform coding model. The transform process can be mainly divided into four parts. This is illustrated in FIG. 9 showing the VAE framework.
[0138] The conversion process can be split into four parts. Figure 9 illustrates a VAE framework including encoder and decoder branches. In Figure 9, the encoder 901 maps the input image x to a latent representation (denoted by y) via the function y = f(x). This latent representation may also be referred to as a part or a point within the "latent space" hereinafter. The function f() is a conversion function that converts the input signal x into a more compressible representation y. 〔以下、^付きの文字は^を前に付けることで表すことがある〕 . The quantizer 902 converts the latent representation y into a quantized latent representation ^y with (discrete) values by ^y = Q(y). Here, Q represents the quantization function. The quantization function may be RDOQ. The entropy model or hyper encoder / decoder (also known as the hyper prior) 903 estimates the distribution of the quantized latent representation ^y in order to obtain the minimum rate achievable by lossless entropy source coding.
[0139] The latent space can be understood as a representation of compressed data (e.g., picture data) where similar data points are closer to each other within the latent space. The latent space is useful for learning data features and for finding a simpler representation of the data for analysis.
[0140] The quantized latent representations T, ^y and the side information ^z of the hyper prior 903 are included (binarized) in the bitstream 2 using arithmetic coding (AE) as shown in Figure 9.
[0141] Furthermore, a decoder 904 is provided that converts the quantized latent representation into the reconstructed image ^x. ^x = g(^y). The signal ^x is an estimate of the input image x. It is desirable for x to be as close as possible to ^x, in other words, for the reconstruction quality to be as high as possible. However, the higher the similarity between ^x and x, the greater the amount of side information that needs to be transmitted. Such side information includes bitstreams 1 and 2 shown in FIG. 9, which are generated by the encoder and transmitted to the decoder. Usually, the greater the amount of side information, the higher the reconstruction quality. However, a large amount of side information means a low compression ratio. Therefore, one objective of the system described in FIG. 9 is to balance the reconstruction quality and the amount of side information transmitted in the bitstream.
[0142] In FIG. 9, component AE 605 is an arithmetic encoding module that converts samples of the quantized latent representation ^y and side information ^z into a binary representation bitstream 1. Samples of ^y and ^z can include, for example, integers or floating-point numbers. One objective of the arithmetic encoding module is to convert the sample values into a string of binary digits (via the binarization process) (the string of binary digits is then included in the bitstream, which may include additional portions corresponding to the encoded image or further side information).
[0143] Arithmetic decoding (AD) 906 is a process that reverses the binarization process, converting the binary digits back to the original sample values. Arithmetic decoding is provided by arithmetic decoding module 906.
[0144] Note that the present disclosure is not limited to this particular framework. Furthermore, the present disclosure is not limited to image or video compression and can also be applied to object detection, image generation, and recognition systems.
[0145] In Figure 9, there are two sub-networks connected to each other. A network in this context is a logical division between parts of the whole network. For example, in Figure 9, the processing units (modules 901, 902, 904, 905, 906) are called an autoencoder / decoder or simply an "encoder / decoder" network. In other words, a network can be defined using processing units (modules) connected to enable functions. For the connected modules 901, 902, 904, 905, 906 in Figure 9, each network performs the function of encoding and decoding the input picture x (for example, an input tensor). Thus, the "encoder / decoder" network (the first network) in the example of Figure 9 is responsible for encoding (generating) and decoding (parsing) the first bitstream "bitstream 1". On the other hand, the connected processing units (modules) 903, 908, 909, 910, and 907 form another network (the second network), which may be called a "hyper encoder / decoder" network. The second network is responsible for encoding (generating) and decoding (parsing) the second bitstream "bitstream 2". In the example of Figure 9, the first bitstream contains the encoded picture data ^y, and the second bitstream contains the side information ^z. Thus, the purposes of the two networks are different. Any of the processing units (modules) of the first and second networks may itself be a network called a sub-network, that is, a particular module is part of a (larger) network. For example, any of the modules 901, 902, 904, 905, and 906 in Figure 9 is a sub-network of the first network. Similarly, any of the modules 903, 908, 909, 910, and 907 is a sub-network of the second network. Within each network, each of the processing units (modules) performs specific functions as needed to realize the processing of the entire first and second networks respectively.In the example of FIG. 9, the functions are encoding-decoding processing of picture data (first network) and encoding-decoding processing of side information (second sub-network). Further, the connected modules 901, 902, and 905 may be regarded as an encoder sub-network (i.e., a sub-network of the encoder-decoder network), and the modules 904 and 906 may be regarded as a decoder sub-network. As is clear from the above description, the first and second networks may each be interpreted as a sub-network with respect to the entire network including all processing units.
[0146] ■The first sub-network is responsible for the following: ●Conversion 901 of the input image x into its latent representation y (more compressible than x) ●Quantizing 902 the latent representation y into a quantized latent representation ŷ ●Using AE by the arithmetic encoding module 905 to compress the quantized latent representation ŷ to obtain a bitstream "bitstream 1" ●Parsing the bitstream 1 via AD using the arithmetic decoding module 906 ●Reconstructing 904 the reconstructed image (x̂) using the parsed data.
[0147] The purpose of the second sub-network is to obtain the statistical characteristics (e.g., mean value, variance, and correlation between samples of the bitstream 1) of the samples of the "bitstream 1" so that the compression of the bitstream 1 by the first sub-network becomes more efficient. The second sub-network generates a second bitstream "bitstream 2" including the information (e.g., mean value, variance, and correlation between samples of the bitstream 1).
[0148] The second network includes an encoding part that converts the quantized latent representation ^y into side information z, quantizes the side information z into quantized side information ^z, and encodes (e.g., binarizes) the quantized side information ^z into bitstream 2. In this example, binarization is performed by arithmetic encoding (AE). The decoding part of the second network includes arithmetic decoding (AD) 910 that converts the input bitstream 2 into decoded quantized side information ^z'. Since arithmetic encoding and decoding are lossless compression methods, ^z' may be the same as ^z. The decoded quantized side information ^z' is then converted 907 into decoded side information ^y'. ^y' represents the statistical characteristics of ^y (e.g., the mean value of ^y samples or the variance of sample values, etc.). The decoded latent representation ^y' is then provided to the arithmetic encoder 905 and arithmetic decoder 906 described above to control the probability model of ^y.
[0149] Figure 9 illustrates an example of a VAE (Variational Autoencoder), and its details may vary in different implementations. For example, in a particular implementation, there may be additional components to more efficiently obtain the statistical characteristics of the samples in bitstream 1. In one such implementation, there may be a context modeler targeted at extracting the mutual correlation information of bitstream 1. The statistical information provided by the second subnetwork may be used by the AE (Arithmetic Encoder) 905 and AD (Arithmetic Decoder) 906 components.
[0150] Figure 9 depicts the encoder and decoder in a single figure. As will be apparent to those skilled in the art, the encoder and decoder may be embedded in different devices, as illustrated in Figures 9A and 9B, and very often are.
[0151] Figure 9A shows the encoder component of the VAE framework alone, and Figure 9B shows the decoder component of the VAE framework alone. As input, the encoder receives a picture (picture data) according to some embodiments. The input picture may include one or more channels such as color channels or other types of channels, such as depth channels or motion information channels, etc. The output of the encoder (as shown in Figure 9A) is bitstream 1 and bitstream 2. Bitstream 1 is the output of the first subnetwork of the encoder, and bitstream 2 is the output of the second subnetwork of the encoder.
[0152] Similarly, in Figure 9B, two bitstreams, bitstream 1 and bitstream 2, are received as input, and the reconstructed (decoded) image ^x is generated at the output.
[0153] As described above, the VAE can be divided into different logical units that perform different actions. This is illustrated in Figures 9A and 9B, where Figure 9A shows the components participating in the encoding of a signal such as a video and the encoded information provided. This encoded information is then received by the decoder component shown in Figure 9B, for example, for decoding. Thus, the same reference numerals in Figures 9, 9A, and 9B indicate that each processing unit (module) performs the same function.
[0154] Specifically, as seen in Figure 9A, the encoder includes an encoder 901 that converts the input x into a signal y, and the signal y is then provided to a quantizer 902. The quantizer 902 provides information to an arithmetic coding module 905 and a hyperencoder 903. The hyperencoder 903 provides the bitstream 2 already described above to a hyperdecoder 907, and the hyperdecoder 907 signals the information to an arithmetic encoding module 605.
[0155] The output of the arithmetic encoding module is bitstream 1. Bitstream 1 and bitstream 2 are the outputs of the encoding of the signal, and these are then provided (sent) to the decoding process.
[0156] Unit 901 is called an "encoder", but it is also possible to call the complete subnetwork described in FIG. 9A an "encoder". The encoding process generally means a unit (module) that converts an input into an encoded (e.g., compressed) output. From FIG. 9A, it can be seen that unit 901 actually converts the input x into y, which is a compressed version of x, so it can be regarded as the core of the entire subnetwork. The compression in encoder 901 may be achieved, for example, by applying a neural network or generally any processing network having one or more layers. In such a network, the compression may be performed by a cascaded process (i.e., sequential processing) that includes downsampling to reduce the size and / or number of input channels. Thus, the encoder may be called, for example, a neural network (NN)-based encoder.
[0157] The remaining parts in the figure (quantization unit, hyperencoder, hyperdecoder, arithmetic encoder / decoder) are all parts responsible for improving the efficiency of the encoding process or converting the compressed output y into a series of bits (bitstream). Quantization may be provided to further compress the output of the NN encoder 901 by lossy compression. AE 905 combined with the hyperencoder 903 and hyperdecoder 907 used to form AE 905 may perform binarization that can further compress the quantized signal by lossless compression. Therefore, it is also possible to call the entire subnetwork in FIG. 9A an "encoder". The same is true for FIG. 9B, and it is also possible to call the entire subnetwork a "decoder".
[0158] Most deep learning (DL)-based image / video compression systems reduce the dimensionality of the signal before converting the signal into binary digits (bits). For example, in the VAE framework, the encoder, which is a non-linear transformation, maps the input image x to y, where y has a smaller width and height than x. Since y has a smaller width and height and thus a smaller size, the dimensionality (size) of the signal is reduced, and thus it is easier to compress the signal y. Note that generally, the encoder does not necessarily need to reduce the size of both (or generally all) dimensions. Rather, some exemplary implementations may provide an encoder that reduces the size only in one dimension (or generally, a subset of dimensions).
[0159] The general principle of compression is illustrated in FIG. 8. The latent space, which is the output of the encoder and the input of the decoder, represents the compressed data. Note that the size of the latent space can be much smaller than the input signal size. Here, the term size can refer to the resolution, e.g., the number of samples of the feature map(s) output by the encoder. The resolution may be given as the product of the number of samples per dimension (e.g., width × height × number of channels of the input image or feature map).
[0160] The reduction in the size of the input signal is illustrated in FIG. 8, which represents a deep learning-based encoder and decoder. In FIG. 8, the input image x corresponds to the input data that is input to the encoder. The transformed signal y corresponds to the latent space, which has a smaller number of dimensions or at least a smaller size in one dimension than the input signal. Each column of circles represents a layer in the processing chain of the encoder or decoder. The number of circles in each layer indicates the size or dimensionality of the signal in that layer.
[0161] From FIG. 8, it can be seen that the encoding operation corresponds to reducing the size of the input signal, and the decoding operation corresponds to reconstructing the original size of the image.
[0162] One way to reduce the signal size is downsampling. As described above, downsampling is a process in which the sampling rate of the input signal is reduced. For example, if the input image has sizes h and w and the output of the downsampling is h2 and w2, at least one of the following holds. h2 < h w2 < w
[0163] The reduction of the signal size usually occurs step by step along the chain of processing layers, rather than all at once. For example, if the input image x has dimensions (indicating height and width) h and w and the latent space y has dimensions h / 16 and w / 16, size reduction may occur in four layers during encoding, and each layer reduces the signal size by a factor of 2 in each dimension.
[0164] Some deep learning-based video / image compression methods use multiple downsampling layers. As an example, the VAE framework shown in Figure 6A utilizes six downsampling layers marked as 601 - 606.
[0165] The layer containing downsampling is indicated by a downward arrow ↓ in the layer description. The layer description "Conv Nx5x5 / 2↓" means that the layer is a convolutional layer with N channels and the convolutional kernel has a size of 5x5. As described above, 2↓ means that downsampling by a factor of 2 is performed in this layer. As a result of the downsampling by a factor of 2, one of the dimensions of the input signal is reduced to half in the output. In Figure 6A, 2↓ indicates that both the width and height of the input image are reduced to half. Since there are six downsampling layers, if the width and height of the input image 814 (also denoted as x) are given by w and h, the output signal ^z 813 has a width and height equal to w / 64 and h / 64, respectively.
[0166] The modules indicated by AE and AD are the arithmetic encoder and arithmetic decoder that have already been described above with respect to FIGS. 9, 9A, and 9B. The arithmetic encoder and decoder are specific implementations of entropy coding. AE and AD (as part of components 613 and 615 in FIGS. 6A and 6B) may be replaced by other means of entropy coding. In information theory, entropy encoding is a lossless data compression method used to convert the value of a symbol into a binary representation, which is a reversible process. Also, "Q" in the figure corresponds to the quantization operation described above in connection with FIGS. 6A and 6B and is further described in the "Quantization" section above. Also, the quantization operation and the corresponding quantization unit as part of component 613 or 615 do not necessarily exist and / or may be replaced by another unit.
[0167] FIGS. 6A and 6B also show a decoder including upsampling layers 607 to 612. Between upsampling layers 611 and 610 in the input processing order, a further layer 620 implemented as a convolutional layer is provided that does not provide upsampling to the received input. The corresponding convolutional layer 620 is also shown for the decoder. Such a layer may be provided in the NN to perform an operation on the input that changes certain characteristics without changing the size of the input. However, it is not necessary to provide such a layer.
[0168] Looking at the processing order of bitstream 2 through the decoder, the upsampling layers are executed in the reverse order, i.e., from upsampling layer 612 to upsampling layer 607. Here, each upsampling layer is shown to provide upsampling with an upsampling ratio 2 indicated by ↑. Of course, it is not necessarily the case that all upsampling layers have the same upsampling ratio, and other upsampling ratios such as 3, 4, 8, etc. can also be used. Layers 607 - 612 are implemented as convolutional layers (conv). Specifically, since they can be intended to provide an operation opposite to that of the encoder for the input, the upsampling layer may apply a deconvolution operation to the received input so that its size is increased by a factor corresponding to the upsampling ratio. However, the present disclosure is generally not limited to deconvolution, and upsampling can be performed in any other way, such as by bilinear interpolation between two neighboring samples, or by nearest neighbor sample copy, or the like.
[0169] In the first subnetwork, after several convolutional layers (601 - 603), generalized divisive normalization (GDN) follows on the encoder side and inverse GDN (IGDN) follows on the decoder side. In the second subnetwork, the activation function applied is ReLU. It should be noted that the present disclosure is not limited to such an implementation, and generally, other activation functions can be used instead of GDN or ReLU.
[0170] Figure 6B shows another example of a VAE-based encoder-decoder structure similar to that of Figure 6A. In Figure 6B, it is shown that the encoder and decoder can include several downsampling layers and upsampling layers. Each layer applies downsampling by a factor of 2 or upsampling by a factor of 2. Further, the encoder and decoder can include additional components such as general divisive normalization (GDN) 650 on the encoder side and inverse GDN (IGDN) 655 on the decoder side. Further, both the encoder and decoder can include one or more ReLUs, specifically leaky ReLUs 660 and 665. Also, a factored entropy model can be provided for the encoder and a Gaussian entropy model 670 for the decoder. Also, a plurality of convolutional masks 680 may be provided. Further, in the embodiment of Figure 6B, the encoder includes a universal quantizer (UnivQuan) and the decoder includes an attention module.
[0171] The total number of strides and downsampling operations defines a condition regarding the input channel size, i.e., the size of the input to the neural network.
[0172] Here, when the input channel size is an integer multiple of 64 = 2×2×2×2×2×2, the channel size remains an integer after all ongoing downsampling operations. By applying corresponding upsampling operations in the decoder during upsampling and applying the same rescaling at the end of the processing of the input through the upsampling layer, the output size becomes the same as the input size in the encoder again.
[0173] Thereby, a reliable reconstruction of the original input is obtained.
[0174] Receptive field: Within the context of a neural network, the receptive field is defined as the size of the region in the input that produces a sample in the output feature map. Basically, this is a measure of the association between the output feature (of any layer) and the input region (patch). Note that the concept of receptive field applies to local operations (i.e., convolution, pooling, etc.). For example, a convolution operation using a kernel of size 3×3 has a receptive field of 3×3 samples in the input layer. In this example, nine input samples are used by the convolution node to obtain one output sample.
[0175] Total receptive field: The total receptive field (TRF) refers to the set of input samples that are used to obtain a specified set of output samples, for example, by applying one or more processing layers of a neural network.
[0176] The total receptive field can be illustrated by Figure 10. Figure 10 illustrates the processing of a one-dimensional input (seven samples on the left side of the figure) having two consecutive transposed convolution (also called deconvolution) layers. The input is processed from left to right, i.e., the "deconv layer 1" processes the input first, and its output is processed by the "deconv layer 2". In this example, the kernel has a size of 3 in both deconvolution layers. This means that three input samples are required to obtain one output sample in each layer. In this example, the set of output samples is marked inside the dashed rectangle and contains three samples. Due to the size of the deconvolution kernel, seven samples in the input are required to obtain an output set of samples containing three output samples. Thus, the total receptive field of the three marked output samples is seven samples in the input.
[0177] In FIG. 10, there are seven input samples, five intermediate output samples, and three output samples. The reduction in the number of samples is due to the fact that there are "missing samples" at the boundaries of the input because the input signal is finite (it does not extend infinitely in each direction). That is, since the deconvolution operation requires three input samples corresponding to each output sample, when the number of input samples is seven, only five intermediate output samples can be generated. In fact, the amount of output samples that can be generated is (k - 1) samples less than the number of input samples, where k is the kernel size. In FIG. 10, since the number of input samples is seven, after the first deconvolution using a kernel size of three, the number of intermediate samples is five. After the second deconvolution using a kernel size of three, the number of output samples is three.
[0178] As can be seen from FIG. 10, the total receptive field of the three output samples is the seven samples in the input. The size of the total receptive field increases by successively applying processing layers with kernel sizes greater than one. In general, the total receptive field of a set of output samples is calculated by starting from the output layer, tracing the connections of each node up to the input layer, and then finding the union of all the samples in the input that are directly or indirectly (through more than one processing layer) connected to the set of output samples. For example, in FIG. 10, each output sample is connected to three samples in the previous layer. The union includes the five samples in the intermediate output layer, which are connected to the seven samples in the input layer.
[0179] Sometimes, it is desirable to keep the number of samples the same after each operation (convolution, deconvolution, or others). In such cases, padding can be applied to the boundaries of the input to compensate for the "missing samples". FIG. 11 shows this case when the number of samples is kept equal. Note that padding is not an essential operation for convolution, deconvolution, or any other processing layer, so the present disclosure is applicable to both cases.
[0180] This should not be confused with downsampling. In the process of downsampling, for every M samples, there are N samples in the output, where N < M. The difference is that M is typically much smaller than the number of inputs. In Figure 10, there is no downsampling; rather, the reduction in the number of samples results from the fact that the input size is not infinite and there are "missing samples" in the input. For example, if the number of input samples was 100, since the kernel size is k = 3, when two convolutional layers are used, the number of output samples would have been 100 - (k - 1) - (k - 1) = 96. In contrast, if both transposed convolutional layers were performing downsampling (at a ratio of M = 2 and N = 1), the number of output samples would have been
Number
[0181] Figure 12 illustrates downsampling using two convolutional layers with a downsampling ratio of 2 (N = 1 and M = 2). In this example, seven input samples become three due to the combined effect of downsampling and "missing samples" at the boundaries. The number of output samples can be calculated after each processing layer using the formula ceil((100 - (k - 1)) / r), where k is the kernel size and r is the downsampling ratio.
[0182] The operations of convolution and transposed convolution (i.e., deconvolution) are identical from a mathematical representation perspective. The difference comes from the fact that the deconvolution operation assumes that the previous convolution operation has been performed. In other words, deconvolution is a process of filtering the signal to compensate for the previously applied convolution. The purpose of deconvolution is to reproduce the signal that existed before the convolution was performed. This disclosure applies to both the convolution operation and the deconvolution operation (in fact, any other operation where the kernel size is greater than 1, as will be explained later).
[0183] FIG. 13 shows another example for explaining how to calculate the overall receptive field. In FIG. 13, a two-dimensional input sample array is processed by two convolutional layers each having a kernel size of 3×3. After applying two deconvolutional layers, an output array is obtained. The set (array) of output samples is marked by a solid rectangle (“output sample”) and contains 2×2 = 4 samples. The overall receptive field of this set of output samples contains 6×6 = 36 samples. The overall receptive field can be calculated as follows. · Each output sample is connected to 3×3 samples in the intermediate output. The union of all samples in the intermediate output connected to the set of output samples contains 4×4 = 16 samples. · Each of the 16 samples in the intermediate output is connected to 3×3 samples in the input. The union of all samples in the input connected to the 16 samples in the intermediate output contains 6×6 = 36 samples. Thus, the overall receptive field of the 2×2 output samples is 36 samples in the input.
[0184] In image and video compression systems, the compression and decompression of input images having a very large size are typically performed by dividing the input image into multiple parts. VVC and HEVC, for example, adopt such a division method by dividing the input image into tiles or wavefront processing units.
[0185] When tiles are used in traditional video coding systems, the input image is typically divided into multiple rectangular-shaped parts. FIG. 14 illustrates one such division. In FIG. 14, part 1 and part 2 may be processed independently of each other, and the bitstream for decoding each part is encapsulated into independently decodable units. As a result, the decoder can independently parse (obtain the syntax elements necessary for sample reconstruction) each bitstream (corresponding to part 1 and part 2) and can also independently reconstruct the samples of each part.
[0186] In the wavefront parallel processing shown in FIG. 15, each part typically consists of one row of coding tree blocks (CTBs). The difference between wavefront parallel processing and tiling is that in wavefront parallel processing, the bitstreams corresponding to each part can be decoded almost independently of each other. However, since the sample reconstruction of each part still has dependencies between parts, the sample reconstruction cannot be executed independently. In other words, wavefront parallel processing makes the parse process independent while keeping the sample reconstruction dependent.
[0187] Both wavefront and tiling are techniques that enable the execution of all or part of the decoding operation independently of each other. The advantages of independent processing are as follows. · Two or more identical processing cores can be used to process the entire image. Therefore, the processing speed can be increased. · If the capabilities of the processing cores are not sufficient to process large images, the image can be divided into multiple parts that require fewer resources for processing. In this case, even if a less capable processing unit cannot process the entire image due to resource limitations, it can process each part.
[0188] To meet the requirements for processing speed and / or memory, HEVC / VVC uses a processing memory large enough to process the encoding / decoding of the entire frame. To achieve this, the top - of - the - line GPU cards are used. In the case of traditional codecs such as HEVC / VVC, since the entire frame is divided into blocks and each block is processed one by one, the memory requirements for processing the entire frame are usually not a major problem. However, the processing speed is a major concern. Therefore, when a single processing unit is used to process the entire frame, the speed of the processing unit must be very high, and thus, the processing unit is usually very expensive.
[0189] On the one hand, the NN-based video compression algorithm considers the entire frame in encoding / decoding instead of the block-based approach in the conventional hybrid coder. The memory requirements are too high to be processed through the NN-based encoding / decoding module.
[0190] In traditional hybrid video encoders and decoders, the amount of memory required is proportional to the maximum allowable block size. For example, in VVC, the maximum block size is 128×128 samples.
[0191] However, the memory required for NN-based video compression is proportional to the size W×H, where W and H represent the width and height of the input / output image. Since typical video resolutions include 3840×2160 picture size (4K video), it can be seen that the memory requirements can be very high compared to hybrid video coders. In image and video compression systems, the compression and decompression of input images with very large sizes are typically performed by dividing the input image into multiple parts. To address memory constraints, the NN-based video coding algorithm can apply tiling in the latent space.
[0192] When tiling is applied without overlap, boundary artifacts may be visible in the reconstructed image. This problem may be partially solved by overlapping the tiles in the latent space or the signal region, where the overlap is large enough to avoid these artifacts. If the overlap is larger than the size of the receptive field of the NN, the operations can be performed in a non-standard way and may cause some overhead in the computational complexity. Next, if the overlap is smaller than the size of the receptive field, the tiling operation is not lossless / transparent and needs to be specified (canonical).
[0193] Furthermore, in the case of a structure having multiple pipelines (e.g., pipelines for processing luma and / or chroma, or generally multiple channels of an input tensor) or multiple sub-networks, a simple approach where tiles always have the same size can result in a performance loss. Rate-Distortion Optimization Quantization (RDOQ) (e.g., unit 908 in FIG. 9) is another computationally complex and memory-intensive operation. It may not be optimal to select the same tile size within RDOQ for all sub-networks. Also, if the scene representation in different pipelines is not sample-aligned during tiling (e.g., CCS), simple tiling may not preserve information regarding the correlations along different components. Sample-alignment means that the size of luma does not match the size of chroma. In such cases, video processing may be less performant as there is no or at least less significant spatial and / or temporal correlation, and as a result, the quality of the reconstructed image may degrade.
[0194] Another problem is that, in order to process a large input with a single processing unit (e.g., a CPU or GPU), the processing unit has to be very fast as it needs to execute a large number of operations per unit time. This requires the unit to have a high clock frequency and a high memory bandwidth, which are expensive design criteria for chip manufacturers. In particular, due to physical limitations, it is not easy to increase the memory bandwidth and the clock frequency.
[0195] Although the latest deep learning-based image and video compression algorithms follow the Variational Autoencoder (VAE) framework, NN-based video coding algorithms for encoding and / or decoding are still in the initial development stage, and there are no consumer devices that include the VAE implementations shown in FIGS. 9, 9A, and 9B. Also, the cost of consumer devices is very sensitive to the memory implemented.
[0196] For an NN-based video coding algorithm to be cost-effective for implementation in consumer devices such as mobile phones, it is therefore necessary to reduce the memory footprint and the required operating frequency of the processing unit. Such optimizations have not yet been performed.
[0197] The present disclosure is applicable to both end-to-end AI codecs and hybrid AI codecs. In a hybrid AI codec, for example, filtering operations (filtering of the reconstructed picture) can be performed by a neural network (NN). The present disclosure applies to such NN-based processing modules. Generally, the present disclosure can be applied to all or part of the video compression and decompression processes when at least part of the process includes an NN and such an NN includes a convolution operation or a transposed convolution operation. For example, the present disclosure is applicable to individual processing tasks that are executed as part of the processing by an encoder and / or a decoder, including in-loop filtering, post-filtering, and / or pre-filtering, and rate-distortion optimization quantization (RDOQ) for the encoder only.
[0198] Some embodiments of the present disclosure can provide a solution to the above problems with respect to enabling a trade-off between memory resources and computational complexity within an NN-based video encoding-decoding framework. In particular, the present disclosure provides the possibility of processing parts of the input independently while still providing that different image components are sample-aligned. This reduces the memory requirements while maintaining the compression performance and has little additional computational complexity.
[0199] The processing can be decoding or encoding. The neural network (NN) in the exemplary implementation described below may be any of the following. ·A network including at least one processing layer in which two or more input samples are used to obtain an output sample (this is a general condition when the problems addressed in the present disclosure occur). ·A network including at least one convolutional (or transposed convolutional) layer. In one example, the convolutional kernel is larger than 1. ·A network including at least one pooling layer (such as max pooling, average pooling, etc.). ·A decode network, a hyper-decoder network, or an encode network. ·A part of the above (sub-network).
[0200] The input may be as follows. ·A feature map. ·The output of a hidden layer. ·A latent space feature map. The latent space can be obtained according to a bitstream. ·An input image.
Example
[0201] The first embodiment Hereinafter, as shown in FIG. 16, a method and apparatus for picture / video encoding-decoding (compression-decompression) are described in which a plurality of sub-networks are used to process an input tensor representing picture data.
[0202] In this exemplary and non - limiting embodiment, a method for encoding an input tensor representing picture data is provided. The input tensor can have a matrix form with width = w, height = h, and a third dimension (e.g., the number of channels) equal to D in the spatial dimension. For example, the input tensor can be the input image itself having D components including one or more color components and possibly additional channels such as depth channels or motion channels. However, the present disclosure is not limited to such inputs. Generally, the input tensor can be a potential representation that can be the result of a previous process (e.g., pre - processing), which is a representation of picture data.
[0203] The input tensor is processed by a neural network including at least a first sub - network and a second sub - network. Examples of the first and / or second sub - networks for the encoding branch are shown in FIG. 16 and include an encoder 1601 and a rate - distortion optimized quantizer (RDOQ) 1602. The processing includes applying the first sub - network to a first tensor, which includes dividing the first tensor into a first plurality of tiles in the spatial dimension and processing the first plurality of tiles by the first sub - network. After applying the first sub - network, the second sub - network is applied to a second tensor, which includes dividing the second tensor in the spatial dimension into a second plurality of tiles and processing the second plurality of tiles by the second sub - network.
[0204] In the example of FIG. 16, the output of the first sub-network 1601 is provided as input to a second sub-network, which is the RDOQ 1602. In this case, the first tensor is the input image x representing picture data, which may be raw picture data. Next, the second tensor input to the second sub-network is the feature tensor in the latent space. However, the present disclosure is not limited to the case where the first sub-network and the second sub-network are directly cascade-connected. Generally, the second sub-network is after the first sub-network, that is, applied after applying the first network, but there may be some additional processing between the first sub-network and the second sub-network. Thus, the above term "after" does not limit the above processing to immediately after in the sense that the output of the first sub-network is directly input to the second sub-network. Rather, "after" means that a plurality of first and second tiles are processed, for example, within the same processing pipeline.
[0205] One or more channels of the first and second tensors are divided into so-called tiles, which basically represent data obtained by dividing the input tensor in one or more spatial dimensions. Like current video coding standards, tiles are intended to provide the possibility of being decoded in parallel, that is, independently of each other. A tile may include one or more samples. A tile may have a rectangular shape, but is not limited to such a regular shape. The rectangular shape may be a square shape. An example of dividing into regularly shaped tiles is shown in FIG. 14. However, the present disclosure is not limited to rectangles, especially squares. For example, it may have a checkerboard or irregular shape. The division of the first and second input tensors may be performed such that the tiles have a triangular or any other shape, which may depend on the type of processing performed by a particular application and / or sub-network.
[0206] In FIG. 16, the general processing of N components is shown in sub-networks 1601 and 1602. Here, the components can be input tensor channels (e.g., color components or latent space representations) that can be processed in parallel. Such processing can include dividing the channels into tiles (including a first plurality of tiles) in the spatial region.
[0207] However, the N components in FIG. 16 may correspond to N tiles (or tile groups) of each of one or more channels. The N tiles (or tile groups) can be processed in parallel. After such processing of the first and / or second plurality of tiles, each sub-network can merge the processed tiles into an output tensor. In FIG. 16, such a merged output tensor is y for sub-network encoder 1601 and ^y for sub-network RDOQ 1602. Note that in FIG. 16, the number of components in the first sub-network 1601 and the second sub-network 1602 is the same (N). However, this is not necessarily the case. The number of parallel processing pipes in the first sub-network may be different from the number of parallel processing pipes in the second sub-network. For example, in the case of post-filtering, the first sub-network is a post-filter that processes each component (e.g., Y, U, and V) separately. Next, in the case of encoding or decoding processing, the encoder or decoder is the second sub-network respectively, and a difference is made between Y and UV. That is, UV is processed identically.
[0208] Furthermore, at least two co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes. Here, the term "co-located" means that two tiles (i.e., one from the first plurality and one from the second plurality) are in corresponding (e.g., at least partially overlapping) positions within the first and second input tensors in the spatial dimensions. In other words, the tiling into tiles may be different for each subnetwork, and thus each subnetwork may use a different tiling (a different partitioning into tiles).
[0209] In one exemplary implementation, the inputs of each subnetwork (i.e., the first and second tensors) can have a smaller size, except for the tiles at the lower and right image boundaries, since the input tensors do not necessarily have a size that is an integer multiple of the tile size, and are (in the spatial region) divided into a grid of tiles of the same size. Such a grid with tiles of the same size can be advantageous due to the possibility of efficient signaling of such a grid in the bitstream and the low processing complexity. On the other hand, a grid that can include tiles of different sizes can result in better performance and content adaptability.
[0210] In the above exemplary implementation, the tiles of the first plurality of tiles adjacent in at least one of the spatial dimensions partially overlap in the at least one spatial dimension. Additionally or alternatively, the tiles of the second plurality of tiles adjacent in at least one of the spatial dimensions partially overlap in the at least one spatial dimension. The term "adjacent" means that the respective tiles are neighboring. Adjacent tile portions 1 and 2 are shown in FIG. 14, which are adjacent but do not overlap. FIG. 17A shows a partial overlap, where the first tensor is, for example, 2D in the x - y plane with 4 tiles (i.e., regions L 1 , L 2 , L 3 , and L 4) is divided into. Similar considerations may be applicable to the second input tensor. The first tile of the first plurality of tiles is L1, and the second tile is L2. L1 and L2 are adjacent to each other in the x-axis direction and have overlapping boundaries along the y-axis. As shown, L1 and L2 partially overlap. Partially overlapping means that tiles L1 and L2 contain one or more of the same tensor elements. The tensor elements may, in some embodiments, correspond to picture samples for the first input tensor.
[0211] FIG. 17B shows a scenario of partial overlap in the x-axis and y-axis directions, the same as FIG. 17A. In both FIGS. 17A and 17B, L1 further overlaps with the adjacent tile on the right (in the y dimension; the boundaries overlap along the x dimension). L1 also slightly overlaps with its diagonally adjacent tile (in both dimensions). Similarly, L2 overlaps with both its directly adjacent tile on the right and its diagonally adjacent tile. FIG. 18 shows another example of partial overlap for tiles L1 and L2 that may be tiles of the first plurality of tiles. L1 and L2 have overlaps with other tiles only in one of their respective dimensions. L1 partially overlaps with the adjacent tile on the right, and the boundary is along the x-axis (i.e., partial overlap in the y dimension). Next, L2 partially overlaps with its adjacent tile L1 at the top. In the examples of FIGS. 17A and 17B, the overlap between L1 and L2 means that L1 includes samples of L2 and L2 includes samples from L1. In FIG. 18, L2 includes samples from L1, but L1 does not include samples from L2. As will be apparent to those skilled in the art, further variations of the overlap may exist. The present disclosure is not limited to a particular manner or extent of overlap. Also, a tile arrangement without overlap may be included.
[0212] FIG. 19 is a further example of L1 and L2 with no overlap with each other or with other tiles. FIG. 20 is the same as FIGS. 17A and 17B, and further details regarding the partial overlap region will be further described below.
[0213] In one exemplary implementation, tiles of the first plurality of tiles (such as L1 and L2) are processed independently by the first subnetwork. Additionally or alternatively, tiles of the second plurality of tiles are processed independently by the second subnetwork. In other words, the processing of the tiles is independent of each other, and thus, mutually independent. The independent processing provides the possibility of parallelization. For example, in some implementations, at least two tiles of the first plurality of tiles are processed in parallel by the first subnetwork and / or at least two tiles of the second plurality of tiles are processed in parallel by the second subnetwork. The parallel processing is shown in FIG. 16 and includes processing 1 to processing N (processing pipelines 1 to N) in the subnetwork encoder 1601 or the quantization RDOQ 1602. Using the encoder subnetwork 1601 as the first subnetwork, the encoder 1601 receives an input tensor x that is divided into N tiles x1 to xN of the first input tensor. Each tile is processed by respective blocks, i.e., processing 1 to processing N that do not need to interact with each other (e.g., wait for each other during processing). The result of each processing is an output tensor y1 to yN for each tile, which may be a feature map in the latent space. The output tensors y1 to yN may be further combined into an output tensor y. This combination may (but does not have to) include cropping as shown in FIGS. 17 to 19. Note that the combination into the tensor y does not have to be executed. It is conceivable that the second subnetwork reuses the tiling of the first subnetwork and simply modifies it (makes the tiling finer by further dividing the tiles or coarser by combining multiple tiles into one). In the example of FIG. 16, processing 1 to N may perform processing on a tile-by-tile basis (i.e., processing i processes tile i). Alternatively, processing i may process component i of a plurality of components of the input tensor. In this case, processing i divides component i into a plurality of tiles and processes the tiles separately or in parallel.
[0214] In the above, for simplicity, parallel processing of all N input tensor tiles by each of the N processing pipes was illustrated. However, the present disclosure is not limited to such parallel processing. There may be more than N tiles in the input tensor, and they are divided into N tile groups that are processed in parallel within each of the N processing pipes (instances of the first subnetwork and / or instances of the second subnetwork). As will be apparent to those skilled in the art, once the tiles are independent of each other, their processing can, in principle, be parallelized. Those skilled in the art can design any number of parallel processing pipes according to their respective performance requirements and / or hardware availability.
[0215] As shown in FIGS. 17A / B to 20, each tile has a certain size, which may have a certain width and a certain height, and they may be different from each other in the case of rectangular-shaped tiles, or the same in the case of square-shaped tiles. In one implementation, dividing the first tensor includes determining the tile sizes within the first plurality of tiles based on a first predetermined condition, and / or dividing the second tensor includes determining the tile sizes within the second plurality of tiles based on a second predetermined condition. For example, the first predetermined condition and / or the second predetermined condition are based on the available decoder hardware resources and / or the motion present in the picture data. As an example of the available hardware resources, the first and / or second predetermined conditions may be the memory resources of the processing device (decoder or encoder). When the available amount of memory is less than a predetermined value (the amount of memory resources), the determined tile size may be smaller than the tile size when the available memory is equal to or greater than the predetermined value. However, the hardware resources are not limited to memory. The first and / or second conditions may be based on the availability of processing power, for example, the number of processors and / or the processing speed of one or more processors.
[0216] Alternatively or additionally, the presence of motion may be used in the first and / or second conditions. For example, the tile size may be determined to be smaller when there is more motion in the input tensor portion corresponding to the tile, when the presence of motion is less prominent, or when there is no motion at all, compared to when there is less motion. Whether the motion is prominent (i.e., fast motion and / or fast / frequent change in motion) is determined by the respective motion vectors with respect to the change in its magnitude and direction, and may be compared with corresponding predetermined values (thresholds) for magnitude and / or direction and / or frequency.
[0217] An alternative or additional condition may be a region of interest (ROI), where the tile size may be determined based on the presence of the ROI. For example, the ROI in the picture data may be a detected object (e.g., vehicle, bicycle, motorcycle, pedestrian, animal, etc.). The objects may have different sizes, move fast or slow, and / or change the direction of their motion fast or slow and / or several times (i.e., more frequently than a predetermined frequency value). In one exemplary implementation, the size or tile in at least one dimension may be smaller for the ROI than for the rest of the input tensor.
[0218] Thus, the tile size may be adapted or optimized to the hardware resources or to the content of the picture data including scene-specific tile sizes. It is also possible to jointly adapt or optimize the tile size to both the hardware resources and the content of the picture data.
[0219] Figures 17A / B to 20 illustrate splitting a first (or second) tensor into a plurality of tiles that partially overlap with adjacent tiles. As a result of the overlap and / or tile processing, the processed tiles corresponding to the region Ri may be subject to cropping. In Figures 17A / B to 20, the subscripts L and R are the same for the corresponding splits in the input and output. For example, L4 corresponds to R 4 The arrangement of Ri follows the same pattern as Li, which means that if L1 corresponds to the upper left corner of the input space, R1 corresponds to the upper left corner of the output space. If L2 is to the right of L1, R2 is to the right of R1. In the example of FIG. 17A, the division of the first tensor is such that each region L i contains the complete receptive field of each R i respectively. Further, the union of R i constitutes the entire target picture R.
[0220] The determination of the full receptive field depends on the kernel size of each processing layer. This can be determined by tracing back the input samples of the first tensor in the reverse direction of processing. The full receptive field consists of the union of the input samples used in the calculation of all output sample sets. Therefore, the full receptive field depends on the connections between each layer and can be determined by tracing all the connections in the direction from the output to the input starting from the output.
[0221] In the example of the convolutional layer shown in FIG. 13, the kernel sizes of convolutional layers 1 and 2 are K1×K1 and K2×K2 respectively, and the downsampling ratios are R1 and R2 respectively. Convolutional layers typically use regular input-output connections (for example, K×K input samples are always used for each output). In this example, the calculation of the size of the full receptive field can be performed as follows. W = ((w × R2) + K2 - 1) × R1 + (K1 - 1) H = ((h × R2) + K2 - 1) × R1 + (K1 - 1) Here, H and W represent the size of the full receptive field, and h and w are the height and width of the output sample set respectively.
[0222] In the above example, the convolution operation is described in a two-dimensional space. If the number of dimensions of the space to which the convolution is applied is higher, a 3D convolution operation may be applied. The 3D convolution operation is a simple extension of the 2D convolution operation, and an additional dimension is added to all of the operations. For example, the kernel size can be represented as K1×K1×N and K2×K2×N, and the total receptive field can be represented as W×H×N based on the previous example, where N represents the size of the third dimension. Since the extension from the 2D convolution operation and the 3D convolution operation is simple, the present invention is applicable to both the 2D convolution operation and the 3D convolution operation. In other words, the size of the third (or even the fourth dimension) may be greater than 1, and the present invention can be applied in the same manner.
[0223] The above equation is an example showing how the size of the total receptive field can be determined. The determination of the total receptive field depends on the actual input-output connections of each layer. The output of the encoding process is R i is. The union of R i constitutes the feature tensor in which the components y 1 to y N are merged (Figure 16). In this example, since R i has overlapping regions, first a cropping operation is applied to obtain R-crop i without overlapping regions. Finally, R-crop i are concatenated to obtain the merged feature tensor y. In this example, L i includes the total receptive field of R i as described above.
[0224] R i and L i (that is, the size of the first and / or second plurality of tiles) can be determined as follows. · First, determine N non-overlapping regions R-crop i . For example, R-crop i are N×M regions of equal size, and N×M is determined by the decoder according to the memory limit. · R-cropi Determine the entire receptive field of L. i L is set equal to the entire receptive field of each R-crop i respectively. · Process each L i to obtain R i . This means that R i is the size of the output sample set generated by NN. Note that actual processing may not be necessary. Once the size of L i is determined, it may be possible to determine the size and position of R i according to a function. Since the structure of NN is known, the relationship between the size of L i and R i is known. Therefore, without actually performing the processing, R i can be calculated by a function according to L i . · If the size of R i is not equal to that of R-crop i , crop R i to obtain R-crop i . ■ If a padding operation is applied to the input sample or intermediate output sample during processing by NN, the size of R-crop i may not be equal to that of R i . Padding can be applied to NN, for example, when some size requirements for the input and intermediate output samples must be met. For example, NN may require (due to its certain structure) that the input size must be a multiple of 16 samples. In this case, if L i is not a multiple of 16 samples in one direction, padding may be applied in that direction to make it a multiple of 16. In other words, the padded samples are dummy samples to ensure the integer multiplicity of each Li with the size requirements of each NN layer. ■ The cropped-out samples can be obtained by determining "the output samples including the padded samples in that calculation". This option may be applicable to the exemplary implementations discussed.
[0225] This is shown in FIG. 17A, where after processing a first plurality of tiles (i.e., regions Li) via a first sub-network (e.g., encoder 1601 of FIG. 16), the size Ri output by the first sub-network may be too large and thus may not have a suitable size for the input to subsequent sub-networks such as the second sub-network RDOQ 1602 of FIG. 16. Thus, the processed tiles R1 and R2 in FIG. 17A are subject to a cropping operation, and then, in this example, the cropped R-crop1 to R-crop4 are merged.
[0226] FIG. 17B shows an example of partial overlap of regions L1 to L4 (i.e., the first and / or second plurality of tiles) where cropping of the processed tiles (i.e., regions R1 to R4) is not involved. R i and L i can be determined as follows and is sometimes referred to as the "simple" non-cropping case shown in FIG. 17B. · First, determine N non-overlapping regions R i . For example, R i may be N×M regions of equal size, where N×M is determined by the decoder according to memory limitations. · Determine the total receptive field of R i . L i is set equal to the total receptive field of each R i . The total receptive field is calculated by tracing each output sample in R i backward to L i . Thus, L i consists of all samples used in the calculation of at least one of the samples in R i . · Process each L i to obtain R i . This means that R i is the size of the output sample set generated by NN.
[0227] Exemplary implementations solve the all-peak memory problem by dividing the input space into multiple, smaller, independently processable regions.
[0228] In the above exemplary implementation, the overlapping regions of L i require additional processing compared to not dividing the input into regions. The larger the overlapping region, the more additional processing is required. In particular, in some cases, the entire receptive field of R i may be too large. In such cases, the total number of calculations to obtain the entire reconstructed picture may increase too much.
[0229] FIG. 18 shows another example of partially overlapping tiles (i.e., regions L1 - L4), which also includes cropping of the processed tiles R1 - R4, similar to FIG. 17A. Compared to FIGS. 17A and 17B, the input regions L i (i.e., tiles) are smaller here, as they each represent only a subset of the entire receptive field of the respective R i . Each region Li is processed independently by its respective sub-network (e.g., encoder 1601 and / or RDOQ 1602 of FIG. 16), thereby obtaining two regions R1 and R2 (i.e., output subsets). Since a subset of the entire receptive field is used to obtain a set of output samples, a padding operation may be required to generate missing samples. In one exemplary implementation, the processing by the first sub-network of the first plurality of tiles and / or by the second sub-network of the second plurality of tiles may include padding before processing using said one or more layers. Thus, samples missing in the input subset may be added by the padding process, which improves the quality of the reconstructed output subset Ri. Thus, after combining the output subsets Ri, the quality of the reconstructed picture is also improved.
[0230] When thinking about it, padding refers to increasing the size of the input (i.e., the input image) by generating new samples at the boundaries of the image (or picture) by using pre-defined sample values or by using the sample values at that position within the input image. This is shown in Figure 11. The generated samples are approximations of the non-existent actual sample values. Thus, the padded samples may be obtained, for example, based on one or more of the nearest neighbor samples of the sample to be padded. For example, the sample is padded by copying the nearest neighbor sample. If there are more adjacent samples at the same distance, the adjacent samples to be used among them may be specified by convention (e.g., by a standard). Another possibility is to interpolate the padding sample from multiple neighboring samples. Alternatively, padding may include using samples of zero values. Intermediate samples generated by the process may also need to be padded. The intermediate samples may be generated based on the samples of the input subset including the one or more padded samples. The padding may be performed before the input of the neural network or within the neural network. However, the padding should be performed before the processing of the one or more layers.
[0231] To be complete, Figure 19 shows another example where, after the output subset Ri has undergone cropping, each cropped region Ri-crop is seamlessly merged without any overlap of the cropped regions. In contrast to the examples of Figures 17A, 17B, and 18, the regions Li do not overlap. Each implementation may be as follows. 1. Determine N non-overlapping regions Li (i.e., a plurality of tiles of the first and second) in the first and second tensors respectively. Here, at least one of those regions includes a subset of the receptive field of one of R i and R iThe union constitutes the complete output of the first or second sub-network (e.g., the feature tensor y). 2. Process each L i independently using the NN to obtain the region R i 3. Merge R i to obtain the merged output tensor.
[0232] Figure 19 shows this case, where Li is selected as the non-overlapping region. This is a special case where the total amount of computation required to obtain the entire reconstructed output (i.e., the output picture) is minimized. However, the reconstruction quality may be compromised.
[0233] As described above, the first and second sub-networks are part of a neural network. A sub-network is itself a neural network and includes at least one processing layer. In one implementation, the first sub-network performs processing by one or more layers including at least one convolutional layer and at least one pooling layer, and / or the second sub-network performs processing by one or more layers including at least one convolutional layer and at least one pooling layer.
[0234] In one implementation example, the first sub-network and the second sub-network perform respective processing that is part of picture or video compression.
[0235] Furthermore, the first sub-network and / or the second sub-network performs one of picture encoding, rate-distortion optimized quantization (RDOQ), and picture filtering by the convolutional sub-network. As described above, the first and / or second sub-network may be the encoding device (or encoding process) 901 or Q / RDOQ 902 of FIG. 9, and the input image x is first processed by the encoder 901 that performs picture encoding. The encoder 901 may be part of a VAE encoder-decoder framework as shown in FIGS. 6A and 6B, and each encoder "ga" performs respective picture encoding by processing the input image through a sequence of convolutional layers 601-604 including processing via a GDN layer. Similarly, the quantizer "Q" in FIGS. 6A and 6B may perform the function of quantization or RDOQ, which is the first and / or second sub-network. The same applies to units Q or RDOQ 902 and 908 in FIG. 9. Picture post-filtering is not further shown in FIGS. 6A / B and 9. In general, some of the processing layers of the neural network in the decoding device can have a post-filtering or general filtering function.
[0236] For example, FIG. 16 shows an example of an encoding device that may correspond to a neural network, including an encoder sub-network 1601 as the first sub-network and an RDOQ sub-network 1602 as the second sub-network. However, the present disclosure is not limited to such embodiments. The neural network may be a network for decoding a picture or a latent representation, and may include a decoding sub-network and a post-filtering sub-network. Further examples of sub-networks are possible, including a pre-processing sub-network or other sub-networks.
[0237] As shown in FIG. 16, the encoding device (or generally the encoding process) may further include a hyper-encoder 1603, a quantizer 1608 for the output z of the hyper-encoder 1603, and an arithmetic encoder 1609 for encoding the quantized information (^z) with a hyper-prior so as to be included in the bitstream 2.
[0238] The decoding device (or decoding process) may correspondingly further include an arithmetic decoder 1610 for hyper-prior information and a subsequent hyper-decoder 1607. The hyper-prior parts of encoding and decoding may correspond to the VAE framework regarding FIGS. 6A and 6B or a modification thereof.
[0239] FIGS. 6A and 6B show an encoder within an NN-based VAE framework having respective convolutional layers 601 to 606, and the respective sizes of the input images (picture data) are then reduced by only 2. For example, a plurality of samples (i.e., samples of picture data) used by a neural network (NN) may depend on the kernel size of the first input layer of the NN. The one or more layers of the neural network (NN) may include one or more pooling layers and / or one or more subsampling layers. The NN may generate one sample of the output subset by pooling a plurality of samples through the one or more pooling layers. Alternatively or additionally, one output sample may be generated by the NN through subsampling (i.e., downsampling) by one or more downsampling convolutional layers (such as convolutional layers 601 to 606 in FIG. 6B). Pooling and downsampling may be combined to generate one output sample. FIG. 12 shows downsampling by two convolutional layers, starting from seven samples of the entire receptive field and providing one sample as the output.
[0240] In FIGS. 6A and 6B, the input image 614 corresponds to picture data. According to one implementation, the input tensor 614 is a picture or a sequence of pictures that includes one or more components at least one of which is a color component. Alternatively, the input tensor may be a latent space representation of a picture that can be an output of preprocessing (e.g., an output tensor). The one or more components are color components and / or depth and / or motion maps and / or other feature maps related to the picture samples.
[0241] Note that the input tensor can represent other types of data (e.g., any type of multi-dimensional data) having one or more spatial components that may be suitable for tiling and / or processing via the subnetwork as described in the first embodiment.
[0242] In one implementation, the input tensor has at least two components, namely a first component and a second component. The first subnetwork divides the first component into a third plurality of tiles and divides the second component into a fourth plurality of tiles, and at least two respective co-located tiles of the third plurality of tiles and the fourth plurality of tiles have different sizes. In principle, the present disclosure is not limited to any particular number of spatial components. There may be one spatial component (e.g., a grayscale picture). However, in this implementation, the input tensor has a plurality of spatial components. The first and second components may be color components. The tiling may be different not only for the subnetwork but also for the components processed by the same subnetwork.
[0243] Thus, additionally or alternatively, the second subnetwork divides the first component into a fifth plurality of tiles, divides the second component into a sixth plurality of tiles, and at least two respective co-located tiles of the fifth plurality of tiles and the sixth plurality of tiles have different sizes. In other words, the tiling of the spatial components of the second input tensor can be different. Further details of the second embodiment with different tilings for different spatial input components are provided below. The second embodiment can be combined with the first embodiment described herein where the tiling is different for different sub-networks of the neural network.
[0244] The encoding of FIG. 16 (similar to FIGS. 6A and 6B) further includes generating a bitstream by including the output of the processing by the neural network in the bitstream. The neural network can further include entropy coding. This is shown in FIG. 6A, where after encoding the input image 614 through the encoder neural network (convolutional layers 601-604), the output y of the encoder NN_ga undergoes quantization (e.g., RDOQ) and arithmetic encoding 613 to provide the bitstream Bitstream 1. Similarly, the hyperprior neural network receives the (bottleneck) feature map y (output of the encoder NN ga), processes it through convolutional layers 605 and 606 and two ReLU layers, and provides it as output z with the statistics of the input image. This is quantized and arithmetic encoded (615) to generate the bitstream Bitstream 2. Similarly, through the processes shown in FIGS. 6B, 9, and 9A, the respective bitstreams Bitstream 1 and Bitstream 2 are generated.
[0245] Once the tile size is determined as described above, generating the bitstream further includes including in the bitstream an indication of the tile size within the first plurality of tiles and / or an indication of the tile size of the second plurality of tiles. This bitstream may be a bitstream including Bitsream 1 and Bitstream 2, i.e., a part of the bitstream generated by the entire encoding device (encoding process).
[0246] Further details and implementation examples regarding the processing of components including color components and the signaling of tile sizes are described in the second embodiment. Here, simply refer to FIG. 20 showing examples of various types of parameters (such as the tile sizes of the first plurality of tiles and / or the second plurality of tiles). Each indication is included in the bitstream.
[0247] The encoding process of the input tensor has a corresponding part of the decoding that shares the functional correspondence in the process. In this exemplary and non-limiting embodiment, a method for decoding a tensor representing picture data is provided. The tensor can have a matrix form with width = w, height = h, and a third dimension (e.g., depth or number of channels) equal to D in two spatial dimensions. The method includes processing an input tensor representing picture data by a neural network including at least a first sub-network and a second sub-network. Note that the width and height of the input tensor of the decoder can be different from the width and height of the input tensor processed by the encoder. Further, as will be apparent to those skilled in the art, the first sub-network and the second sub-network of the decoder may perform functions that are completely or partially (e.g., functionally) inverse to the first sub-network and the second sub-network of the encoder. However, the inverse function may not be interpreted in a strict mathematical way. Rather, the term "inverse" refers to the process for the purpose of decoding the tensor to reconstruct the original picture data. It is understood by those skilled in the art that the encoding compression and decoding decompression may include additional processes that may not be necessary for decoding and / or encoding. For example, RDOQ shown in FIG. 16 is a process only for the encoder. Further, note that the terms "first" and "second" sub-networks are merely labels for distinguishing between the sub-networks of the decoder (and for that purpose, also between the sub-networks of the encoder described above).
[0248] Examples of the first and / or second subnetwork for the encoding branch are shown in FIG. 16 and include decoder 1604 and post-filter 1611. In this method, the process includes applying the first subnetwork to a first tensor, which includes dividing the first tensor in the spatial dimension into a first plurality of tiles and processing the first plurality of tiles by the first subnetwork; and after applying the first subnetwork, applying the second subnetwork to a second tensor, which includes dividing the second tensor in the spatial dimension into a second plurality of tiles and processing the second plurality of tiles by the second subnetwork. In the example of FIG. 16, the first subnetwork is decoder 1604, and its output is provided as input to the second subnetwork, which is post-filter 1611. Here, the first tensor is the quantized feature tensor ^y in the latent space. This is pre-decoded from bitstream Bitstream 1 by arithmetic decoder 1606. Next, the second tensor ^x' input to the second subnetwork 1611 for post-filtering is a feature tensor, such as a feature, a feature map, or a feature map in the latent space. Thus, similar to the encoder's process, the type of the input tensor (e.g., the first and second tensors) can depend on the process executed by the preceding subnetwork. In the example of FIG. 16, the preceding subnetwork is encoder 1604. The term "after" above does not limit the above decoding process immediately in the sense that the output of the first subnetwork is directly input to the second subnetwork. Rather, "after" means that the first and second pluralities of tiles are processed in the same pipeline in a certain temporal order, which may not be immediate in time.
[0249] Furthermore, the type of input may also depend on a layer (e.g., a layer of a neural network NN or a layer of a non-trained network) where the processed input data can be branched to be used as input to another (e.g., subsequent) subnetwork. Such a layer may be, for example, the output of a hidden layer. Similar to the encoding process, the first and second tensors are split into the so-called tiles already defined above in the decoding process.
[0250] In FIG. 16, decoder 1604 can be a first subnetwork having processes 1 to N (processing pipelines 1 to N). As shown in FIG. 16, decoder 1604 receives input feature tensor ŷ and splits it into N tiles ŷ 1 ~ŷ N (the first plurality of tiles). Each tensor tile is then processed by each block, i.e., processes 1 to N that do not need to interact with each other (e.g., wait for each other during processing). The result of each process provides N tiles ^x 1 '~^x N ' (tensors). After processing the first plurality of tiles, decoder subnetwork 1604 may merge the processed tiles into a first output tensor ^x'. In FIG. 16, the first output tensor may be a second tensor used as input by post-filter 1611. Post-filter 1611 can be a second subnetwork having processes 1 to N (processing pipelines 1 to N). Similar to decoder 1604, post-filter 1611 splits input tensor ^x' into a second plurality of tiles ^x 1 '~^x N '. These are processed by respective processes 1 to N of post-filter 1611. Processes 1 to N output respective tiles ^x 1 ~^x Nis provided. This can be merged into tile^x. In the example of FIG. 16, the merged^x refers to the decoded tensor representing the reconstructed picture data. The merge can include (but does not have to include) cropping as shown in FIGS. 17-19. Note that the merge / combination into tensor^x' or^x does not have to be performed. It is conceivable that the second subnetwork reuses the tiling of the first subnetwork and simply modifies it (refining the tiling by further dividing the tiles or coarsening the tiling by combining multiple tiles into one). In the example of FIG. 16, processes 1 to N may perform the processing in tile units. That is, process i processes tile i. Alternatively, process i may process component i among the plurality of components of the input tensor. In this case, process i divides component i into a plurality of tiles and processes those tiles separately or in parallel.
[0251] Furthermore, at least two respective co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes. In other words, the subdivision of the tiles may be different for each subnetwork, and thus each subnetwork may use different tile sizes. However, the inputs of each subnetwork (i.e., the first and second tensors) are divided into a grid of tiles of the same size, except for the tiles at the lower and right image boundaries, which can have a smaller size.
[0252] Otherwise, the characteristics and / or features of the first and second plurality of tiles used in the decoding process are similar to one of the encoding processes described above. In particular, in the above exemplary implementation, the tiles of the first plurality of tiles adjacent in at least one of the spatial dimensions overlap partially, and / or the tiles of the second plurality of tiles adjacent in at least one of the spatial dimensions overlap partially. Examples of adjacent tiles with partial overlap are shown in FIGS. 17A, 17B, 18, 19, and 20.
[0253] Furthermore, in some exemplary implementations, the tiles of the first plurality of tiles are processed independently by the first subnetwork and / or the tiles of the second plurality of tiles are processed independently by the second subnetwork. In other words, the processing of the tiles is independent of each other and thus can be parallelized. In one example, at least two tiles of the first plurality of tiles are processed in parallel by the first subnetwork and / or at least two tiles of the second plurality of tiles are processed in parallel by the second subnetwork. FIG. 16 shows parallel processing on the decoder side, where the components of the first input tensor (e.g., tile and / or spatial components) ^y 1 ~^y N are processed by the decoder 1604 without interaction between the processing for components 1 to N. The result of each processing is the output tensor component ^x 1 '~^x N '. These may be further combined into the output tensor ^x'. The second subnetwork may be the post-filter 1611 of FIG. 16, which takes as input the tensor ^x' from the decoder 1604 (the first subnetwork). In this example, the input of the post-filter 1611 is directly connected to the output of the decoder 1604. Alternatively, there may be additional processing between the post-filter and the decoder of FIG. 16. The input tensor ^x' is split by the post-filter into N tiles ^x 1 '~^x N '. The post-filter then processes each tile independently or in parallel. The parallel processing is reflected in FIG. 16 in such a way that the processing of components 1 to N does not interact with each other. The result of the parallel processing of the post-filter is the reconstructed tile ^x 1 ~^x N . This can be merged into one tensor ^x representing the reconstructed picture data.
[0254] In addition, the splitting of the first tensor includes determining the size of a tile among the first plurality of tiles based on a first predetermined condition, and / or the splitting of the second tensor includes determining the size of a tile among the second plurality of tiles based on a second predetermined condition. For example, the first predetermined condition and / or the second predetermined condition are based on other characteristics as already described above with respect to available decoder hardware resources and / or motion or encoding present in the picture data.
[0255] The first and second sub-networks for decoding processing may have the same configuration as the sub-network for encoding processing. Specifically, the first sub-network performs processing by one or more layers including at least one convolutional layer and at least one pooling layer, and / or the second sub-network performs processing by one or more layers including at least one convolutional layer and at least one pooling layer. FIGS. 6A and 6B show a decoder within an NN-based VAE framework having respective convolutional layers 607 to 6012, and the size of each input tensor (feature map) ^y is then doubled (upsampled). Also, the first sub-network and the second sub-network perform respective processes that are part of decompressing a picture or video. Such processes may be provided by the VAE encoder shown in FIGS. 6A and 6B, taking the feature tensor ^y as input, decompressing the output image, and reconstructing it as the output of convolutional layer 607. That represents the reconstructed picture data (i.e., the decoded tensor). For example, the first sub-network and / or the second sub-network performs one of picture decoding by a convolutional sub-network and picture filtering. As described above, the first and / or second sub-networks for decoding processing may be the decoder 904 of FIG. 9 that processes the feature map tensor ^y to generate reconstructed image data ^x that approximates the original picture data x. As shown in FIG. 9, the decoder 904 may be part of a VAE encoder-decoder framework as shown in FIGS. 6A and 6B, and each decoder "gs" performs respective picture decoding by processing the feature tensor ^y through a sequence of convolutional layers 610 to 607 including processing via an inverse IGDN layer. Picture filtering (e.g., post-filtering) is not further shown in FIGS. 6A / B and 9.
[0256] In one implementation, the input tensor is a picture or sequence of pictures that includes one or more components where at least one is a color component. Alternatively, the input tensor may be a latent space representation of a picture, which can be the output of a pre - processing (e.g., an output tensor). The one or more components are color components and / or depth and / or motion maps and / or other feature maps (s) related to the picture samples. The input tensor has at least two components, namely a first component and a second component. The first sub - network divides the first component into a third plurality of tiles and divides the second component into a fourth plurality of tiles, and at least two co - located tiles of the third plurality of tiles and the fourth plurality of tiles have different sizes, and / or the second sub - network divides the first component into a fifth plurality of tiles and divides the second component into a sixth plurality of tiles, and at least two co - located tiles of the fifth plurality of tiles and the sixth plurality of tiles have different sizes.
[0257] The decoding method described above further includes extracting an input tensor from a bitstream for processing by a neural network. The neural network may further include entropy decoding. FIGS. 6A and 6B show decoding an input tensor ^y for a decoder sub - network gs from a bitstream Bitstream 1 via arithmetic decoding. Entropy encoding is performed by a sub - network hs. Here, an entropy tensor ^z can be decoded by arithmetic decoding from bitstream Bitstream 2. ^z is processed to obtain statistical information about the encoded picture data that is used for the decoding process of ^y. FIG. 6B shows further details of entropy decoding, the mean μ and variance σ 1 , σ 2The information about the distribution given for ^y is further input into a Gaussian entropy model 670 and used for the arithmetic decoding of ^y. As described above, a subnetwork hs (FIG. 6A) having various upsampling convolutional layers 611 and 612, possibly including a leaky ReLU layer 660, may be part of the hyper-decoder 907 shown in FIGS. 9 and 9B. In one exemplary implementation, a second subnetwork performs picture post-filtering, and for at least two tiles of a second plurality of tiles, one or more parameters of the post-filtering are different and are extracted from the bitstream. Further, the decoding process further includes parsing from the bitstream an indication of the size of the tiles of a first plurality of tiles and / or an indication of the size of the tiles of a second plurality of tiles.
[0258] In this illustrative and non - limiting embodiment, a computer program stored on a non - transitory medium is provided that, when executed on one or more processors, includes code to perform any of the steps of the encoding and decoding methods described above. Flowcharts of the encoding and the encoding process are shown in FIGS. 21 and 22 respectively. Regarding the encoding process of FIG. 21, in step 2110, a first sub - network processes a first tensor. This includes dividing the first tensor into a plurality of tiles, which are then processed by the first sub - network. Note that in FIG. 21, picture data (i.e., an input tensor representing picture data) is input to the first sub - network shown by the dashed line. This means that the raw picture data does not necessarily have to be directly input to the first sub - network, and this may depend on the processing performed by the first sub - network regarding the order of the processing sequence. In other words, the first tensor, which is the input for the first sub - network, is derived from the picture data. If the first sub - network is the encoder 901 shown in FIG. 9, the first tensor may be the input tensor x. In step S2120, the first plurality of tiles are further processed by the first sub - network. After the processing by the first sub - network, a second sub - network processes a second tensor in step S2130. This includes dividing the second tensor into a second plurality of tiles. Here too, the second tensor input to the second sub - network shown by the dashed line does not necessarily have to be a direct output from the first sub - network. In the implementation example of FIG. 9, the output of the feature tensor y of the encoder 901 (the first sub - network) is directly input to the RODQ 902 (the second sub - network). However, there may be additional processing between the encoder 901 and the RODQ 902. In step S2140, the second plurality of tiles are further processed. The output of the processing of the second sub - network, which may include further processing (not shown in FIG. 21), may include generating a bitstream (e.g., bitstream 1 and / or bitstream 2 in FIG. 9).Also, the process for determining the size of the tiles among the first and second plurality of tiles, and / or including an indication of said size in the bitstream, may be a processing step before providing the bitstream as an output of neural network processing. Regarding the decoding process of FIG. 22, the flowchart depicts reverse processing steps starting from an input tensor that can be a direct or indirect input to the first subnetwork, as indicated by the dashed line. Here too, in the exemplary implementation of FIG. 9, the feature tensor ^y is used as a direct input to the decoder 904, which is the first subnetwork in this case. In step S2220, as shown in FIG. 16, the first plurality of tiles are processed by the first subnetwork (e.g., decoder 1604). After the processing of the first subnetwork, the second tensor is processed by the second subnetwork in step S2230 by splitting the second tensor into a plurality of tiles, which is then further processed by the second subnetwork in step S2240. In the exemplary implementation of FIG. 16, the output ^x of the first subnetwork of decoder 1604 is the second tensor, which is processed by the post-filter 1611 corresponding to the second subnetwork. The processing of the second plurality of tiles in the example of FIG. 16 outputs ^x corresponding to the reconstructed picture data as an output. It should be noted that there may be additional processing between the processing of decoder 1604 and post-filter 1611 in FIG. 16. In that case, the output ^x of decoder 1604 may not be directly input to post-filter 1611.
[0259] Furthermore, as already described, the present disclosure also provides a device (apparatus) configured to execute the steps of the above-described method.
[0260] In this exemplary and non - limiting embodiment, a processing apparatus for encoding an input tensor representing picture data is provided. FIG. 23 shows a processing apparatus 2300 having respective modules for performing steps of an encoding process, including a processing circuit 2310. The processing circuit is configured to process the input tensor by a neural network including at least a first sub - network and a second sub - network realized by an NN processing module - sub - network 1 2311 and an NN processing module - sub - network 2 2312. The NN processing module - sub - network 1 2311 may have a separate splitting module 1 2313 that splits a first tensor in the spatial dimension into a first plurality of tiles and processes the first plurality of tiles by the first sub - network. Alternatively, the respective modules for processing the first input tensor and / or the first plurality of tiles may be implemented in a single module that may be included in a single circuit or separate circuits. Similarly, the NN processing module - sub - network 2 2312 and the splitting module 2314 apply the second sub - network to a second tensor, which includes splitting the second tensor in the spatial dimension into a second plurality of tiles and processing the second plurality of tiles by the second sub - network. It should be noted that the processing of the second tensor after the processing of the first tensor may be implemented by wiring the respective modules such that the signals (from the perspective of those input and output signals) are input in their respective chronological order (not necessarily immediate chronological order). Alternatively or additionally, the respective order of signaling may be implemented by configuring the modules by software. The modules 2313 and 2314 may further provide a function of determining the tile size of the first and / or second plurality of tiles. The processing apparatus may further have a bit - stream module 2315 that provides a function of generating a bit - stream, and the output of the neural network processing is included in the bit - stream.Furthermore, module 2315 provides a function of including in the bitstream the display of the tile sizes of the first plurality of tiles and / or the second plurality of tiles.
[0261] In this exemplary and non-limiting embodiment, a processing device for encoding an input tensor representing picture data is provided, the processing device comprising one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming configuring an encoder to execute a method according to the encoding method described above when executed by the one or more processors.
[0262] In this exemplary and non - limiting embodiment, a processing device for decoding a tensor representing picture data is provided. FIG. 24 shows a processing device 2400 having respective modules for performing steps of a decoding process, comprising a processing circuit 2410. The processing circuit is configured to process an input tensor by a neural network including at least a first sub - network and a second sub - network realized by an NN processing module - sub - network 1 2411 and an NN processing module - sub - network 2 2412. The NN processing module - sub - network 1 2411 may have a separate splitting module 1 2413 that splits a first tensor in the spatial dimension into a first plurality of tiles and processes the first plurality of tiles by the first sub - network. Alternatively, each module for processing the first input tensor and / or the first plurality of tiles may be implemented in a single module that may be included in a single circuit or separate circuits. Similarly, the NN processing module - sub - network 2 2412 and the splitting module 2414 apply a second sub - network to a second tensor, which includes splitting a second tensor in the spatial dimension into a second plurality of tiles and processing the second plurality of tiles by the second sub - network. Modules 2413 and 2414 may further provide a function of determining the tile size of the first and / or second plurality of tiles. Here again, the processing of the second tensor after the processing of the first tensor may be implemented, for example, by wiring the respective modules such that signals (with respect to those input and output signals) are input in their respective temporal (not necessarily immediate) order. Alternatively or additionally, the respective order of signaling may be implemented by configuring the modules by software. The processing device may further have a parsing module 2415 that provides a function of parsing an indication of the tile size of the first plurality of tiles and / or the second plurality of tiles from a bitstream. Module 2415 may further provide a function of extracting the input tensor from the bitstream.
[0263] In this exemplary and non-limiting embodiment, a processing apparatus for decoding a tensor representing picture data, comprising: one or more processors; and a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, wherein when the programming is executed by the one or more processors, the encoder is configured to execute the method according to the above-described decoding method.
[0264] The encoding and decoding processes performed by the VAE encoder-decoder described above and shown in FIG. 16 may be implemented within the coding system 10 of FIG. 1A. Thereby, the source device 12 represents the encoding side and provides compression of the input picture data 21 including the input tensor x of FIG. 16. In particular, the encoder 20 of FIG. 1A may include modules for the encoding process according to the present disclosure, such as the encoder 1601, the quantizer or RDOQ 1602, and the arithmetic encoder 1605. The encoder 20 may further include modules of the hyperprior, such as the hyperdecoder 1603, the quantizer or RDOQ 1608, and the arithmetic encoder 1609. Similarly, the destination device 14 of FIG. 1A represents the decoding side that provides decompression of the input tensor representing the picture data. In particular, the decoder 30 of FIG. 1A may include modules for the decoding process according to the present disclosure, such as the decoder 1604 and the postfilter 1611 of FIG. 16, together with the arithmetic decoder 1606. In addition, the decoder 30 may further include a hyperprior for decoding ^z, the arithmetic decoder 1610, the hyperdecoder 1607, and the arithmetic decoder 1606. In other words, the encoder 20 and the decoder 30 of FIG. 1A may be implemented and configured to include any of the modules of FIG. 16 to implement the encoding or decoding process according to the present disclosure, and a plurality of sub-networks process the input tensor divided into a first plurality of tiles and a second plurality of tiles, which are then processed as described in the first embodiment. FIG. 1A shows the encoder 20 and the decoder 30 separately, but they may be implemented via the processing circuit 46 of FIG. 1B. In other words, the processing circuit 46 may provide the functionality of the encoding-decoding process of the present disclosure by implementing respective circuits for each of the modules of FIG. 16.
[0265] Similarly, the video coding device 200 having the coding module 270 of the processor 230 in FIG. 2 can execute the encoding process or the decoding process of the present disclosure. For example, the video coding device 200 may be an encoder or a decoder having each module in FIG. 16, and executes the encoding or decoding process as described above.
[0266] The apparatus 300 in FIG. 3 can be implemented as an encoder and / or a decoder having an arithmetic encoder 1605, 1609 and an arithmetic decoder 1606, 1610, an encoder 1601, a quantizer or RDOQ 1602, a decoder 1604, a post filter 1611, a hyper encoder 1603, and a hyper decoder 1607 so as to execute the tile processing as described according to the first embodiment. For example, the processor 302 in FIG. 3 may have respective circuits for executing the encoding and / or decoding process according to the foregoing methods.
[0267] The exemplary implementations of the encoder 20 shown in FIG. 4 and the decoder 30 shown in FIG. 5 can also implement the encoding and decoding functions of the present disclosure. For example, the splitting unit 452 in FIG. 4 can perform splitting the first and / or second tensors into the first and / or second plurality of tiles, which is executed by the encoder 1601 in FIG. 16. Thereby, the syntax element 465 can include an indication (singular or plural) about the size and position of the tile, together with an indication such as a filter index. Similarly, the quantization unit 408 may perform quantization or RDOQ of the RDOQ module 1602, and the entropy encoding unit 470 can implement the functions of the hyper priors (i.e., modules 1603, 1605, 1607, 1608, 1609). Next, the entropy decoding unit 504 in FIG. 5 can perform the functions of the decoder 1604 in FIG. 16 by splitting the encoded picture data 21 (input tensor) into tiles and parsing an indication about the tile size or position, etc. from the bitstream as the syntax element 566. The entropy decoding unit 504 can further implement the hyper priors (i.e., modules 1606, 1607, 1610). The post filtering 1611 in FIG. 16 can be executed, for example, also by the entropy decoding unit 504. Alternatively, the post filter 1611 may be implemented within the mode application unit 560 as an additional unit (not shown in FIG. 5).
Embodiment
[0268] Second Embodiment In an exemplary implementation prior to the first embodiment, the tiles of the first plurality of tiles and / or the second plurality of tiles are processed by the first and second sub-networks respectively, as described previously. Here, the encoding and decoding processes of the components in the plurality of pipelines are discussed. Each pipeline processes one or more picture / image components, which can be color planes. As can be seen from the following description, the first embodiment and the second embodiment partially share similar or the same processes.
[0269] In this illustrative and non - limiting embodiment, a method for processing an input tensor representing picture data is provided. The input tensor can have a matrix format with width = w, height = h, and a third dimension (e.g., depth or number of channels) equal to D. Further, the input tensor may be processed by a neural network. In this method, a plurality of components of the input tensor are processed, and the components include a first component and a second component in the spatial dimension. The input tensor is a picture or a sequence of pictures including one or more components of which at least one of the plurality of components is a color component. In one implementation, the first component represents the luma component of the picture data, and the second component represents the chroma component of the picture data. For example, the plurality of components of the input tensor may be in the YUV format, where Y is the luma component and UV are the chroma components. The luma component may sometimes be called the primary component, and the chroma component may sometimes be called the secondary component. Note that the terms "primary" and "secondary" are labels for placing more weight and / or importance on the primary component(s) than on the secondary component(s). As will be apparent to those skilled in the art, even though the luma component is typically the component with higher importance, there may be cases where a different component other than luma is prioritized for use. Such prioritization often involves using information about the primary component (e.g., luma) as auxiliary information for processing the secondary component (i.e., less important than luma). The plurality of components may be in the RGB format, or any other format suitable for processing the components in their respective processing pipelines. This method includes processing the first component, which includes dividing the first component in the spatial dimension into a first plurality of tiles and processing the tiles of the first plurality of tiles separately; and processing the second component, which includes dividing the second component in the spatial dimension into a second plurality of tiles and processing the tiles of the second plurality of tiles separately. Note that separately does not mean independently, i.e., there may still be auxiliary information shared for processing the first component and / or the second component.Examples of tiles having a rectangular shape are shown in FIGS. 14, 17A, 17B, and 18-20. In this method, at least two co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes. In other words, the tiles of the first plurality of tiles and the second plurality of tiles have different sizes. Regarding the size of the tiles within each plurality of tiles, all the tiles within the first plurality of tiles have the same size and / or all the tiles within the second plurality of tiles have the same size. Thus, each pipeline processes tiles of the same size that are different between pipelines. In particular, Y and UV can have a block partitioning different from VVC. Tiles are introduced to solve memory problems and can contain multiple blocks. Tiles can be coded independently of each other and decoded in parallel. A block is part of a tile / picture / slice, and each block uses its own coding method and the decoding is sequential, thus an encoding tree (block structure) is required for better local adaptation. That is, in this exemplary and non-limiting embodiment, a separate chroma coding tree may be used.
[0270] In one implementation, at least two tiles out of the first plurality of tiles are processed independently or in parallel, and / or at least two tiles out of the second plurality of tiles are processed independently or in parallel. FIG. 25 shows an example of encoder-decoder processing when the first component is luma and the second component is one of chroma U or V, and these are processed in separate pipelines. It should be noted that, as shown in FIG. 25, chroma components U and V can be jointly processed as one chroma component. In one exemplary implementation, the processing of the input tensor includes processing that is part of picture or video compression. For example, the processing of the first component and / or the second component includes one of picture encoding by a neural network, rate-distortion optimized quantization (RDOQ), and picture filtering. The compression processing is shown in FIG. 25 by respective modules, encoders 2501, 2508 and RDOQ 2502, 2509. Encoders 2501 and 2508 perform picture encoding of the input picture x by a neural network (NN) for luma Y and chroma component(s) UV. Modules RDOQ 2502 and 2509 quantize the feature tensor y by optimizing the rate distortion and provide the quantized feature tensor ŷ for each component as an output. Modules, hyper-encoders 2503, 2510 and RDOQ 2504, 2511 form parts of the hyper-prior network on the encoding side that generate statistical information ẑ of the picture data. Further, a bitstream is generated by including the outputs of the processing of the first component and the second component in the bitstream. In the example of FIG. 25, the quantized feature tensors ŷ of the first (luma) component and the second (chroma) component are each arithmetic encoded (1605) and included in bitstreams Bitstream Y1 and Bitstream UV1, respectively. Similarly, the quantized statistical information ẑ of luma and chroma is arithmetic encoded (1609) and included in bitstreams Bitstream Y2 and Bitstream UV2, respectively.In the example of FIG. 25, two bitstreams are generated, allowing the encoded picture data to be separated from the encoded statistics of the picture data.
[0271] In another exemplary implementation, the processing of the input tensor includes processing that is part of decompressing a picture or video. For example, the processing of the first component and / or the second component includes one of neural network-based picture decoding and picture filtering. The decompression process is shown in FIG. 25 by respective modules, decoders 2506, 2513 and postfilters 2507, 2514. Decoders 2506 and 2513 perform picture decoding of the input picture x by a neural network (NN) for the luma Y and chroma components UV by decoding the quantized feature tensors ŷ for the luma and chroma components from the respective bitstreams Bitstream Y1 and Bitstream UV1. As described above, the hyperprior also requires a decoding module that includes hyperdecoders 2505, 2512 that decode the quantized statistical information ẑ of the picture data from the bitstreams Bitstream Y2 and Bitstream UV2 for the luma and chroma components. This information is input to the arithmetic decoder 1606 and its output is provided to the decoders 2505 and 2513. In one implementation, the processing of the first component and / or the second component includes picture postfiltering. As shown in FIG. 25, the decoder output is input to the postfilters 2507 and 2514, and these postfilters perform picture postfiltering of the pictures of their components and provide as output the reconstructed picture data ^x for the reconstructed luma ^Y and chroma ^UV, respectively. The functional blocks (modules) shown in FIG. 25 are similar in structure arrangement to the respective modules of the VAE encoder-decoder shown in FIG. 16 of the first embodiment.
[0272] In the example of FIG. 25, the input tensor is YUV having Y, U, and V as multiple components. In particular, the first component here is the luma Y input to the encoder 2501 of the luma pipeline. Similarly, as described above, U and V may be collectively considered as the second components input to the encoder 2508 of the chroma pipeline. Note that different reference numerals are assigned to the respective functional modules (units) of the VAE encoder-decoder of the luma and chroma pipelines so that each tile having different sizes for the luma and chroma pipelines can be processed in different ways. Still, those modules (e.g., encoder 2501 and encoder 2508) may perform the same / similar functions regarding encoding, RDOQ, decoding, post-filtering, etc. The said functions are the same as those performed by the respective modules of the VAE encoder-decoder of FIG. 16 and have already been described above. In one implementation, processing the second component includes decoding the chroma components of the picture based on the representation of the luma components of the picture. This is shown in FIG. 25, where the luma component is the primary component and is thus considered to have a higher importance than the chroma components UV. As shown, in the decoding process of luma and chroma, the luma feature tensor ^y is obtained from the luma bitstream Bitstream Y1 by arithmetic decoding 1606, which is used as an input for the decoder 2513 for decoding the UV components.
[0273] Similar to the first embodiment, the tiles of the first plurality of tiles adjacent in at least one dimension of the spatial dimension partially overlap, and / or the tiles of the second plurality of tiles adjacent in at least one dimension of the spatial dimension partially overlap. Examples of adjacent tiles that can partially overlap are described above with reference to FIGS. 17A, 17B, and 18-20. In one exemplary implementation, the splitting of the first component includes determining the size of the tiles in the first plurality of tiles based on a first predefined condition, and / or the splitting of the second component includes determining the size of the tiles in the second plurality of tiles based on a second predefined condition. For example, the first predefined condition and / or the second predefined condition are based on available decoder hardware resources and / or motion present in the picture data. The first condition and / or the predefined condition may be the memory resources of the decoder and / or encoder. If the available memory is less than a predefined value of the memory resources, the determined tile size may be smaller than the tile size when the available memory is greater than or equal to the predefined value. Alternatively or additionally, if the presence of motion is significant, the tile size may be determined to be smaller than when the presence of motion is less significant. Whether the motion is significant (i.e., fast motion and / or fast / frequent change in motion) is determined by the respective motion vector(s) with respect to the change in its magnitude and direction and may be compared with corresponding predefined values (thresholds) for magnitude and / or direction and / or frequency. Additionally, another predefined condition may be a region of interest (ROI), where if the ROI size is smaller than the ROI reference size, the tile size may be determined to be small.For example, the ROI in the picture data may be based on the detected objects (such as vehicles, bicycles, motorcycles, pedestrians, animals, etc.), and the objects may have different sizes, move fast or slowly, and / or change the direction of their movements fast or slowly and / or multiple times (i.e., more frequently based on a predefined frequency value). The ROI may further be based on the smoothness of the regions within the picture data. For example, the tile size may be large for regions having a large smoothness (measured against a predefined value of smoothness), while a small tile size may be used for less smooth regions, i.e., regions having very prominent, e.g., spatial variations. That is, the tile size may be determined based on the degree of texture of the regions within the picture data. Further, for primary components such as, for example, luminance, a larger tile size may be determined, while for secondary components (singular or plural) such as chroma (singular or plural), a smaller tile size may be determined. Thus, the tile size can be optimized for the hardware resources together with the content of the picture data including scene-specific tile sizes. In a further implementation, the step of determining the size of the tiles within the second plurality of tiles includes scaling the tiles of the first plurality of tiles. In other words, the tile size of the first component is used as a reference to derive the tile size of the second component by scaling the tile size of the first component. The scaling may include enlarging (scaling up) such that the scaled tiles of the second plurality of tiles are larger than the size of the first plurality of tiles. Alternatively, the scaling may include shrinking (scaling down) such that the scaled tiles of the second plurality of tiles are smaller than the size of the first plurality of tiles. Whether to scale up or scale down may depend on the importance of the first and / or second components.When at least one of the plurality of components being processed has at least one color component, the use of upscaling or downscaling may also depend on a specific color (such as a color component indicating "danger / warning", etc.).
[0274] In the following, tile size information can be signaled and decoded from the bitstream. Various aspects of the signaling and decoding of tile size and / or other suitable parameters will be described. In one exemplary implementation, an indication of the determined size of tiles within a first plurality of tiles and / or within a second plurality of tiles is encoded in the bitstream. Further, the indication further includes the position of the tiles within the first plurality of tiles and / or the second plurality of tiles. For example, the first component is a luma component, an indication of the tile size of the first plurality of tiles is included in the bitstream, the second component is a chroma component, an indication of a scaling factor is included in the bitstream, and the scaling factor associates the tile size of the first plurality of tiles with the tile size of the second plurality of tiles.
[0275] Examples of signaling the tile size and / or position of tiles for the various components are shown in the following table, along with an excerpt of the code syntax for implementing the signaling. The syntax table is an example and may not be limited to this particular syntax. In particular, the syntax examples are for illustrative purposes only and may apply to the signaling of each respective indication for the tile size and / or tile position for the luma components and / or chroma components included in the bitstream. The first syntax table refers to an auto-decoder and the second table refers to a post-filter. As is apparent from Tables 1 and 2, the same / similar indications may be used and may be included in the bitstream, but may be included separately for their use in the auto-decoder (Table 1) and the post-filter (Table 2). In one implementation, for at least two tiles of the first plurality of tiles, one or more parameters of the post-filtering are different and are extracted from the bitstream, and for at least two tiles of the second plurality of tiles, one or more parameters of the post-filtering are different and are extracted from the bitstream. In other words, different tile-based post-filter parameters may be signaled in the bitstream, and thus, the post-filtering of the tiles of each of the first plurality of tiles and / or the second plurality of tiles can be performed with high accuracy.
[0276] In the following, examples of the various parts of the encoder-decoder VAE module of FIG. 25 and how the module uses each display will be described with reference to Tables 1 and 2.
[0277] [Table 1] TIFF2025516914000008.tif135170
[0278] [Table 2] TIFF2025516914000010.tif253170 TIFF2025516914000011.tif215170
[0279] A. Decoder In this implementation example, as described above with reference to FIG. 25 showing separate pipelines for decoding the luma component and the chroma component, the luma and chroma are decoded separately. However, as indicated by the dashed line pointing from the arithmetic decoder 1606 of the luma pipeline to the decoder 2513 of the chroma pipeline, the decoding of the chroma component requires the latent spaces of both chroma and luma (i.e., CCS).
[0280] Derivation of tile map: As shown in Table 1, several ways to obtain a tile map are possible. 1. The tile map is explicitly signaled only for the primary component. Other components use the same tile map (for YUV420 and CCS, the primary component is luma / Y and the other components are chroma / UV). Using the same tile map includes using the same tile map in a literal way. In one implementation, the same tile map (e.g., for the luma component) is used to derive a tile map for the secondary component (e.g., chroma UV) by scaling the tile size of the luma component. For that purpose, an indication of the scaling factor is included in the bitstream. In Table 1, such a scaling factor is "chroma_tile_scaling_factor". 2. The tile map is explicitly signaled for each component of the image.
[0281] In the signaling example of Table 1, whether the tiles of each component are signaled or not depends on the indications "tiles_enabled_for_chroma" and "tiles_enabled_for_luma". The indication may be by a simple flag "0" or "1" indicating that it can be turned off or on.
[0282] When the tile map is signaled explicitly, this can be done in one of the following ways.
[0283] 1. A regular grid of tiles of the same size (excluding the bottom and right boundaries) is used. Offset and size values are signaled, from which the tile map can then be derived (see below). In this case, the tiles of the first component (luma) have the same size, where the tile size is defined using the width and height when the tile is rectangular. Each indication of the tile size is "tile_width_luma" and "tile_height_luma" in Table 1. Similar indications for the same chroma tile size can be used, except that the chroma tile size is different from the luma tile size.
[0284] 2. A regular grid of tiles of the same size (excluding the bottom and right boundaries) is used. Offset and size values are used but not signaled directly. Instead, the size and offset values are derived from the already decoded level definition. The tile map can be derived from the size and offset values (see below).
[0285] 3. Any grid of tiles is used. First, the number of tiles is signaled, and then for each tile, its position and size are signaled (duplication will be implicitly included in this signaling). In Table 1, any grid is reflected in that for a given number "num_luma_tiles" of luma tiles, each tile can have an individual size with respect to "tile_width[i]" and "tile_height[i]". Further, an indication about the tile position (here the start position) is also included in the bitstream, and the indication about the tile position is "tile_start_x[i]" and "tile_start_y[i]". In this example, the tiles are in a 2D x - y space.
[0286] When chroma signaling is enabled and the chroma components are processed independently of the luma components (Table 1: "!use_dependent_chroma_tiles", i.e., chroma tiles are not dependent on luma tiles), a similar signaling with each indication included in the bitstream is used.
[0287] Further indications about the overlap of tiles of luma components and / or chroma components may be included in the bitstream and thus signaled to the decoder.
[0288] When the tile map is signaled via the tile size (tile width is equal to height) and the value for duplication (value in the signal space size / coordinates), N overlapping regions are derived as follows: for tile_start_y in range(0, image_height - overlap, tile_height - overlap): for tile_start_x in range(0, image_width - overlap, tile_width - overlap): height=min(tile_height, image_height - tile_start_y) width = min(tile_width, image_width - tile_start_x) im_tile i = (tile_start_x, tile_start_y, width, height)
[0289] Decoding of luma components: There are N overlapping regions (tiles) within the image. Each tile (im_tile i ) has a corresponding tile within the latent space (lat_tile i ). Furthermore, im_tile i covers only a subset of the entire receptive field of lat_tile i . For each im_tile in the signal space, the matching lat_tile in the latent space is derived as follows. Here, alignment_size is a power of 2 that depends on the number of downsampling layers within the subnetwork: image_tile = im_tilei lat_tile_start_y = image_tile.position.y / / alignment_size lat_tile_start_x = image_tile.position.x / / alignment_size if image_tile.size.height % alignment_size: height = math.ceil(image_tile.size.height / alignment_size) else: height = image_tile.size.height / / alignment_size if image_tile.size.width % alignment_size: width = math.ceil(image_tile.size.width / alignment_size) else: width = image_tile.size.width / / alignment_size lat_tile i = (lat_tile_start_x, lat_tile_start_y, width, height)
[0290] Using lat_tile, the corresponding region of the latent space is extracted and processed by the decoder sub-network. The decoder sub-network also has im_tile as an auxiliary input, which is required to correctly pad at the image boundaries, especially when the tile size may not be a multiple of alignment_size. The output of the decoder is assigned to the region of the image specified by im_tile. This step can include a cropping operation to remove the overlapping parts of the reconstruction that overlap with the reconstruction of another tile.
[0291] The need for cropping of regions R1 - R4 for the processed tiles L1 - L4 is explained with reference to FIGS. 17A and 18 to ensure seamless merging.
[0292] Decoding of chroma components: There are N overlapping regions (tiles) within the image. In the example shown in FIGS. 17A and 18, there are four partially overlapping tiles. Each tile (im_tile i ) has a corresponding tile within the latent space (lat_tile i ). Further, im_tile i covers only a subset of the full receptive field of lat_tile i . The N overlapping regions are derived based on the signaling parameters for tile size and tile overlap (tile width is equal to height). The size is signaled in signal space units (i.e., not in the latent space). Then, the position and size of the tiles in signal space are derived as follows. for tile_start_y in range(0, image_height - overlap, tile_height - overlap): for tile_start_x in range(0, image_width - overlap, tile_width - overlap): height=min(tile_height, image_height - tile_start_y) width=min(tile_width, image_width - tile_start_x) im_tile i =(tile_start_x, tile_start_y, width, height)
[0293] For each im_tile in the signal space, the matching lat_tile in the latent space is derived as follows. Here, alignment_size is a power of 2 that depends on the number of downsampling layers in the subnetwork. image_tile=im_tile i lat_tile_start_y=image_tile.position.y / / alignment_size lat_tile_start_x=image_tile.position.x / / alignment_size if image_tile.size.height % alignment_size: height=math.ceil(image_tile.size.height / alignment_size) else: height=image_tile.size.height / / alignment_size if image_tile.size.width % alignment_size: width = math.ceil(image_tile.size.width / alignment_size) else: width = image_tile.size.width / / alignment_size lat_tile i =(lat_tile_start_x, lat_tile_start_y, width, height)
[0294] Using lat_tile, the corresponding region of the chroma latent space lat_UV is extracted. Further, the corresponding region of the luma latent space lat_Y is determined. For the YUV420 example, a possible way to do this is to downsample the luma latent space by a factor of 2 and then use the same lat_tile as for the chroma to extract lat_Y. Both lat_Y and lat_UV are then processed by the decoder subnetwork. The decoder subnetwork also has im_tile as an auxiliary input, which is required to correctly pad at the image boundaries, especially when the tile size may not be a multiple of alignment_size. The output of the decoder is assigned to the region of the image specified by im_tile. This step may include a cropping operation to remove the part of the reconstruction that overlaps with the reconstruction of another tile.
[0295] B. Post-processing filter Derivation of tile map: As shown in Table 2, several ways to obtain the tile map are possible.
[0296] 1. The same tile map as the decoder is used. For the example of YUV420, this can be implemented such that for filtering the Y component, the same tile as for luma / Y in the decoder is used, and for U and V, the same tile as used by the decoder for chroma (UV) is used. Such behavior can be signaled with a single flag. 2. The tile map is explicitly signaled only for the primary component (e.g., the luma component). Other components (e.g., the chroma components) use the same tile map. As shown in Table 2, the signaling of the tile map (i.e., tile size, position, and overlap) can be done in a similar way as in Table 1 for the auto decoder. 3. The tile map is explicitly signaled for each component of the image. Again, the signaling of the tile map for each component can be done by using the same instructions included in the bitstream as in Table 1.
[0297] Note that when the filter selection uses multi-scale structural similarity (MS-SSIM) as the distortion criterion (encoder), the tiles should be large enough for MS-SSIM (this includes several downsampling steps). If the tiles at the bottom and right boundaries of the image are too small, those tiles are enlarged by removing areas from their respective neighboring regions, i.e., adjacent regions. Since the decoder has to perform the same processing, this is normative. When MSE and / or peak signal-to-noise ratio (PSNR) are used as the distortion criterion, there is no problem regarding the possibility of the tile size being too small.
[0298] When the tile map is explicitly signaled, this can be done in one of the following ways. 1. A regular grid of tiles of the same size (except for the bottom and right boundaries) is used. Offset and size values are signaled from which the tile map can be derived (see the following explanation). 2. A regular grid of tiles of the same size (excluding the lower and right boundaries) is used. Offset and size values are used but not signaled directly. Instead, the size and offset values are derived from already decoded level definitions. The tile map can be derived from the size and offset values (see the following description). 3. Any grid of tiles is used. First, the number of tiles is signaled, and then for each tile, its position and size are signaled (overlap will be implicitly included in this signaling).
[0299] Derivation of filter used per tile tile I For a specific image component of tile, the filter specified by filter index filter I is used. In other words, the tile-based filter can be specified with one parameter, namely the filter index. The selection of filter I for each tile can be signaled via one of the following ways as shown in Table 2.
[0300] 1. The same model / filter is used for all tiles of the image component. This can be signaled with a single flag together with the filter index filter I In Table 2, such a flag is "same_model_for_all_luma" with a filter index of "model_idx_luma".
[0301] 2. The filter model is signaled for each tile of the image component. a. A default filter index is signaled. In Table 2, such a default index is "default_model_idx_luma". For each tile IFor (the number of tiles is already known from the tile map), a flag is signaled to indicate whether the filter index is different from the default. In Table 2, such a flag for each tile is "use_default_idx[i]" with an index "i" that labels each tile within the first or second plurality of tiles. If it is different, the filter index filter I is signaled. In Table 2, each filter index is "model_idx_luma[i]". In other words, one or more parameters of the filter are different for at least two tiles out of the first plurality of tiles. The same also applies to one or more parameters regarding the second plurality of tiles. b. Each tile I For (the number of tiles is already known from the tile map), the filter index filter for the tile I is signaled. In Table 2, each filter index is "model_idx_luma[i]".
[0302] The filter index for each component can also be explicitly signaled. Alternatively, the filter index for a component can be derived from those filter indexes signaled for another component, for example, the luma component.
[0303] When the selected model of the filter is signaled, the post-filter can use the same tile map as that used by the decoder. This can reduce the overhead for signaling the tile size.
[0304] Filtering of components: There are N overlapping regions (tiles) in the image. Each tile is processed independently (possibly in parallel). tile IRegarding it, its position, size, overlap, and the filter index used are known from the signaling described above. The reconstructed region I [Region I] is extracted from the reconstructed image based on the position and size values. Then, filtering based on the filter index is performed on this region. The output is cropped using the position, size, and overlap values and assigned to the filtered output image at the position described by the position and size values.
[0305] C. RDOQ In this implementation, RDOQ is applied separately for the luma and chroma components, i.e., the latent spaces for each component luma and chroma are optimized separately. Figure 25 shows the respective RDOQ processing of luma Y and chroma UV in separate pipelines by RDOQ modules 2502 and 2509, respectively. However, the decoding of the chroma component requires both the chroma and luma latent spaces (i.e., CCS). Thus, the processing by RDOQ 2502 is first performed for luma, and a new optimized luma latent ^y is obtained. Then, the luma latent is kept fixed and used as additional input (i.e., auxiliary information) for the optimization of the chroma latent space. In Figure 25, this is shown by the dashed line.
[0306] When the same tile map is used for luma and chroma in the RDOQ process, the optimization of the chroma component does not need to wait until the luma RDOQ optimization is performed for all luma tiles. Instead, when the optimization for a particular luma tile is completed, it can be directly used as input for the optimization of the corresponding chroma latent tile.
[0307] Derivation of tile map: Since RDOQ is an encoder-only process, the tiles used need not be signaled. In this implementation, a regular tile grid was used for the luma and chroma components. Each grid is described by offset and size values (encoder arguments). Using these, the tile map for a given component can be derived (see below).
[0308] Optimization of luma components: RDOQ sequentially and iteratively optimizes the cost of the tiles given by Cost = R + λD Here, R is an estimate of the number of bits required to encode the tile, D is a metric (peak signal-to-noise ratio (PSNR), multi-scale structural similarity (MS-SSIM), etc.) for the distortion of the reconstruction (of the tile) compared to the original. λ is a parameter set according to the operating point of the encoder. For im_tile
[0309] R and D are obtained as follows. i There are N overlapping regions (tiles) in the image. Each tile (im_tile
[0310] ) has a corresponding tile in the latent space (lat_tile i ). Further, im_tile i covers only a subset of the full receptive field of lat_tile i . For each im_tile in the signal space, the respective matching lat_tile in the latent space is derived as follows. Here, alignment_size is a power of 2 that depends on the number of downsampling layers in the subnetwork. image_tile = im_tilei lat_tile_start_y = image_tile.position.y / / alignment_size lat_tile_start_x = image_tile.position.x / / alignment_size if image_tile.size.height % alignment_size: height = math.ceil(image_tile.size.height / alignment_size) else: height = image_tile.size.height / / alignment_size if image_tile.size.width % alignment_size: width = math.ceil(image_tile.size.width / alignment_size) else: width = image_tile.size.width / / alignment_size lat_tilei = (lat_tile_start_x, lat_tile_start_y, width, height)
[0311] Using the lat_tile, the corresponding region of the latent space is extracted and processed by the decoder sub-network to obtain the reconstructed tile. The decoder sub-network also has im_tile as an auxiliary input, which is required to correctly pad, especially at the image boundaries where the tile size may not be a multiple of alignment_size. Using im_tile, the corresponding region is extracted from the original image. D can be calculated as a function (peak signal-to-noise ratio (PSNR), multi-scale structural similarity (MS-SSIM), etc.) of the original tile and the reconstructed tile. R is obtained by calling the encoding function for the extracted latent tile and measuring / estimating the amount of bits required to encode it. The RDOQ process is sequential and iterative, i.e., it is performed iteratively over a number of iteration steps (typically 10 to 30 iteration steps).
[0312] Optimization of chroma components: RDOQ sequentially and iteratively optimizes the cost for the tiles given by Cost = R + λD Here, R is an estimated value of the number of bits required to encode the tile, D is a metric (such as peak signal-to-noise ratio (PSNR), multi-scale structural similarity (MS-SSIM), etc.) for the distortion of the reconstruction (of the tile) compared to the original. λ is a parameter set according to the operating point of the encoder. In the proposed implementation, chroma is encoded conditionally with respect to luma. Thus, R here also includes the number of bits required to encode the corresponding luma tile:
[0313] R = R luma + R chroma R luma remains constant throughout the chroma RDOQ optimization process. R chroma and D for im_tile i are obtained as follows.
[0314] There are N overlapping regions (tiles) in the image. Each tile (im_tile i ) has a corresponding tile in the latent space (lat_tile i ). Furthermore, im_tile i covers only a subset of the full receptive field of lat_tile i . For each im_tile in the signal space, the matching lat_tile in the latent space is derived as follows. Here, alignment_size is a power of 2 that depends on the number of downsampling layers in the subnetwork. image_tile = im_tile i lat_tile_start_y = image_tile.position.y / / alignment_size lat_tile_start_x = image_tile.position.x / / alignment_size if image_tile.size.height % alignment_size: height = math.ceil(image_tile.size.height / alignment_size) else: height = image_tile.size.height / / alignment_size if image_tile.size.width % alignment_size: width = math.ceil(image_tile.size.width / alignment_size) else: width = image_tile.size.width / / alignment_size lat_tile i = (lat_tile_start_x, lat_tile_start_y, width, height)
[0315] Using the lat_tile, the corresponding region of the chroma latent space lat_UV is extracted. Further, the corresponding region of the luma latent space lat_Y is determined. For the YUV420 example, one way to do this is to downsample the luma latent space by a factor of 2 and then extract lat_Y using the same lat_tile as for the chroma. Both lat_Y and lat_UV are then processed by the decoder subnetwork to obtain the reconstructed chroma tile. The decoder subnetwork also has the im_tile as an auxiliary input, which is needed to correctly perform padding, especially at the image boundaries where the tile size may not be a multiple of the alignment_size. Using the im_tile, the corresponding region is extracted from the original image. Then, D can be calculated as a function (peak signal-to-noise ratio (PSNR), multi-scale structural similarity (MS-SSIM), etc.) of the original chroma tile and the reconstructed chroma tile. R chroma is obtained by calling the encoding function for the extracted chroma latent tile and measuring / estimating the amount of bits required to encode it.
[0316] The RDOQ process is sequential and iterative, i.e., it is performed iteratively over a number of iteration steps (typically 10 to 30 times).
[0317] In this exemplary and non-limiting embodiment, there is provided a non-transitory medium storing a computer program including code that, when executed on one or more processors, performs steps of a method for processing an input tensor representing picture data described in the second embodiment. Each flowchart is shown in FIG. 26. In step S2610, a first component of the input tensor is processed, which includes dividing the first component into a first plurality of tiles. Similarly, in step S2630, a second component of the input tensor is processed, which includes dividing the second component into a second plurality of tiles. Then, in steps S2620 and S2640, each of the first and second pluralities of tiles is processed separately. As suggested by the flowchart of FIG. 26, steps S2610 and S2620 are executed separately from steps S2630 and S2640, which reflects processing the first and second components in two separate pipelines as shown in FIG. 25 for the luma Y component (the first component) and the chroma components UV (the second component). The processing steps S2610 and / or S2630 of dividing the first and second tensors into the first and second pluralities of tiles may be part of the encoding process of encoder 2501 and / or encoder 2508. Further processing the first and second pluralities of tiles in different ways may include any of encoding, quantization of RDOQ, decoding, and post-filtering, as executed by the respective modules shown in FIG. 25. At the end of the separate processing of the first and second pluralities of tiles, a reconstructed first component and a reconstructed second component are provided. The reconstructed components may be the reconstructed picture data of the luma component ^y and the chroma component ^UV, as shown in FIG. 25. In the processing of FIG. 26, the dashed horizontal arrows indicate that the processing of the first and second pluralities of tiles may interact. May interact, for example, in that auxiliary information from the processing of the first / second plurality of tiles may be used for the processing of the second / first plurality of tiles.As described above in connection with FIG. 25, the processing of the UV chroma component in the chroma pipeline can use the information of the luma component of the processing of the luma pipeline as auxiliary information.
[0318] Furthermore, as already described, the present disclosure also provides a device configured to execute the steps of the methods described for the second embodiment.
[0319] In this exemplary and non - limiting embodiment, an apparatus for processing an input tensor representing picture data is provided. FIG. 27 shows an apparatus 2700 comprising a processing circuit 2710 having respective modules for performing the method steps of FIG. 26. Modules 2711 and 2712 are configured to perform processing of first and second components, including processing of respective first and second pluralities of tiles. Further, two modules 2713 and 2714 are used to divide first and second components of the input tensor in the spatial dimension into respective first and second pluralities of tiles. FIG. 27 shows separate modules 2711 and 2712, and modules 2713 and 2714, but it should be noted that the modules may be combined into one module, but the first and second components may be configured to be processed separately. This includes separate processing of the first and second pluralities of tiles to enable pipeline - based processing of each component. In particular, modules 2711 and 2712 may provide the functions of the individual modules shown in FIG. 25 for respective pipelines, including encoders 2501 and 2508, quantization units RDOQ 2502 and 2509, decoders 2506 and 2513, and post - filters 2507 and 2514. Further, the functions of the encoder - decoder hyperprior may be performed by respective modules 2711 and 2712 for respective pipelines. In FIG. 27, module 2715 performs processing including generating bitstreams Bitstream Y / UV 1 and Bitstream Y / UV 2, where an indication of the tile size of the first and / or second pluralities of tiles may also be included in respective bitstreams. Parse module 2716 parses Bitstream Y / UV 1 and Bitstream Y / UV 2. This includes extracting an indication (s) of tile size and / or position from Bitstream Y / UV 1.
[0320] In this exemplary and non - limiting embodiment, an apparatus for processing an input tensor representing picture data is provided. The apparatus includes one or more processors and a non - transitory computer - readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors. When the programming is executed by the one or more processors, the apparatus is configured to execute the method described in the second embodiment.
[0321] Further details regarding signaling Instructions regarding the tile size of the first and / or second plurality of tiles are included in the bitstream Bitstream Y / UV 1, as discussed above. This also applies to instructions for tile position and / or tile duplication, and / or scaling and / or filter index, etc. Alternatively, all or part of the instructions may be included in side information. The side information may then be included in the bitstream Bitstream 1, from which the decoder parses the side information to determine the tile size, position, etc. required for the decoding process (decompression). For example, instructions related to tiles (e.g., tile size, position, and / or duplication) may be included in the first side information. Also, instructions related to scaling may be included in the second side information, and instructions related to the filter model (e.g., filter index, etc.) may be included in the third side information. In other words, the instructions may be grouped and included in group - specific side information (e.g., the first to third side information). Thereby, the groups refer here to groups such as "tile", "filter", etc.
[0322] In the following, with reference to FIG. 20 showing some signaling examples on the decoder side, further details regarding signaling instructions are provided for an example of a tile. It should be noted that the same applies to the encoder side. FIG. 20 shows various parameters such as the sizes of regions Li, Ri, and overlapping regions that can be included in (and parsed from) the bitstream. Region Li refers to a tile of the first and / or second tensor to be processed (e.g., a tile of the input tensor x of a certain component in FIG. 25), and region Ri refers to the corresponding tile generated as output after processing tile Li. For example, tile Ri may be a tile of output y in FIG. 25 after processing the input tensor x.
[0323] For example, the side information includes one or more of the following instructions. · The number of the input subsets · The size of the input set, i.e., the number of tiles (e.g., luma components and / or chroma components) · The size (h1, w1) of each of the two or more input subsets, i.e., the size of the tile to be processed · The size (H, W) of the reconstructed picture (R) · The size (H1, W1) of each of the two or more output subsets, i.e., the size of the tile after processing · The overlap between the two or more input subsets (L1, L2), i.e., the amount of overlap between the tiles to be processed · The overlap between the two or more output subsets (R1, R2), i.e., the amount of overlap between the tiles after processing.
[0324] Therefore, signaling of various parameters through side information can be performed in a flexible manner. Therefore, the signaling overhead can be adapted depending on which of the above parameters are signaled in the side information, while other parameters should be derived from those parameters being signaled. The respective sizes of the two or more input subsets may be different. Alternatively, the input subsets may have a common size.
[0325] In one example, the position and / or amount of samples to be cropped are determined according to a neural network size change parameter of a neural network that specifies the relationship between the size of the input subset indicated in the side information, the size of the input to the network, and the size of the output from the network. Therefore, the position and / or cropping amount can be determined more accurately by considering both the size of the input subset and the characteristics of the neural network (i.e., its size change parameter). Therefore, the cropping amount and / or position may be adapted to the characteristics of the neural network, which further improves the quality of the reconstructed picture data.
[0326] The size change parameter may be an additive term subtracted from the input size to obtain the output size. In other words, the output size of the output subset is related to its corresponding input subset by exactly an integer. Alternatively, the size change parameter may be a ratio. In this case, the size of the output subset is associated with the size of the input subset by multiplying the size of the input subset by the ratio to obtain the size of the output subset.
[0327] As described above, the determination of Li, Ri, and the cropping amount can be obtained from the bitstream according to a predetermined rule or a combination of the two. · An indication of the amount of cropping may be included in (parsed from) the bitstream. In this case, the side information includes the amount of cropping (duplication amount). · The amount of cropping may be a fixed number. For example, such a number may be predefined by a standard, or may be fixed once the relationship between the input and output sizes (dimensions) is known. · The amount of cropping may be related to horizontal, vertical, or both-direction cropping.
[0328] Cropping can be performed according to a preconfigured rule. After the amount of cropping is obtained, the cropping rule may be as follows. · It follows the position of Ri (such as upper left, center, etc.) in the output space. If the side of Ri does not coincide with the output boundary, cropping can be applied to that side (top, left, bottom, or right).
[0329] The size and / or coordinates of Li (i.e., the tile) can be included in the bitstream. Alternatively, the number of partitions can be indicated in the bitstream, and the size of each Li can be calculated based on the input size and the number of partitions.
[0330] The duplication amount of each input subset Li can be as follows. · An indication of the amount of duplication may be included in (parsed from or derived from) the bitstream. · The amount of duplication may be a fixed number. As described above, "fixed" in this context means known by a convention such as a standard or a proprietary configuration, or preconfigured as part of the encoding parameters or neural network parameters. · The amount of duplication may be related to horizontal, vertical, or both-direction cropping. · The amount of duplication can be calculated based on the amount of cropping.
[0331] The following provides several numerical examples to show which parameters can be signaled via the side information included in (and parsed from) the bitstream, and then how these signaled parameters are used to derive the remaining parameters. These examples are for illustrative purposes only and do not limit the present disclosure.
[0332] For example, the bitstream can include the following information related to the signaling of Li. · Number of partitions on the vertical axis = 2. This corresponds to an example where the space L is vertically divided into two parts in FIG. 20. · Number of partitions on the horizontal axis = 2. This corresponds to an example where the space L is horizontally divided into two parts in FIG. 20. · Equal - size partition flag = true. This is illustrated in the figure by showing L1, L2, L3, and L4 having the same size. · Size of the input space L (wL = 200, hL = 200). The width w and height h are measured in units of the number of samples in these embodiments. · Duplication amount = 10. In this example, the duplication is measured in units of the number of samples.
[0333] According to the above information, since the duplication amount is 10 and the partitions are shown to be of equal size, the size of the partition can be obtained as w=(200 / 2 + 10)=110, h=(200 / 2 + 10)=110.
[0334] Also, since the number of partitions on each axis is 2 and the size of the partition is (110, 110), the upper - left coordinates of the partitions can be obtained as follows. · Upper - left coordinates for the first partition L1(x = 0, y = 0) · Upper - left coordinates for the second partition L2(x = 90, y = 0) · Upper - left coordinates for the third partition L3(x = 0, y = 90) · The upper left coordinates L4(x = 90, y = 90) for the fourth partition.
[0335] The following examples illustrate various options for signaling all or some of the above parameters, with reference to FIG. 20. FIG. 20 shows how various parameters related to the input subset Li, the output subset Ri, the input picture, and the reconstructed picture are linked.
[0336] Note that the above-signaled parameters do not limit the present disclosure. As will be described below, there are many possible ways to signal information from which the sizes, cropping, or padding of the input space and output space and subspaces can be derived. Some further examples are shown below.
[0337] The first signaling example: FIG. 20 shows a first example in which the following information is included in the bitstream. · The number of regions in the latent space (corresponding to the input space on the decoder side). This is equal to 4. · The total size (height and width) of the latent space. This is equal to (h, w) (referred to as wL and hL above). · h1 and w1 used to derive the size of the regions, i.e., the size of the input subsets (here the size of the four Lis). · The total size (H, W) of the reconstructed output R. · H1 and W1. H1 and W1 represent the size of the output subsets.
[0338] Next, the following information is predefined or determined in advance. · The amount of overlap X of the regions Ri. For example, X also determines the cropping amount. · The amount of overlap y between the regions Li.
[0339] According to the information included in the bitstream and the predefined information, the sizes of Li and Ri can be determined as follows. ·L1 = (h1 + y, w1 + y) ·L2 = ((h - h1) + y, w1 + y) ·L3 = (h1 + y, (w - w1) + y) ·L4 = ((h - h1) + y, (w - w1) + y) ·R1 = (H1 + X, W1 + X) ·R2 = ((H - H1) + X, W1 + X) ·R3 = (H1 + X, (W - W1) + X) ·R4 = ((H - H1) + X, (W - W1) + X)
[0340] As can be recognized from the first signaling example, the size (h1, w1) of the input subset L1 is used to derive the sizes of all the remaining input subsets L2 - L4. This is possible because the same amount of overlap y is used for the input subsets L1 - L4, as shown in FIG. 20. In this case, only a few parameters need to be signaled. The same argument applies to the output subsets R1 - R4, where only the signaling of the size (H1, W1) of the output subset R1 is required to derive the sizes of the output subsets R2 - R4.
[0341] In the above example, h1 and w1, and H1 and W1 are the intermediate coordinates in the input space and the output space respectively. Thus, in this first signaling example, a single coordinate (h1, w1) and (H1, W1) is used to calculate the four - way partitioning of the input space and the output space respectively. Alternatively, the sizes of two or more input subsets and / or output subsets can be signaled.
[0342] In another example, if the structure of the NN that processes Li, i.e., how the output size will be when the input size is Li, is known, it may be possible to calculate Ri from Li. In this case, the size (Hi, Wi) of the output subset Ri may not need to be signaled through side information. However, in some other implementations, since the determination of the size Ri may not be possible until the actual NN operation is executed, it may be desirable to signal the size Ri in the bitstream (as in this case).
[0343] Second signaling example: The second example of signaling involves determining H1 and W1 according to a formula based on h1 and w1. The formula may be, for example, as follows: ·H1 = (h1 + y) * scalar - X ·W1 = (w1 + y) * scalar - X Here, the scalar is a positive number. The scalar is related to the size change ratio of the encoder and / or decoder network. For example, the scalar may be an integer such as 16 for the decoder and a fraction such as 1 / 16 for the encoder. Thus, in the second signaling example, H1 and W1 are not signaled in the bitstream but rather are derived from the signaled sizes of the respective input subsets L1. Also, the scalar is an example of a size change parameter.
[0344] Third signaling example: In the third example of signaling, the amount of overlap y between the regions Li is not determined in advance but rather is signaled in the bitstream. The amount of cropping X of the output subset is then determined according to the following formula based on the amount of cropping y of the input subset. ·X = y * scalar Here, the scalar is a positive number. The scalar is related to the size change ratio of the encoder and / or decoder network. For example, the scalar is an integer such as 16 for the decoder and a fraction such as 1 / 16 for the encoder.
[0345] It should be noted that the present disclosure is not limited to a specific framework. Furthermore, the present disclosure is not restricted to image or video compression and can also be applied to object detection, image generation, and recognition systems.
[0346] The present invention can be implemented in hardware (HW) and / or software (SW). Furthermore, the HW-based implementation may be combined with the SW-based implementation.
[0347] For clarity, any of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0348] The encoding and decoding processes performed by the VAE encoder-decoder described above and shown in FIG. 25 may be implemented within the coding system 10 of FIG. 1A. Thereby, the source device 12 represents the encoding side and provides compression of the input picture data 21 including the input tensor x of FIG. 25, which may be the components Y and UV respectively. In particular, the encoder 20 of FIG. 1A may include modules for processing (e.g., compressing and / or decompressing) according to the present disclosure for processing multiple components independently. For example, the encoder 20 of FIG. 1A may include an encoder 2501, a quantizer or RDOQ 2502, and an arithmetic encoder 1605 for processing the luma component Y. The encoder 20 may further include modules of hyper priors such as a hyper decoder 2503, a quantizer or RDOQ 2504, and an arithmetic encoder 1609. Further, the encoder 20 of FIG. 1A may include an encoder 2508, a quantizer or RDOQ 2509, and an arithmetic encoder 1605 for processing the chroma components UV. The encoder 20 may further include modules of hyper priors such as a hyper decoder 2510, a quantizer or RDOQ 2511, and an arithmetic encoder 1609.
[0349] Similarly, the destination device 14 of FIG. 1A represents the decoding side that provides decompression of the input tensor representing picture data. In particular, the decoder 30 of FIG. 1A may include modules for decompression processing of the luma component Y, such as the decoder 2506 and the post-filter 2507 of FIG. 25, together with the arithmetic decoder 1606. Additionally, the decoder 30 may further include hyper priors for decoding ^z, such as the arithmetic decoder 1610, the hyper decoder 2505, and the arithmetic decoder 1606. The decoder 30 may further include hyper priors for decoding ^z, such as the arithmetic decoder 1610, the hyper decoder 2505, and the arithmetic decoder 1606. To process the chroma components UV, the decoder 30 may include the decoder 2513 and the post-filter 2514 of FIG. 25, together with the arithmetic decoder 1606. Additionally, the decoder 30 may further include hyper priors for decoding ^z for chroma, such as the arithmetic decoder 1606, the hyper decoder 2512, and the arithmetic decoder 1610. In other words, the encoder 20 and the decoder 30 of FIG. 1A may be implemented and configured to include any of the modules of FIG. 25 to implement the encoding or decoding processing of multiple components in their respective pipelines (e.g., luma and chroma pipelines), where the input tensor is divided into multiple tiles and has multiple components processed as described in the second embodiment. Although FIG. 1A shows the encoder 20 and the decoder 30 separately, they may be implemented via the processing circuit 46 of FIG. 1B. In other words, the processing circuit 46 can provide the encoding-decoding processing function of the present disclosure by implementing circuits for the modules of FIG. 25 for each pipeline or both pipelines (luma and / or chroma).
[0350] Similarly, the video coding device 200 having the coding module 270 of the processor 230 in FIG. 2 can perform the functions of the processes (compression and decompression) of the present disclosure. For example, the video coding device 200 can be an encoder or a decoder having the respective modules in FIG. 25 to perform the encoding or decoding process as described above.
[0351] The apparatus 300 in FIG. 3 may be implemented as an encoder and / or decoder having the encoders 2501, 2508, quantizers or RDOQs 2502, 2509, decoders 2506, 2513, post filters 2507, 2514, hyper encoders 2503, 2510, and hyper decoders 2505, 2512, together with arithmetic encoders 2505, 2509, and arithmetic decoders 2506, 2510, thereby performing tile processing for each component as described according to the second embodiment. For example, the processor 302 in FIG. 3 may have respective circuits for performing compression and / or decompression processing according to the aforementioned methods.
[0352] The exemplary implementations of the encoder 20 shown in FIG. 4 and the decoder 30 shown in FIG. 5 can also implement the encoding and decoding functions of the present disclosure. For example, the splitting unit 452 in FIG. 4 may perform splitting the first tensor and / or the second tensor into a first plurality of tiles and / or a second plurality of tiles for the first component and the second component, respectively, as executed by the encoders 2501 and 2508 in FIG. 25. Thereby, the syntax element 465 may include an indication of the size and position of the tiles, together with an indication such as a filter index. Similarly, the quantization unit 408 may perform quantization or RDOQ of the RDOQ modules 2502 and 2509 in FIG. 25, and the entropy encoding unit 470 may implement the functions of the hyper priors (i.e., modules 2503, 2505, 2510, 2511, 1605, 1608, 1609). Also, the entropy decoding unit 504 in FIG. 5 can perform the functions of the decoders 2506 and 2513 in FIG. 25 by splitting the encoded picture data 21 (input tensor) into tiles and parsing an indication of the size or position of the tiles, etc. from the bitstream as the syntax element 566. The entropy decoding unit 504 may further implement the hyper prior modules (i.e., modules 2505, 2512, 1610). The post filtering 2507 and 2514 in FIG. 25 can be performed, for example, also by the entropy decoding unit 504. Alternatively, the post filters 2507 and 2514 may be implemented within the mode application unit 560 as additional units (not shown in FIG. 5).
[0353] Further embodiments According to an aspect of the present disclosure, a method for processing an input tensor representing picture data is provided, the method including: processing a plurality of components of the input tensor including a first component and a second component in a spatial dimension, the processing including processing the first component in the spatial dimension by dividing the first component into a first plurality of tiles and separately processing the tiles of the first plurality of tiles; and processing the second component in the spatial dimension by dividing the second component into a second plurality of tiles and separately processing the tiles of the second plurality of tiles, at least two co-located tiles of the first plurality of tiles and the second plurality of tiles being of different sizes. As a result, the input tensor representing picture data can be efficiently processed component-by-component by using tiles in a sample-aligned manner within a plurality of pipelines. Thus, memory requirements are reduced while improving processing performance (e.g., compression and decompression) without increasing the complexity of the calculation.
[0354] In some exemplary implementations, at least two tiles of the first plurality of tiles are processed independently or in parallel; and / or at least two tiles of the second plurality of tiles are processed independently or in parallel. Thus, the components of the input tensor can be processed at high speed and the processing efficiency can be improved.
[0355] In a further implementation, the first component represents the luma component of the picture data and the second component represents the chroma component of the picture data. Thus, both the luma component and the chroma component can be processed via a plurality of pipelines within the same processing framework.
[0356] In one example, the tiles of the first plurality of tiles adjacent in at least one dimension of the spatial dimension partially overlap; and / or the tiles of the second plurality of tiles adjacent in at least one dimension of the spatial dimension partially overlap. Thus, the quality of the reconstructed picture can be improved, especially along the boundaries of the tiles. Thus, picture artifacts can be reduced.
[0357] According to one implementation, the dividing of the first component includes determining the size of tiles within a first plurality of tiles based on a first predetermined condition, and / or the dividing of the second component includes determining the size of tiles within a second plurality of tiles based on a second predetermined condition. For example, the first predetermined condition and / or the second predetermined condition are based on available decoder hardware resources and / or motion present in the picture data. Thus, the tile size can be adapted and optimized according to available decoder resources and / or motion, enabling a content-based tile size. In a further example, determining the size of tiles in the second plurality of tiles includes scaling the tiles of the first plurality of tiles. As a result, the tile size of the second plurality of tiles can be determined quickly and the efficiency of tile processing can be improved.
[0358] In an exemplary implementation, an indication of the determined size of tiles within the first plurality of tiles and / or within the second plurality of tiles is encoded in the bitstream. Thus, the representation of the tile size is efficiently included in the bitstream and the processing required is at a low level.
[0359] In another implementation, the size of all tiles within the first plurality of tiles is the same, and / or the size of all tiles within the second plurality of tiles is the same. As a result, the tiles can be processed efficiently without additional processing for handling different tile sizes, and tile processing can be accelerated.
[0360] In a second example, the indication further includes the position of tiles within the first plurality of tiles and / or within the second plurality of tiles.
[0361] According to one implementation, the first component is a luma component, and the indication of the tile size of the first plurality of tiles is included in the bitstream; the second component is a chroma component, and the indication of the scaling factor is included in the bitstream, and the scaling factor associates the tile size of the first plurality of tiles with the tile size of the second plurality of tiles. Thus, the tile size of the chroma component can be quickly obtained by a fast operation of scaling the tile size of the luma component. Further, the overhead for signaling the tile size for chroma can be reduced by using the scaling factor as an indication.
[0362] In an exemplary implementation, the processing of the input tensor includes processing that is part of picture or video compression. For example, the processing of the first component and / or the second component includes one of picture encoding by a neural network, rate distortion optimization quantization (RDOQ), and picture filtering. Thus, the compression processing can be performed in a flexible manner that includes various types of processing (encoding, RDOQ, filtering).
[0363] A further exemplary implementation includes generating a bitstream by including the outputs of the processing of the first component and the second component in the bitstream. Thus, the processing output can be quickly included in the bitstream, and the processing required is at a low level.
[0364] In an exemplary implementation, processing of the input tensor includes processing that is part of decompressing a picture or video. For example, processing of the first component and / or the second component includes one of picture decoding by a neural network and picture filtering. Thus, the decompression process can be performed in a flexible manner that includes various types of processing (encoding, filtering). For example, processing of the second component includes decoding the chroma component of a picture based on the representation of the luma component of the picture. Thus, the luma component can be used as auxiliary information for decoding the chroma component. This can improve the quality of the decoded chroma. In a further example, processing of the first component and / or the second component includes picture post-filtering; for at least two tiles of a first plurality of tiles, one or more parameters of the post-filtering are different and are extracted from the bitstream; for at least two tiles of a second plurality of tiles, one or more parameters of the post-filtering are different and are extracted from the bitstream. Thus, the filter parameters can be efficiently signaled via the bitstream. Further, the post-filtering is performed using filter parameters adapted to the tile size to improve the quality of the reconstructed picture data.
[0365] In an exemplary implementation, the input tensor is a picture or sequence of pictures that includes one or more of a plurality of components, and at least one of the components is a color component.
[0366] According to an aspect of the present disclosure, there is provided a computer program stored in a non-transitory medium that, when executed on one or more processors, includes code for performing any of the steps of the foregoing aspects of the present disclosure.
[0367] According to an aspect of the present disclosure, there is provided an apparatus for processing an input tensor representing picture data, the apparatus having a processing circuit configured to process a plurality of components of the input tensor including a first component and a second component in a spatial dimension, the processing including processing the first component in the spatial dimension by dividing the first component into a first plurality of tiles and processing the tiles of the first plurality of tiles separately; and processing the second component in the spatial dimension by dividing the second component into a second plurality of tiles and processing the tiles of the second plurality of tiles separately, wherein at least two co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes.
[0368] According to an aspect of the present disclosure, there is provided an apparatus for processing an input tensor representing picture data, the apparatus comprising one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming configuring the processing apparatus to perform a method according to any of the foregoing aspects of the present disclosure when executed by the one or more processors.
[0369] In summary, the present disclosure relates to neural network-based picture encoding and decoding of tile-based image regions. An input tensor representing picture data is processed by a neural network including at least a first and a second sub-network. The first sub-network is applied to a first tensor, where the first tensor is divided in the spatial dimension into a first plurality of tiles. The first tile is then further processed by the first sub-network. After application of the first sub-network, a second sub-network is applied to a second tensor, which is divided in the spatial dimension into a second plurality of tiles. The second tile is then further processed by the second sub-network. Among the first and second pluralities of tiles, there are at least two respective co-located tiles of different sizes. In the case of encoding, the first and second sub-networks perform a part of the compression including picture encoding, rate distortion optimized quantization, and picture filtering. In the case of decoding, the first and second sub-networks perform a part of the decompression including picture decoding and picture filtering.
[0370] Furthermore, the present disclosure relates to picture encoding and decoding of tile-based image regions. In particular, a plurality of components of an input tensor including first and second components in the spatial dimension are processed in a plurality of pipelines. Processing the first component includes dividing the first component in the spatial dimension into a first plurality of tiles. Similarly, processing the second component includes dividing the second component in the spatial dimension into a second plurality of tiles. Then, each of the first and second pluralities of tiles is processed separately. Among the first and second pluralities of tiles, there are at least two respective co-located tiles of different sizes. In the case of compression, the processing of the first and / or second component includes picture encoding, rate distortion optimized quantization, and picture filtering. In the case of decompression, the processing includes picture decoding and picture filtering.
Claims
1. A method for encoding an input tensor representing picture data, the method comprising: processing the input tensor by a neural network comprising at least a first sub-network and a second sub-network, the processing comprising: applying the first sub-network to a first tensor by dividing the first tensor into a first plurality of tiles in a spatial dimension and processing the first plurality of tiles by the first sub-network; after applying the first sub-network, applying the second sub-network to a second tensor by dividing the second tensor into a second plurality of tiles in the spatial dimension and processing the second plurality of tiles by the second sub-network, wherein at least two co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes; A method.
2. tiles of the first plurality of tiles adjacent in at least one dimension of the spatial dimension overlap partially, and / or tiles of the second plurality of tiles adjacent in at least one dimension of the spatial dimension overlap partially; The method according to claim 1.
3. tiles of the first plurality of tiles are processed independently by the first sub-network, and / or tiles of the second plurality of tiles are processed independently by the second sub-network; The method according to claim 1 or 2.
4. at least two tiles of the first plurality of tiles are processed in parallel by the first sub-network, and / or at least two tiles of the second plurality of tiles are processed in parallel by the second sub-network; The method according to claim 3.
5. dividing the first tensor includes determining the size of tiles in the first plurality of tiles based on a first predetermined condition, and / or dividing the second tensor includes determining the size of tiles in the second plurality of tiles based on a second predetermined condition; The method according to any one of claims 1 to 4.
6. The method according to claim 5, wherein the first predetermined condition and / or the second predetermined condition are based on available decoder hardware resources and / or motion present in the picture data.
7. The first subnetwork performs processing by one or more layers including at least one convolutional layer and at least one pooling layer; and / or The second subnetwork performs processing by one or more layers including at least one convolutional layer and at least one pooling layer. The method according to any one of claims 1 to 6.
8. The method according to any one of claims 1 to 7, wherein the first subnetwork and the second subnetwork perform respective processing that is part of picture or video compression.
9. The first subnetwork and / or the second subnetwork - Picture encoding by a convolutional subnetwork; - Rate-Distortion Optimization Quantization (RDOQ); - Picture filtering The method according to claim 8, which performs one of the above.
10. The method according to any one of claims 1 to 9, wherein the input tensor is a picture or a sequence of pictures including one or more components at least one of which is a color component.
11. The input tensor has at least two components, namely a first component and a second component; The first subnetwork divides the first component into a third plurality of tiles and divides the second component into a fourth plurality of tiles, and at least two co-located tiles of the third plurality of tiles and the fourth plurality of tiles have different sizes; and / or The second subnetwork divides the first component into a fifth plurality of tiles and divides the second component into a sixth plurality of tiles, and at least two co-located tiles of the fifth plurality of tiles and the sixth plurality of tiles have different sizes. The method according to claim 10.
12. The method according to any one of claims 1 to 11, including generating a bitstream by including the output of the processing by the neural network in the bitstream.
13. The method according to claim 12, further comprising including in the bitstream an indication of the tile size in the first plurality of tiles and / or an indication of the tile size in the second plurality of tiles. **Claim 14** A method for decoding a tensor representing picture data, the method comprising: processing an input tensor representing the picture data by a neural network including at least a first sub-network and a second sub-network, the processing comprising: applying the first sub-network to a first tensor by dividing the first tensor into a first plurality of tiles in a spatial dimension and processing the first plurality of tiles by the first sub-network; after applying the first sub-network, applying the second sub-network to a second tensor by dividing the second tensor into a second plurality of tiles in the spatial dimension and processing the second plurality of tiles by the second sub-network, wherein at least two co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes. Method. **Claim 15** tiles of the first plurality of tiles adjacent in at least one dimension of the spatial dimension overlap partially, and / or tiles of the second plurality of tiles adjacent in at least one dimension of the spatial dimension overlap partially, The method according to claim 14. **Claim 16** tiles of the first plurality of tiles are processed independently by the first sub-network, and / or tiles of the second plurality of tiles are processed independently by the second sub-network, The method according to claim 14 or 15. **Claim 17** at least two tiles of the first plurality of tiles are processed in parallel by the first sub-network, and / or at least two tiles of the second plurality of tiles are processed in parallel by the second sub-network, The method according to claim 16. **Claim 18** Dividing the first tensor includes determining the tile size in the first plurality of tiles based on a first predetermined condition, and / or Said splitting of the second tensor includes determining the size of tiles in said second plurality of tiles based on a second predetermined condition. The method according to any one of claims 14 to 17.
19. The method according to claim 18, wherein the first predetermined condition and / or the second predetermined condition are based on available decoder hardware resources and / or motion present in the picture data.
20. The first subnetwork performs processing by one or more layers including at least one convolutional layer and at least one pooling layer, and / or The second subnetwork performs processing by one or more layers including at least one convolutional layer and at least one pooling layer. The method according to any one of claims 14 to 19.
21. The method according to any one of claims 14 to 20, wherein the first subnetwork and the second subnetwork perform respective processing that is part of decompressing a picture or video.
22. The first subnetwork and / or the second subnetwork · Picture decoding by a convolutional subnetwork, · Picture filtering performs one of them. The method according to claim 21.
23. The method according to any one of claims 14 to 22, wherein the input tensor is a picture or a sequence of pictures including one or more components at least one of which is a color component.
24. The input tensor has at least two components, namely a first component and a second component. The first subnetwork divides the first component into a third plurality of tiles and divides the second component into a fourth plurality of tiles, and at least two co-located tiles of the third plurality of tiles and the fourth plurality of tiles have different sizes; and / or The second subnetwork divides the first component into a fifth plurality of tiles and divides the second component into a sixth plurality of tiles, and at least two co-located tiles of the fifth plurality of tiles and the sixth plurality of tiles have different sizes. The method according to claim 23.
25. The method according to any one of claims 14 to 24, further comprising extracting the input tensor from the bitstream for processing by the neural network.
26. The second subnetwork performs picture post-filtering, for at least two of the second plurality of tiles, one or more parameters of the post-filtering are different and are extracted from the bitstream, The method according to claim 25.
27. The method according to claim 25 or 26, further comprising parsing an indication of the tile size in the first plurality of tiles and / or an indication of the tile size in the second plurality of tiles from the bitstream.
28. A computer program stored in a non-transitory medium, including code that, when executed on one or more processors, executes the steps of the method according to any one of claims 1 to 27.
29. A processing device is provided for encoding an input tensor representing picture data, the processing device comprising: a processing circuit configured to process the input tensor by a neural network including at least a first subnetwork and a second subnetwork, the processing comprising: applying the first subnetwork to a first tensor by dividing the first tensor into a first plurality of tiles in a spatial dimension and processing the first plurality of tiles by the first subnetwork; after applying the first subnetwork, applying the second subnetwork to a second tensor by dividing the second tensor into a second plurality of tiles in the spatial dimension and processing the second plurality of tiles by the second subnetwork, wherein at least two co-located tiles of the first plurality of tiles and the second plurality of tiles are of different sizes, processing device.
30. A processing device for encoding an input tensor representing picture data, the processing device comprising: one or more processors; A non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming, when executed by the one or more processors, configuring the encoder to perform the method according to any one of claims 1 to 13. Processing device.
31. A processing device for decoding a tensor representing picture data, the processing device comprising: A processing circuit configured to process an input tensor representing the picture data by a neural network including at least a first sub-network and a second sub-network, the processing comprising: - Applying the first sub-network to a first tensor by dividing the first tensor into a first plurality of tiles in a spatial dimension and processing the first plurality of tiles by the first sub-network; - After applying the first sub-network, applying the second sub-network to a second tensor by dividing the second tensor into a second plurality of tiles in the spatial dimension and processing the second plurality of tiles by the second sub-network, wherein at least two co-located tiles of the first plurality of tiles and the second plurality of tiles have different sizes. Processing device.
32. A processing device for decoding a tensor representing picture data, the processing device comprising: One or more processors; A non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming, when executed by the one or more processors, configuring the decoder to perform the method according to any one of claims 14 to 29. Processing device.
Citation Information
Patent Citations
Implementation of a neural network in multicore hardware
GB2599910A
Tile partitions including subtiles in video coding
JP2021528003A
Data compression / decompression system and method
JP2022187683A
Neural network-based bitstream decoding and encoding
JP2023547941A
A Front-End Architecture for Neural Network-Based Video Coding
JP2023553369A