Resampling in image compression
Neural networks are employed for image and video coding, addressing compression performance limitations by parsing bitstreams and resampling with integer factors, resulting in efficient and accurate image reconstruction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-06-27
- Publication Date
- 2026-07-17
AI Technical Summary
Conventional image and video coding methods face challenges with high costs and practical limitations in compression performance, particularly in terms of memory and processing speed, necessitating improved techniques for efficient compression and decompression with minimal quality sacrifice.
A method utilizing neural networks for image and video coding, involving parsing bitstreams to obtain latent tensors, resampling using integer factors, and processing secondary components based on primary components, allowing for parallel processing and reduced processing load, while maintaining high accuracy in image reconstruction.
The method enables efficient compression and decompression with improved accuracy and reduced memory requirements, achieving high compression performance and ease of use by leveraging neural networks for reliable coding and decoding.
Smart Images

Figure 2026524097000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to International Application PCT / EP2023 / 067439, filed on 27 June 2023. The disclosure of the above patent application is incorporated herein by reference in its entirety.
[0002] This disclosure relates generally to the field of image and video coding, and more particularly to image and video coding with resampling in image compression. [Background technology]
[0003] Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the internet and mobile networks, real-time conversational applications like video chat, video conferencing, DVD and Blu-ray discs, video content acquisition and editing systems, and camcorders for security purposes.
[0004] Even the amount of video data required to create a relatively short video can be substantial, which can pose difficulties when the data is to be streamed or otherwise transmitted across communication networks with limited bandwidth. Therefore, video data is generally compressed before being transmitted across modern telecommunications networks. Video size can also be a problem when video is stored on a storage device, as memory resources can be limited. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Compression techniques are also preferably applied in the context of still image encoding.
[0005] With limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques that improve compression ratios with little or no sacrifice in image quality are desirable.
[0006] Neural network (NN) and deep learning (DL) techniques, which utilize artificial neural networks, have been used for some time now in the field of encoding and decoding video, images (for example, still images), and other similar technologies.
[0007] It is desirable to further improve the efficiency of such image coding (video coding or still image coding) based on a trained network that takes into account limitations in available memory and / or processing speed.
[0008] In detail, conventional image compression methods suffer from high costs and practical limitations in compression performance. [Overview of the Initiative]
[0009] The present invention relates in particular to a method and apparatus for coding image or video data using a neural network, for example, a neural network described in the embodiments for carrying out the invention described below. The use of a neural network may enable reliable coding and decoding in a self-learning manner as well as estimation of an entropy model, resulting in high accuracy of the image reconstructed from compressed input data.
[0010] The above and other objectives are achieved by the subject matter of the independent claims. Further implementations are evident from the dependent claims, description, and drawings. [Means for solving the problem]
[0011] According to a first embodiment, a method is provided for reconstructing at least a portion of an image, the method comprising: parsing a first bitstream to obtain a first latent tensor (for at least a portion of the image); and processing the first latent tensor to obtain a first tensor representing the primary component of the image. Furthermore, the method comprises: parsing a second bitstream different from the first bitstream to obtain a second latent tensor different from the first latent tensor (for at least a portion of the image); resampling the first latent tensor based on an integer factor to obtain a resampled first latent tensor; and obtaining a second tensor representing at least one secondary component of the image based on the second latent tensor and the resampled first latent tensor.
[0012] Since integer factors are used to obtain the resampled first latent tensor, this is extremely beneficial for both compression performance and ease of use.
[0013] In principle, an image may be a still image or an intraframe of a video sequence. Here, and in the following description, it should be understood that an image has components, specifically luminance and color components. The components may be considered as the dimensions of an orthogonal basis representing a full-color image. For example, when an image is represented in YUV space, the components are lumens Y, chroma U, and chroma V. One of the components of the image is selected as the primary component, and one or more other components are selected as secondary (non-primary) components. The terms “secondary component” and “non-primary component” are used interchangeably herein and refer to components that are coded using the auxiliary information provided by the primary component. Encoding and decoding the secondary component using the auxiliary information provided by the primary component results in a high accuracy of the reconstructed image obtained after the decoding process.
[0014] The first latent tensor can be processed independently of the processing of the second latent tensor. In effect, even if data for the quadratic component is lost, the encoded primary component can be recovered. The compressed original image data can be reliably and quickly reconstructed by this method due to the possible parallel processing of the first and second bitstreams.
[0015] In one implementation, the primary component of an image is the lumen component, and at least one secondary component of the image is the chromen component. For example, two secondary components of an image, one of which is a chromen component and the other is a different chromen component, are coded in parallel. In another implementation, the primary component of an image is the chromen component, and at least one secondary component of the image is the lumen component. This results in a high degree of flexibility in the actual conditioning of one component by another component.
[0016] The overall coding may, in detail, include processing in latent space, which may enable the processing of downsampled input data, and therefore fixed processing with less processing load. Note that in this specification, the terms “downsampling” and “upsampling” are used to mean reducing and expanding the size of the tensor representation of the data, respectively.
[0017] At least one of the height or width dimensions of the first latent tensor may be smaller than the corresponding height or width dimension of the first tensor, and / or the height or width dimension of the second latent tensor may be smaller than the corresponding height or width dimension of the connected tensor. For example, a scaling factor of 16 or 32 in the height and / or width dimensions may be used.
[0018] It may be a fact that the sample size or subpixel offset of the second tensor in at least one of the height and width dimensions of the tensor is different from the sample size or subpixel offset of the first tensor in at least one of the height and width dimensions.
[0019] In this case as well, at least one of the height or width dimensions of the first latent tensor may be smaller than the corresponding height or width dimension of the first tensor, and / or the height or width dimension of the second latent tensor may be smaller than the corresponding height or width dimension of the connected tensor. For example, a reduction factor of 16 or 32 in the height and / or width dimensions may be used. Adjusting the sample location of the first tensor to match the sample location of the second tensor may involve, for example, downsampling the width and height of the first tensor by a factor of 2.
[0020] In one implementation, the first latent tensor has a channel dimension, and the second latent tensor has a channel dimension, where the size of the first latent tensor in the channel dimension is one of being greater than, less than, or equal to the size of the second latent tensor in the channel dimension. If the first-order component is considered to be of greater importance than the second-order component, which is usually the case, the channel length of the first-order component may be longer than that of the second-order component. If the signal of the first-order component is relatively clear and the signal of the non-first-order component is relatively noisy, the channel length of the first-order component may be shorter than that of the second-order component. Numerical experiments have shown that shorter channel lengths compared to those in the art can be used without significant degradation of the quality of the reconstructed image, and therefore memory requirements can be reduced.
[0021] Generally, a first tensor may be transformed into a first latent tensor using a first neural network, and a connected tensor may be transformed into a second latent tensor using a second neural network different from the first neural network. In this case, the first and second neural networks may be trained collaboratively to determine the size of the first latent tensor in the channel dimension and the size of the second latent tensor in the channel dimension. The determination of the channel length may be performed by exhaustive search or in a content-adaptive manner. A set of models may be trained, each model based on a different number of channels for coding the first and second-order components. Thereafter, the neural network may be able to optimize the channel lengths involved.
[0022] The determined channel length must also be used by the decoder used to reconstruct the encoded components. Thus, according to one implementation, the size of a first latent tensor in the channel dimension may be signaled in the first bitstream, and the size of a second latent tensor in the channel dimension may be signaled in the second bitstream. The signaling may be performed explicitly or implicitly, allowing the decoder to be directly informed of the channel length in a bit-saving manner.
[0023] In one implementation, a first bitstream is parsed based on a first entropy model, and a second bitstream is parsed based on a second entropy model different from the first. Such entropy models allow for the reliable estimation of the statistical properties used in the process of converting the tensor representation of data into a bitstream.
[0024] The method of disclosure may be advantageously implemented in the context of a hyper-prior architecture that sequentially provides secondary information useful for coding parts of an image, thereby improving the accuracy of parts of the reconstructed image.
[0025] The resampling process can be performed without interpolation (which is a computationally expensive operation). This can lead to improved coding efficiency.
[0026] The integer factors may be calculated based on a first scaling factor for the linear component and a second scaling factor for at least one quadratic component. For example, the integer factors are s UV / s Y It is calculated as, where, s Y represents the first scaling factor, s UV This represents the second scaling factor.
[0027] s Y When s exists in the bitstream,Y is obtained by parsing the bitstream. s Y When s does not exist in the bitstream, s Y is set to a default value. s Y The default value of may be 1.
[0028] s UV When s exists in the bitstream, s UV is obtained by parsing the bitstream. s UV When s does not exist in the bitstream, s UV is set to a default value. s UV The default value of may be 2.
[0029] As another example, the integer factor r may exist in the bitstream or may be set to a default value.
[0030] When r exists in the bitstream, r is obtained by parsing the bitstream. When r does not exist in the bitstream, r is set to a default value. The default value of r may be 2.
[0031] s Y When s exists in the bitstream, s Y is obtained by parsing the bitstream. s Y When s does not exist in the bitstream, s Y is set to a default value. s Y The default value of may be 1.
[0032] Then, s UV where s UV = r·s Y is derived, provided that s UV represents a scaling factor for at least one second-order component.
[0033] According to one implementation of the first embodiment, the first scaling factor may include a first horizontal scaling factor and a first vertical scaling factor. Similarly, the second scaling factor may include a second horizontal scaling factor and a second vertical scaling factor.
[0034] Resampling is performed by using nearest neighbor upsampling or nearest neighbor downsampling.
[0035] For nearest-near-side upsampling, This is denoted as s↑. This layer has a size [C, h in , w in It receives a tensor input of size [C, s·h]. in , s·w in Outputs the tensor output of ].
[0036] Interpolation is not required for this downsampling; it simply involves copying. Output[c, i, j]=Input[c, s·i, s·j], i=0,...,h in -1, j=0,...,w in -1, c=0,...,C-1 This is the case where s represents an integer factor.
[0037] For nearest-neighbor downsampling, This is represented as s↓. This layer has a size [C, s·h out , s·w out It receives a tensor input of size [C, h out , w out This outputs a tensor output of ], and this procedure reduces the spatial resolution of each tensor channel.
[0038] Interpolation is not required for this downsampling; it simply involves copying. Output[c, s·i, s·j]=Input[c, i, j], i=0,...,h in -1, j=0,...,w in-1, c=0,...,C-1 This is the case where s represents an integer factor.
[0039] According to one implementation of the method according to the first embodiment, the processing of the first latent tensor comprises converting the first latent tensor into a first tensor.
[0040] According to one implementation of the method according to the first aspect, obtaining a second tensor representing at least one quadratic component of an image is: The process comprises concatenating a second latent tensor with a resampled first latent tensor to obtain a connected tensor, and transforming the connected tensor into a second tensor. At least one of these transformations may involve upsampling. Thus, the processing in the latent space may be performed at a lower resolution, such as is necessary for the accurate reconstruction of the components in YUV space or any other space preferably used for image representation.
[0041] According to another implementation of the method according to the first embodiment, each of the first and second latent tensors has height and width dimensions, processing the first latent tensor comprises converting the first latent tensor to the first tensor, and processing the second latent tensor comprises determining whether the sample size or subpixel offset of the second latent tensor in at least one of the height and width dimensions is different from the sample size or subpixel offset of the first latent tensor in at least one of the height and width dimensions. When it is determined that the sample size or subpixel offset of the second latent tensor is different from the sample size or subpixel offset of the first latent tensor, the sample location of the first latent tensor is adjusted to match the sample location of the second latent tensor. Thereafter, the adjusted first latent tensor is obtained. Furthermore, when it is determined that the sample size or subpixel offset of the second latent tensor is different from the sample size or subpixel offset of the first latent tensor, the second latent tensor and the adjusted first latent tensor are concatenated to obtain a connected latent tensor; otherwise, the second latent tensor and the first latent tensor are concatenated to obtain a connected latent tensor, and the connected latent tensor is transformed into the second tensor.
[0042] A first bitstream may be processed by a first neural network, and a second bitstream may be processed by a second neural network different from the first neural network. A first latent tensor may be transformed by a third neural network different from the first and second networks, and a connected latent tensor may be transformed by a fourth neural network different from the first, second, and third networks.
[0043] According to another implementation of the method according to the first embodiment, the first latent tensor has a channel dimension, and the second latent tensor has a channel dimension, where the size of the first latent tensor in the channel dimension is one of being greater than, less than, or equal to the size of the second latent tensor in the channel dimension. Information about the sizes of the first and second latent tensors in the channel dimension may be obtained from information signaled in the first and second bitstreams, respectively.
[0044] According to another implementation of the method according to the first embodiment, the first tensor represents the first residual component of the residual for the first component of the image, and the second tensor represents at least one second residual component of the residual for at least one second component of the image, using information from the first latent tensor.
[0045] A second embodiment provides a method for reconstructing at least a portion of an image, the method comprising: parsing a first bitstream, for example, based on a first entropy model, to obtain a first latent tensor (for at least a portion of the image); and processing the first latent tensor to obtain a first tensor representing a first residual component of the residual for a first component of the image. Furthermore, the method comprises: parsing a second bitstream different from the first bitstream, for example, based on a second entropy model different from the first entropy model, to obtain a second latent tensor different from the first latent tensor (for at least a portion of the image); resampling the first latent tensor based on an integer factor to obtain a resampled first latent tensor; and obtaining a second tensor representing at least one second residual component of the residual for at least one second component of the image, based on the second latent tensor and the resampled first latent tensor.
[0046] Therefore, a residual is obtained having a first residual component for the primary component and a second residual component for at least one secondary component. In principle, the image can be a still image or an interframe of a video sequence.
[0047] The first and second entropy models may be provided by the hyperplier pipeline described above.
[0048] The first latent tensor may be processed independently of the processing of the second latent tensor.
[0049] The first-order component of the image may be a lumen component, and at least one of the second-order components of the image may be a chromen component. In this case, the second tensor may represent two residual components for two second-order components, one of which is a chromen component and the other is a different chromen component. Alternatively, the first-order component of the image may be a chromen component, and at least one of the second-order components of the image may be a lumen component.
[0050] Since integer factors are used to obtain the resampled first latent tensor, this is extremely beneficial for both compression performance and ease of use.
[0051] In principle, an image may be a still image or an intraframe of a video sequence. Here, and in the following description, it should be understood that an image has components, specifically luminance and color components. The components may be considered as the dimensions of an orthogonal basis representing a full-color image. For example, when an image is represented in YUV space, the components are lumens Y, chroma U, and chroma V. One of the components of the image is selected as the primary component, and one or more other components are selected as secondary (non-primary) components. The terms “secondary component” and “non-primary component” are used interchangeably herein and refer to components that are coded using the auxiliary information provided by the primary component. Encoding and decoding the secondary component using the auxiliary information provided by the primary component results in a high accuracy of the reconstructed image obtained after the decoding process.
[0052] The first latent tensor can be processed independently of the processing of the second latent tensor. In effect, even if data for the quadratic component is lost, the encoded primary component can be recovered. The compressed original image data can be reliably and quickly reconstructed by this method due to the possible parallel processing of the first and second bitstreams.
[0053] In one implementation, the primary component of an image is the lumen component, and at least one secondary component of the image is the chromen component. For example, two secondary components of an image, one of which is a chromen component and the other is a different chromen component, are coded in parallel. In another implementation, the primary component of an image is the chromen component, and at least one secondary component of the image is the lumen component. This results in a high degree of flexibility in the actual conditioning of one component by another component.
[0054] The overall coding may, in detail, include processing in latent space, which may enable the processing of downsampled input data, and therefore fixed processing with less processing load. Note that in this specification, the terms “downsampling” and “upsampling” are used to mean reducing and expanding the size of the tensor representation of the data, respectively.
[0055] At least one of the height or width dimensions of the first latent tensor may be smaller than the corresponding height or width dimension of the first tensor, and / or the height or width dimension of the second latent tensor may be smaller than the corresponding height or width dimension of the connected tensor. For example, a scaling factor of 16 or 32 in the height and / or width dimensions may be used.
[0056] It may be a fact that the sample size or subpixel offset of the second tensor in at least one of the height and width dimensions of the tensor is different from the sample size or subpixel offset of the first tensor in at least one of the height and width dimensions.
[0057] In this case as well, at least one of the height or width dimensions of the first latent tensor may be smaller than the corresponding height or width dimension of the first tensor, and / or the height or width dimension of the second latent tensor may be smaller than the corresponding height or width dimension of the connected tensor. For example, a reduction factor of 16 or 32 in the height and / or width dimensions may be used. Adjusting the sample location of the first tensor to match the sample location of the second tensor may involve, for example, downsampling the width and height of the first tensor by a factor of 2.
[0058] According to one implementation of the second embodiment, the first latent tensor has a channel dimension, and the second latent tensor has a channel dimension, where the size of the first latent tensor in the channel dimension is one of being greater than, less than, or equal to the size of the second latent tensor in the channel dimension. If the first-order component is considered to be of greater importance than the second-order component, which is usually the case, the channel length of the first-order component may be longer than the channel length of the second-order component. If the signal of the first-order component is relatively clear and the signal of the non-first-order component is relatively noisy, the channel length of the first-order component may be shorter than the channel length of the second-order component. Numerical experiments have shown that shorter channel lengths compared to those in the art can be used without significant degradation of the quality of the reconstructed image, and therefore memory requirements can be reduced.
[0059] Generally, a first tensor may be transformed into a first latent tensor using a first neural network, and a connected tensor may be transformed into a second latent tensor using a second neural network different from the first neural network. In this case, the first and second neural networks may be trained collaboratively to determine the size of the first latent tensor in the channel dimension and the size of the second latent tensor in the channel dimension. The determination of the channel length may be performed by exhaustive search or in a content-adaptive manner. A set of models may be trained, each model based on a different number of channels for coding the first and second-order components. Thereafter, the neural network may be able to optimize the channel lengths involved.
[0060] The determined channel length must also be used by the decoder used to reconstruct the encoded components. Thus, according to one implementation, the size of a first latent tensor in the channel dimension may be signaled in the first bitstream, and the size of a second latent tensor in the channel dimension may be signaled in the second bitstream. The signaling may be performed explicitly or implicitly, allowing the decoder to be directly informed of the channel length in a bit-saving manner.
[0061] According to one implementation of the second aspect, a first bitstream is parsed based on a first entropy model, and a second bitstream is parsed based on a second entropy model different from the first entropy model. Such entropy models allow for the reliable estimation of the statistical properties used in the process of converting the tensor representation of the data into a bitstream.
[0062] The method of disclosure may be advantageously implemented in the context that a hyperplier architecture that sequentially provides secondary information useful for coding (parts of) an image improves the accuracy of (parts of) the reconstructed image.
[0063] The resampling process can be performed without interpolation (which is a computationally expensive operation). This can lead to improved coding efficiency.
[0064] The integer factors may be calculated based on a first scaling factor for the linear component and a second scaling factor for at least one quadratic component. For example, the integer factors are s UV / s Y It is calculated as, where, s Y represents the first scaling factor, s UV This represents the second scaling factor.
[0065] s Y When s exists in the bitstream, YThis is obtained by parsing the bitstream. Y When s does not exist in the bitstream, Y This is set to the default value. Y The default value for this can be 1.
[0066] s UV When s exists in the bitstream, UV This is obtained by parsing the bitstream. UV When s does not exist in the bitstream, UV This is set to the default value. UV The default value can be 2.
[0067] As another example, an integer factor r may exist in the bitstream, or it may be set to a default value.
[0068] If r exists in the bitstream, it is obtained by parsing the bitstream. If r does not exist in the bitstream, it is set to its default value. The default value of r may be 2.
[0069] s Y When s exists in the bitstream, Y This is obtained by parsing the bitstream. Y When s does not exist in the bitstream, Y This is set to the default value. Y The default value for this can be 1.
[0070] Next, s UV ga s UV =r·s Y It is derived as, however, s UV This represents the scaling factor for at least one quadratic component.
[0071] According to one implementation of the second embodiment, the first scaling factor may include a first horizontal scaling factor and a first vertical scaling factor. Similarly, the second scaling factor may include a second horizontal scaling factor and a second vertical scaling factor.
[0072] Resampling is performed by using nearest neighbor upsampling or nearest neighbor downsampling.
[0073] For nearest-near-side upsampling, This is denoted as s↑. This layer has a size [C, h in , w in It receives a tensor input of size [C, s·h]. in , s·w in Outputs the tensor output of ].
[0074] Interpolation is not required for this downsampling; it simply involves copying. Output[c, i, j]=Input[c, s·i, s·j], i=0,...,h in -1, j=0,...,w in -1, c=0,...,C-1 This is the case where s represents an integer factor.
[0075] For nearest-neighbor downsampling, This is represented as s↓. This layer has a size [C, s·h out , s·w out It receives a tensor input of size [C, h out , w out This outputs a tensor output of ], and this procedure reduces the spatial resolution of each tensor channel.
[0076] Interpolation is not required for this downsampling; it simply involves copying. Output[c, s·i, s·j]=Input[c,i,j], i=0,...,h in -1, j=0,...,w in -1, c=0,...,C-1 This is the case where s represents an integer factor.
[0077] According to one implementation of the method according to the second embodiment, the processing of the first latent tensor comprises converting the first latent tensor into a first tensor.
[0078] According to one implementation of the method according to the second aspect, obtaining a second tensor representing at least one quadratic residual component of the residual for at least one quadratic component of the image is: The process comprises concatenating a second latent tensor with a resampled first latent tensor to obtain a connected tensor, and transforming the connected tensor into a second tensor. At least one of these transformations may involve upsampling. Thus, the processing in the latent space may be performed at a lower resolution, such as is necessary for the accurate reconstruction of the components in YUV space or any other space preferably used for image representation.
[0079] According to another implementation of the method according to the second embodiment, each of the first and second latent tensors has height and width dimensions, processing the first latent tensor comprises transforming the first latent tensor into a first tensor, and processing the second latent tensor comprises determining whether the sample size or subpixel offset of the second latent tensor in at least one of the height and width dimensions is different from the sample size or subpixel offset of the first latent tensor in at least one of the height and width dimensions. When it is determined that the sample size or subpixel offset of the second latent tensor is different from the sample size or subpixel offset of the first latent tensor, the sample location of the first latent tensor is adjusted to match the sample location of the second latent tensor. Thereafter, the adjusted first latent tensor is obtained. Furthermore, when it is determined that the sample size or subpixel offset of the second latent tensor is different from the sample size or subpixel offset of the first latent tensor, the second latent tensor and the adjusted first latent tensor are concatenated to obtain a connected latent tensor; otherwise, the second latent tensor and the first latent tensor are concatenated to obtain a connected latent tensor, and the connected latent tensor is transformed into the second tensor.
[0080] A first bitstream may be processed by a first neural network, and a second bitstream may be processed by a second neural network different from the first neural network. A first latent tensor may be transformed by a third neural network different from the first and second networks, and a connected latent tensor may be transformed by a fourth neural network different from the first, second, and third networks.
[0081] According to another implementation of the method in the second aspect, the first latent tensor has a channel dimension, and the second latent tensor has a channel dimension, where the size of the first latent tensor in the channel dimension is greater than, less than, or equal to the size of the second latent tensor in the channel dimension. Information about the sizes of the first and second latent tensors in the channel dimension may be obtained from information signaled in the first and second bitstreams, respectively.
[0082] Any of the exemplary implementation configurations described above may be combined in any way deemed appropriate. Methods according to any of the above-described embodiments and implementation configurations may be implemented within the apparatus.
[0083] According to a third aspect, an apparatus is provided for reconstructing at least a portion of an image, the apparatus comprising one or more processors and a non-temporary computer-readable storage medium coupled to one or more processors and storing a program for execution by one or more processors, wherein the apparatus is configured such that, when the program is executed by one or more processors, it performs a method according to any one of the first and second aspects and the corresponding implementations described above.
[0084] According to a fourth aspect, a processing apparatus is provided for reconstructing at least a portion of an image, the processing apparatus comprising a processing circuit configuration configured to perform the method according to any one of the first and second aspects and the corresponding implementations described above.
[0085] Furthermore, according to a fifth aspect, a computer program stored on a non-temporary medium is provided, which, when executed on one or more processors, performs the steps of the method according to any of the above-described aspects and implementations.
[0086] The technical background and embodiments of the present invention will be described in more detail below with reference to the accompanying drawings and illustrations. [Brief explanation of the drawing]
[0087] [Figure 1] This is a schematic diagram showing the channels processed by the layers of a neural network. [Figure 2] This is a schematic diagram showing the encoder types of neural networks. [Figure 3] This is a schematic diagram showing a network architecture that includes a hyperplier model. [Figure 4] This is a block diagram showing the structure of a cloud-based solution for machine-based tasks such as machine vision tasks. [Figure 5] A cylinder showing the structure of an end-to-end trainable video compression framework. [Figure 6] This is a block diagram showing a network for motion vector (MV) compression. [Figure 7] This is a block diagram showing learned image compression configurations in the field of technology. [Figure 8] This is a block diagram showing another learned image compression configuration in the field of technology. [Figure 9] This is a diagram illustrating the concept of conditional coding. [Figure 10] This diagram illustrates the concept of residual coding. [Figure 11] This is a diagram illustrating the concept of residual conditional coding. [Figure 12] This figure shows conditional intracoding using one embodiment of the present invention. [Figure 13] This figure shows conditional residual coding according to one embodiment of the present invention. [Figure 14] This figure shows conditional coding according to one embodiment of the present invention, which means that the sampling coefficients in latent space sample location adjustment are constrained to being integers. [Figure 15]This figure shows conditional intracoding for input data in YUV420 format according to one embodiment of the present invention. [Figure 16] This figure shows conditional intracoding for input data in YUV444 format according to one embodiment of the present invention. [Figure 17] This figure shows conditional residual coding for input data in YUV420 format according to one embodiment of the present invention. [Figure 18] This figure shows conditional residual coding for input data in YUV444 format according to one embodiment of the present invention. [Figure 19] This figure shows conditional residual coding for input data in YUV420 format according to another embodiment of the present invention. [Figure 20] This figure shows conditional residual coding for input data in YUV444 format according to another embodiment of the present invention. [Figure 21] This is a flowchart illustrating an exemplary method for reconstructing at least a portion of an image according to one embodiment of the present invention. [Figure 22] This figure shows a decoder architecture according to one embodiment of the present invention. [Figure 23] This figure shows a decoder architecture according to another embodiment of the present invention. [Figure 24] This figure shows a signal decoder architecture according to one embodiment of the present invention. [Figure 25] This figure shows a processing apparatus configured to perform a method for encoding or reconstructing at least a portion of an image, according to one embodiment of the present invention. [Figure 26] This is a block diagram showing an example of a video coding system configured to implement embodiments of the present invention. [Figure 27] This block diagram shows another example of a video coding system configured to carry out embodiments of the present invention. [Figure 28]This is a block diagram showing an example of an encoding or decoding device. [Figure 29] This is a block diagram showing another example of an encoding or decoding device. [Modes for carrying out the invention]
[0088] In the following description, accompanying drawings forming part of this disclosure are referenced, which, for example, illustrate specific aspects of embodiments of the present invention, or specific aspects in which embodiments of the present invention may be used. It is understood that embodiments of the present invention may be used in other aspects, and which may include structural or logical modifications not shown in the drawings. Accordingly, the embodiments for carrying out the invention described below should not be understood in an restrictive sense, and the scope of the invention is defined by the appended claims.
[0089] For example, disclosures relating to a method described may also apply to a corresponding device or system configured to perform the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units, e.g., functional units (e.g., one unit performing one or more steps, or multiple units each performing one or more of the steps), even if such one or more units are not explicitly described or illustrated in the figures, for performing the one or more method steps described. On the other hand, if a particular device is described based on one or more units, e.g., functional units, the corresponding method may include one step for performing the functionality of one or more units (e.g., one step performing the functionality of one or more units, or multiple steps each performing one or more of the functionality of the units), even if such one or more steps are not explicitly described or illustrated in the figures. Furthermore, unless otherwise specifically noted, features of the various exemplary embodiments and / or aspects described herein may be combined with each other.
[0090] The following provides an overview of some of the technical terms used.
[0091] Artificial neural networks An artificial neural network (ANN), or connectionist system, is a computing system vaguely conceived from the biological neural networks that make up the brains of animals. Such systems generally "learn" to perform tasks by examining examples, without being programmed with task-specific rules. For example, in image recognition, such a system might learn to identify images containing cats by analyzing exemplary images that have been manually labeled as "cat" or "not a cat," and then using the results to identify cats in other images. Such a system does this without any prior knowledge of cats, such as that cats have fur, tails, whiskers, and cat-like faces. Instead, such a system automatically generates the characteristics to identify from the examples they process.
[0092] ANNs are based on a collection of connected units or nodes called artificial neurons, which roughly model neurons in the biological brain. Each connection, like a synapse in the biological brain, can transmit signals to other neurons. The artificial neurons that receive the signals can then process them and signal to the neurons connected to them.
[0093] In ANN implementations, the "signals" at connections are real numbers, and the output of each neuron is calculated by some nonlinear function of the sum of its inputs. Connections are called edges. Neurons and edges typically have weights that are adjusted as learning progresses. These weights increase or decrease the intensity of the signal at the connection. Neurons may have thresholds such that they only send a signal if the integrated signal exceeds that threshold. Neurons are typically integrated within layers. Different layers may perform different transformations on their inputs. Signals travel from the first layer (input layer) to the last layer (output layer), sometimes traversing multiple layers.
[0094] The initial goal of ANN methods was to solve problems in a similar way to how the human brain would solve them. Over time, the focus shifted to performing specific tasks, leading to a deviation from biology. ANNs are used in a wide range of tasks, including computer vision, speech recognition, machine translation, social network filtering, playing board and video games, medical diagnosis, and even in activities traditionally considered to be reserved for humans, such as painting.
[0095] Convolutional Neural Network The name "Convolutional Neural Network" (CNN) indicates that the network employs a mathematical operation called convolution. Convolution is a specialized type of linear operation. A convolutional network is simply a neural network that uses convolution in at least one of its layers instead of general matrix multiplication.
[0096] Figure 1 schematically illustrates the general concept of processing by neural networks such as CNNs. A convolutional neural network consists of input and output layers, as well as several hidden layers. The input layer is the layer to which the input (such as a portion of an image, as shown in Figure 1) is provided for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that convolve using multiplication or other dot products. The result of the layers is one or more feature maps (f.maps in Figure 1), sometimes called channels. There may be subsampling that involves some or all of the layers. As a result, the feature maps may be smaller, as shown in Figure 1. The activation function in a CNN is usually a RELU (Normalized Linear Unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, and normalization layers, which are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. These layers are colloquially called convolutions, but this is merely a convention. Mathematically, it is technically a sliding dot product or cross-correlation. This has importance to indices in a matrix in that it affects how weights are determined at specific index points.
[0097] As shown in Figure 1, when programming a CNN to process images, the input is a tensor with shape (number of images) × (image width) × (image height) × (image depth). Then, after passing through the convolutional layer, the image is abstracted into a feature map with shape (number of images) × (feature map width) × (feature map height) × (feature map channels). The convolutional layer in the neural network should have the following attributes: a convolutional kernel defined by width and height (hyperparameters); the number of input and output channels (hyperparameters); and the depth of the convolutional filter (input channels) should be equal to the number of channels (depth) of the input feature map.
[0098] In the past, conventional multilayer perceptron (MLP) models have been used for image recognition. However, due to the fully connected nature of the nodes, they suffer from high dimensionality and do not scale well with higher resolution images. A 1000x1000 pixel image with RGB color channels has 3 million weights, which is too many to efficiently handle at scales involving fully connected nodes. Furthermore, such network architectures do not take into account the spatial structure of the data, treating distant input pixels as if they were close to each other. This ignores locality of reference in image data, both computationally and semantically. Therefore, for purposes such as image recognition, which are governed by spatially local input patterns, the fully connected nature of neurons is wasteful.
[0099] Convolutional neural networks (CNNs) are a biologically inspired variation of multilayer perceptrons, specifically designed to mimic the behavior of the visual cortex. These models mitigate the challenges posed by MLP architectures by leveraging the strong, spatially localized correlations present in natural images. Convolutional layers are the core building blocks of a CNN. The layer parameters consist of a set of learnable filters (the kernels described above) that have small receptive fields but extend throughout the full depth of the input volume. During the forward pass, each filter is convolved across the width and height of the input volume to compute the dot product between the filter's entry and the input, generating a two-dimensional activation map of that filter. As a result, the network learns filters that activate when the network detects certain types of features at certain spatial locations in the input.
[0100] Stacking activation maps for all filters along the depth dimension forms the total output volume of the convolutional layer. Every entry in the output volume can be interpreted as the output of a neuron that shares parameters with neurons in the same activation map, as well as looking at a small region in the input. A feature map or activation map is the output activation for a given filter. Feature map and activation have the same meaning. In some papers, it is called an activation map because it is a mapping corresponding to the activations of different parts of an image, and also called a feature map because it is a mapping of where certain types of features can be found in the image. High activation means that several features have been found.
[0101] Another important concept in CNNs is pooling, which is a form of nonlinear downsampling. There are several nonlinear functions for performing pooling, the most common of which is max pooling. It divides the input image into sets of non-overlapping rectangles and outputs the maximum value for each such sub-region.
[0102] Intuitively, the precise location of a feature is less important than its rough location relative to other features. This is the idea behind the use of pooling in convolutional neural networks. Pooling layers work to gradually reduce the spatial size of representations in order to reduce the number of parameters, the memory footprint, and the amount of computation in the network, and therefore also to control overfitting. It is common to periodically insert pooling layers between consecutive convolutional layers in a CNN architecture. Pooling operations result in another form of translational invariance.
[0103] A pooling layer operates independently across all depth slices of the input, spatially resizing them. The most common form is a pooling layer with a 2x2 filter applied with a stride of 2, downsampling by 2 along both width and height at every depth slice in the input, discarding 75% of the activations. In this case, every max operation spans four numbers. The depth dimension remains unchanged.
[0104] In addition to max pooling, the pooling unit can use other functions such as mean pooling or L2 norm pooling. Mean pooling has historically been frequently used, but recently it has become less preferred compared to max pooling, which actually performs better. There is a recent trend toward using smaller filters or even discarding the pooling layer entirely, due to the aggressive reduction of representation size. Region of interest pooling (also called ROI pooling) is a variation of max pooling where the output size is fixed and the input rectangle is parameterized. Pooling is a key component of convolutional neural networks for object detection based on fast R-CNN architectures.
[0105] The aforementioned ReLU is an abbreviation for Normalized Linear Unit, which applies a non-saturated activation function. It effectively removes negative values from the activation map by setting them to 0. It enhances the nonlinear properties of the decision function and the network as a whole without affecting the receptive fields of the convolutional layers. Other functions, such as saturated hyperbolic tangent and sigmoid functions, are also used to enhance nonlinearity. ReLU is often preferred over other functions because it trains the neural network several times faster without a significant penalty to generalization accuracy.
[0106] After several convolutional and max-pooling layers, high-level inference in the neural network is performed via fully connected layers. Neurons in the fully connected layers have connections to all activations in the previous layer, as seen in typical (non-convolutional) artificial neural networks. Their activations can thus be computed as affine transformations, which involve matrix multiplication followed by bias offsets (vector addition of learned or fixed bias terms).
[0107] The "loss layer" specifies how training penalizes the deviation between the predicted (output) label and the true label, and is usually the final layer of a neural network. Different loss functions may be used, depending on the task. The softmax loss is used to predict a single class out of K mutually exclusive classes. The sigmoid cross-entropy loss is used to predict K independent probability values in [0, 1]. The Euclidean loss is used to regress to real-valued labels.
[0108] In summary, Figure 1 shows the data flow in a typical convolutional neural network. First, the input image is passed through a convolutional layer and abstracted into a feature map having several channels corresponding to the number of filters in the set of learnable filters in that layer. The feature map is then subsampled, for example, using a pooling layer, which reduces the dimensionality of each channel in the feature map. The next data comes to another convolutional layer, which may have a different number of output channels connected to a different number of channels in the feature map. As mentioned above, the number of input and output channels are hyperparameters of the layer. To establish network connectivity, these parameters need to be synchronized between two connected layers such that the number of input channels for the current layer should be equal to the number of output channels for the previous layer. For the input data, for example, in the first layer that processes the input data, the number of input channels is usually equal to the number of channels in the data representation, for example, 3 channels for an RGB or YUV representation of an image or video, or 1 channel for a grayscale image or video representation.
[0109] Autoencoders and unsupervised learning An autoencoder is a type of artificial neural network used to learn efficient data coding in an unsupervised manner. A schematic diagram of it is shown in Figure 2. The goal of an autoencoder is to learn a representation (encode) for a set of data, usually for dimensionality reduction, by training the network to ignore signal "noise". Along with the reduction side, the reconstruction side is learned, where the autoencoder attempts to generate a representation as close as possible to its original input from the reduced encoding, hence the name. In its simplest form, given one hidden layer, the encoder stage of the autoencoder takes input x and maps it to h. h = σ(Wx + b)
[0110] This image h is usually called the code, latent variable, or latent representation. Here, σ is an element-wise activation function, such as a sigmoid function or normalized linear unit. W is the weight matrix, and b is the bias vector. The weights and biases are typically initialized randomly and then iteratively updated during training through backpropagation. The decoder stage of the autoencoder then maps h to a reconstructed x' of the same shape as x, i.e., x'=σ'(W' h'+b') Here, σ', W', and b' for the decoder may be independent of the corresponding σ, W, and b for the encoder.
[0111] Variational autoencoder models make strong assumptions about the distribution of latent variables. They use variational methods for latent representation learning, which results in an additional loss component and a specific estimator for the training algorithm called the Stochastic Gradient Variational Bayes (SGVB) estimator. It is a directed graphical model p θ The data is generated by (x|h), and the encoder is the posterior distribution p θ Approximation q to (h|x) φ It is assumed that (h|x) is being trained, where φ and θ represent the parameters of the encoder (recognition model) and decoder (generative model), respectively. The probability distribution of the latent vector of a VAE usually fits the probability distribution of the training data much better than a standard autoencoder. The purpose of a VAE has the following form:
[0112]
number
[0113] Here, D KL This represents the Kullback-Leibler divergence. The pryor across latent variables is typically a central isotropic multivariate Gaussian p. θ(h) is set to N(0,I). Typically, the shapes of the variational and likelihood distributions are chosen such that they are factorized Gaussians, i.e., q φ (h|x)=N(ρ(x),ω 2 (x)I) p φ (x|h)=N(μ(h),σ 2 (h)I) And, where ρ(x) and ω 2 (x) is the encoder output, and μ(h) and σ 2 (h) is the decoder output.
[0114] Recent advances in the field of artificial neural networks, particularly in convolutional neural networks, have enabled researchers to apply neural network-based techniques to the tasks of image and video compression. For example, end-to-end optimized image compression using a network based on variational autoencoders has been proposed. Thus, data compression is considered a fundamental and well-considered problem in engineering and is usually formulated with the goal of designing a code for a given set of discrete data with minimal entropy. Its solution relies heavily on knowledge of the probabilistic structure of the data, and therefore the problem is closely related to probabilistic source modeling. However, since all practical codes must have a finite entropy, continuous data (such as a vector of image pixel intensity) must be quantized into a finite set of discrete values, which introduces errors. In this context, known as the irreversible compression problem, one must balance two competing costs: the entropy (rate) of the discretized representation and the errors (distortion) resulting from quantization. Different compression applications, such as data storage or transmission over channels with limited capacity, require different rate-distortion tradeoffs. Simultaneous optimization of rate and distortion is difficult. Without further constraints, the general problem of optimal quantization in high-dimensional spaces is cumbersome. For this reason, most existing image compression methods operate by linearly transforming the data vector into a suitable continuous-value representation, independently quantizing its elements, and then encoding the resulting discrete representation using a reversible entropy code. This method is called transform coding due to the central role of the transformation. For example, JPEG uses a discrete cosine transform for blocks of pixels, and JPEG2000 uses multiscale orthogonal wavelet decomposition. Typically, the three components of a transform coding method—the transform, the quantizer, and the entropy code—are optimized separately (often through manual parameter tuning).Modern video compression standards such as HEVC, VVC, and EVC also use transformed representations to encode residual signals after prediction. For this purpose, several transformations are used, including discrete cosine and sine transforms (DCT, DST), as well as the low-frequency non-separable manually optimized transform (LFNST).
[0115] Variational image compression In "Density Modeling of Images Using a Generalized Normalization Transformation" (hereinafter referred to as "Balle"), published in arXiv e-prints in 2015 by J. Balle, L. Valero Laparra, and EP Simoncelli at the 4th International Conference on Representation Learning, 2016, the authors proposed a framework for end-to-end optimization of image compression models based on nonlinear transformations. Previously, the authors demonstrated that a model consisting of linear nonlinear block transformations optimized for a measure of perceptual distortion exhibits visually superior performance compared to a model optimized for mean squared error (MSE). Here, the authors optimize for MSE but use a more flexible transformation constructed from a cascade of linear convolution and nonlinearity. In detail, the authors demonstrate the effectiveness of Gaussianizing image density using generalized divisive normalization (GDN) coupled nonlinearity, inspired by models of neurons in the biological visual system. This cascaded transformation is followed by uniform scalar quantization (i.e., each element is rounded to the nearest integer), which effectively performs a parametric form of vector quantization on the original image space. Using an approximate parametric nonlinear inverse transform, the compressed image is reconstructed from these quantized values.
[0116] For any desired point along the rate-distortion curve, the parameters for both the analysis and composite transformation are jointly optimized using stochastic gradient descent. To achieve this in the presence of quantization (which generates a zero gradient almost everywhere), the authors substitute the quantization step with additive uniform noise using a proxy loss function based on the continuous relaxation of a stochastic model. The relaxed rate-distortion optimization problem has some similarities to those used to fit generating image models, particularly variational autoencoders, but the constraints the authors impose to ensure it approximates all discrete problems along the rate-distortion curve are different. Finally, rather than reporting difference or discrete entropy estimates, the authors implement the entropy code and report the performance using actual bitrates, thus demonstrating the feasibility of the solution as a complete lossy compression method.
[0117] In J. Balle, an end-to-end trainable model for image compression based on variational autoencoders is described. The model incorporates a hyperplier to effectively capture spatial dependency in the latent representation. This hyperplier relates to secondary information also sent to the decoder, a concept virtually universal to all modern image codecs but largely unexplored in image compression using ANNs. Unlike existing autoencoder compression methods, this model trains a complex plier in cooperation with the underlying autoencoder. The authors demonstrate that this model delivers state-of-the-art image compression when measuring visual quality using a common MS-SSIM index, and when evaluated using a more archaic metric based on squared error (PSNR), it provides rate-distortion performance that surpasses published ANN-based methods.
[0118] Figure 3 shows a network architecture including the hyperplier model. The left side (ga, gs) shows the image autoencoder architecture, and the right side (ha, hs) corresponds to the autoencoders that perform the hyperplier. The factorized plier model uses the same architecture for analysis ga and synthetic transformation gs. Q represents quantization, and AE and AD represent the arithmetic encoder and arithmetic decoder, respectively. The encoder applies ga to the input image x to produce a response y (latent representation) with a spatially varying standard deviation. The encoded ga includes multiple convolutional layers with subsampling, and Generalized Piecewise Normalization (GDN) as the activation function.
[0119] The response is fed into ha, and z outlines the distribution of the standard deviation. z is then quantized, compressed, and transmitted as secondary information. The encoder then processes the quantized vector
[0120]
number
[0121] Using
[0122]
number
[0123] That is, we estimate the spatial distribution of the standard deviation used to obtain probability values (or frequency values) for arithmetic coding (AE), and use it to represent quantized images.
[0124]
number
[0125] The (or latent) representation is compressed and transmitted. The decoder then processes the compressed signal.
[0126]
number
[0127] First, restore the data. The decoder then uses HS.
[0128]
number
[0129] Obtain,
[0130]
number
[0131] It also provides it with an accurate probability estimate for successfully reconstructing it. The decoder then obtains the reconstructed image.
[0132]
number
[0133] It will be supplied to GS.
[0134] Further work has improved the hyperplier-based probabilistic modeling by introducing an autoregressive model based on the PixelCNN++ architecture, which allows us to leverage the context of already decoded symbols in the latent space for better probability estimation of further symbols to be decoded, for example, as shown in Figure 2 of L. Zhou, Zh. Sun, X. Wu, J. Wu, End-to-end Optimized Image Compression with Attention Mechanism, CVPR, 2019 (hereinafter referred to as "Zhou").
[0135] Cloud solutions for machine tasks Machine video coding (VCM) is another computer science trend that is common today. The main idea behind this technique is to transmit coded representations of image or video information that are targeted for further processing by computer vision (CV) algorithms, such as object segmentation, detection, and recognition. In contrast to traditional image and video coding that targets human perception, its quality characteristics are not reconstructed quality but performance on computer vision tasks, such as object detection accuracy. This is illustrated in Figure 4.
[0136] Video coding for machines, also known as collaborative intelligence, is a relatively new paradigm for the efficient deployment of deep neural networks across mobile and cloud infrastructures. By splitting the network between mobile and the cloud, it is possible to distribute the computational workload in such a way that the overall energy and / or latency of the system is minimized. In general, collaborative intelligence is a paradigm in which the processing of a neural network is distributed among two or more different computing nodes, e.g., devices, but generally any functionally defined node. Here, the term “node” does not refer to the neural network nodes described above. Rather, a (computational) node here refers to a separate device / module (physically or at least logically) that implements a part of the neural network. Such devices may be different servers, different end-user devices, a mixture of servers and / or user devices and / or the cloud and / or processors, etc. In other words, computing nodes may be considered nodes that belong to the same neural network and communicate with each other to transmit coded data within / for the neural network. For example, to enable the execution of complex calculations, one or more layers may run on a first device and one or more layers may run on another device. However, the distribution may also be finer-grained, with a single layer running on multiple devices. In this disclosure, the term “multiple” refers to two or more. In some existing solutions, part of the neural network functionality runs on a device (such as a user device or edge device) or multiple such devices, and the output (feature map) is then passed to the cloud. The cloud is a collection of processing or computing systems located outside the device that powers part of the neural network. The concept of collaborative intelligence has also been extended to model training.In this case, data flows in both directions: from the cloud to mobile during backpropagation in training, and from mobile to the cloud during the forward path in training and during inference.
[0137] Several studies have proposed semantic image compression by encoding deep features and then reconstructing the input image from them. Compression based on uniform quantization has been demonstrated, followed by context-based adaptive arithmetic coding (CABAC) from H.264. In some scenarios, it may be more efficient to send the output of the hidden layer (deep feature map) from the mobile part to the cloud rather than sending the compressed natural image data to the cloud and performing object detection using the reconstructed image. Efficient compression of feature maps benefits image and video compression and reconstruction for both human perception and machine vision. Entropy coding methods, such as arithmetic coding, are common techniques for compressing deep features (i.e., feature maps).
[0138] Today, video content contributes more than 80% of internet traffic, and that percentage is expected to increase further. Therefore, it is crucial to create efficient video compression systems that can produce higher-quality frames within a given bandwidth budget. In addition, most video-related computer vision tasks, such as video object detection or video object tracking, are sensitive to the quality of compressed video, and efficient video compression can provide advantages over other computer vision tasks. Meanwhile, techniques in video compression are also useful for action recognition and model compression. However, over the past few decades, video compression algorithms have relied on handcrafted modules, such as block-based motion estimation and discrete cosine transform (DCT), to reduce redundancy in video sequences, as mentioned above. While each module may be well-designed, the overall compression system is not end-to-end optimized. It is desirable to further improve video compression performance by collectively optimizing the entire compression system.
[0139] End-to-end image or video compression Recently, deep neural network (DNN)-based autoencoders for image compression have achieved performance equal to or even better than traditional image codecs such as JPEG, JPEG2000, or BPG. One possible interpretation is that DNN-based image compression methods can leverage large-scale end-to-end training and highly nonlinear transformations not used in conventional methods. However, it is not trivial to directly apply these techniques to create an end-to-end learning system for video compression. Firstly, learning how to generate and compress the edited motion information for video compression remains an unresolved problem. Video compression methods rely heavily on motion information to reduce temporal redundancy in video sequences. A simple solution would be to use learning-based optical flow to represent the motion information. However, current learning-based optical flow methods aim to generate the most accurate flow field possible. A precise optical flow is often not optimal for a particular video task. In addition, the data volume of optical flow is significantly larger compared to motion information in conventional compression systems, and directly applying existing compression methods to compress optical flow values significantly increases the number of bits required to store motion information. Secondly, it is unclear how to construct a DNN-based video compression system while minimizing the rate-distortion-based objectives for both residuals and motion information. Rate-distortion optimization (RDO) aims to achieve higher quality (i.e., less distortion) of the reconstructed frame given a number of bits (or bitrate) for compression. RDO is crucial for video compression performance. To leverage the end-to-end training capability for learning-based compression systems, an RDO policy is needed to optimize the entire system.
[0140] In "DVC: An End-to-end Deep Video Compression Framework," Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11006-11015, the authors propose an end-to-end deep video compression (DVC) model that collaboratively learns motion estimation, motion compression, and residual coding.
[0141] Such an encoder is shown in Figure 5. In detail, Figure 5 shows the overall structure of an end-to-end trainable video compression framework. A CNN is specified to transform the optical flow into a corresponding representation suitable for better compression in order to compress the motion information. In detail, an autoencoder-style network is used to compress the optical flow. The motion vector (MV) compression network is shown in Figure 6. The network architecture is somewhat similar to ga / gs in Figure 3. In detail, the optical flow is fed into a series of convolutional operations and nonlinear transformations, including GDN and IGDN. The number of output channels for convolution (deconvolution) is 128, except for the last deconvolution layer which is equal to 2. Given an optical flow of size M × N × 2, the MV encoder generates a motion representation of size M / 16 × N / 16 × 128. The motion representation is then quantized, entropy coded, and sent to a bitstream. The MV decoder receives the quantized representation and uses the MV encoder to reconstruct the motion information.
[0142] In detail, the following definition holds true.
[0143] Picture size (the terms "image" and "picture" are used interchangeably herein): Refers to the width or height of a picture, or a width-height pair. The width and height of an image are typically measured in terms of the number of lumens.
[0144] Downsampling: Downsampling is the process of reducing the sampling rate (sampling interval) of a discrete input signal.
[0145] Upsampling: Upsampling is the process of increasing the sampling rate (sampling interval) of a discrete input signal.
[0146] Cropping: The process of cutting off the outer edges of a digital image. Cropping can be used to make an image smaller (in terms of the number of samples) and / or to change the aspect ratio (length to width) of the image.
[0147] Padding: Padding refers to increasing the size of an input image (or image) by generating new samples at the boundaries of the image, either by using predefined sample values or by using sample values at that location within the input image.
[0148] Convolution: Convolution is given by the following general formula. Below, f() may be defined as an input signal, and g() may be defined as a filter.
[0149]
number
[0150] NN module: A neural network module, a component of a neural network. It may be a layer or subnetwork within a neural network. A neural network is a sequence of NN modules. In the context of this specification, it is assumed that a neural network is a sequence of K NN modules.
[0151] Latent space: An intermediate step in neural network processing. The latent space representation includes the outputs of the input or hidden layers, which are not intended to be seen.
[0152] Irreversible NN Module: Information processed by an irreversible NN module results in information loss, and the irreversible module makes the processed information unrecoverable.
[0153] Reversible NN module: Information processed by a reversible NN module does not result in information loss, and reversibility makes the processed information recoverable.
[0154] Bottleneck: Latent space tensor proceeding to the reversible coding module.
[0155] Autoencoder: A model that converts a signal into a (compressed) latent space and then converts it back into the original signal space.
[0156] Encoder: Downsamples an image using a convolutional layer with nonlinearity and / or residuals to obtain a latent tensor (y).
[0157] Decoder: The latent tensor (y) is upsampled to the original image size using a convolutional layer with nonlinearity and / or residuals.
[0158] Hyperencoder: Further convolutional layers with nonlinearity and / or residuals are used to downsample the latent tensor to a smaller latent tensor (z).
[0159] Hyperdecoder: For entropy estimation, a smaller latent tensor (z) is upsampled using a convolutional layer with nonlinearity and / or residuals.
[0160] AE / AD (Arithmetic Encoder / Decoder): Encodes a latent tensor into a bitstream, or decodes a latent tensor from a bitstream with a given statistical plier.
[0161] Autoregressive entropy estimation: A process of estimating the statistical pryor of a latent tensor in a continuous manner. Q: Quantization block.
[0162]
number
[0163] ,
[0164]
number
[0165] : The quantized version of the corresponding latent tensor.
[0166] Masked Convolution (MaskedConv): A type of convolution that masks some latent tensor elements so that the model can only make predictions based on latent tensor elements that have already been observed.
[0167] H, W: Height and width of the input image.
[0168] Block / Patch: A subset of latent tensors on a rectangular grid.
[0169] Information sharing: A collaborative process of sharing information from different patches. P: Size of the rectangular patch. K: The kernel size that defines the numbered adjacent patches included in the information sharing. L: The kernel size that determines how many of the previously coded latent tensor elements are included in the information sharing.
[0170] Masked Convolution (MaskedConv): A type of convolution that masks some latent tensor elements so that the model can only make predictions based on latent tensor elements that have already been observed.
[0171] PixelCNN: A convolutional neural network that includes one or more layers of masked convolution.
[0172] Component: One dimension of the orthogonal basis representing the full-color image.
[0173] Channel: A layer within a neural network.
[0174] Intracodec: The first or main frame of a video is processed as an intraframe, and it is usually treated as an image.
[0175] Intercodec: After the intracodec, the video compression system performs inter-prediction. The first motion estimation tool calculates the motion vector of the object, and then the motion compensation tool uses the motion vector to predict the next frame.
[0176] Residual codec: The predicted frame is not always identical to the current frame; the difference between the current frame and the predicted frame is the residual. A residual codec compresses the residual, similar to how an image is compressed.
[0177] Signal conditioning: An additional signal is used to aid in NN inference, however, the additional signal is not present in the output and is significantly different from the output during the training procedure.
[0178] Conditional codecs: Codecs that use signal conditioning to assist (guide) compression and reconstruction. Since the auxiliary information required for conditioning is not part of the input signal, state-of-the-art conditional codecs are used for compressing video streams rather than images.
[0179] FIG. 7 is a block diagram showing a particular learned image compression configuration comprising an autoencoder and hyperprior components in the art that can be improved in accordance with the present disclosure. The input image to be compressed is represented as a 3D tensor having a size of H×W×C or C×H×W, where H and W are the height and width (dimensions) of the image, respectively, and C is the number of components (e.g., a luminance component and two chroma components). The input image is passed through an encoder 71. The encoder downsamples the input image by applying a plurality of convolutions and non-linear transformations to generate a latent tensor y. Note that in the context of deep learning, the terms "downsampling" and "upsampling" do not refer to resampling in the classical sense, but rather are general terms for changing the size of the H and W dimensions of a tensor. Resampling may also be spelled as resample. Similarly, resampling and resampled may also be spelled as resampling and resampled, respectively.
[0180] The latent tensor y output by the encoder 71 represents an image in the latent space,
[0181]
Number
[0182] having a size of, D e is the downsampling factor of the encoder 71, and C e is the number of channels (e.g., the number of neural network layers involved in the transformation of the tensor representing the input image).
[0183] The latent tensor y is further downsampled by the hyper encoder 72 using convolution and non - linear transformation to become the hyper latent tensor z. The hyper latent tensor z has a size
[0184]
Number
[0185] and is quantized by block Q to obtain the quantized hyper latent tensor
[0186]
Number
[0188]
Number
[0189] are estimated using a factorization entropy model. The arithmetic encoder AE uses these statistical characteristics to create a bit - stream representation of the tensor
[0190]
Number
[0191] Without requiring an autoregressive process, all elements of the tensor
[0192]
Number
[0193] are written into the bit - stream.
[0194] The factored entropy model functions as a codebook whose parameters are available on the decoder side. The arithmetic decoder AD uses the factored entropy model to recover the hyper-latent tensor
[0195]
Number
[0196] from the bitstream. The recovered hyper-latent tensor
[0197]
Number
[0198] is upsampled by the hyper-decoder 73 by applying a plurality of convolutional operations and non-linear transformations. The upsampled recovered hyper-latent tensor is denoted by Ψ. The entropy of the quantized latent tensor
[0199]
Number
[0200] is autoregressively estimated based on the upsampled recovered hyper-latent tensor Ψ. The autoregressive entropy model thus obtained is used to estimate the statistical characteristics of the quantized latent tensor
[0201]
Number
[0202] .
[0203] The arithmetic encoder AE uses the quantized latent tensor
[0204] [[ID=S5]]
Number
[0205] These estimated statistical properties are used to create a bitstream representation of the image. In other words, the arithmetic encoder AE of the autoencoder component compresses the image information in latent space by entropy coding based on secondary information provided by the hyperplier component. The latent tensor y is reconstructed from the bitstream by the arithmetic decoder AD at the receiver side using an autoregressive entropy model. The reconstructed latent tensor y is upsampled by decoder 74 by applying multiple convolution operations and nonlinear transformations to obtain a tensor representation of the reconstructed image.
[0206] Figure 8 shows a modified form of the architecture shown in Figure 7. The processing of encoder 81 and decoder 84 of the autoencoder component is similar to the processing of encoder 71 and decoder 74 of the autoencoder component shown in Figure 7, and the processing of encoder 82 and decoder 83 of the hyperplier component is similar to the processing of encoder 72 and decoder 73 of the hyperplier component shown in Figure 7. Note that each of these encoders 71, 81, 72, 82 and decoders 73, 83, 74, 84 may each have or be connected to a neural network. Furthermore, a neural network may be used to provide the entropy model involved.
[0207] Unlike the configuration shown in Figure 7, the configuration shown in Figure 8 involves a quantized latent tensor.
[0208]
number
[0209] teeth,
[0210]
number
[0211] To obtain a tensor Φ with a reduced number of elements compared to the original, it undergoes a masked convolution. The entropy model is obtained based on the connected tensor Φ and Ψ (a restored hyperlatent tensor that is upsampled). The entropy model thus obtained is a quantized latent tensor.
[0212]
number
[0213] It is used to estimate the statistical properties of [the subject].
[0214] Conditional coding represents a specific type of coding in which auxiliary information is used to improve the quality of the reconstructed image. Figure 9 illustrates the principle of conditional coding. Auxiliary information A is concatenated with the input frame x and processed jointly by encoder 91. The quantized coded information in latent space is written into a bitstream by an arithmetic encoder and reconstructed from the bitstream by an arithmetic decoder AD. The reconstructed coded information in latent space must be decoded by decoder 92 to obtain the reconstructed frame X. In this decoding stage, the latent representation a of auxiliary information A must be provided to the input of decoder 92. The latent representation a of auxiliary information A is provided by another encoder 93 and concatenated with the output of decoder 92.
[0215] In the context of video compression, a conditional codec is implemented to compress the residuals used for interpretation of the current block of the current frame, as shown in Figure 10. The residuals are calculated by subtracting the current block from its predicted version. The residuals are encoded by encoder 101 to obtain a residual bitstream. The residual bitstream is decoded by decoder 102. The predicted block is obtained by prediction unit 103 by using information from the previous frame / block. The predicted block is processed in a similar manner as it has the same size and dimensions as the current block. The reconstructed residuals are added to the predicted block to provide the reconstructed block.
[0216] A conditional residual coding (CodeNet) in this field is shown in Figure 11. The configuration is similar to the one shown in Figure 9. The conditional encoder 111 uses the predicted frame as auxiliary information to condition the codec.
[0217]
number
[0218] The current frame xt is encoded using information from the latent space. The quantized encoded information in the latent space is written into the bitstream by the arithmetic encoder and restored from the bitstream by the arithmetic decoder AD. The restored encoded information in the latent space is decoded by decoder 112 to obtain the reconstructed frame Xt. In this decoding stage, auxiliary information
[0219]
number
[0220] latent expression
[0221]
number
[0222] This needs to be applied to the input of decoder 112. (Supplementary information)
[0223]
number
[0224] latent expression
[0225]
number
[0226] This is provided by another encoder 113 and coupled with the output of decoder 112.
[0227] CodeNet uses predicted frames but does not use explicit differences (residuals) between the predicted frames and the current frames. Coding the current frame while extracting all information from the predicted frames can be advantageous compared to residual coding in that it brings to light the relatively small amount of information that needs to be sent.
[0228] However, CodeNet, due to the entropy prediction involved, does not enable high levels of parallelism and requires a large memory space. According to this disclosure, it is possible to reduce memory requirements and improve overall processing execution time.
[0229] This disclosure provides conditional coding in which a primary component of an image is encoded independently of one or more non-primary components, and one or more non-primary components are encoded using information from the primary component. Hereafter, the primary component may be a lumen component and one or more non-primary components may be chroma components, or the primary component may be a chroma component and a single non-primary component may be a lumen component. The primary component can be encoded and decoded independently of the non-primary components. Therefore, the primary component can be decoded even if the non-primary components are lost for some reason. One or more non-primary components can be encoded jointly and in parallel, and they can be encoded in parallel with the primary component. Decoding of one or more non-primary components utilizes information from the latent representation of the primary component. This type of conditional coding can be applied to intra-predictive and inter-predictive processing of video sequences. Furthermore, it can be applied to still image coding.
[0230] Figure 12 illustrates the basics of conditional intra prediction according to an exemplary embodiment. The tensor representation x of the input image / frame i is quantized and supplied to the encoding device 121. Note here, and in the following description, the entire image or only a portion of the image, such as one or more blocks, slices, tiles, etc., may be coded.
[0231] Prior to the encoding device 121, the tensor representation x is separated into a primary intra component and at least one non-primary (secondary) intra component, the primary intra component is converted into a primary intra component bitstream, and at least one non-primary intra component is converted into at least one non-primary intra component bitstream. The bitstreams represent compressed information about the component used by the decoding device 122 for component reconstruction. The two bitstreams can be interleaved with each other. The encoding device 121 may be treated as a conditional color separation (CCS) encoding device. The encoding of at least one non-primary intra component is based on information from the primary intra component, as will be described in detail later. Each bitstream is decoded by the decoding device 122 to reconstruct the image / frame. The decoding of at least one non-primary intra component is based on information from the latent representation of the primary intra component, as will be described in detail later.
[0232] Figure 13 illustrates the basics of residual coding according to an exemplary embodiment. The tensor representation x' of the input image / frame i' is quantized, the residuals are calculated and fed to the coding device 131. Prior to the coding device 131, the residuals are separated into first-order residual components and at least one non-first-order residual component, the first-order residual component is converted into a first-order residual component bitstream, and at least one non-first-order residual component is converted into at least one non-first-order residual component bitstream. The coding device 131 may be treated as a conditional color separation (CCS) coding device. The coding of at least one non-first-order residual component is based on information from the first-order residual component, as will be described in detail later. Each bitstream is decoded by the decoding device 132 to reconstruct the image / frame. The decoding of at least one non-first-order residual component is based on information from the latent representation of the first-order residual component, as will be described in detail later. The calculation of residuals and the predictions required for the reconstructed image / frame are provided by the prediction unit 133.
[0233] In the configurations shown in Figures 12 and 13, the encoding devices 121 and 131 and the decoding devices 131 and 132 may each have their own neural network or be connected to each other's neural networks. The encoding devices 121 and 131 may have variational autoencoders. A different number of channel / neural network layers may be involved in processing the primary component compared to processing at least one non-primary component. The encoding devices 121 and 131 may determine the appropriate number of channel / neural network layers by performing a brute-force search or in a content-adaptive manner. A set of models may be trained, each model based on a different number of channels for encoding primary and non-primary components. During processing, the encoding devices 121 and 131 may determine the filters to best execute. The neural networks of the encoding devices 121 and 131 may be trained collaboratively to determine the number of channels used to process primary and non-primary components. In some applications, the number of channels used to process the first-order components may be greater than the number of channels used to process the non-first-order components. In other applications, for example, if the first-order component signal is less noisy than the non-first-order component signal, the number of channels used to process the first-order components may be less than the number of channels used to process the non-first-order components. In principle, the choice of the number of channels may be derived from optimizations relating to processing rate on the one hand and signal distortion on the other. Extra channels can reduce distortion but may result in a greater processing load. Experiments have shown that a suitable number of channels may be, for example, 128 for the first-order components and 64 for the non-first-order components, or 128 for both first-order and non-first-order components, or 192 for the first-order components and 64 for the non-first-order components.
[0234] The number of channels / neural network layers used in the encoding process may be implicitly or explicitly signaled to the decoding devices 122 and 132, respectively.
[0235] Figure 14 shows, to some extent, in more detail an embodiment of conditional coding of an image (a frame of a video sequence or a still image). Encoder 141 receives a tensor representation of the primary component P of the image having a size of H P ×W P ×C P where H P represents the height dimension of the image, W P represents the width dimension of the image, and C P represents the input channel dimension. Hereinafter, a tensor having a size of A×B×C is usually simply abbreviated and referred to as tensor A×B×C. Similarly, a tensor having a size of C×A×B is usually simply abbreviated and referred to as tensor C×A×B.
[0236] Exemplary sizes in the height, width, and channel dimensions of the tensor output by encoder 141 are H P / 16×W P / 16×128.
[0237] Note that encoders 141 and 142 may be provided within encoding devices 121 and 131.
[0238] Based on the output of encoder 141, that is, the representation of the tensor representation of the primary component of the image in the latent space, a bitstream is generated and converted back to the latent space to obtain the restored tensor
[0239]
Number
[0240] in the latent space.
[0241] H NP represents the height dimension of the image, W NP represents the width dimension of the image, and C NP represents the input channel dimension, a tensor representation H NP ×WNP ×C NP is the tensor representation H of the primary component P P ×W P ×C P and the concatenation with (thus, the tensor H NP ×W NP ×(C NP +C P )) is input into another encoder 142 and then input into that other encoder 142. Exemplary sizes in the height, width, and channel dimensions of the tensor output by encoder 142 are H P / 16 × W P / 16 × 64 or H P / 32 × W P / 32 × 64.
[0242] Before the concatenation, the sample locations of the tensor representation H P ×W P ×C P of the primary component P may have to be adjusted to the sample locations of the tensor representation H NP ×W NP ×C NP of at least one non-primary component NP if the sample sizes or sub-pixel offsets of the samples of the tensor are different from each other. Based on the output of the other encoder 142, i.e., the representation of the concatenated tensor image in the latent space, a bitstream is generated and converted back to the latent space to obtain the restored concatenated tensor
[0243]
Number
[0244] in the latent space.
[0245] Constraint s UV / s Y ∈ integer, the sub-pixel offset in the resampling operation is equal to 0, and the resampling process can be performed without using interpolation (which is a computationally expensive operation). Correspondingly, the coding efficiency can be improved.
[0246] On the first side, the reconstructed tensor in the latent space
[0247]
number
[0248] However, the reconstructed tensor representation H P ×W P ×C P Based on this, the first-order component P of the image is input into the decoder 143 for reconstruction.
[0249] Furthermore, in latent space, tensor
[0250]
number
[0251] Tensor with
[0252]
number
[0253] The concatenation is performed. In this case as well, if the sample sizes or subpixel offsets of these tensors to be concatenated are different from each other, some adjustment of the sample locations is required. On the non-primary side, the tensor obtained from this concatenation
[0254]
number
[0255] However, the reconstructed tensor representation H NP ×W NP ×C NP Based on this, at least one non-linear component NP of the image is input into another decoder 144 for reconstruction.
[0256] The coding described above may be performed on the primary component P independently of at least one non-primary component NP. For example, the coding of the primary component P and at least one non-primary component NP may be performed in parallel. Compared to the present art, overall processing parallelism can be increased. Furthermore, numerical experiments have shown that shorter channel lengths can be used compared to the present art without significant degradation of the quality of the reconstructed image, and therefore memory requirements can be reduced.
[0257] Below, exemplary implementations of conditional coding of image components represented in YUV space (one lumen component Y and two chroma components U and V) are described with reference to Figures 15-20. It goes without saying that the disclosed conditional coding is also applicable to any other (color) space that may be used for image representation.
[0258] In the embodiment shown in Figure 15, input data in YUV420 format is processed, where Y represents the lumen component of the current image to be processed, UV represents the chroma components U and V of the current image to be processed, and 420 indicates that the size of the lumen component Y in the height and width dimensions is four times larger than the size of the chroma component UV (twice the height and twice the width). In the embodiment shown in Figure 15, Y is selected to be a first-order component processed independently of UV, and UV is selected to be a non-first-order component. The UV components are processed together.
[0259] The YUV representation of the image to be processed is separated into (primary) Y components and (non-primary) UV components. The encoder 151, which includes a neural network,
[0260]
number
[0261] The encoder receives a tensor representing the Y component of the image to be processed using the following size, where H and W are the height and width dimensions, and the input depth (i.e., the number of channels) is 1 (for one lumen component). The output of encoder 151 is,
[0262]
number
[0263] This is a latent tensor with the size C, where C y C is the number of channels assigned to the Y component. In this embodiment, the four downsampling layers in the encoder 151 reduce both the height and width of the input tensor by a coefficient of 16 (downsampling), and the number of channels C y There are 128 of them. The obtained latent representations of the Y component are processed by the hyperprior Y pipeline.
[0264] The UV component of the image to be processed is a tensor
[0265]
number
[0266] This is represented by , where H and W are the height and width dimensions, and the number of channels is 2 (for the two chroma components). Conditional coding of the UV component requires auxiliary information from the Y component. If the plane size (H and W) of the Y component is different from the size of the UV component, a resampling unit is used to align the position of the sample in the tensor representing the Y component with the position of the sample in the tensor representing the UV component. Similarly, if there is an offset between the position of the sample in the tensor representing the Y component and the position of the sample in the tensor representing the UV component, alignment must be performed.
[0267] The aligned tensor representation of the Y component is a tensor
[0268]
number
[0269] To obtain the latent tensor, it is concatenated with the tensor representation of the UV component. The encoder 152, which has a neural network, uses this concatenated tensor as a latent tensor.
[0270]
number
[0271] Convert to C uv is the number of channels assigned to the UV component. In this embodiment, five downsampling layers in encoder 152 reduce both the height and width of the input tensor by a coefficient of 32 (downsampling), resulting in 64 channels. The obtained latent representation of the UV component is processed by a hyperplier UV pipeline similar to the hyperplier Y pipeline (see the explanation of Figure 7 above for the operation of the pipelines). Note that both the hyperplier UV pipeline and the hyperplier Y pipeline may include neural networks.
[0272] The hyperplier Y pipeline provides an entropy model used for entropy coding of the (quantized) latent representation of the Y component. The hyperplier Y pipeline comprises a (hyper)encoder 153, an arithmetic encoder, an arithmetic decoder, and a (hyper)decoder 154.
[0273] To obtain the hyperlatent tensor which is converted into a bitstream by the arithmetic-encoded AE (in some cases, after quantization not shown in Figure 15, in effect, any quantization performed by the quantization unit Q is optional here and below), the latent tensor representing the Y component in the latent space
[0274]
number
[0275] However, it is further downsampled by the (hyper)encoder 153 using convolution and nonlinear transformations. The statistical properties of the (quantized) hyperlatent tensor are estimated using an entropy model, such as a factorized entropy model, and the arithmetic encoder AE of the hyperplier Y pipeline uses these statistical properties to create a bitstream. All elements of the (quantized) hyperlatent tensor may be written into the bitstream without requiring an autoregressive process.
[0276] The (factorized) entropy model functions as a codebook for which its parameters are available on the decoder side. The arithmetic decoder AD of the hyperprior Y pipeline reconstructs the hyperlatent tensor from the bitstream by using the (factorized) entropy model. The reconstructed hyperlatent tensor is upsampled by the (hyper)decoder 154 by applying multiple convolution operations and nonlinear transformations. The latent tensor representing the Y component in latent space.
[0277]
number
[0278] The quantized latent tensor is then quantized by the quantization unit Q of the hyperplier Y pipeline, and the entropy of the quantized latent tensor is autoregressively estimated based on the upsampled, reconstructed hyperlatent tensor output by the (hyper)decoder 154.
[0279] The latent tensor representing the Y component in the latent space
[0280]
number
[0281] It is also quantized by another arithmetic encoder AE, provided by the hyperplier Y pipeline, which uses the estimated statistical properties of the tensor, before being converted into a bitstream (which may be transmitted from the transmitter to the receiver).
[0282]
number
[0283] The latent tensor is reconstructed from the bitstream by another arithmetic decoder AD using an autoregressive entropy model provided by the hyperplier Y pipeline.
[0284]
number
[0285] teeth,
[0286]
number
[0287] The image, having a size such as , is upsampled by decoder 155 by applying multiple convolution operations and nonlinear transformations to obtain a reconstructed tensor representation of the Y component.
[0288] The hyperplier UV pipeline takes the output of encoder 152, i.e., the latent tensor.
[0289]
number
[0290] This latent tensor is processed. This latent tensor is further downsampled by the (hyper)encoder 156 of the hyperplier UV pipeline using convolution and nonlinear transformations to obtain a hyperlatent tensor that is converted into a bitstream by the arithmetic encoded AE of the hyperplier UV pipeline (possibly after quantization not shown in Figure 15). The statistical properties of the (quantized) hyperlatent tensor are estimated using an entropy model, e.g., a factorized entropy model, and the arithmetic encoder AE of the hyperplier Y pipeline uses these statistical properties to create the bitstream. All elements of the (quantized) hyperlatent tensor may be written into the bitstream without requiring an autoregressive process.
[0291] The (factorized) entropy model functions as a codebook for which its parameters are available on the decoder side. The arithmetic decoder AD of the hyperplier UV pipeline reconstructs the hyperlatent tensor from the bitstream by using the (factorized) entropy model. The reconstructed hyperlatent tensor is upsampled by the (hyper)decoder 157 of the hyperplier UV pipeline by applying multiple convolution operations and nonlinear transformations. The latent tensor representing the UV component
[0292]
number
[0293] The latent tensor is quantized by the quantization unit Q of the hyperplier UV pipeline, and the entropy of the quantized latent tensor is autoregressively estimated based on the upsampled, reconstructed hyperlatent tensor output by the (hyper)decoder 157.
[0294] Latent tensor representing the UV component in latent space
[0295]
number
[0296] It is also quantized by another arithmetic encoder AE, provided by the hyperplier UV pipeline, which uses the estimated statistical properties of that tensor, before being converted into a bitstream (which may be transmitted from the transmitter to the receiver). The latent tensor representing the UV component in latent space.
[0297]
number
[0298] However, it is reconstructed from the bitstream by another arithmetic decoder AD using an autoregressive entropy model provided by the hyperplier UV pipeline.
[0299] Reconstructed latent tensor representing the UV component in latent space
[0300]
number
[0301] This is a tensor that is input into decoder 158 on the UV processing side.
[0302]
number
[0303] To obtain the reconstructed latent tensor after further downsampling,
[0304]
number
[0305] This is linked to, i.e., the restored latent tensor.
[0306]
number
[0307] However, (as auxiliary information required for decoding the UV component) tensor
[0308]
number
[0309] It is connected to,
[0310]
number
[0311] The image is upsampled by its decoder 158 by applying multiple convolution operations and nonlinear transformations to obtain a tensor representation of the reconstructed UV component of the image having a size such that the image is upsampled. The tensor representation of the reconstructed UV component of the image is coupled with the tensor representation of the reconstructed Y component of the image to obtain a reconstructed image in YUV space.
[0312] Figure 16 shows an embodiment similar to that shown in Figure 15, but for processing input data in YUV444 format, where the sizes of the tensors representing the Y and UV components are the same in the height-width dimension, respectively. The encoder 161 is the tensor representing the Y component of the image to be processed.
[0313]
number
[0314] This is transformed into latent space. The auxiliary information does not need to be resampled according to this embodiment and is therefore a tensor representing the UV component of the image to be processed.
[0315]
number
[0316] This is a tensor representing the Y component.
[0317]
number
[0318] It can be directly connected to and a connected tensor
[0319]
number
[0320] This is converted to latent space by encoder 162 on the UV side. The hyperplier Y pipeline, comprising (hyper)encoder 163 and (hyper)decoder 164, and the hyperplier UV pipeline, comprising (hyper)encoder 166 and (hyper)decoder 167, operate in the same manner as described above with reference to Figure 15. Since the restored latent representations of the U and UV components have the same size in height and width, they can be concatenated together in latent space without resampling. Restored latent representation of the U component
[0321]
number
[0322] This is upsampled by decoder 165, and the restored concatenated latent representation of the Y and UV components.
[0323]
number
[0324] The image is upsampled by decoder 168, and the outputs of decoders 165 and 168 are combined to obtain a reconstructed image in YUV space.
[0325] Figures 17 and 18 illustrate embodiments in which conditional residual coding is provided. Conditional residual coding may be used for interpretation of the current frame of a video sequence or for still image coding. Unlike the embodiments shown in Figures 15 and 16, residuals with residual components in YUV space are processed. The residuals are separated into residual Y components for the Y component and residual UV components for the UV component. Processing of the residual components is similar to processing of the Y and UV components as described above with reference to Figures 15 and 16. According to the embodiment shown in Figure 17, the input data is in YUV420 format. Therefore, the residual Y component must be downsampled before concatenation with the residual UV component. Encoders 171 and 172 provide their respective latent representations. A hyperplier Y pipeline with (hyper)encoder 173 and (hyper)decoder 174, and a hyperplier UV pipeline with (hyper)encoder 176 and (hyper)decoder 177, operate as described above with reference to Figure 15. On the residual Y component side, decoder 175 outputs the restored representation of the residual Y component. On the residual UV side, decoder 178 outputs the restored representation of the residual UV component based on the auxiliary information provided in the latent space, and downsampling of the restored latent representation of the residual Y component is required. The outputs of decoders 175 and 178 are combined to obtain the restored residuals in YUV space, which can be used to obtain (a portion of) the restored image.
[0326] According to the embodiment shown in Figure 18, the input data is in YUV444 format. Downsampling of auxiliary information is not required. Processing of residual Y and UV components is similar to the processing of Y and UV components described above with reference to Figure 16. The encoder 181 is a tensor representing the residual Y component of the image to be processed.
[0327]
number
[0328] Transforms into latent space. Tensor representing the residual UV component of the image to be processed.
[0329]
number
[0330] This is a tensor representing the residual Y component.
[0331]
number
[0332] It can be directly connected to and a connected tensor
[0333]
number
[0334] This is converted to latent space by encoder 182 on the residual UV side.
[0335] The hyperplier Y pipeline, comprising (hyper)encoder 183 and (hyper)decoder 184, and the hyperplier UV pipeline, comprising (hyper)encoder 186 and (hyper)decoder 187, operate in the same manner as described above with reference to Figure 15.
[0336] Since the restored latent representations of the residual U component and residual UV component have the same height and width, they can be concatenated without resampling. Restored latent representation of the residual U component
[0337]
number
[0338] This is upsampled by decoder 185, and the restored concatenated latent representation of the residual Y and residual UV components.
[0339]
number
[0340] The outputs of decoders 185 and 188 are combined to obtain the restored residual of the image in YUV space, which can be used to obtain the restored image (or portion thereof) that has been upsampled by decoder 188.
[0341] Figure 19 shows an alternative embodiment to the embodiment shown in Figure 17. The only difference in the configuration shown in Figure 19 is that the autoregressive entropy model is not employed. Tensor
[0342]
number
[0343] The representation of the residual Y component, expressed by , is transformed in latent space by encoder 191. The residual Y component is a tensor
[0344]
number
[0345] Using encoder 192 which outputs a tensor
[0346]
number
[0347] It is used as auxiliary information for coding the residual UV component represented by . The hyperplier Y pipeline, comprising a (hyper)encoder 193 and a (hyper)decoder 194, is used for the latent representation of the residual Y component
[0348]
number
[0349] It provides secondary information used for coding. Decoder 195 is a tensor
[0350]
number
[0351] The reconstructed residual Y component is output by (hyper)encoder 196 and (hyper)decoder 197. The hyperplier UV pipeline, comprising (hyper)encoder 196 and (hyper)decoder 197, outputs the tensor output by encoder 192.
[0352]
number
[0353] The latent representation of, i.e., a tensor
[0354]
number
[0355] This provides secondary information used for coding. Decoder 198 is a connected tensor in latent space.
[0356]
number
[0357] Received, Tensor
[0358]
number
[0359] It outputs the reconstructed residual UV component represented by [the specified formula].
[0360] Figure 20 shows an alternative embodiment with respect to the embodiment shown in Figure 18. In this case as well, the only difference is that the configuration shown in Figure 20 does not employ an autoregressive entropy model.
[0361] tensor
[0362]
number
[0363] The representation of the residual Y component, expressed by , is transformed in latent space by encoder 201. The residual Y component is a tensor
[0364]
number
[0365] Tensor using encoder 202 which outputs
[0366]
number
[0367] It is used as auxiliary information for coding the residual UV component represented by . The hyperplier Y pipeline, comprising a (hyper)encoder 203 and a (hyper)decoder 204, is used for the latent representation of the residual Y component
[0368]
number
[0369] It provides secondary information used for coding. Decoder 205 is a tensor
[0370]
number
[0371] The reconstructed residual Y component, represented by the (hyper)encoder 206 and (hyper)decoder 207, outputs the tensor output by encoder 202. The hyperplier UV pipeline, comprising (hyper)encoder 206 and (hyper)decoder 207, outputs the tensor output by encoder 202.
[0372]
number
[0373] The latent representation of, i.e., a tensor
[0374]
number
[0375] It provides secondary information used for coding. Decoder 208 is a reconstructed representation of the residual UV component in latent space.
[0376]
number
[0377] Received, Tensor
[0378]
number
[0379] It outputs the reconstructed residual UV component represented by [the specified formula].
[0380] Processing that does not involve the adoption of an autoregressive entropy model may reduce the overall complexity of the process and, depending on the actual application, may further improve the accuracy of the reconstructed image.
[0381] According to the embodiment shown in Figure 21, a method for reconstructing at least a portion of an image is provided. In S231, a first bitstream is parsed based on, for example, a first entropy model, in order to obtain a first latent tensor, and the first latent tensor is processed to obtain a first tensor representing the first-order components of the image, S233. Furthermore, a second bitstream different from the first bitstream is parsed based on, for example, a second entropy model different from the first entropy model, in order to obtain a second latent tensor different from the first latent tensor, S235. In S237, the first latent tensor is resampled based on integer factors to obtain a resampled first latent tensor. Then, in S239, a second tensor representing at least one second-order component of the image is obtained based on the second latent tensor and the resampled first latent tensor. The integer factors are sometimes called integer scaling factors.
[0382] Since integer factors are used to obtain the resampled first latent tensor, this is extremely beneficial for both compression performance and ease of handling. The resampling process can be performed without interpolation (which is a computationally expensive operation). Correspondingly, coding efficiency can be improved.
[0383] The integer factors may be calculated based on a first scaling factor for the linear component and a second scaling factor for at least one quadratic component. For example, the integer factors are s UV / s Y It is calculated as, where, s Y represents the first scaling factor, s UV This represents the second scaling factor.
[0384] s Y When s exists in the bitstream, Y This is obtained by parsing the bitstream. Y When s does not exist in the bitstream, Y This is set to the default value. Y The default value for this can be 1.
[0385] s UV When s exists in the bitstream, UV This is obtained by parsing the bitstream. UV When s does not exist in the bitstream, UV This is set to the default value. UV The default value can be 2.
[0386] As another example, an integer factor r may exist in the bitstream, or it may be set to a default value.
[0387] If r exists in the bitstream, it is obtained by parsing the bitstream. If r does not exist in the bitstream, it is set to its default value. The default value of r may be 2.
[0388] s Y When s exists in the bitstream, Y This is obtained by parsing the bitstream. Y When s does not exist in the bitstream, Y This is set to the default value. Y The default value for this can be 1.
[0389] Next, s UV ga s UV =r·s Y It is derived as, however, s UV This represents the scaling factor for at least one quadratic component.
[0390] In the sample above, s Y This is the scaling factor for the lumens (both horizontally and vertically), s UV This is the scaling factor for chroma (both horizontally and vertically).
[0391] In other implementations, different scaling factors may be used, for example, horizontally and vertically. s Y _hor is the scaling factor for the rumor (horizontally), s Y _ver is the scaling factor for the lumens (vertically), s UV _hor is the scaling factor for the chroma (horizontally), s UV _ver is the scaling factor for the chroma (vertically).
[0392] The integer factors may also include horizontal and vertical coefficients.
[0393] Resampling is performed by using nearest neighbor upsampling or nearest neighbor downsampling.
[0394] For nearest-near-side upsampling, This is denoted as s↑. This layer has a size [C, h in , w in It receives a tensor input of size [C, s·h]. in , s·w in Outputs the tensor output of ].
[0395] Interpolation is not required for this downsampling; it simply involves copying. Output[c, i, j]=Input[c, s·i, s·j], i=0,...,h in -1, j=0,...,w in -1, c=0,...,C-1 This is the case where s represents an integer factor.
[0396] For nearest-neighbor downsampling, This is represented as s↓. This layer has a size [C, s·h out , s·w out It receives a tensor input of size [C, h out , w out This outputs a tensor output of ], and this procedure reduces the spatial resolution of each tensor channel.
[0397] Interpolation is not required for this downsampling; it simply involves copying. Output[c, s·i, s·j]=Input[c, i, j], i=0,...,h in -1, j=0,...,w in -1, c=0,...,C-1 This is the case where s represents an integer factor.
[0398] The decoder architecture shown in Figure 22 is an example of how to implement the method shown in Figure 21. The data (tensors and streams) are shown inside the "white" box, the neural network modules required for decoding are shown in the gray shaded box, and the switchable tools are shown in the purple shaded box.
[0399] For the primary and secondary color components, the code streams may be parsed independently and reconstructed using modules consisting of the same sequence of neural network layers, with only differences in size and number of tensor channels in the input tensor.
[0400] The first stream z may be parsed by a reversible entropy decoder (me-tANS decoder).
[0401]
number
[0402] The probability distribution for reversible coding is assumed to be a Gaussian with pre-trained parameters (part of the trained model), and the cumulative distribution function calculated based on those pre-trained parameters is used in the reversible entropy decoder.
[0403] Decoded hyperplier tensor
[0404]
number
[0405] It is used as input for two different processes, namely, a hyperdecoder and a hyperscale decoder.
[0406] Next, stream y may be parsed by a reversible decoder (me-tANS decoder).
[0407]
number
[0408] The probability distribution for parsing is assumed to be a Gaussian with a mean and standard deviation of 0, given as the output when the hyperscale decoder outputs a tensor σ[C,h4,w4], which is then scaled according to a rate control parameter β inside the sigma scale to produce σ', and then masked and scaled according to an RVS parameter inside the adaptive sigma scale to produce σ''. Finally, the tensor σ'' value is quantized (converted to an index in the probability distribution table). Some elements of the residual tensor are skipped (not encoded / decoded) and replaced with 0 in the decoder skip module, which then reshapes the output into a parsed set of syntax elements {s} from the tANS decoder, mask_sigma from the skip mask generation module, and a reconstructed residual tensor in 3D shape.
[0409]
number
[0410] Receive.
[0411] On the decoder side, residual
[0412]
number
[0413] It is scaled by the inverse gain unit according to parameter β.
[0414]
number
[0415] This generates the residual tensor. Then, the residual tensor is scaled within the invRVS (inverse residual and variance scale) module.
[0416]
number
[0417] This forms the reconstructed latent tensor.
[0418]
number
[0419] It is used for that purpose.
[0420] The hyperdecoder generates explicit_prediction which is input to the multi-stage context model MCM, and the multi-stage context model MCM then reconstructs the residuals.
[0421]
number
[0422] It also takes the same as input, and is a latent space tensor
[0423]
number
[0424] This is an 8-stage neural network process that outputs the following: Latent Scaling Before Synthesis (LSBS) followed by the reconstructed latent space tensor.
[0425]
number
[0426] The signal is ready for reconstruction. Latent tensor reconstructions for the first and second-order components are independent of each other.
[0427] Reconstructed latent space tensor
[0428]
number
[0429] is the input to the composite transformation. Another input to the composite transformation is the auxiliary tensor.
[0430]
number
[0431] In the case of quadratic component synthesis, the auxiliary tensor is generated from the first-order component reconstructed latent tensor. In the case of first-order components, the auxiliary tensor is not used. Y ) and secondary (s UVThe size of the tensor is shown in Table 1, depending on the input picture height H and width W for the component, as well as the scaling factor.
[0432] [Table 1]
[0433] For the first-order component, parameter C d = 0. This means that the composite transformation of the first-order components does not receive auxiliary information (it is reconstructed independently). In the case of the second-order components, C d = 128, and the auxiliary for quadratic transformation composition is the integer factor s of the linear component. UV / s Y Reconstructed latent space tensor
[0434]
number
[0435] Resampled by
[0436]
number
[0437] Resampling may be performed using nearest neighbor upsampling or nearest neighbor downsampling as described above.
[0438] In one embodiment, resampling is performed after the composite transformation and before the filter. The filter may be any filter disclosed in this application. In one example, resampling is performed by receiving the output of the composite transformation and resampling the output.
[0439] Resampling may be performed using nearest neighbor upsampling or nearest neighbor downsampling as described above.
[0440] In one embodiment, if the size of the quadratic component of the coded picture parameter is not equal to the color sampling mode of the output picture, the quadratic component is coded at a lower resolution compared to the output size, and an upsampling process is applied. Supported color sampling modes for input and output images having corresponding ratios between the primary component size and the quadratic component size are listed in Table 2. In one example, the resampling (S237) step can be performed by resampling the first latent tensor based on an integer factor to obtain the resampled first latent tensor, the integer factor being c in Table 2. hor Or c ver Either or c in Table 2. hor Or c ver It is obtained based on this.
[0441] In one embodiment, the decoding process begins with parsing the picture header, which contains the picture size, quality parameter β, model identifier, tiling, tool, and color sampling mode (s) of the output picture. ver s hor ), the size of the quadratic component of the coded picture parameter (c ver c hor This includes information about tools, etc.
[0442] In one embodiment, the next step is z-stream decoding (for the first order z Y - z for streams and quadratic components UV The z-stream is a probability distribution table ("z-table") for the z-stream arithmetic decoder, which is part of the trained model. This process is performed for the first and second components, respectively, and for the three-dimensional hypertensor.
[0443]
number
[0444] and
[0445]
number
[0446] Generates.
[0447] In one embodiment, the decoded hypertensor
[0448]
number
[0449] It is used as input for two different processes, namely, a hyperdecoder and a hyperscale decoder.
[0450] In one embodiment, an arithmetic decoder for quality maps takes a residual stream "q-stream" as input and outputs a quality_map of size [h4, w4].
[0451] In one embodiment, the entropy and residual decoding operations are quantized, using multipliers of 8 bits or less, to generate 16-bit data at every layer. The quantized operation ensures no overflow of the 32-bit integer register, thus guaranteeing bit-strict behavior. The hyperscale decoder has a residual tensor size ([C for the first order). p [C for h4, w4] and the second-order component s A residual signal variable (I) having a logarithmic scale, with [h4, w4]) σ The hyperscale decoder is followed by a sigma scale ("VarScale"). The sigma scale operation is logarithmic scale I'. σ It is part of the variable rate support that outputs I'. This tensor is used in the mask generation process for SKIP, RVS, and LSBS. σIt passes through the adaptive sigma scale ("RVS Scale"). Finally, the variance I'' takes on a logarithmic scale. σ This is quantized and shows a probability distribution table for residual decoding by an arithmetic decoder (size [C] for the first order). p [C for h4, w4] and the second-order component s Provides SigmaIdx for [h4, w4].
[0452] In one embodiment, the arithmetic decoder for residuals takes the residual stream "r-stream", SigmaIdx (for deriving the probability distribution table), and SkipMask as input. Tensor elements to be skipped for coding by SkipMask are replaced with 0. The arithmetic decoder is the residual tensor (for the first order).
[0453]
number
[0454] and (for the second-order component)
[0455]
number
[0456] Outputs.
[0457] In one embodiment, the residual is descaled within the inverse gain unit module for variable rate support and further modified within the inverse RVS.
[0458] In one embodiment, the hyperdecoder uses a prediction tensor (
[0459]
number
[0460] and
[0461]
number
[0462] ) generates residuals inside the latent tensor reconstruction (
[0463]
number
[0464] and
[0465]
number
[0466] ) is combined with the first-order component. In the case of the first-order component, the reconstructed latent representation of the image is, i.e., the first-order component
[0467]
number
[0468] The multi-stage context model used to generate it. For quadratic components,
[0469]
number
[0470] The residual generated is simply the prediction plus the residual.
[0471] In one embodiment, primary and secondary color component code streams can be parsed, and latent tensor (
[0472]
number
[0473] ,
[0474]
number
[0475] ) can be reconstructed independently.
[0476] In one embodiment, the composite transformation is preceded by the LSBS process (if the LSBS tool is enabled). Several composite transformations with different architectures and parameters are defined. The composite transformations for the first and second-order components are specified by the DecoderID. None of the composite transformation networks are latent tensors.
[0477]
number
[0478] and
[0479]
number
[0480] It can be used to reconstruct an image from. Sub-clause 10.3 specifies three different synthesis networks (DecoderID=0, 1, 2). The synthesis transformation for the first-order component is as input.
[0481]
number
[0482] Only is received (the primary component can be reconstructed independently). The composite transformation for the secondary component is as follows:
[0483]
number
[0484] and
[0485]
number
[0486] It receives both. The output of the composite transformation network is the output picture for the first order.
[0487]
number
[0488] and
[0489]
number
[0490] It has the following dimensions.
[0491] In one embodiment, s ver / c ver ↑ s hor / c hor ↑'). Those color components enter an optional filter module (specified in Annex I) consisting of several chroma filters and one lumar edge filter.
[0492] In one embodiment, the analytical transformation network generates a tensor y having sizes [Cp, h4, w4] for the first-order component and [Cs, h4, w4] for the second-order component. The hypertensor z has sizes [Cp, h6, w6] for the first-order component and [Cs, h6, w6] for the second-order component. Supported color sampling modes for input and output images having corresponding ratios between the first-order and second-order component sizes are listed in Table 2.
[0493] [Table 2]
[0494] In the example shown in Figure 22,
[0495]
number
[0496] is the first latent tensor,
[0497]
number
[0498] This is the resampled first latent tensor.
[0499]
number
[0500] or
[0501]
number
[0502] This is the first tensor representing the first-order component of the image.
[0503]
number
[0504] This is the second latent tensor.
[0505]
number
[0506] or
[0507]
number
[0508] This is a second tensor representing at least one quadratic component of the image.
[0509] The composite transformation for first-order and second-order components consists of the same neural network layers, the only difference being the size of the input tensor and the number of tensor channels.
[0510]
number
[0511] The output is shown (the tensor sizes are listed in Table 1).
[0512] As shown in Figure 22, after reconstruction, the first and second-order components are resampled using their own scaling factors (in Figure 22, the upsampling module is indicated by "s↑"). After resampling to the original picture size, all three color components pass through an Inter-Channel Correlation Information (ICCI) filter.
[0513] The reconstruction process is completed by an inverse color transformation ("invColorTr" in Figure 22).
[0514] The decoder architecture shown in Figure 23 is another example of implementing the method shown in Figure 21. The difference between Figure 23 and Figure 22 is that in Figure 22,
[0515]
number
[0516] Nearest neighbor downsampling is performed on the result, but in Figure 23
[0517]
number
[0518] This involves performing nearest neighbor upsampling on the data.
[0519] An example of a signal decoder in Figures 22 and 23 is shown in Figure 24. Signal decoders are sometimes called synthetic transforms. Learning-based reconstruction (also called synthetic transform) involves two pipelines with identical neural network architectures except for the input size and number of channels.
[0520] The input for analysis and transformation is as follows: - Auxiliary information tensor
[0521]
number
[0522] The reconstructed latent space tensor of the connected shape [C, h4, w4]
[0523]
number
[0524] , - Operating point indicator opIdx, - Size H of the input / output tensor in , W in , - Model parameters for the composite transformation network, defined by the pair (modelIdx, opIdx).
[0525] The output of the analysis transformation is size [C in , H in , W in ] Reconfigured color components
[0526]
number
[0527] It is a tensor.
[0528] The sizes of the tensors for the first and second-order components are listed in Table 1.
[0529] Composite transformations are the main (
[0530]
number
[0531] ) and support (
[0532]
number
[0533] The process begins with connection to the input. Depending on the operating point indicator (opIdx), the decoder executes the following sequence of steps:
[0534] With respect to the base operating point (opIdx=0), the number of channels is C+C. d One lightweight residual block is followed by a series of two transposed convolutions with kernel size 4×4, coupled with a residual activation unit having a cropping layer (stride 2, corresponding depths 4 and 3) and kernel size 3×3. The number of output channels in the transposed convolutions are correspondingly C1 and C2. The stride for both transposed convolutions is 2. The next step in the process is a normal convolution with kernel size 3×3, stride 1, and an unchanged number of channels C2 coupled with a residual activation unit (kernel size 3×3). Then the number of channels is reduced from C2 to 16C in There is a 3x3 convolution with a stride of 1 to increase the number of channels. The next step is a pixel shuffle with a stride of 4 outputs, which has a number of channels C. inThis is done to ensure that it will have the desired properties. The process concludes with a cropping layer (stride 4, depth 1).
[0535] For a high operating point (opIdx=1), the number of channels is C+C. d The two residual blocks are followed by a series of two transposed convolutions with kernel size 3x3, coupled with a residual activation having a cropping layer (stride 2, corresponding depths 4 and 3) and kernel size 3x3. The number of output channels in both transposed convolutions is C. The stride for both transposed convolutions is 2. The next step in the process is a normal convolution with kernel size 3x3 and stride 1, with output channels 4C. This is done to ensure that the next step, a pixel shuffle with stride 2 output, will have number of channels C. Then, a residual nonlocal attention block (∝=1) is executed, coupled with a residual activation having a cropping layer (stride 2, depth 2) and kernel size 3x3. The process has kernel size 3x3, stride 2, and number of output channels C. in It terminates with a transposed convolution having a crossed layer, followed by a cropping layer (stride 2, depth 1).
[0536] A processing device 250 for reconstructing at least a portion of an image is provided in Figure 25, and the processing device 250 comprises a processing circuit configuration 255 configured to perform a method as shown in Figure 21.
[0537] Several exemplary implementation forms in hardware and software A corresponding system in which the encoder-decoder processing chain described above can be deployed is shown in Figure 26. Figure 26 is a schematic block diagram showing exemplary coding systems, e.g., video, image, audio, and / or other coding systems (or short coding systems) that may utilize the techniques of this application. The video encoder 20 (or short encoder 20) and video decoder 30 (or short decoder 30) of the video coding system 10 represent examples of devices that may be configured to perform the techniques described in the various examples in this application. For example, video coding and decoding may employ a neural network, such as those shown in Figures 1 to 6, which may be distributed and may apply the bitstream parsing and / or bitstream generation described above to transmit feature maps between distributed computing nodes (two or more).
[0538] As shown in Figure 26, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 in order to decode encoded picture data 13.
[0539] The source device 12 includes an encoder 20 and may additionally, i.e., optionally, include a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.
[0540] The picture source 16 may comprise, or may comprise, any type of picture capture device, e.g., a camera for capturing real-world pictures, and / or any type of picture generation device, e.g., a computer graphics processor for generating computer-animated pictures, or any other type of device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may comprise any type of memory or storage for storing any of the pictures described above.
[0541] In contrast to the processing performed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 is sometimes referred to as the raw picture or raw picture data 17.
[0542] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain a preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising. It can be understood that the preprocessing unit 18 may consist of optional components. It should also be noted that the preprocessing may employ a neural network (as shown in any of Figures 1 to 7) that uses presence indicator signaling.
[0543] The video encoder 20 is configured to receive pre-processed picture data 19 and provide encoded picture data 21.
[0544] The communication interface 22 of the source device 12 may be configured to receive encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) to another device, such as the destination device 14 or any other device, via the communication channel 13 for storage or direct reconstruction.
[0545] The destination device 14 includes a decoder 30 (for example, a video decoder 30) and may additionally, i.e., optionally, include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0546] The communication interface 28 of the destination device 14 is configured to receive encoded picture data 21 (or any further processed version thereof) from, for example, the source device 12 directly or from any other source, for example, a storage device, for example, an encoded picture data storage device, and to provide the encoded picture data 21 to the decoder 30.
[0547] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 13 via a direct communication link between the source device 12 and the destination device 14, for example, via a direct wired or wireless connection, or via any type of network, for example, a wired or wireless network or any combination thereof, or any type of private and public network, or any type of combination thereof.
[0548] The communication interface 22 may be configured, for example, to package the encoded picture data 21 into an appropriate format, such as a packet, and / or to process the encoded picture data using any kind of transmit encoding or processing for transmission over a communication link or communication network.
[0549] The communication interface 28, which forms the counterpart to the communication interface 22, may be configured, for example, to receive the transmitted data and process the transmitted data using any kind of corresponding transmit decoding or processing and / or depackaging to obtain the encoded picture data 21.
[0550] Both communication interfaces 22 and 28 may be configured as unidirectional or bidirectional communication interfaces, as indicated by the arrows to the communication channel 13 in Figure 26, pointing from the source device 12 to the destination device 14, and may be configured to send and receive messages, for example, to set up a connection, to confirm and exchange any other information relating to the communication link and / or data transmission, such as the transmission of encoded picture data. The decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (for example, by employing a neural network based on one or more of Figures 1 to 7).
[0551] The post-processor 32 of the destination device 14 is configured to post-process decoded picture data 31 (also called reconstructed picture data), for example, decoded picture 31, in order to obtain post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare, for example, the decoded picture data 31 for display by the display device 34.
[0552] The display device 34 of the destination device 14 is configured to receive post-processed picture data 33 for displaying the picture to, for example, a user or viewer. The display device 34 may be any type of display for representing the reconstructed picture, for example, an integrated or external display or monitor, or may comprise such a display. The display may comprise, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a microLED display, a liquid crystal on silicon (LCoS), a digital optical processor (DLP), or any other type of display.
[0553] Figure 26 shows the source device 12 and the destination device 14 as separate devices, but the device embodiment may also have both or both functionalities, the source device 12 or its corresponding functionality, and the destination device 14 or its corresponding functionality. In such embodiments, the source device 12 or its corresponding functionality and the destination device 14 or its corresponding functionality may be implemented using the same hardware and / or software, or by separate hardware and / or software, or in any combination thereof.
[0554] As will become apparent to those skilled in the art based on the description, the functionality of different units or the presence and (strict) division of functionality within the source device 12 and / or destination device 14 as shown in Figure 26 may vary depending on the actual device and application.
[0555] The encoder 20 (e.g., video encoder 20) or the decoder 30 (e.g., video decoder 30), or both the encoder 20 and the decoder 30, may be implemented via a processing circuit configuration, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof dedicated to video coding. The encoder 20 may be implemented via a processing circuit configuration 46 to embody various modules, including neural networks, such as those shown in any or part of Figures 1 to 6. The decoder 30 may be implemented via a processing circuit configuration 46 to embody various modules, such as those described with respect to Figures 1 to 7 and / or any other decoder systems or subsystems described herein. The processing circuit configuration may be configured to perform various operations, as described later. If the technique is partially implemented in software, the device may store instructions for the software in a suitable non-temporary computer-readable storage medium, and one or more processors may be used to execute the instructions in hardware to perform the technique of this disclosure. Both the video encoder 20 and the video decoder 30 may be integrated as part of a composite encoder / decoder (codec) within a single device, for example, as shown in Figure 27.
[0556] The source device 12 and destination device 14 may comprise any of a wide range of devices, including any type of handheld or stationary device, such as a notebook or laptop computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video game console, video streaming device (such as a content service server or content distribution server), broadcast receiver device, broadcast transmitter device, etc., and may or may not have an operating system. In some cases, the source device 12 and destination device 14 may be equipped for wireless communication. Therefore, the source device 12 and destination device 14 may be wireless communication devices.
[0557] In some cases, the video coding system 10 shown in Figure 26 is merely an example, and the techniques of this application may be applied to video coding configurations (e.g., video coding or video decoding) that do not necessarily involve any data communication between the coding device and the decoding device. In other examples, data may be retrieved from local memory and streamed over a network. The video coding device may code the data and store it in memory, and / or the video decoding device may retrieve the data from memory and decode it. In some examples, coding and decoding are performed by devices that do not communicate with each other, but simply code the data into memory and / or retrieve the data from memory and decode it.
[0558] Figure 28 is a schematic diagram of a video coding device 2000 according to one embodiment of the present disclosure. The video coding device 2000 is suitable for carrying out the disclosed embodiments as described herein. In one embodiment, the video coding device 2000 may be a decoder, such as the video decoder 30 in Figure 26, or an encoder, such as the video encoder 20 in Figure 26.
[0559] The video coding device 2000 comprises an inlet port 2010 (or input port 2010) and a receiver unit (Rx) 2020 for receiving data, a processor, logic unit, or central processing unit (CPU) 2030 for processing data, a transmitter unit (Tx) 2040 and an exit port 2050 (or output port 2050) for transmitting data, and memory 2060 for storing data. The video coding device 2000 may also comprise an optical-to-electrical (OE) component and an electric-to-optical (EO) component coupled to the inlet port 2010, a receiver unit 2020, a transmitter unit 2040, and an exit port 2050 for the exit or input of optical or electrical signals.
[0560] The processor 2030 is implemented by hardware and software. The processor 2030 may be implemented as one or more CPU chips, cores (e.g., a multi-core processor), FPGAs, ASICs, and DSPs. The processor 2030 communicates with the input port 2010, the receiver unit 2020, the transmitter unit 2040, the output port 2050, and the memory 2060. The processor 2030 includes a coding module 2070. The coding module 2070 implements the embodiments disclosed above. For example, the coding module 2070 performs, processes, prepares, or provides various coding operations. Thus, the inclusion of the coding module 2070 results in a significant improvement in the functionality of the video coding device 2000 and affects the transition of the video coding device 2000 to different states. Alternatively, the coding module 2070 is implemented as instructions stored in the memory 2060 and executed by the processor 2030.
[0561] Memory 2060 may comprise one or more disks, tape drives, and solid-state drives, and may be used as an overflow data storage device to store the program when such a program is selected for execution, and to store instructions and data read during program execution. Memory 2060 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random-access memory (RAM), ternary associative memory (TCAM), and / or static random-access memory (SRAM).
[0562] Figure 29 is a simplified block diagram of the apparatus 800, which may be used as one or both of the source device 12 and destination device 14 from Figure 26, according to an exemplary embodiment.
[0563] The processor 2102 in the device 2100 may be a central processing unit. Alternatively, the processor 2102 may be any other type of device, existing or to be developed in the future, capable of manipulating or processing information. The disclosed implementation may be practiced using a single processor, e.g., processor 2102, as shown in the figure, but advantages in speed and efficiency may be realized by using two or more processors.
[0564] The memory 2104 in the device 2100 may be a read-only memory (ROM) device or a random access memory (RAM) device in one implementation. Any other suitable type of storage device may be used as memory 2104. Memory 2104 may contain code and data 2106 accessed by the processor 2102 using the bus 2112. Memory 2104 may further contain an operating system 2108 and an application program 2110, the application program 2110 including at least one program that enables the processor 2102 to perform the method described herein. For example, the application program 2110 may include applications 1 to N, and applications 1 to N further include a video coding application that performs the method described herein.
[0565] The device 2100 may also include one or more output devices, such as a display 2118. The display 2118 may, in one example, be a touch-sensitive display that combines the display with a touch-sensitive element capable of operating to sense touch input. The display 2118 may be coupled to the processor 2102 via a bus 2112.
[0566] Although illustrated here as a single bus, the bus 2112 of device 2100 may consist of multiple buses. Furthermore, secondary storage may be directly coupled to other components of device 2100 or accessible via a network, and may comprise a single integrated unit such as a memory card or multiple units such as multiple memory cards. Device 2100 can be implemented in a wide variety of configurations.
[0567] Furthermore, the processing unit 250 shown in Figure 25 may include a source device 12 or destination device 14 shown in Figure 26, a video coding system 40 shown in Figure 27, a video coding device 2000 shown in Figure 28, or a device 2100 shown in Figure 29. [Explanation of Symbols]
[0568] 10 Video Coding Systems 12 Source Devices 13 Communication Channels 14 Destination device 16 Picture Sources 17. Pictures, picture data, unprocessed pictures, unprocessed picture data 18 Preprocessors, Picture Preprocessors, Preprocessing Units 19 Pre-processed pictures, pre-processed picture data 20 Video Encoders 21 Encoded Picture Data 22 Communication interface, communication unit 28 Communication interface, communication unit 30 video decoders 31 Decrypted picture data, decrypted picture 32 Post-processors, Post-processing Units 33 Post-processed pictures 34 Display Devices 40 Video Coding Systems 46 Processing Circuit Configuration 71 Encoders 72 Hyper Encoders 73 Hyper Decoder 74 Decoders 81 Encoders 82. Encoders of Hyperplier Components 83 Hyperplier Component Decoder 84 Decoders 91 encoders 92 Decoder 93 Another encoder 101 Encoder 102 Decoder 103 prediction units 111 Conditional Encoder 112 Decoder 113 Another encoder 121 Encoding Devices 122 Decryption Devices 131 Encoding Devices 132 Decryption Devices 133 prediction units 141 encoders 142 Another encoder 143 Decoder 144 Another Decoder 151 encoders 152 encoders 153 Hyper Encoder 154 Hyperdecoder 155 Decoder 156 Hyper Encoder 157 Hyper Decoder 158 Decoder 161 encoders 162 encoders 163 Hyper Encoder 164 Hyper Decoder 165 Decoder 166 Hyper Encoders 167 Hyper Decoder 168 Decoders 171 Encoder 172 encoders 173 Hyper Encoder 174 Hyper Decoder 175 Decoder 176 Hyper Encoder 177 Hyper Decoder 178 Decoder 181 Encoder 182 encoders 183 Hyper Encoder 184 Hyper Decoder 185 Decoder 186 Hyper Encoder 187 Hyper Decoder 188 Decoders 191 encoders 192 encoders 193 Hyper Encoder 194 Hyper Decoder 195 Decoder 196 Hyper Encoder 197 Hyper Decoder 198 Decoder 201 Encoder 202 encoders 203 Hyper Encoder 204 Hyper Decoder 205 Decoder 206 Hyper Encoder 207 Hyper Decoder 208 Decoder 250 Processing Units 255 Processing Circuit Configuration 2000 Video Coding Devices 2010 Inlet port, Input port 2020 Receiver Unit 2030 Processor, Logical Unit, Central Processing Unit 2040 Transmitter Unit 2050 Exit port, output port 2060 memory 2070 Coding Module 2100 equipment 2102 Processor 2104 memory 2106 Data 2108 Operating Systems 2110 Application Program 2112 Bus 2118 Display
Claims
1. A method for reconstructing at least a portion of an image, The first step is to parse the first bitstream in order to obtain the first latent tensor (S231), The steps include processing the first latent tensor to obtain a first tensor representing the first-order component of the aforementioned image (S233), In order to obtain a second latent tensor different from the first latent tensor, the process involves parsing a second bitstream different from the first bitstream (S235), Step (S237) of resampling the first latent tensor based on integer factors in order to obtain a resampled first latent tensor, Step (S239) of obtaining a second tensor representing at least one quadratic component of the image based on the second latent tensor and the resampled first latent tensor. A method for providing this.
2. The method according to claim 1, wherein the resampling step (S237) is performed based on the integer factor without using interpolation.
3. The method according to claim 1 or 2, wherein the integer factor is calculated based on a first scaling coefficient for the first-order component and a second scaling coefficient for at least one second-order component.
4. The integer factor is s UV / s Y It is calculated as s Y represents the first scaling factor, s UV The method according to claim 3, wherein represents the second scaling factor.
5. s Y If it exists in the bitstream or is set to the first default value, UV The method according to claim 4, wherein is present in the bitstream or is set to a second default value.
6. s Y The first default value of is 1, and s UV The method according to claim 5, wherein the second default value of is 2.
7. The method according to claim 1 or 2, wherein the integer factor is present in the bitstream or set to a first default value.
8. s Y is present in the bitstream or set to a second default value, s UV where s UV =r · s Y is derived as, s Y represents a first scaling factor for the primary component, s UV represents a second scaling factor for the at least one secondary component, and r represents the integer factor, the method according to claim 7.
9. The first default value of r is 2, and s Y The method according to claim 8, wherein the second default value of is 1.
10. The method according to any one of claims 1 to 9, wherein the resampling is performed by using nearest neighbor upsampling or nearest neighbor downsampling.
11. The nearest neighbor upsampling tensor input is of size [C, h in , w in ] and the tensor output of the nearest neighbor upsampling is of size [C, s・h in , s・w in The method according to claim 10, wherein s represents the integer factor and interpolation is not performed for the nearest neighbor upsampling.
12. The tensor input for the nearest neighbor downsampling is of size [C, s・h out , s・w out ] and the tensor output of the nearest neighbor downsampling is of size [C, h out , w out The method according to claim 10, wherein s represents the integer factor and interpolation is not performed for the nearest neighbor downsampling.
13. The method according to any one of claims 1 to 12, wherein the primary component of the image is a lumen component, and the at least one secondary component of the image is a chroma component.
14. The method according to any one of claims 1 to 13, wherein the second tensor represents two quadratic components, one of which is a chromatic component and the other is a different chromatic component.
15. The processing of the first latent tensor (S234) includes the step of converting the first latent tensor into the first tensor. The method according to any one of claims 1 to 14.
16. The step of obtaining the second tensor representing at least one quadratic component of the aforementioned image is: The method comprises the steps of concatenating the second latent tensor with the resampled first latent tensor to obtain a connected tensor, and converting the connected tensor to the second tensor. The method according to claim 15.
17. The method according to any one of claims 1 to 16, wherein the first bitstream is parsed by a first neural network, and the second bitstream is parsed by a second neural network different from the first neural network.
18. The method according to claim 16, wherein the first latent tensor is transformed by a third neural network, and the connected latent tensor is transformed by a fourth neural network different from the third neural network.
19. The first tensor represents the first residual component of the residual with respect to the first component of the image, The second tensor uses information from the first latent tensor to represent at least one quadratic residual component of the residual for at least one quadratic component of the image, The method according to any one of claims 1 to 18.
20. The method according to any one of claims 3 to 6, 8, and 9, wherein the first scaling factor includes a first horizontal scaling factor and a first vertical scaling factor.
21. The method according to any one of claims 3 to 6, 8, 9, and 20, wherein the second scaling factor includes a second horizontal scaling factor and a second vertical scaling factor.
22. A computer program stored on a non-temporary medium, comprising code that, when executed on one or more processors, performs a step of the method according to any one of claims 1 to 21.
23. Processing apparatus (40, 250, 2000, 2100) for reconstructing at least a portion of an image, One or more processors (43, 255, 2030, 2102) The apparatus comprises a non-temporary computer-readable storage medium coupled to one or more processors and storing a program for execution by the one or more processors, wherein the apparatus is configured such that when the program is executed by the one or more processors, it performs the method according to any one of claims 1 to 21. Processing units (40, 250, 2000, 2100).
24. A processing device (250) for reconstructing at least a portion of an image, wherein the processing device (40, 250, 2000, 2100) comprises a processing circuit configuration (255) configured to perform the method described in any one of claims 1 to 21.