Signaling of Feature Map Data

By using presence indicators and side information signaling, the method addresses inefficiencies in feature map data compression, enhancing computational efficiency and reducing latency in neural network processing.

JP7710034B2Active Publication Date: 2025-07-17HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023516786
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-20
Filing Date
2021-10-20
Publication Date
2025-07-17
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

Existing methods for compressing feature map data in neural networks are inefficient, leading to high data transfer requirements and increased computational complexity, especially in distributed systems where resource constraints are prevalent.

Method used

A method for efficiently encoding and decoding feature map data by using presence indicators to determine whether to parse or skip regions of the feature map, along with signaling side information, thereby reducing the amount of data transferred and simplifying entropy decoding.

Benefits of technology

This approach reduces the amount of data transmitted and simplifies entropy decoding, improving computational efficiency and reducing latency in neural network processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710034000036
    Figure 0007710034000036
  • Figure 0007710034000037
    Figure 0007710034000037
  • Figure 0007710034000038
    Figure 0007710034000038
Patent Text Reader

Abstract

The present disclosure relates to efficient signaling of feature map information for systems employing neural networks. In particular, at the decoder side, a presence indicator is obtained based on information parsed from a bitstream. Based on the value of the obtained presence indicator, further data related to the feature map region is analyzed or the analysis is bypassed. The presence indicator may be, for example, a region presence indicator indicating whether the bitstream contains feature map data, or a side information presence indicator indicating whether the bitstream contains side information related to the feature map data. Similarly, encoding methods and even encoding and decoding devices are also provided. Thus, feature map data may be processed more efficiently, including reducing the decoding complexity and even the amount of transmitted data by applying a bypass.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of signaling information related to neural network processing, such as feature maps or other side information.

Background Art

[0002] Hybrid image and video codecs have been used for decades to compress image and video data. In such codecs, a signal is typically encoded with respect to a block by predicting the block and further encoding only the difference between the original block and its prediction. In particular, such encoding may include transformation, quantization, and bitstream generation, but usually includes some form of entropy encoding. Typically, the three components of a hybrid encoding method, transformation, quantization, and entropy encoding, are optimized separately. Recent video compression standards such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC) also use transform representations to encode the residual signal after prediction.

[0003] In recent years, machine learning has been applied to image and video coding. In general, machine learning can be applied to image and video coding in a variety of different ways. For example, some end-to-end optimized image or video coding schemes have been discussed. Additionally, machine learning has been used to determine or optimize some parts of end-to-end coding, such as the selection or compression of prediction parameters or the like. These applications are common in that they generate some feature map data to be transmitted between the encoder and the decoder. An efficient structure of the bitstream can significantly contribute to reducing the number of bits used to encode the image / video source signal.

[0004] Efficient signaling of feature map data is also beneficial for other machine learning applications or for transmitting feature map data or related information between layers of a machine learning system that may be distributed.

[0005] Collaborative intelligence is one of several new paradigms for efficiently deploying deep neural networks on mobile-cloud infrastructure. By splitting the network, for example, between a (mobile) device and the cloud, it is possible to distribute the computational workload so that the overall energy and / or latency of the system is minimized. In general, distributing the computational workload enables resource-constrained devices to be used in the deployment of neural networks. Additionally, computer vision tasks or image / video encoding for machines are applications that can operate in a distributed manner and utilize collaborative intelligence.

[0006] A neural network typically includes two or more layers. A feature map is the output of a layer. In a neural network split between devices, such as between a device and the cloud or between different devices, the feature map at the output of the split location (e.g., the first device) is compressed and transmitted to the remaining layers of the neural network (e.g., to the second device).

[0007] Since transmission resources are typically limited, it is desirable to reduce the amount of data transferred while still providing configurability to support various quality requirements. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0008] The present invention relates to a method and apparatus for compressing data used in neural networks. Such data may include features (e.g., feature maps), without limitation.

[0009] The present invention is defined by the scope of the independent claims. Some of the advantageous embodiments are provided in the dependent claims.

[0010] In particular, some embodiments of the present disclosure relate to the signaling of feature map data that can be used by a neural network. Efficiency can be improved by carefully designing the bitstream structure. In particular, a presence indicator is signaled to indicate the presence or absence of feature map data or generally side information. Depending on the presence indicator, encoding and / or decoding of the feature map or regions thereof is either performed or skipped.

[0011] According to one aspect, a method for decoding a feature map for processing by a neural network based on a bitstream is provided, the method including obtaining a region presence indicator based on information from the bitstream for a region of the feature map, and decoding the region including parsing data from the bitstream to decode the region when the region presence indicator has a first value and bypassing parsing data from the bitstream to decode the region when the region presence indicator has a second value. Thus, the feature map data can be efficiently encoded within the bitstream. In particular, by skipping the parsing, the complexity of entropy decoding can be reduced and further the rate can be reduced (reducing the amount of data to be included within the bitstream).

[0012] In some embodiments, the region presence indicator is decoded from the bitstream.

[0013] In some embodiments, the region presence indicator is derived based on information signaled within the bitstream. For example, when the region presence indicator has a second value, decoding the region further includes setting the region according to a predetermined rule. The ability to have a rule indicating how to fill in the missing parts of the feature map data enables the decoder to appropriately compensate for non-signaled data.

[0014] For example, the predetermined rule specifies setting the features of the region to a constant. Constant embedding is a simple and efficient way to fill in missing parts without involving entropy decoding of all feature elements. In some embodiments, the constant is zero.

[0015] This method may include the step of decoding the constant from the bitstream. Signaling the constant can provide flexibility and improve the reconstructed result.

[0016] In an exemplary implementation, the method includes obtaining a side information presence indicator from the bitstream, parsing the side information from the bitstream when the side information presence indicator has a third value, and bypassing parsing the side information from the bitstream when the side information presence indicator has a fourth value, where the side information may further include at least one of the region presence indicator and information regarding being processed by a neural network to obtain an estimated probability model for use in entropy decoding of the region. Additional side information may allow for more refined settings. The presence indicator for the side information can make its encoding mode efficient.

[0017] For example, when the side information presence indicator has a fourth value, the method includes setting the side information to a predetermined side information value. Providing the function of filling in the side information enables handling of missing (non-signaled) data.

[0018] In some embodiments, the region presence indicator is a flag that can take on only one of two values formed by a first value and a second value. The binary flag provides a particularly rate-efficient indication.

[0019] In some embodiments, the region is a channel of the feature map. The channel level is a natural layer granularity. Since the channels are already separately (separable) available, no further partitioning is necessary. Thus, utilizing per-channel granularity is an efficient way to scale rate and distortion. However, it is understood that the region of a channel can have any other granularity, such as a sub-region or unit of a channel, values of different channels corresponding to the same spatial arrangement of sub-regions or units, a single value of the feature map, values at the same spatial arrangement in different channels.

[0020] This method may further include obtaining a significance order indicating the significance of a plurality of channels of the feature map, obtaining a last significant channel indicator, and obtaining a region presence indicator based on the last significant channel indicator. Ordering the channels, or generally regions of the feature map, can help simplify the signaling of region presence. In some embodiments, the indication of the last significant channel corresponds to the index of the last significant channel within the significance order. Signaling the last significant region as side information in the bitstream, instead of signaling a flag for each region, can be more efficient.

[0021] The indication of the last significant channel corresponds to a quality indicator decoded from the bitstream and can indicate the quality of the encoded feature map resulting from the compression of the region of the feature map. Thus, region channel presence can be derived without additional signaling since quality is typically also indicated as side information in the bitstream for other purposes.

[0022] For example, obtaining the significance order includes decoding an indication of the significance order from the bitstream. Explicit signaling may provide complete flexibility in setting / configuring the order. However, the present disclosure is not limited to explicit signaling of the significance order.

[0023] In an exemplary implementation, obtaining the significance order includes deriving the significance order based on previously decoded information (or side information) regarding the source data from which the feature map was generated. One advantage of this implementation is that the additional signaling overhead is low.

[0024] For example, obtaining the significance order includes deriving the significance order based on previously decoded information regarding the type of source data from which the feature map was generated. The type of source data may also provide a good indication about the purpose and characteristics of the data and thus the quality required for its reconstruction.

[0025] For example, from the bitstream, it includes decoding channels sorted within the bitstream according to the significance order from the most significant channel to the least significant channel. This will enable the use of a bitrate scalability feature that allows reducing the bitrate without re-encoding by dropping (cutting) the least significant channels according to a desired quality level.

[0026] In an exemplary implementation, the method decodes region division information from a bitstream that instructs to divide the region of the feature map into units, and further includes a step of decoding or not decoding a unit presence indication that indicates whether the feature map data should be parsed from the bitstream to decode the units of the region according to the division information. By further dividing (partitioning) a region such as a channel into units (partitions), the granularity of indicating the presence of these units is made finer, and thus the overhead can be scaled more flexibly.

[0027] For example, the region division information for a region includes a flag indicating whether the bitstream contains unit information specifying the dimensions and / or positions of the units of the region, and the method includes decoding a unit presence indication for each unit of the region from the bitstream, and parsing or not parsing the feature map data for the unit from the bitstream according to the value of the unit presence indication for the unit. By signaling the parameters of the division, the rate and distortion relationship can be controlled with additional degrees of freedom.

[0028] In some embodiments, the unit information specifies a hierarchical division of a region including at least one of a quadtree, a binary tree, a ternary tree, or a triangular division. These divisions provide the advantage of efficient signaling, such as those developed for hybrid codecs.

[0029] In some exemplary implementations, decoding of regions from a bitstream includes extracting from the bitstream a last significant coefficient indicator that indicates the position of the last coefficient among the coefficients of the region, decoding the significant coefficients of the region from the bitstream, setting coefficients following the last significant coefficient indicator according to a predetermined rule, and obtaining feature data of the region by inverse-transforming the coefficients of the region. Application of the transform and encoding of the coefficients after the transform enable easy identification of the significant coefficients since the significant coefficients are typically located within a region of coefficients with low indices. This can further enhance the signaling efficiency.

[0030] For example, the inverse transform is an inverse discrete cosine transform, an inverse discrete sine transform, or an inverse transform obtained by modifying the inverse discrete cosine transform or the inverse discrete cosine transform, or a convolutional neural network transform. These exemplary transforms are already used for the purpose of residual encoding and can be efficiently implemented, for example, with existing hardware / software / algorithm approaches.

[0031] This method may further include decoding from the bitstream a side information presence flag that indicates whether the bitstream contains any side information for the feature map, where the side information includes information regarding being processed by a neural network to obtain an estimated probability model for use in entropy decoding of the feature map. The presence indication for the side information can provide more efficient signaling, similar to the feature map region presence indication.

[0032] For example, decoding of the region presence indicator includes decoding by a context adaptive entropy decoder. Further encoding / decoding by adaptive entropy coding makes the bitstream more compact.

[0033] According to one aspect, a method for decoding a feature map for processing by a neural network from a bitstream is provided, the method including obtaining from the bitstream a side information indicator regarding the feature map, and decoding the feature map including, when the side information indicator has a fifth value, parsing side information for decoding the feature map from the bitstream, and when the side information indicator has a sixth value, bypassing parsing side information for decoding the feature map from the bitstream. An existence indication regarding the side information may provide more efficient signaling, similar to the feature map region existence indication.

[0034] The method may include entropy decoding, which is based on the decoded feature map processed by the neural network. This enables efficient adaptation of the entropy coder to the content.

[0035] For example, when the side information indicator has a sixth value, the method includes setting the side information to a predetermined side information value. Thus, signaling of the side information value is omitted, which may lead to rate improvement. For example, the predetermined side information value is zero.

[0036] In some embodiments, the method includes decoding the predetermined side information value from the bitstream. Providing a value to be used by default increases the flexibility of its setting.

[0037] According to one aspect, a method for decoding an image is provided. The method includes a method according to any of the methods described above for decoding a feature map for neural network processing from a bitstream, and obtaining a decoded image including processing the decoded feature map by a neural network. The application of the above method for machine learning-based image or video encoding can significantly improve the efficiency of such encoding. For example, the feature map represents encoded image data and / or encoded side information for decoding the image data.

[0038] In another aspect, a method for computer vision is provided. The method includes a method according to any of the methods described above for decoding a feature map for neural network processing from a bitstream, and performing a computer vision task including processing the decoded feature map by a neural network. The application of the above method to a machine vision task can help reduce the rate required by the transfer of information within a machine learning-based model such as a neural network, especially when implemented in a distributed manner. For example, the computer vision task is object detection, object classification, and / or object recognition.

[0039] According to one aspect, a method for encoding a feature map for neural network processing in a bitstream is provided. The method includes, for a region of the feature map, obtaining a region presence indicator based on an obtained region presence indicator of the feature map, and determining to encode the region of the feature map when the region presence indicator has a first value, and bypassing encoding the region of the feature map when the region presence indicator has a second value.

[0040] In some embodiments, the region presence indicator is indicated within the bitstream.

[0041] In some embodiments, the region presence indicator is derived based on information already signaled within the bitstream.

[0042] The encoding methods referred to herein impart a specific structure to the bitstream and enable the use of the advantages described above for the corresponding decoding methods.

[0043] For example, determining (or obtaining the region presence indicator) includes evaluating the values of the characteristics of the region. Context / content-based adaptation of signaling utilizes certain correlations and makes it possible to reduce the rate of the encoded stream as much as possible. In such a way, determining (or obtaining the region presence indicator) may include evaluating values in the context of the region, in other words, evaluating values that are spatially adjacent to the region in the feature map.

[0044] In some embodiments, determining is based on the impact of the region on the quality of the result of the neural network processing. Such a determination leads to a rate reduction that is sensitive to the distortion resulting from the information cut.

[0045] According to an exemplary implementation, determining includes stepwise determining the total number of bits required for transmission of the feature map starting from the bits of the most significant region and continuing to regions with decreasing significance until the total exceeds a pre-configured threshold, encoding regions where the total does not exceed the pre-configured threshold, and encoding a region presence indicator having a first value for the encoded regions and a second value for non-encoded regions (also referred to as unencoded regions). This approach provides a computationally efficient approach that meets the constraints imposed on the rate and distortion of the resulting stream.

[0046] According to one aspect, a method is provided for encoding a feature map for processing by a neural network within a bitstream, the method comprising obtaining the feature map, determining whether to indicate side information regarding the feature map, and indicating within the bitstream either a side information indicator that indicates a third value and the side information, or a side information indicator that indicates a fourth value without side information. Thus, the encoding may be able to signal only the information that is practically necessary to achieve the required quality / rate.

[0047] In some embodiments, the region presence indicator and / or the side information indicator is a flag that can take on only one of two values formed by a first value and a second value. The binary flag provides a particularly rate-efficient indication. In some embodiments, the region is a channel of the feature map.

[0048] It is understood that the regions of the feature map can have different granularities, for example, different channel values corresponding to sub-regions or units of a channel, a single value of the feature map, values at the same spatial location in different channels.

[0049] The method may further comprise obtaining a significance order that indicates the significance of a plurality of channels of the feature map, obtaining a last significant channel indicator, and obtaining a region presence indicator based on the last significant channel indicator. In some embodiments, the indication of the last significant channel corresponds to the index of the last significant channel within the significance order.

[0050] The indication of the last significant channel corresponds to a quality indicator that is encoded (inserted) into the bitstream as side information and may indicate the quality of the encoded feature map resulting from the compression of the region of the feature map.

[0051] For example, the significance order indication is inserted (encoded) into the bitstream as side information.

[0052] In an exemplary implementation, obtaining the significance order includes deriving a significance order based on previously encoded information regarding the source data from which the feature map was generated. For example, obtaining the significance order includes deriving a significance order based on previously encoded information regarding the type of source data from which the feature map was generated.

[0053] For example, it includes encoding channels sorted within the bitstream according to the significance order from the most significant channel to the least significant channel within the bitstream.

[0054] In an exemplary implementation, the method further includes encoding region division information in the bitstream that instructs to divide the region of the feature map into units, and depending on the division information, encoding (inserting into the bitstream) or not encoding a unit presence indication that indicates whether feature map data is included in the bitstream for decoding the units of the region.

[0055] For example, the region division information for a region includes a flag indicating whether the bitstream contains unit information specifying the dimensions and / or positions of the units of the region, and the method includes decoding a unit presence indication for each unit of the region from the bitstream and, depending on the value of the unit presence indication for the unit, parsing or not parsing the feature map data for the unit from the bitstream. By signaling the parameters of the division, the rate and distortion relationship can be controlled with additional degrees of freedom.

[0056] In some embodiments, the unit information specifies a hierarchical division of a region including at least one of a quadtree, a binary tree, a ternary tree, or a triangular division.

[0057] In some exemplary implementations, encoding a region into a bitstream includes including in the bitstream a last significant coefficient indicator that indicates the position of the last coefficient among the coefficients of the region, encoding the significant coefficients of the region in the bitstream, and transforming the feature data of the region by a transform to thereby obtain the coefficients of the region.

[0058] For example, the inverse transform is an inverse discrete cosine transform, an inverse discrete sine transform, or an inverse transform obtained by modifying an inverse discrete cosine transform or an inverse discrete sine transform, or a convolutional neural network transform.

[0059] This method may further include encoding in the bitstream a side information presence flag that indicates whether the bitstream includes any side information for the feature map, where the side information includes information regarding being processed by a neural network to obtain an estimated probability model for use in entropy decoding of the feature map.

[0060] For example, encoding the region presence indicator and / or the side information indicator includes encoding by a context adaptive entropy encoder.

[0061] According to one aspect, a method for encoding a feature map for processing by a neural network in a bitstream is provided, the method including including in the bitstream a side information indicator regarding the feature map, where encoding the feature map includes inserting a fifth value of the side information indicator in the bitstream, the side information being information for decoding the feature map, and including the side information in the bitstream or including a sixth value of the side information indicator in the bitstream and not inserting side information for the feature map region in the bitstream.

[0062] This method may include entropy encoding, which is based on an encoded feature map processed by a neural network.

[0063] For example, when the side information indicator has a sixth value, the encoder controls it to set the side information to a predetermined side information value on the decoder side. In some embodiments, this method includes encoding the predetermined side information value into the bitstream.

[0064] According to one aspect, a method for encoding an image is provided, which includes obtaining an encoded image by processing an input image with a neural network, thereby obtaining a feature map, and a method according to any of the methods described above for encoding the feature map into a bitstream. For example, the feature map represents encoded image data and / or encoded side information for decoding the image data.

[0065] In another aspect, a method for computer vision is provided, which includes a method according to any of the methods described above for encoding a feature map processed by a neural network into a bitstream, and a step of processing an image to thereby obtain a feature map with a neural network and providing the decoder with the feature map for performing a computer vision task based thereon. For example, the computer vision task is object detection, object classification, and / or object recognition.

[0066] According to one aspect, a computer program stored in a non-transitory medium includes code for performing the steps of any of the methods presented above when executed on one or more processors.

[0067] According to one aspect, a device for decoding a feature map for neural network processing based on a bitstream is provided. The device includes a region existence indicator acquisition module configured to acquire a region existence indicator for a region of the feature map based on information from the bitstream, and a decoding module configured to, when the region existence indicator has a first value, analyze data from the bitstream to decode the region, and when the region existence indicator has a second value, bypass analyzing data from the bitstream to decode the region, for decoding the region.

[0068] According to one aspect, a device for decoding a feature map for neural network processing from a bitstream is provided. The device includes a side information indicator acquisition module configured to acquire a side information indicator regarding the feature map from the bitstream, and a decoding module configured to decode the feature map, where decoding includes analyzing side information for decoding the feature map from the bitstream when the side information indicator has a fifth value, and bypassing analyzing side information for decoding the feature map from the bitstream when the side information indicator has a sixth value.

[0069] According to one aspect, a device for encoding a feature map for neural network processing in a bitstream is provided. The device includes a feature map acquisition module configured to acquire the feature map, and an encoding control module configured to, based on the acquired feature map, acquire a region presence indicator and, based on the acquired region presence indicator, determine whether to encode a region of the feature map when the region presence indicator has a first value or bypass encoding the region of the feature map when the region presence indicator has a second value.

[0070] According to one aspect, a device for encoding a feature map for neural network processing in a bitstream is provided. The device includes a feature map acquisition module configured to acquire the feature map, and an encoding control module configured to determine whether to indicate side information regarding the feature map, the encoding control module indicating, in the bitstream, either a side information indicator indicating a third value and side information or a side information indicator indicating a fourth value without side information.

[0071] According to one aspect, a device for decoding a feature map for neural network processing based on a bitstream is provided. The device includes a processing circuit configured to perform steps of a method according to any of the above methods.

[0072] According to one aspect, a device for encoding a feature map for neural network processing in a bitstream is provided. The device includes a processing circuit configured to perform steps of a method according to any of the above methods.

[0073] According to one aspect, the decoder device is implemented by the cloud. In such a scenario, some embodiments may provide a good trade-off relationship between the rate required for transmission and the accuracy of the neural network.

[0074] Any of the devices described above may be embodied on an integrated chip.

[0075] Any of the above-described embodiments and exemplary implementations may be combined.

[0076] Hereinafter, embodiments of the present invention will be described in more detail with reference to the attached drawings and figures.

Brief Description of the Drawings

[0077]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

DETAILED DESCRIPTION OF THE INVENTION

[0078] In the following description, reference is made to the accompanying drawings which form a part hereof and which illustrate, by way of example, specific aspects of embodiments of the invention or specific aspects in which embodiments of the invention may be used. It is understood that embodiments of the invention may be used in other aspects and may include structural or logical changes not shown in the figures. Accordingly, the following detailed description should not be construed in a limiting sense, and the scope of the invention is defined by the appended claims.

[0079] For example, the disclosure related to the described method may also apply to the corresponding device or system configured to perform the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units, such as functional units (e.g., one unit performing one or more steps, or multiple units each performing one or more of the multiple steps), even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific device is described based on one or more units, such as functional units, the corresponding method may include one step of performing the functions of the one or more units (e.g., one step of performing the functions of one or more units, or multiple steps each performing the functions of one or more of the multiple units), even if such one or more steps are not explicitly described or illustrated in the figures. Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specified.

[0080] Some embodiments are directed to providing low-complexity compression of data for neural networks. For example, the data may include feature maps, or other data used in neural networks such as weights or other parameters. In some exemplary implementations, signaling is provided that may enable generation of an efficient bitstream for a neural network. By a collaborative intelligence paradigm, a mobile device or edge device may have feedback from the cloud if needed. Note, however, that the present disclosure is not limited to the framework of a collaborative network that includes the cloud. This may be employed in any distributed neural network system. Additionally, it may be employed in neural networks that are not necessarily distributed to store feature maps.

[0081] Next, an overview of some of the terms used is provided.

[0082] Artificial Neural Network An artificial neural network (ANN) or connectionist system is a computational system that is vaguely inspired by the biological neural networks that make up animal brains. Such systems generally "learn" to perform tasks by considering examples, without being programmed with task-specific rules. For example, in image recognition, it may be possible to learn to identify images that contain cats by analyzing exemplary images that are manually labeled as "cat" or "no cat" and using the results to identify cats in other images. This is done without prior knowledge about cats, such as, for example, fur, tail, whiskers, and a cat-like face. Instead, discriminative features are automatically generated from the examples being processed.

[0083] ANN is based on a collection of interconnected units or nodes called artificial neurons, which roughly model the neurons of the biological brain. Each connection, like a synapse in the biological brain, can transmit signals to other neurons. Next, the artificial neuron that receives the signal can process it and send the signal to the connected neurons.

[0084] In the implementation form of ANN, the "signal" at the connection is a real number, and the output of each neuron is calculated by some non-linear function of the sum of its inputs. The connection is called an edge. Neurons and edges typically have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal at the connection. A neuron may have a threshold such that a signal is transmitted only when the aggregated signal exceeds its threshold. Typically, neurons are aggregated into layers. Different layers can perform different transformations on their inputs. Signals propagate from the first layer (input layer) to the last layer (output layer), sometimes traversing multiple layers in between.

[0085] The original goal of the ANN approach was to solve problems in the same way as the human brain does. As time went by, the interest shifted to performing specific tasks and deviated from biology. ANN is used in various tasks in activities that were traditionally thought to be possible only for humans, including computer vision, speech recognition, machine translation, social network filtering, playing board games and video games, medical diagnosis, and even drawing pictures.

[0086] Convolutional Neural Network The name "Convolutional Neural Network" (CNN) indicates that the network employs a mathematical operation called convolution. Convolution is a specialized type of linear operation. A convolutional network is simply a neural network that uses convolution instead of general matrix multiplication in at least one of its layers.

[0087] FIG. 1 illustrates an overview of the general concept of processing by a neural network such as a CNN. A convolutional neural network consists of an input and output layer, as well as a plurality of hidden layers. The input layer is the layer where the input (such as a part of the image shown in FIG. 1) is provided for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that are convolved with multiplication or other dot products. The result of the layer is one or more feature maps (f.maps in FIG. 1), sometimes also referred to as channels. Some or all of the layers may be accompanied by subsampling. As a result, the feature maps can become smaller, as illustrated in FIG. 1. The activation function in a CNN is usually a RELU (Rectified Linear Unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, and normalization layers, which are referred to as hidden layers because their inputs and outputs are masked by the activation function and the final convolution. These layers are referred to as convolutions in colloquial terms, but this is just a convention. Mathematically speaking, this is technically a sliding dot product or cross-correlation. This is significant in that it affects how the weights are determined at a particular index point with respect to the indices within the matrix.

[0088] When programming a CNN for image processing, as shown in FIG. 1, the input is a tensor having a shape (number of images)×(image width)×(image height)×(image depth). Then, after passing through the convolutional layer, the image is abstracted into a feature map having a shape (number of images)×(feature map width)×(feature map height)×(feature map channels). The convolutional layer within the neural network should have, as attributes, a convolutional kernel (hyperparameter) defined by the width and height, the number of input and output channels (hyperparameters). The depth of the convolutional filter (input channels) should be equal to the number of channels (depth) of the input feature map.

[0089] Previously, conventional multi-layer perceptron (MLP) models have been used for image recognition. However, due to the full connectivity between nodes, these have the problem of high dimensionality and could not handle high-resolution images well. A 1000×1000 pixel image with RGB color channels has 3 million weights, which is too high to be processed efficiently enough for full connectivity. Also, such network architectures do not consider the spatial structure of the data and treat distant input pixels the same as nearby pixels. This ignores the locality of references within the image data, both computationally and semantically. Therefore, the full connectivity of neurons is wasted for purposes such as image recognition, which is dominated by spatially local input patterns.

[0090] Convolutional neural networks are biologically inspired variants of multi-layer perceptrons specifically designed to emulate the behavior of the visual cortex. These models reduce the problems imposed by the MLP architecture by exploiting the strong spatial local correlations present in natural images. Convolutional layers are the core building blocks of CNNs. The parameters of a layer consist of a set of learnable filters (the kernels described above), which have small receptive fields but span the full depth of the input volume. In the forward pass, each filter is convolved over the width and height of the input volume, computing the dot product between the entries of the filter and the input, and generating a 2D activation map for that filter. As a result, the network learns filters that activate when detecting some particular kind of feature at a certain spatial position within the input.

[0091] By stacking the activation maps for all filters along the depth dimension, the full output volume of the convolutional layer is formed. Thus, each entry in the output volume can also be interpreted as looking at a small region in the input and being the output of neurons that share neurons and parameters within the same activation map. The feature map, or activation map, is the output activation for a given filter. Feature maps and activations have the same meaning. In some papers, this is called the activation map because it is a mapping corresponding to the activation of different parts of the image, and it is also called the feature map because it is a mapping of where certain features are found within the image. High activation means that a certain feature has been found.

[0092] Another important concept in CNNs is pooling, which is a form of non-linear downsampling. There are several non-linear functions for implementing the most common form of pooling, which is max pooling. This partitions the input image into a set of non-overlapping rectangles and outputs the maximum value for each such sub-region.

[0093] Intuitively, the exact placement of features is less important compared to the general placement with respect to other features. This is the thinking behind the use of pooling in convolutional neural networks. The pooling layer gradually reduces the spatial size of the representation, reducing the number of parameters, memory footprint, and amount of computation within the network, and thus also playing a role in controlling overfitting. In CNN architectures, it is common to periodically insert a pooling layer between consecutive convolutional layers. The pooling operation provides another form of translational invariance.

[0094] The pooling layer operates independently on all depth slices of the input and spatially resizes. The most common form is a pooling layer where a 2x2 filter is applied with a stride of 2 downsampling for each depth slice within the input along both width and height, discarding 75% of the activations. In this case, all max operations take four numbers. The depth dimension remains unchanged.

[0095] In addition to max pooling, the pooling unit can use other functions such as average pooling and l2-norm pooling. Average pooling has been historically used often but has recently become less popular compared to max pooling which actually has better performance. Due to the aggressive reduction of the size of the representation, there is a tendency recently to use smaller filters or to discard the pooling layer altogether. "Region of interest" pooling (also called ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important component of convolutional neural networks for object detection based on the Fast R-CNN architecture.

[0096] The aforementioned ReLU is short for rectified linear unit, applying a non-saturating activation function. This effectively removes negative values from the activation map by setting them to 0. This enhances the decision function and the overall non-linear characteristics of the network without affecting the receptive field of the convolutional layer. Other functions such as the saturated hyperbolic tangent function and the sigmoid function are also used to enhance non-linearity. ReLU is often preferred over other functions as it can train neural networks several times faster without a significant penalty to generalization accuracy.

[0097] After several convolutional layers and max-pooling layers, high-level inferences of the neural network are made through fully-connected layers. Neurons in a fully-connected layer have connections to all activations in the previous layer, as seen in a normal (non-convolutional) artificial neural network. Thus, its activations can be well considered as being computed as an affine transformation, performing a bias offset (vector addition of learned or fixed bias terms) after matrix multiplication.

[0098] The "loss layer" specifies how to penalize the deviation between the predicted label (output) of the training and the true label, and is usually the final layer of the neural network. Various loss functions suitable for different tasks can be used. The softmax loss is used to predict a single class out of K mutually exclusive classes. The sigmoid cross-entropy loss is used to predict K independent probability values within [0, 1]. The Euclidean loss is used for regression to real-valued labels.

[0099] In summary, Figure 1 shows the data flow of a typical convolutional neural network. First, the input image is passed through a convolutional layer and is abstracted into a feature map containing multiple channels corresponding to the number of filters within the set of learnable filters of this layer (e.g., one channel for each filter). Next, the feature map is subsampled, for example, using a pooling layer, thereby reducing the dimensions of each channel within the feature map. Since the next data comes to another convolutional layer that may have a different number of output channels, the number of channels within the feature map may also vary. As described above, the number of input channels and output channels are hyperparameters of the layer. To establish the connectivity of the network, those parameters need to be synchronized between two connected layers. For example, the number of input channels for the current layer should be equal to the number of output channels of the previous layer. For the first layer that processes input data, such as an image, the number of input channels is usually equal to the number of channels of the data representation. For example, it is 3 channels for the RGB or YUV representation of an image or video, or 1 channel for a grayscale image or video representation.

[0100] Autoencoders and Unsupervised Learning An autoencoder is a type of artificial neural network used to learn efficient data encoding in unsupervised learning. Its schematic diagram is shown in Figure 2. The purpose of an autoencoder is to learn a representation (encoding) for a set of data, typically for dimensionality reduction, by training the network to ignore the signal "noise". Along with the reduction side, a reconstruction side is learned, and the autoencoder attempts to generate a representation as close as possible to the original input, thus its name, from the reduced encoding. In the simplest case, when given one hidden layer, the encoder stage of the autoencoder receives the input x and maps it to h, that is h = σ(Wx + b) where. This image h is typically referred to as a code, latent variable, or latent representation. Here, σ is an element-wise activation function such as the sigmoid function or the rectified linear unit. W is a weight matrix, and b is a bias vector. The weights and biases are usually initialized randomly and then updated iteratively when training through backpropagation. Subsequently, the decoder stage of the autoencoder maps h to a reconstruction x' of the same shape as x, i.e., x' = σ'(W'h' + b') where. Here, σ', W', and b' for the decoder may be independent of the corresponding σ, W, and b for the encoder.

[0101] Variational autoencoder (VAE) models make strong assumptions about the distribution of latent variables. These use a variational approach to latent representation learning, resulting in additional loss components and a specific estimator for a training algorithm called the stochastic gradient variational Bayes (SGVB) estimator. The data is generated by a directed graphical model p θ (x|h), and the encoder is assumed to be learning an approximation q θ (h|x) of the posterior distribution p φ (h|x). Let φ and θ represent the parameters of the encoder (recognition model) and decoder (generative model), respectively. The probability distribution of the latent vectors of a VAE typically matches the probability distribution of the training data much more closely than that of a standard autoencoder. The objective variable of a VAE is

[0102] [Number]

[0103] takes the form of. Here, D KL represents the Kullback-Leibler Divergence. The prior distribution over the latent variables is usually a centered isotropic multivariate Gaussian distribution p θ(h) is set to N(0, I). Generally, the shapes of the variational distribution and the likelihood distribution are selected to be factorized Gaussians as follows. q φ (h|x) = N(ρ(x), ω 2 (x)I) p φ (x|h) = N(μ(h), σ 2 (h)I) Here, ρ(x) and ω 2 (x) are the encoder outputs, and μ(h) and σ 2 (h) are the decoder outputs.

[0104] Recent advancements in the field of artificial neural networks, particularly convolutional neural networks, have enabled researchers to focus on applying neural network-based techniques to the tasks of image and video compression. For example, end-to-end optimized image compression has been proposed, which uses a network based on variational autoencoders. Thus, data compression is considered a fundamental well-studied problem in engineering and is generally formulated in terms of the goal of designing a code for a given discrete data ensemble with minimum entropy. This solution heavily depends on knowledge of the probabilistic structure of the data, and thus the problem is closely related to probabilistic source modeling. However, since all practical codes must have a finite entropy, continuous-valued data (such as a vector of image pixel intensities) must be quantized into a finite set of discrete values, introducing an error. In this context, it is known as the non-reversible compression problem, and a trade-off relationship between two competing costs, the entropy (rate) of the discretized representation and the error (distortion) resulting from quantization, must be considered. Different compression applications, such as data storage or transmission over a limited-capacity channel, require consideration of different rate-distortion trade-off relationships. Simultaneous optimization of rate and distortion is difficult. Without additional constraints, the general problem of optimal quantization in high-dimensional spaces is intractable. For these reasons, most existing image compression methods operate by linearly transforming the data vector into a suitable continuous-valued representation, independently quantizing its elements, and then encoding the resulting discrete representation using a reversible entropy code. This approach is called transform coding due to the central role of the transformation. For example, JPEG uses the discrete cosine transform for blocks of pixels, and JPEG2000 uses multi-scale orthogonal wavelet decomposition. Typically, the three components of a transform coding method - the transform, the quantizer, and the entropy code - are optimized separately (often through manual parameter tuning). The latest video compression standards such as HEVC, VVC, and EVC also use the transformed representation to encode the residual signal after prediction.Several transforms, such as the discrete cosine transform and discrete sine transform (DCT, DST), and even the low-frequency non-separable manual optimization transform (LFNST), are used for that purpose.

[0105] Variational image compression In J. Balle, L. Valero Laparra, and E. P. Simoncelli (2015). "Density Modeling of Images Using a Generalized Normalization Transformation". In: arXiv e-prints. Presented at the 4th Int. Conf. for Learning Representations, 2016 (hereinafter referred to as "Balle"), the authors proposed a framework for the end-to-end optimization of an image compression model based on a non-linear transform. Previously, the authors demonstrated that a model consisting of a linear-nonlinear block transform optimized for a measure of perceptual distortion exhibits visually superior performance compared to a model optimized for the mean squared error (MSE). Here, the authors optimize for the MSE but use a more flexible transform constructed from a cascade of linear convolution and non-linearity. In particular, the authors use a generalized divisive normalization (GDN) coupled non-linearity inspired by the model of neurons in the biological visual system and proven to be effective for gaussianization of image density. Following this cascade transform, uniform scalar quantization (i.e., each element is rounded to the nearest integer) follows, which effectively implements a parametric form of vector quantization in the original image space. The compressed image is reconstructed from these quantized values using an approximate parametric non-linear inverse transform.

[0106] For any desired point along the rate-distortion curve, the parameters of both the analysis transform and the synthesis transform are simultaneously optimized using stochastic gradient descent. To achieve this in the presence of quantization (which gives rise to gradients that are mostly zero everywhere), the authors use a surrogate loss function based on a continuous relaxation of the probabilistic model and replace the quantization step with additive uniform noise. The relaxed rate-distortion optimization problem is somewhat similar to the problem used to fit generative image models, particularly variational autoencoders, but differs in the constraints imposed by the authors to ensure that discrete problems are approximated along the entire rate-distortion curve. Finally, rather than reporting differential or discrete entropy estimates, the authors implement an entropy coder and report performance using the actual bitrate, thereby demonstrating the feasibility of the solution as a complete lossless compression method.

[0107] J. Balle's paper describes an end-to-end trainable model for image compression based on variational autoencoders. This model incorporates a hyperprior distribution to effectively capture the spatial dependencies in the latent representation. This hyperprior distribution is related to the side information that is also transmitted on the decoder side and is a concept that is practically universal in all modern image codecs but has not been extensively investigated in image compression using ANNs. Unlike existing autoencoder compression methods, this model is trained together with an autoencoder based on a complex prior distribution. The authors demonstrate that this model provides state-of-the-art image compression when measuring visual quality using the popular MS-SSIM index and achieves rate-distortion performance that exceeds that of published ANN-based methods when evaluated using more conventional metrics based on peak signal-to-noise ratio (PSNR).

[0108] Figure 3 shows the network architecture including the hyperprior distribution model. On the left (g a , g s ), it shows the image autoencoder architecture, and on the right (h a , h s) corresponds to an autoencoder that implements a hyperprior distribution. The factored prior distribution model uses the same architecture for the analysis transform g a and the synthesis transform g s . Q represents quantization, and AE and AD represent an arithmetic encoder and an arithmetic decoder, respectively. When the encoder passes x, the input image, through g a , it produces a response y (latent representation) with a spatially varying standard deviation. The encoder g a includes multiple convolutional layers with subsampling and, as an activation function, generalized divisive normalization (GDN).

[0109] The response is fed into h a which summarizes the distribution of the standard deviation in z. Then z is quantized, compressed, and transmitted as side information. Next, the encoder uses the quantized vector

[0110]

Number

[0111] to obtain the spatial distribution of the standard deviation used to acquire probability values (or frequency values) for arithmetic coding (AE), and uses it to compress and transmit the quantized image representation

[0112]

Number

[0113] (or latent representation). The decoder first extracts from the compressed signal

[0114]

Number

[0115] (or latent representation) for compression and transmission. The decoder first extracts from the compressed signal

[0116]

Number

[0117] Restore it. Then, h s is used to

[0118]

Number

[0119] obtain, but this is on top of that

[0120]

Number

[0121] also gives the correct probability estimate value for normal restoration. Then, this is

[0122]

Number

[0123] is supplied to g s to obtain the reconstructed image.

[0124] In further work, probability modeling with a hyper-prior distribution was further improved by introducing, for example, an autoregressive model based on the PixelCNN++ architecture, which, as illustrated in Figure 2 of, for example, L. Zhou, Zh. Sun, X. Wu, J. Wu, "End-to-end Optimized Image Compression with Attention Mechanism", CVPR 2019 (hereinafter referred to as "Zhou"), enables the use of the context of the already decoded symbols in the latent space for better probability estimation of the further symbols to be decoded.

[0125] Cloud Solutions for Machine Tasks Video Coding for Machines (VCM) is a recently popular direction in other computer science fields. The main idea behind this approach is to transmit an encoded representation of image or video information that is to be further processed by computer vision (CV) algorithms such as object segmentation, detection, and recognition. In contrast to the conventional encoding of images and videos for human perception, the quality characteristic is not the reconstructed quality but rather the performance of computer vision tasks, such as object detection accuracy. This is illustrated in Figure 4.

[0126] Video Coding for Machines, also referred to as Collaborative Intelligence, is a relatively new paradigm for efficiently deploying deep neural networks in a mobile-cloud infrastructure. By splitting the network between the mobile and the cloud, it is possible to distribute the computational workload so that the overall energy and / or latency of the system is minimized. In general, Collaborative Intelligence is a paradigm in which the processing of neural networks is distributed among two or more different computing nodes, e.g., devices, but generally any functionally defined nodes. Here, the term "node" does not refer to the nodes of the neural network described above. Rather, a (computing) node here refers to a separate device / module (physically or at least logically) that implements a part of the neural network. Such devices may be different servers, different end-user devices, a mixture of servers and / or user devices and / or the cloud and / or processors, or the like. In other words, computing nodes can be considered as nodes that belong to the same neural network and communicate with each other to transfer encoded data within / for the neural network. For example, one or more layers can be executed on a first device and one or more layers can be executed on another device to enable complex computations. However, the distribution can also be finer, with a single layer being executed on multiple devices. In the present disclosure, the term "plurality" refers to two or more. In some existing solutions, a part of the neural network functionality is executed in a device (user device or edge device or the like) or multiple such devices, and then the output (feature map) is passed to the cloud. The cloud is an aggregate of processing or computing systems located outside the device that operates a part of the neural network. The concept of Collaborative Intelligence has also been extended to the training of models.In this case, data flows both from the cloud to the mobile during backpropagation in training and from the mobile to the cloud during the forward pass in training, and the same is true for inference.

[0127] Some research has presented semantic image compression by encoding deep features and then reconstructing the input image from them. Compression based on uniform quantization has been shown, followed by context-based adaptive arithmetic coding (CABAC) from H.264. In some scenarios, it may be more efficient to transmit the output of the hidden layer (deep feature map) from the mobile part to the cloud rather than sending the compressed natural image data to the cloud and performing object detection using the reconstructed image. Efficient compression of the feature map is beneficial for both human perception and machine vision in terms of image and video compression and reconstruction. Entropy coding methods, such as arithmetic coding, are a popular approach for compressing deep features (i.e., feature maps).

[0128] Currently, video content contributes more than 80% of Internet traffic, and that proportion is expected to increase even further. Therefore, it is important to build an efficient video compression system to generate higher-quality frames with a given bandwidth budget. In addition, most video-related computer vision tasks, such as video object detection or video object tracking, are sensitive to the quality of the compressed video, and efficient video compression can also bring benefits to other computer vision tasks. On the other hand, techniques in video compression are also useful for action recognition and model compression. However, over the past few decades, video compression algorithms have relied on handcrafted modules, such as block-based motion estimation and discrete cosine transform (DCT), to reduce redundancy in the video sequence as described above. Each module is well-designed, but the overall compression system is not end-to-end optimized. It is desirable to further improve video compression performance by simultaneously optimizing the entire compression system.

[0129] End-to-End Image or Video Compression In recent years, deep neural network (DNN)-based autoencoders for image compression have achieved performance equal to or better than that of conventional image codecs such as JPEG, JPEG2000, or BPG. One possible explanation is that DNN-based image compression methods can utilize large-scale end-to-end training and highly non-linear transformations that are not used in conventional approaches. However, it is not obvious to directly apply these techniques to build an end-to-end learning system for video compression. First, how to generate and compress motion information tailored for video compression remains an unsolved problem. Video compression methods rely heavily on motion information to reduce the temporal redundancy of video sequences. A direct solution is to use a learning-based optical flow to represent motion information. However, current learning-based optical flow approaches aim to generate the most accurate flow field possible. An accurate optical flow is often not optimal for a specific video task. In addition, the data volume of optical flow increases significantly when compared with motion information in conventional compression systems. Directly applying existing compression approaches to compress optical flow values significantly increases the number of bits required to store motion information. Second, the method of building a DNN-based video compression system by minimizing rate-distortion-based objectives for both residual information and motion information is not clear. Rate-distortion optimization (RDO) aims to achieve a higher-quality reconstructed frame (i.e., less distortion) when a given number of bits (or bitrate) for compression is provided. RDO is important for video compression performance. To utilize the power of end-to-end training of learning-based compression systems, an RDO strategy for optimizing the entire system is needed.

[0130] Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, Zhiyong Gao, "DVC: An End-to-end Deep Video Compression Framework", Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11006 - 11015, where the authors proposed an end-to-end deep video compression (DVC) model that jointly learns motion estimation, motion compression, and residual coding.

[0131] Such an encoder is illustrated in Figure 5. In particular, Figure 5 shows the overall structure of an end-to-end trainable video compression framework. To compress motion information, a CNN is designated to convert the optical flow into a corresponding representation better suited for compression. Specifically, an autoencoder-style network is used to compress the optical flow. This motion vector (MV) compression network is shown in Figure 6. The network architecture is somewhat similar to ga / gs in Figure 3. In particular, the optical flow is fed into a series of convolutional operations as well as non-linear transformations including GDN and IGDN. The number of output channels for the convolution (deconvolution) is equal to 2, except for the last deconvolution layer which is 128. Given an optical flow of size M×N×2, the MV encoder generates a motion representation of size M / 16×N / 16×128. The motion representation is quantized, entropy-coded, and sent to the bitstream. The MV decoder receives the quantized representation and reconstructs the motion information using the MV encoder.

[0132] Figure 7 shows the structure of the motion compensation unit. Here, the previous reconstructed frame x t-1By using the reconstructed motion information, the warping unit generates the warped frame (usually with the help of an interpolation filter such as a bilinear interpolation filter). Then, a separate CNN with three inputs generates the predicted picture. The architecture of the motion compensation CNN is also shown in FIG. 7.

[0133] The residual information between the original frame and the predicted frame is encoded by the residual encoder network. A highly non-linear neural network is used to convert the residuals into corresponding latent representations. Compared with the discrete cosine transform in conventional video compression systems, this approach can utilize the power of non-linear transformation more appropriately and achieve higher compression efficiency.

[0134] From the above overview, it can be seen that the CNN-based architecture can be applied to both image compression and video compression by considering different parts of the video framework including motion estimation, motion compensation, and residual coding. Entropy coding is a popular method used for data compression, widely adopted in the industry, and applicable to feature map compression for either human perception or computer vision tasks.

[0135] Improvement of Encoding Efficiency Channel information is not equally important for the final task. It has been observed that information in certain channels can be dropped without significantly degrading the final image or video reconstruction quality, or object detection accuracy, for example, not being transmitted to the decoder. At the same time, the amount of bits saved by dropping unimportant information can improve the overall rate-distortion trade-off relationship.

[0136] Furthermore, encoding and decoding latency is one of the important parameters of a compression system for a practical implementation of the compression system. Due to the nature of artificial neural networks, operations within one layer can be performed in parallel, and the overall network latency is usually not very high, determined by the amount of subsequent connected layers. By leveraging modern graphics processing units (GPUs) or network processing units (NPUs) that can support super parallel processing, an acceptable execution time can be achieved. However, entropy encoding methods such as arithmetic encoding and range encoding imply sequential operations for range interval calculation, normalization, and probability interval matching. There is little possibility that these operations can be parallelized without sacrificing compression efficiency. Therefore, entropy encoding can become a bottleneck that limits the overall system latency. Reduction of the amount of data passed through entropy encoding is desirable.

[0137] Some embodiments of the present disclosure may improve compression efficiency. Furthermore, they may speed up bypassing the transmission of some feature map data from the latent space of a CNN-based image and video codec. Based on such an approach, speeding up the entropy encoding and decoding process can be achieved, which may be considered important for practical implementations.

[0138] In particular, the transfer of information of some feature map regions or channels may be skipped. In particular, that feature map data may be skipped, and it is determined that its absence does not lead to a significant deterioration of the reconstructed quality. Therefore, the compression efficiency can be improved. Still further, the amount of data passed through entropy encoding and decoding is reduced, thereby further shortening the encoding and decoding time and reducing the overall end-to-end compression system latency.

[0139] According to one embodiment, a method for decoding a feature map for processing by a neural network based on a bitstream is provided. This method is illustrated in FIG. 8. This method includes step S110 of obtaining a region presence indicator based on information from the bitstream for a region of the feature map. After the obtaining step S110, a decoding step S150 of decoding the region follows. The decoding step S150 includes step S130 of analyzing data from the bitstream to decode the region when the region presence indicator has a first value. Whether the region presence indicator has a first value can be determined in a determination step S120. On the other hand, when the region presence indicator has a second value (for example, in step S120), the decoding S150 may include step S140 of bypassing the analysis of data from the bitstream to decode the region.

[0140] The obtaining S110 of the region presence indicator may correspond to analyzing the presence indicator from the bitstream. In other words, the bitstream may include the indicator. The obtaining may include, for example, decoding the presence indicator by an entropy decoder. The region presence indicator may be a flag that can take one of a first value and a second value. This may be encoded by, for example, a single bit. However, the present disclosure is not limited thereto. Generally, the obtaining step S110 may correspond to an indirect obtaining based on the value of other parameters analyzed from the bitstream.

[0141] The decoding S150 may refer to a generalized processing of the bitstream, which may include one or more of information analysis, entropy decoding information, bypass (skip) of information analysis, derivation of information used for further analysis based on one or more of the already analyzed elements of the information, or similar operations.

[0142] According to one embodiment, a method is provided for encoding a feature map for processing by a neural network into a bitstream. This method is illustrated in FIG. 9 and can provide a bitstream portion that can be easily decoded by the decoding method described with reference to FIG. 8. The method includes step S160 of obtaining a region presence indicator for a region of the feature map. This region of the feature map may be obtained from a feature map generated by one or more layers of the neural network. The feature map may be stored in a memory or other storage device. Further, the method includes step S170 of determining whether to indicate or not indicate the region of the feature map in the bitstream based on the obtained region presence indicator. Thus, in step S170, if it is determined to indicate the region in the bitstream, the bitstream shall include the region of the feature map. In step S170, if it is determined not to indicate the region in the bitstream, the bitstream shall not include the region of the feature map. Steps S170 to S190 can be regarded as part of a general step S165 of generating (also referred to as encoding) the bitstream.

[0143] In one possible implementation, the region presence indicator is indicated within the bitstream. Thus, in step S170, if it is determined to indicate the region in the bitstream, the bitstream shall include the region presence indicator having a first value and the region of the feature map. In step S170, if it is determined not to indicate the region in the bitstream, the bitstream shall include the region presence indicator having a second value without the region of the feature map.

[0144] The phrase "including a region presence indicator having a certain value" substantially means including a certain value in a bit stream, for example, in a binary form or, in some cases, in an entropy-coded form. The certain value gives the meaning of the region presence indicator using semantics defined by a convention, for example, by a standard.

[0145] Handling of non-existent feature map regions According to an exemplary implementation, when the region presence indicator has a second value, decoding the region further includes setting the region according to a predetermined rule. For example, the predetermined rule specifies setting the features of the region to constants. However, it should be noted that the present disclosure is not limited to the predetermined rule being a rule that specifies that the feature amount should be set to a constant. Rather, the rule may define a method for determining / calculating the feature amount based on, for example, previously decoded features or information from the bit stream.

[0146] FIG. 10 is a diagram illustrating a method for decoding a bitstream. In this example, the regions of the feature maps correspond to the CNN channels. In particular, the method for decoding image or video information includes parsing syntax elements from the bitstream that indicate whether corresponding CNN channel information is present in the bitstream. The CNN channel information may include the feature maps of the corresponding channels or other information related to a particular channel. In the exemplary implementation shown in FIG. 10, the decoder iterates S205 over the input channels of the generation model (or reconstruction network). In each iteration S205, a channel presence flag (corresponding to the region presence flag described above) is read S210 from the bitstream. If the channel presence flag is equal to true, the corresponding channel information is read S235 from the bitstream. This data may be further decoded S240 by an entropy decoder, such as an arithmetic decoder. Otherwise, if the channel presence flag is equal to false, reading the channel information is bypassed and the corresponding channel information is initialized S230 according to a predefined rule, for example, by setting all channel values to a constant. For example, the constant is 0. Then, steps S210 - S240 are repeated for each channel. Thereafter, in step S244, the reconstructed channels are input S244 into an appropriate layer of a neural network (which may be, for example, one of the hidden or output layers of the neural network). The neural network then further processes the input channel information.

[0147] The following table provides an exemplary implementation of the bitstream syntax.

[0148] [Table 1]

[0149] The variable channels_num defines the number of channels to be iterated (see step S205 in FIG. 10). The variables latent_space_height and latent_space_width can be derived based on high-level syntactic information regarding the width and height of the picture and the architecture of the generative model (e.g., the reconstruction network). Other implementations leading to the same logic are possible. For example, the conditional check of channel_presence_flag can be performed outside the loop that iterates over the width and height of the latent space. In that case, the input channel value y_cap[i] can be initialized in another loop other than the channel data reading loop.

[0150] Alternatively, the channel presence indicator may have the opposite interpretation, e.g., a channel_skip_flag (or channel bypass flag), indicating that the reading of the corresponding channel information should be skipped. In other words, the channel (or generally region) indication may be a presence indication or an absence indication. Such an indication may be taken as one of two values, i.e., a first value and a second value. One of these values indicates the presence of the channel (feature map) data, and the other of these values indicates the absence of the channel (feature map) data. It should be further noted that in FIG. 10, a channel indicator indicating the presence / absence of all channels is signaled. However, the present disclosure is not limited thereto, and an indicator may be signaled for a part of the feature map (region) or for a group of channels organized into groups in some predefined way.

[0151] One of the technical advantages of this embodiment may be the reduction of signaling overhead by excluding the transmission of information that is unnecessary or not important or of low importance for image or video reconstruction. Another technical advantage is to speed up the entropy decoding process, which is known as a bottleneck of image and video compression systems, by excluding the processing of unnecessary or not important information by entropy decoding.

[0152] The decode_latent_value() subprocess (the parsing process corresponding to the syntax part) may include entropy decoding. The channel_presence_flag may be encoded as an unsigned integer 1-bit flag by using context-adaptive entropy coding (denoted as ae(v)) or without context-adaptive entropy coding (denoted as u(1)). Using context-adaptive entropy coding makes it possible to reduce the signaling overhead introduced by the channel_presence_flag and further improve the compression efficiency.

[0153] In the above syntax, when the channel_presence_flag indicates that the channel data is not signaled, the input channel value is set to a constant, which is 0 here (y_cap[i][y][x]=0). However, the present disclosure is not limited to the constant value of 0. The constant may take different values or may even be preset by the encoder and then signaled in the bitstream. In other words, in some embodiments, the method further includes the step of decoding the constant from the bitstream.

[0154] Corresponding to the decoding method being described with reference to FIG. 10, as shown in FIG. 11, an encoding method can be provided. This encoding method can generate a bitstream according to the syntax described above. Correspondingly, an image or video encoder can include a unit for determining whether to transmit or bypass the corresponding channel information to the receiving side. For each output channel of the encoding model (assumed to be the input channel of the generation model), the encoder makes a decision regarding the importance of the corresponding channel for image or video reconstruction or machine vision tasks.

[0155] Such an encoding method in FIG. 11 includes a loop over all channels. In each step S208, a channel is taken, and in step S250, a decision regarding the presence of the channel is made. The encoder can make a decision based on some prior knowledge or metric that enables the evaluation of the importance of the channel.

[0156] For example, the sum of the absolute values of the feature maps of the corresponding channels in the latent representation

[0157]

Number

[0158] can be used as a metric for determination on the encoder side, where ChannelPresenceFlag i is a flag representing the encoder's decision regarding the presence of the i-th CNN channel in the bitstream, and w, h correspond to the i-th channel accordingly

[0159]

Number

[0160] are the width and height of the potential representation, and the threshold is some predefined value, e.g., 0. In other implementations, the sum of absolute values can be

[0161]

Number

[0162] normalized by the number of elements in the channel as follows.

[0163] In other possible implementations, instead of the sum of absolute values, the sum of squared values can be used. Also, another possible criterion is the variance of the corresponding channel's feature map in the potential representation defined as the value obtained by dividing the sum of the squared distances of each feature map element in the channel from the average value of the feature map elements in the channel by the number of feature map elements in the channel.

[0164]

Number

[0165] Here, μ is the average value of the feature map elements in the channel, and threshold_var is the threshold. The present disclosure is not limited to any specific decision algorithm or metric. Further alternative implementations are also possible.

[0166] After decision step S250, the determined channel presence flag is written into the bitstream at S255. Next, based on the value of the channel presence flag, at step S260, it is determined whether to write the channel data (the region of the feature map data) into the bitstream or not. In particular, if the channel presence flag is true, at step 270, the channel data is written into the bitstream, which may further include encoding of the channel data by an entropy encoder such as an arithmetic encoder at S280. On the other hand, if the channel presence flag is false, at step 270, the channel data is not written into the bitstream, that is, the writing is bypassed (or skipped).

[0167] FIG. 12 shows an exemplary implementation of an encoder-side method for making a decision regarding the presence of a specific channel in a bitstream based on the relative importance of the channel with respect to the quality of the reconstructed picture. The steps of the encoding method are performed for each channel. In FIG. 12, the loop over those channels is illustrated by step S410 that sets the channel for which the next step is to be performed.

[0168] As a first step S420, for each channel, a significance metric is determined. The metric can be determined, for example, by calculating the sum of the absolute values of all feature map elements within the channel or as the variance of the feature map elements within the channel. The value distribution of the feature map can be an indication of the impact the feature map has on the reconstructed data. Another possible metric for sorting can be an estimate of the number of bits required to transmit the channel information corresponding to the rate (in other words, the number of bits required to transmit the feature map data of a particular channel). Another possible metric for sorting can be, for example, in dB or the contribution of a particular channel to the reconstructed picture quality evaluated with another quality metric such as the multi-scale structural similarity index measure (MS-SSIM) or any other objectively or perceptually weighted / designed metric. The above and other metrics can be combined. There can also be other performance criteria suitable for machine vision tasks, such as object detection accuracy, which is evaluated, for example, by comparing the overall channel reconstruction performance with all but a particular channel reconstruction performance subtracted. In other words, some estimates of the contribution of channel data to the quality / accuracy of a machine vision task can be used as sorting criteria.

[0169] In the next step, all channels are sorted or ranked according to the calculated significance metric, for example, from the most significant (index = 0 or index = 1 depending on the start of the count index) to the least significant.

[0170] In the next steps (S430 to S470), by having as input parameters the contribution of each channel into the bitstream (the amount of bits necessary to transmit the channel feature map value) and the estimation of the desired target bitrate, the encoder makes a decision regarding putting specific channel information into the bitstream. Specifically, in step S430, the number of bits numOfBits of the resulting bitstream portion is initialized to 0. Further, the channel presence flag is initialized to 0 for each channel, indicating that the channel is not included in the bitstream. Then, in step S440, a loop over several channels is started. In the loop, the channels are scanned in order from the most significant to the least significant. For each channel within the loop, in step S450, the number of bits channelNumOfBits necessary to encode the data of the channel is obtained, and the total number of bits numOfBits is incremented by the number of bits channelNumOfBits necessary to encode the data of the obtained channel, i.e., numOfBits += channelNumOfBits (which means numOfBits = numOfBits + channelNumOfBits).

[0171] In step S460 of the loop, the total number of bits numOfBits is compared with the required number of bits requiredNumOfBits. requiredNumOfBits is a parameter that can be set according to the desired rate. If the result of the comparison numOfBits >= requiredNumOfBits is TRUE, this means that the total number of bits has reached or exceeded the required number of bits, and in that case, the method ends. This means that for the current channel i, the channel presence flag remains 0 as initialized in step S430, and the channel data is not included in the bitstream. In step S460, if the result of the comparison numOfBits >= requiredNumOfBits is FALSE, this means that the total number of bits has not reached the required number of bits when the data of channel i is added to the bitstream. Therefore, in step S470, the channel presence flag for the current channel i is set to 1, and the data of that channel is included in the bitstream.

[0172] In summary, the encoder calculates the incremental sum of the bits required for the transmission of the channels starting from the most significant channel, and sets the channel_presence_flag[i] to a value equal to TRUE. After the total number of bits has reached the required number of bits, the remaining least significant channels are determined not to be transmitted, and the channel_presence_flag[i] is set to be equal to FALSE for these channels.

[0173] Alternatively, a desired reconstructed picture quality level (e.g., in dB or other metric such as MS-SSIM) can be used as a criterion for determining the number of channels to transmit. Starting from the most significant channel, the encoder evaluates the reconstructed quality due to the contribution of the most significant channels and, after the desired quality is achieved, the remaining channels are determined to be unnecessary and their corresponding channel_presence_flag[i] are set to be equal to FALSE. In other words, iterations S440 to S470 can be performed by accumulating a quality counter that increases with each channel i and stopping the iteration when the addition of the current channel i reaches or exceeds the desired quality level. As will be apparent to those skilled in the art, the bitrate and quality are merely illustrative and both combinations can be used, or additional criteria such as complexity can be added or used instead.

[0174] The method described with reference to FIG. 12 is merely illustrative. Another alternative method for determining the channels to transmit is to use a rate-distortion optimization (RDO) procedure that minimizes a cost value calculated as follows. Cost = Distortion + Lambda * Rate, or Cost = Rate + Beta * Distortion where Lambda and Beta are the Lagrange multipliers of the constrained optimization method. One of the technical advantages of the solution described above is that it can match the desired target bitrate for the desired reconstructed quality, which is an important aspect of the practical use of the compression system.

[0175] In the embodiments described with reference to FIGS. 10 and 11, the region of the feature map is the entire feature map. However, as will be shown later, the present disclosure is not limited to such data granularity.

[0176] Side information signaling As described with reference to FIG. 3, the hyper-prior distribution proposed in Balle is used to generate side information that is transmitted along with the latent representation of the input signal, enabling the encoding system to capture the statistical characteristics specific to a given input signal and obtain the probability (or frequency) estimates required for the arithmetic decoder. As further demonstrated by Zhou, probability estimation can be further improved by incorporating context based on the already decoded symbols of the latent representation. From FIG. 3, it can be seen that the side information z is based on the output of the convolutional layer. In a general way, the feature map can be considered to represent side information itself.

[0177] According to an embodiment of the present disclosure, a method for decoding a feature map for processing by a neural network from a bitstream is provided. The method is illustrated in FIG. 13 and includes step S310 of obtaining, from the bitstream, a side information indicator indicating whether side information exists in the bitstream. The method further includes S350 of decoding the feature map. The decoding S350 of the feature map further includes step S320, where the value of the side information indicator is determined and acted upon. In particular, when the side information indicator has a fifth value, the method includes step S330 of parsing, from the bitstream, side information for decoding the feature map. Otherwise, when the side information indicator has a sixth value, the method further includes S340 of bypassing parsing, from the bitstream, side information for decoding the feature map. The fifth value indicates that the bitstream contains side information, while the sixth value indicates that the bitstream does not contain side information for a particular part (e.g., region) of the feature map, or for the entire feature map or the like.

[0178] The hyper-prior distribution model may include autoregressive context modeling based on the already decoded symbols, so the transmission of side information may not be necessary, and the hyper-prior distribution network can efficiently model the distribution based only on the context. At the same time, the efficiency of context modeling strongly depends on the statistical characteristics of the input image or video content. Some content may be highly predictable, while some may not. For flexibility, it is beneficial to have the option of transmitting and using side information that can be determined by the encoder based on the statistical characteristics of the input content. This further increases the possibility for the compression system to adapt to the input content characteristics.

[0179] Corresponding to the decoder process of FIG. 13, an embodiment of the bitstream may include a syntax element side_information_available that controls the presence of side information (z_cap) in the bitstream. If the side information is not available (for example, if determined by the encoder not to be necessary to obtain symbol probabilities for arithmetic coding), the decoder skips reading the side information value,

[0180]

Number

[0181] initialize it to some value, for example 0, and then

[0182]

Number

[0183] send it to the hyper-prior generation unit (g s ). This makes it possible to further optimize the signaling. An exemplary bitstream is shown below.

[0184]

Table 2

[0185] Corresponding to the decoding method, according to one embodiment, a method for encoding a feature map for processing by a neural network into a bitstream is provided. The method is illustrated in FIG. 14 and may include step S360 of obtaining a feature map (or at least one region of the feature map). The method further includes an encoding step S365 that may include step S370 of determining whether to indicate side information regarding the feature map in the bitstream. If the determination is affirmative, the method may further include S380 of inserting into the bitstream a third value (e.g., corresponding to the fifth value described above, and thus the encoder and decoder may understand each other or be part of the same system) and a side information indicator indicating the side information. If the determination is affirmative, the method may further include S390 of inserting into the bitstream a side information indicator indicating a fourth value (e.g., corresponding to the sixth value described above, and thus the encoder and decoder may understand each other or be part of the same system) without side information. The side information may correspond to a region of the feature map or to the entire feature map.

[0186] Note that the side information may correspond to the side information as described with reference to the autoencoder of FIG. 3. For example, the decoding method may further include entropy decoding, and the entropy decoding is based on the decoded feature map processed by the neural network. Correspondingly, the encoding method may further include entropy encoding based on the encoded feature map processed by the neural network. In particular, the feature map is the hyper-prior distribution h as shown in FIG. 3 a / h scan correspond to the feature map from. The side information based on the feature map can correspond to the distribution of the standard deviation summarized by z. z can be quantized, further compressed (e.g., entropy coding), and transmitted as side information. Then, the encoder (and further the decoder) uses the quantized vector

[0187] [Number]

[0188] to obtain the spatial distribution of the standard deviation actually used to obtain the probability values (or frequency values, occurrence counts) for arithmetic coding (or generally, other types of entropy coding).

[0189] [Number]

[0190] to estimate it and use it to encode the quantized image representation

[0191] [Number]

[0192] (or latent representation). The decoder first restores

[0193] [Number]

[0194] from the compressed signal and decodes the latent representation accordingly. As described above, the distribution modeling network can be further enhanced by context modeling based on the already decoded symbols of the latent representation. In some cases, the transmission of z may not be necessary, and the input to the generation part of the hyper-prior distribution network (h s ) can be initialized by values following rules, e.g., constants.

[0195] Note that the encoders and decoders in FIG. 3 are merely examples. In general, side information can convey other information (

[0196]

Number

[0197] different from that). For example, the probability model for entropy encoding can be signaled directly or derived from other parameters signaled in the side information. Entropy encoding does not have to be arithmetic encoding and can be other types of entropy or variable-length encoding, which can be, for example, context-adaptive and controllable by side information.

[0198]

[0199] In an exemplary implementation, when the side information indicator has a sixth value, the decoding method includes setting the side information to a predetermined side information value. For example, the predetermined side information value is 0. The decoding method may further include the step of decoding the predetermined side information value from the bitstream.Correspondingly, in an exemplary implementation of the encoding method, when the side information indicator has a fourth value, the encoding method may be applied to an encoding parameter derived based on a predetermined side information value such as 0. The encoding method may further include encoding the predetermined side information value into the bitstream. In other words, the bitstream may convey the predetermined side information once, for example, for a plurality of regions of the feature map, or for the entire feature map or a plurality of feature maps. In other words, a predetermined value (which may be a constant) may be signaled less frequently than the side information indicator. However, this is not necessarily the case, and in some embodiments, the predetermined value may be signaled individually, for each channel (a part of the feature map data), or for each channel region, or in a similar manner.

[0200] The present disclosure is not limited to signaling a predetermined value, and instead, it may be specified by a standard or derived from some other parameter included in the bitstream. Further, instead of signaling a predetermined value, information indicating the rules to follow when a predetermined information is determined may be signaled.

[0201] Therefore, the described method is also applicable for optimizing the transmission of the potential representation of the side information

[0202]

Number

[0203] as described above, the side information

[0204]

Number

[0205] (FIG. 3) is based on the output of a convolutional layer that includes a plurality of output channels. In a general approach, the feature map may be considered to represent side information by itself. Thus, the method described above and illustrated in FIG. 10 is also applicable to the optimization of side information signaling and can also be combined with a side information presence indicator.

[0206] The following is an exemplary syntax table illustrating how these two methods can be combined.

[0207] [Table 3]

[0208] A method for analyzing a bitstream may include obtaining a side information presence indicator (side_information_available) from the bitstream. The method may further include parsing side information (z_cap) from the bitstream when the side information presence indicator has a third value (e.g., side_information_available is TRUE) and bypassing completely parsing side information from the bitstream when the side information presence indicator has a fourth value (e.g., side_information_available is FALSE). Further, the side information indicated in the bitstream may include at least one of a region presence indicator (e.g., side_inf or mation_channel_presence_flag shown in the above syntax and described with reference to FIGS. 8 to 11) and information (z_cap) about being processed by a neural network to obtain an estimated probability model for entropy decoding of the region. Depending on the value of the region presence indicator (side_inf or mation_channel_presence_flag), the feature map data (z_cap) may be included in the bitstream.

[0209] In other words, in this example, a method similar to that described for the latent space y_cap is applied to the side information z_cap, which may also be a feature map including the same channel. z_cap is decoded by decode_latent_value. This may utilize the same decoding process as y_cap, but not necessarily so.

[0210] The corresponding encoding method and decoding method may include a combination of steps as described with reference to FIGS. 8 to 14.

[0211] Furthermore, the encoder may determine which constant value should be used to initialize the values in the latent representation of the hyper-prior distribution. In that case, the value (side_information_init_value) is transmitted in the bitstream as exemplified in the following syntax.

[0212] [Table 4]

[0213] In other words, in some embodiments, the side information presence indicator has a fourth value, and the method includes setting the side information in the decoder to a predetermined side information value (which can also be done in the encoder if the side information is used in encoding). In the above exemplary syntax, when side_information_channel_presence_flag is FALSE, the feature map value (z_cap) is set to 0 (or to another value preset or transmitted in the bitstream as the side_information_init_value syntax element, in which case it should be pre-parsed before assignment). However, the constant 0 is merely exemplary. As already explained in the above embodiments, the feature map data can be predefined by a standard, signaled in the bitstream (e.g., the side_information_init_value syntax element from the above example should be pre-parsed before assignment in that case), or set to predetermined feature map data that can be derived according to predefined or signaled rules.

[0214] In some embodiments, the region presence indicator is a flag that can take only one of two values formed by a first value and a second value. In some embodiments, the side information presence indicator is a flag that can take only one of two values formed by a third value and a fourth value. In some embodiments, the side information presence indicator is a flag that can take only one of two values formed by a fifth value and a sixth value. These embodiments may enable efficient signaling using only 1 bit.

[0215] By skipping and transferring less important deep features or feature maps, improved efficiency can be provided with respect to the encoding and decoding rates and complexities. The skip can be on a per-channel basis, or for each region of a channel, or in a similar manner. The present disclosure is not limited to any particular granularity of the skip.

[0216] It should be further understood that the method described above is applicable to any feature map transmitted and obtained from a bitstream, which is assumed to be input to a generative model (or reconstruction network). The generative model can be used, for example, for image reconstruction, motion information reconstruction, residual information reconstruction, obtaining probability values (or frequency values) for arithmetic coding (e.g., according to the hyper-prior distribution described above), object detection and recognition, or further applications.

[0217] Sorting order Another embodiment of the present disclosure is to reorder channels according to relative importance. In some exemplary implementations, it is possible to convey information regarding the order in the bitstream, simplify signaling, and enable the use of bitrate scalability features. That is, in this embodiment, channel importance sorting is applied, and the sorting order is known on both the encoder and decoder sides either by being transmitted in the bitstream or by referring to predetermined information.

[0218] In particular, a method for decoding feature map data (channel data) may include obtaining a significance order that indicates the importance of a plurality of channels (a part of the feature map data) of a specific layer. The term channel significance here refers to a measure of the importance of a channel with respect to the quality of a task performed by a neural network. For example, if the task is video encoding, the significance may be a metric that measures the reconstructed quality of the encoded video and / or the resulting rate. If the task is machine vision such as object recognition, the significance may be something that measures the contribution of the channel to recognition accuracy or the like. Obtaining the significance order can generally be done by any means. For example, the significance order may be explicitly signaled in the bitstream, or predefined in a standard or other convention, or may be implicitly derived based on other signaled parameters, for example.

[0219] The method may further include obtaining a last significant channel indicator and obtaining a region presence indicator based on the last significant channel indicator. The last significant channel indicator indicates the last significant channel among specific layers or network portions or channels of a network. The term "last" should be understood within the context of the significance order. According to some embodiments, such a last significant channel indication can also be interpreted as the least significant channel indication. The last significant channel is the channel among the channels ordered according to significance, after which there follow channels of low significance that are considered not important for the task of the neural network. For example, if channels 1 to M are ordered in descending order according to significance (from most significant to least significant), channel k is the last significant channel if channels k + 1 to M are considered not important (not significant) for the neural network task. Such channels k + 1 to M then need not be included in the bitstream, for example. In this case, the last significant channel indicator is used to indicate the absence of these channels k + 1 to M within the bitstream, whereby their encoding / decoding can be bypassed. The data related to channels 1 to k is then transmitted in the bitstream such that their encoding / decoding is performed.

[0220] In an exemplary implementation, the sort order is predefined at the design stage and is known at the decoder side (and further at the encoder side) without the need to be transmitted. The index of the last significant coefficient is inserted into the bitstream at the encoder. FIG. 15 illustrates an example of the method at the decoder. The index of the last significant coefficient is obtained from the bitstream at step S510 at the decoder. The last significant channel index can be signaled in the bitstream in a direct form, i.e., the syntax element last_significant_channel_idx can be included in the bitstream (and optionally entropy-coded). Alternative approaches and indirect signaling are described below.

[0221] In step S520, a loop is performed over all channels of the neural network layer or layers. The current channel is the channel at the current iteration i, i.e., the i-th channel. Then, for each i-th channel, the decoder determines channel_presence_flag[i] by comparing the index i with the last significant channel index in S530. If i is higher than the last significant index, channel_presence_flag[i] is set equal to FALSE in step S540. Otherwise, channel_presence_flag[i] is set equal to TRUE in step S550. It should be noted that in this embodiment, with respect to the indication of the last significant channel, the channel presence flag is not an actual flag included in the bitstream. Rather, the channel presence flag in this case is merely an internal variable derived from the signaled last significant channel index as shown above in steps S520 - S550.

[0222] A corresponding exemplary syntax table illustrating both the generation and parsing processes of the bitstream is presented below.

[0223]

Table 5

[0224] Next, in the channel information (data) analysis stage, at step S560, the decoder uses the derived channel_presence_flag[i] to determine whether to analyze the channel information from the bit stream at step S570 (when channel_presence_flag[i] = TRUE), or instead whether to skip (bypass) the analysis at step S580 (when channel_presence_flag[i] = FALSE). As described in the previous embodiment, bypassing may include initializing the channel information with a predefined value such as 0 at step S580. The flowchart of FIG. 15 showing the derivation of the channel presence flag is for illustrative purposes only and represents only an example. Generally, step S530 may directly determine whether to proceed to step S570 or step S580 without intermediate derivation of the channel presence flag. In other words, the channel presence flag is only implicitly derived and may be stored if it is beneficial for further use by other parts of the analysis process or semantic interpretation process depending on the implementation. This also applies to the above-described syntax.

[0225] In this example, decoding is performed by decoding the channels sorted within the bit stream in order of significance from the most significant channel to the last significant channel from the bit stream. The channel information (data) may be entropy encoded, in which case the above method may include step S590 of entropy decoding the channel data. After collecting all the channels, at step S595, the channels are supplied to the corresponding layer of the neural network.

[0226] However, the present disclosure is not limited to cases where a significance order is derived or known on both sides of the encoder and decoder. According to an exemplary implementation, obtaining the significance order includes decoding a significance order indication from the bitstream.

[0227] In particular, in an exemplary implementation, the decoder obtains the significance order from the bitstream by analyzing the corresponding syntax element. The significance order enables establishing a correspondence between each significance index and a channel design index. Here, under the channel design index, the index followed when the channel is indexed in the model design is understood. For example, the channel design order may be the order in which the channels of one layer are numbered according to a pre - defined convention independent of the content of the channels in the neural network processing.

[0228] An example of significance order information including significance indices and corresponding design indices is given in part A) of FIG. 18. As can be seen from the figure, design indices having values from 0 to 191 represent channels, for example, channels of a specific layer within a neural network. The assignment of indices to channels, as described above, is referred to as the significance order in certain cases where the indexing is done according to a given significance. In part A) of FIG. 18, the least significant channel index is 190, which means that the data of 191 channels from 0 to 190 are signaled and the data of the 192nd channel with index 191 is not signaled.

[0229] Similarly, FIG. 17 illustrates how, in step S610, the last significant channel indicator is obtained from the bitstream. Next, step S620 represents a loop over the channels. For each i-th channel, in the significance order obtained from the bitstream, the decoder determines the design index d in S630 and determines the corresponding channel_presence_flag[d] by comparing the index i with the last significant channel index in S635. If i is higher than the last significant index, channel_presence_flag[d] is set equal to FALSE in step S645. Otherwise, channel_presence_flag[d] is set equal to TRUE in step S640. Next, in the channel information analysis stage, the decoder uses channel_presence_flag[d] to analyze the channel information from the bitstream in step S670 (if channel_presence_flag[d] = TRUE) or to initialize the channel information to a predefined value, e.g., 0, in step S660 (if channel_presence_flag[d] = FALSE) (e.g., in decision step S650). Further, in step S680, the channel information can be further decoded by an entropy decoder such as an arithmetic decoder or the like. Finally, in step S690, the decoded channel information for all channels is provided to the corresponding neural network layer.

[0230] An exemplary syntax table illustrating the process of generating and analyzing a bitstream for an embodiment in which the significance order is signaled within the bitstream is presented below.

[0231]

Table 6

[0232] An exemplary syntax for obtaining significance order elements is as follows.

[0233] [Table 7]

[0234] As can be seen from the syntax of obtain_significance_order(), the significance order is obtained by iterating over the channel index i and assigning to each channel i the design channel index design_channel_idx that has the i-th highest significance. This syntax can also be regarded as a list of design channel indices ordered according to the significance of the corresponding respective channels. However, it should be noted that this syntax is merely exemplary and does not limit the ways in which the significance order can be transmitted between the encoder and the decoder.

[0235] One of the technical advantages of this approach (signaling the last significant channel indicator) is content adaptability. In fact, different channels of a CNN can represent different features of natural images or videos. For generality, during the design and training phases, the CNN can be targeted to cover a wider range of variations of possible input signals. In some embodiments, the encoder has the flexibility to identify and optimize the amount of information for signaling, for example, by using one of the methods as described above. It should be noted that the significance order information (indicator) can be transmitted at different levels of granularity, for example, once per picture, once per part of a picture, once per sequence or group of pictures. The granularity can be preset or predefined by a standard. Alternatively, the encoder can have the flexibility to determine when and how many times to signal the significance order information based on the content characteristics. For example, it may be beneficial to transmit or update the significance order information after a new scene is detected within a sequence, due to a change in the statistical characteristics.

[0236] In some embodiments, in the decoder, obtaining the significance order includes deriving the significance order based on previously decoded information regarding the source data from which the feature map was generated. In particular, the decoder may determine the significance order using supplementary information regarding the type of the encoded content transmitted in the bitstream. For example, the supplementary information may distinguish between expert-generated content, user-generated content, camera-captured content or computer-generated content, video games, or the like. Alternatively, or in addition, the supplementary information may distinguish between different chroma sampling rate types, such as YUV420, YUV422, or YUV444. Further source description information may be used in addition to, or alternatively to, this.

[0237] According to an exemplary implementation, obtaining the significance order includes deriving the significance order based on previously decoded information regarding the type of the source data from which the feature map was generated.

[0238] In any of the above-described embodiments and implementations, the encoded bitstream may be arranged in a way that provides some bitrate scalability support. This can be achieved, for example, in that the CNN channel information is placed in the bitstream according to the relative importance of the channels, with the most important channels coming first and the least important channels coming last. The relative channel importance defining the channel order in the bitstream may be defined at the design stage as described above and may be known to both the encoder and the decoder such that it does not need to be transmitted in the bitstream. Alternatively, the significance order information may be defined or adjusted in the encoding process and may be included in the bitstream. The latter option provides some greater flexibility and adaptability with respect to specific features of the content, which may enable the use of higher compression efficiency.

[0239] According to an exemplary implementation, the indication of the last significant channel corresponds to a quality indicator decoded from the bitstream and indicates the quality of the encoded feature map resulting from the compression of the region of the feature map. The last significant channel index can be signaled indirectly, for example, by using a quality indicator transmitted in the bitstream and a look-up table that defines the correspondence between the quality level used to reconstruct a picture with a desired quality level and the channel. For example, parts A and B of FIG. 16 show such a look-up table, where each quality level is associated with the design index of the corresponding last significant channel. In other words, for an input (desired) quality level, the look-up table provides the last significant channel index. Therefore, when the quality level of the bitstream is signaled or derivable at the decoder side, the indication of the last significant channel need not be signaled in the bitstream.

[0240] Note that the amount of quality gradation may not correspond to the amount of channels in the model. For example, in a given example of (FIG. 16, part B), the amount of quality levels is 100 (from 0 to 99) and the amount of channels is 192 (from 0 to 191). For example, in a given example, for quality level 0, the last significant channel index is 1, and for quality level 98, the last significant channel index is 190.

[0241] FIG. 18 illustrates a further example of a look-up table. As described above, for example, referring to FIG. 17, the last significant channel index can be signaled in the bitstream in a direct form or, indirectly, by using, for example, a quality indicator transmitted in the bitstream and a look-up table that defines the correspondence between the quality level used to reconstruct a picture having a desired quality level and the channels. Also in FIG. 18, the amount of quality levels is 100 (from 0 to 99) and the amount of channels is 192 (from 0 to 191). In that case, the last significant channel is defined as the last channel in the order of significance corresponding to the specified (desired) quality level. For example, in a given example, for quality level 0, the last significant channel index is 1, and for quality level 98, the last significant channel index is 190. In part B) of FIG. 18, the relationship between the quality level, the significance index, and the design index is shown. As can be seen, the significance index can be substantially different from the design index.

[0242] Generally, the indication of the last significant channel corresponds to the index of the last significant channel within the order of significance. In other words, M channels are ordered and indexed in an order of significance from 1 to M.

[0243] For example, in an application including image or video encoding, there may be quality settings selected by an encoder (e.g., by a user or an application or the like). Such quality settings (indicators) can then be associated with specific values of the last significant channel as shown in FIGS. 16 and 18. For example, the higher the desired image / video quality after reconstruction, the more channels will be signaled, i.e., the last significant channel index (among the indexes of all channels ordered according to their significance in descending order) will be higher. These are just examples, and it should be noted that the channel ordering can be done in ascending order, or the ascending / descending order can even be selectable and signaled.

[0244] Depending on the desired level of granularity defined, for example, by application requirements, the scalability level can be defined per CNN channel as depicted in FIG. 19, or by grouping several channels into one scalability level. This provides additional flexibility that can better fit application-specific tasks, thereby enabling reduction of the bitrate without re-encoding by dropping the least important feature map channels corresponding to the scalability level.

[0245] Other criteria for channel ordering could be the channel similarity. This similarity could be estimated, for example, as the amount of bits required to encode the channel information. A similar amount of bits required to represent the channel information could be considered as a similar level of the channel information. The amount of information is used as a criterion (from the maximum amount of bits to the minimum amount of bits, or from the minimum amount of bits to the maximum amount of bits) for sorting channels, and similar channels can be successively placed within the bitstream. This results in an additive advantage in compression efficiency due to stabilizing the probability model, which is updated along with symbol decoding when using context-adaptive arithmetic coding (CABAC). It provides more accurate probability estimation by the model and improves compression efficiency.

[0246] Presence of a portion of the feature map Signaling of region presence indicators and / or side information presence indicators can be performed for regions of the feature map. In some embodiments, the regions are channels of the feature map. The feature map generated by the convolutional layer can (at least partially) preserve the spatial relationship between the feature map values as well as the spatial relationship of the input picture samples. Thus, the feature map channels can have regions of different importance for the reconstruction quality, for example, flat regions and non-flat regions including the edges of objects. Dividing the feature map data into units makes it possible to capture this structure of the regions and utilize it for compression efficiency, for example, by skipping the transmission of the feature values of flat regions.

[0247] According to an exemplary implementation, the channel information including the channel feature map has one or more indicators indicating whether a specific spatial region of the channel exists within the bitstream. In other words, the CNN region is an arbitrary region of the CNN channel, and the bitstream further includes information for defining the arbitrary region. To specify a region, known methods for partitioning a two-dimensional space can be used. For example, a quadtree partitioning or other hierarchical tree partitioning method can be applied. Quadtree partitioning is a partitioning method that enables partitioning a two-dimensional space by recursively subdividing it into four quadrants or regions.

[0248] For example, a decoding method (such as the method according to any of the above-described embodiments and exemplary implementations) may further include a step of decoding, from the bitstream, region division information indicating to divide the region of the feature map into units, and decoding a unit existence indication indicating whether feature map data is to be parsed from the bitstream that is not for decoding the unit of the region according to the division information (parsing from the bitstream) or not decoding (bypassing the parsing).

[0249] Correspondingly, the encoding method determines whether the region of the feature map should be further divided into units. If affirmed, division information indicating to divide the region of the feature map into units is inserted into the bitstream. Otherwise, division information indicating not to divide the region of the feature map into units is inserted into the bitstream. It should be noted that the division information may also be a flag that can be further entropy-encoded. However, the present disclosure is not limited thereto, and the division information may also simultaneously signal the parameters of the division when the division is applied.

[0250] When segmentation is applied, the encoder may further determine whether a particular unit exists within the bitstream. Correspondingly, a unit presence indication is provided that indicates whether the feature map data corresponding to the unit is included in the bitstream. The unit presence indicator may be regarded as a special case of the region presence indicator described above, and it should be noted that the above description also applies to this embodiment.

[0251] In some embodiments, the region segmentation information for a region includes a flag indicating whether the bitstream contains unit information specifying the dimensions and / or position of the units of the region. The decoding method includes decoding a unit presence indication for each unit of the region from the bitstream. Depending on the value of the unit presence indication for a unit, the decoding method includes parsing or not parsing the feature map data for the unit from the bitstream.

[0252] In some embodiments, the unit information specifies a hierarchical segmentation of a region including at least one of a quadtree, a binary tree, a ternary tree, or a triangular segmentation.

[0253] A quadtree is a tree data structure in which each internal node has exactly four child nodes. Quadtrees are most often used to partition a two-dimensional space by recursively subdividing it into four quadrants or regions. The data associated with leaf cells varies depending on the application, but leaf cells represent "units of interesting spatial information." In the quadtree partitioning method, the subdivided regions are squares, each of which can be further divided into four child regions. Also, binary and ternary tree methods can include rectangular-shaped units, which correspondingly have two or three children for each internal node. Additionally, any partitioning shape can be achieved using masks or geometric rules, for example, by obtaining a triangular unit shape. This data structure is named a partition tree. The partitioning method can include combinations of different trees such as quadtrees, binary trees, and / or ternary trees within one tree structure. Having different unit shapes and partitioning rules enables more appropriately capturing the spatial structure of feature map data and signaling it to the bitstream in the most efficient way. All forms of the partition tree share some common features, namely, they decompose space into adaptable cells, each cell (or bucket) has a maximum capacity, and when the maximum capacity is reached, the bucket is divided and the tree directory follows the spatial decomposition of the partition tree.

[0254] Figure 20 is a schematic diagram illustrating the quadtree partitioning of a block (a unit of a two-dimensional image or feature map) 2000. Each block can be divided at each step of hierarchical division into four, but it may not be divided. In the first step, block 2000 is divided into four equal-sized blocks, one of which is 2010. In the second step, two of the four blocks (the right blocks in this example) are further partitioned, and each of them is again divided into four equal-sized blocks including block 2020. In the third step, one of the eight blocks is further divided into four equal-sized blocks including block 2030.

[0255] It is worth noting that other partitioning methods, such as binary tree and ternary tree partition splitting, can be used to achieve the same purpose of analyzing any region definition and the corresponding presence flags. Shown below is a syntax table that illustrates possible implementations of the syntax and the corresponding analysis process (and the corresponding bitstream generation process). For each CNN channel, first the parse_quad_tree_presence function is called to read information (split_qt_flag) regarding the arbitrary region definition, fill in information regarding the region size (width, height) and position (x, y), and analyze the presence_flag corresponding to each arbitrary region of the channel.

[0256] Note that other variations of the implementation that produce the same result are also possible. For example, the analysis of the region presence flags can be combined with the analysis of the channel information in one iteration of the analysis loop, or the region splitting information can be shared within a channel group of all channels to reduce the signaling overhead regarding the partition information. In other words, the region splitting information can specify the splitting that should be applied to the splitting of multiple (two or more or all) channels (regions) of the feature map.

[0257] [Table 8]

[0258] In the following syntax table, one implementation of analyzing the region splitting information is presented. The quadtree method is used as an example.

[0259] [Table 9]

[0260] The following syntax table illustrates an example of sharing region partition information (channel_regions_info) within a group of channels, which can also be all channels of the feature map.

[0261] [Table 10]

[0262] The presence information can be hierarchically organized. For example, if the channel_presence_flag for a specific channel (described in some of the above examples and embodiments) is equal to FALSE in the bitstream (indicating that all channel-related information is omitted), further analysis of the channel-related information is not performed for that channel. If the channel_presence_flag is equal to TRUE, the decoder analyzes the region presence information corresponding to the specific channel and then extracts the region-wise channel information from the bitstream that may be related to the unit of the channel. This is illustrated by the exemplary syntax shown below.

[0263] [Table 11]

[0264] Alternatively, or in combination with the previous embodiments, additional syntax elements can be used to indicate whether an arbitrary region presence signaling mechanism should be enabled. In particular, an enable flag can be used to signal whether channel partitioning is enabled (permitted).

[0265] [Table 12]

[0266] The element enable_channel_regions_flag[i] is signaled for each channel i when the channel is present in the bitstream (controlled by the channel_presence_flag[i] flag), and indicates whether channel partitioning is enabled. If the partitioning is enabled, the decoder calls the parse_quad_tree_presence() function to parse the region partitioning information from the bitstream and the presence flags corresponding to the partition units. After the unit presence flag (channel_regions_info[i][n].presence_flag) is obtained, the decoder iterates over the feature map elements of the partition unit. If the unit presence flag is equal to TRUE, the decoder parses the feature map value (decode_latent_value(i,x,y)). Otherwise, the feature map value y_cap[i][y][x] is set to be equal to a constant such as 0 in this given example.

[0267] Application of transformation According to this embodiment, the (CNN) channel feature map representing the channel information is transformed before being signaled in the bitstream. The bitstream further includes a syntax element that defines the position of the last significant transformed channel feature map coefficient. The transformation (also referred to as transform herein) can be any suitable transformation that can, for example, result in some energy compression. After completing the forward transformation of the feature map data, the encoder performs a zigzag scan on the last non-zero coefficient starting from the upper left corner of the transformed coefficient matrix (the upper left corner is the starting point where x = 0, y = 0). In other possible implementations, the encoder may decide to drop some non-zero coefficients considering them to be the least important. The position (x,y) of the last significant coefficient is signaled within the bitstream.

[0268] On the receiving side, for example, the decoder analyzes the corresponding syntax element to define the position of the last significant transform coefficient of the corresponding channel's feature map. After the position is defined, the decoder analyzes the transformed feature map data belonging to the region starting from the upper left corner (x = 0, y = 0) and ending at the position of the last significant coefficient (x = last_significant_x, y = last_significant_y). The remaining transformed feature map data is initialized with a constant value, for example 0. When the transformed coefficients of the corresponding CNN channel are known, an inverse transform can be performed. Based on this, the channel feature map data can be obtained. Such a process can be completed for all input channels. Then, the feature map data may be supplied to a reconstruction network representing the generation model.

[0269] Generally, the decoding of a region from a bitstream is illustrated in FIG. 21 as if the region were a channel. In particular, in step S710, the channel is fetched from among all the channels to be processed. Step S710 corresponds to a loop over the channels to be processed. The decoding includes extracting in S720 a last significant coefficient indicator that indicates the position of the last coefficient among the coefficients of the region (here, exemplarily, the channel) from the bitstream. Step S730 corresponds to a loop over the coefficients. In particular, the loop goes over each of the coefficients from the most significant coefficient to the last significant coefficient. In the loop, the method implements decoding (including parsing S750 and optionally entropy decoding S760) of the significant coefficients from the bitstream (for the region of the S710 loop) and setting in S740 the coefficients following the last significant coefficient indicator according to a predefined rule. In the exemplary flowchart of FIG. 21, the predefined rule is to set the coefficients following the last significant coefficient in significance order to 0. However, as previously explained for the phrase "predefined rule", the rule may define the setting of the coefficients to different values, or that the value to be set should be obtained from the bitstream, or by a particular method (such as derivation from previously used values), or by a similar method. Generally, the present embodiment is not limited to any particular predefined rule.

[0270] Regarding the method steps, in step S735, it is tested (determined, checked) whether the current coefficient (the coefficient given by the steps of loop S730) is a significant coefficient or a non-significant coefficient. Here, a significant coefficient is a coefficient where the sample position is below and equal to the last significant coefficient. A non-significant coefficient is a coefficient above (exceeding) the last significant coefficient. In this example, it is assumed that the significance order corresponds to the order of the samples after transformation. This generally assumes that DC and low-frequency coefficients are more significant than higher-frequency coefficients. The sample position is the position (index) of the transformed (optionally scanned in advance to be 1D) coefficients obtained as a result of the 1D transformation of the region feature map. Note that the present disclosure is not limited to the application of 1D transformation or any specific data format. The significance order of the coefficients can be defined in different ways - for example, by a standard or convention, or even provided by a bitstream or the like.

[0271] After all the coefficients have been obtained (after loop S730 has ended), the decoding method includes S770 of obtaining the region feature data by inverse-transforming the coefficients of the region. As shown in FIG. 21, when the data of all channels are decoded, the decoding method may further include step S780 of supplying the data of all channels to the reconstruction network.

[0272] According to an exemplary implementation, the inverse transformation is an inverse discrete cosine transform, an inverse discrete sine transform, or an inverse transformation obtained by modifying the inverse discrete cosine transform or the inverse discrete cosine transform, or a convolutional neural network transform. For example, a convolutional neural network layer may be regarded as a transformation, and if trained with the last significant coefficient signaling method as described above, energy compression can be performed as desired. Further, the last layer of the analysis part of the autoencoder (which generates the latent space representation y) and the first layer of the generation part of the autoencoder (which receives the quantized latent representation) may be regarded as the forward transformation and the inverse transformation correspondingly. Therefore, if the last significant coefficient signaling method as described above is used during the feature map encoding stage, the desired energy compression of the forward convolutional layer transformation obtained during the joint training will be performed. The present disclosure is not limited to these types of transformations. Rather, other transformations such as KLT, Fourier, Hadamard, or other orthogonal and in some cases unitary transformations may be used.

[0273] The following shows a syntax table exemplifying an analysis process when a transformation is applied to two-dimensional channel feature map data whose dimensions are represented by indices x and y. It should be noted that the present disclosure is not limited to any specific dimensionality of the feature map data. The transformation can be applied after scanning or after assigning positions within some predefined significance order to the data.

[0274] [Table 13]

[0275] In the above example, the position of the last significant coefficient is represented by two numerical values corresponding to the x and y coordinates of the position of the last significant coefficient in the 2D feature map (or its region). In other possible implementations, the position of the last significant coefficient may be syntactically represented by a single number corresponding to the order of the zigzag scan of the 2D feature map (its region).

[0276] FIG. 22 shows a flowchart illustrating an exemplary method that can be performed on the encoder side. Note that the encoding of FIG. 22 is compatible with the decoding of FIG. 21, and thus the bitstream generated by the encoding method of FIG. 22 can be parsed / decoded by the decoding method of FIG. 21.

[0277] The encoding method includes a loop S810 over the channels of the feature map (data output by a layer of a machine learning model, such as a neural network). In step S820, the data of the current channel (the channel at the current step of loop S810) is transformed by a forward transform. As described above, the forward transform may here include scanning the channel features into a 1D sample sequence and then transforming it with a 1D transform to obtain a 1D sequence of coefficients. Such a 1D sequence of coefficients may be considered to follow a significance order. As described above, this example is not limiting. The transform may have more dimensions and the scanning and / or significance order may be different.

[0278] In step S830, the last significant position is determined. The determination of the last significant position can be made based on the desired reconstructed quality and / or rate. For example, similar to the method shown while referring to FIG. 12, the quality and / or rate for the coefficients are accumulated, and as soon as the quality and / or rate or a combination of both exceeds the desired threshold, the current coefficient is regarded as the last significant coefficient. However, other approaches are also possible. In step S840, the position (its indication) of the last significant coefficient determined in step S830 is inserted into the bitstream. Before inserting into the bitstream, the indication of the position can be encoded by an entropy encoding such as arithmetic coding or other variable-length coding methods. Step S850 represents a loop for each of the transformed coefficients. If the coefficient position is lower than or the same as the position of the last significant coefficient, in step S880, the current coefficient is inserted into the bitstream (optionally encoded by an arithmetic encoder in step S870). After encoding the coefficients of all channels, the bitstream may be supplied (or transmitted) to the reconstruction network.

[0279] As will be apparent to those skilled in the art, this embodiment may be combined with any of the above-described embodiments. In other words, the last significant coefficient indication may be signaled together with the last significant channel. The signaling may be performed for an area smaller than the channel (obtained, for example, by partitioning) or for a larger area. Also, side information may be provided. Generally, the above-described embodiments may be combined to provide more flexibility.

[0280] Furthermore, as already described, the present disclosure also provides a device configured to perform the steps of the methods described above. FIG. 23 shows a device 2300 for decoding a feature map for neural network processing based on a bitstream. The device includes a region presence indicator acquisition module 2310 configured to acquire a region presence indicator based on information from the bitstream for a region of the feature map. Further, the device 2300 may further include a decoding module configured to analyze data from the bitstream to decode a region when the region presence indicator has a first value, and to bypass analyzing data from the bitstream to decode a region when the region presence indicator has a second value, so as to decode the region.

[0281] Furthermore, FIG. 24 shows a device 2400 for decoding a feature map for neural network processing from a bitstream. This device includes a side information indicator acquisition module 2410 configured to acquire a side information indicator regarding the feature map from the bitstream, and a decoding module 2420 configured to decode the feature map, the decoding module 2420 including analyzing side information for decoding the feature map from the bitstream when the side information indicator has a fifth value, and bypassing analyzing side information for decoding the feature map from the bitstream when the side information indicator has a sixth value.

[0282] Corresponding to the above-described decoding device 2300, a device 2350 for encoding a feature map for neural network processing in a bit stream is shown in FIG. 23. The device includes a feature map region presence indicator acquisition module 2360 configured to acquire a region presence indicator for a region of the feature map. Further, the device is configured to, based on the acquired feature map region presence indicator, when the region presence indicator has a first value, determine whether to indicate the region of the feature map and indicate it in the bit stream, or when the region presence indicator has a second value, bypass indicating the region of the feature map, and includes an encoding control module 2370 configured as such.

[0283] Corresponding to the above-described decoding device 2400, a device 2450 for encoding a feature map for neural network processing in a bit stream is shown in FIG. 24. This device may include a feature map acquisition module 2460 configured to acquire a feature map. Further, the device 2450 may be configured to determine whether to indicate side information regarding the feature map, and further include an encoding control module 2470 configured to indicate either a side information indicator that indicates a third value and side information in the bit stream, or a side information indicator that indicates a fourth value without side information.

[0284] Note that these devices can be further configured to perform any of the additional features, including the exemplary implementations described above. For example, a device for decoding a feature map for neural network processing based on a bitstream is provided, and this device includes a processing circuit configured to perform any of the steps of the decoding methods described above. Similarly, a device for encoding a feature map for neural network processing in a bitstream is provided, and this device includes a processing circuit configured to perform any of the steps of the encoding methods described above.

[0285] Additional devices are provided, which utilize devices 2300, 2350, 2400, and / or 2450. For example, a device for image or video encoding may include encoding devices 2400 and / or 2450. In addition, this may also include decoding device 2300 and / or 2350. A device for image or video decoding may include decoding devices 2300 and / or 2350.

[0286] In summary, in some embodiments of the present disclosure, a decoder (video or image or feature map) may include an artificial neural network configured to obtain feature map data from a bitstream. The decoder reads a region presence flag from the bitstream, reads the feature map data corresponding to the region if the region presence flag is TRUE, initializes the region of the feature map data with a predefined value (the predefined value is illustratively 0) if the region presence flag is FALSE, and may be configured to supply data to the artificial neural network based on the obtained feature map data. The artificial neural network may be a convolutional neural network. The region may be a channel output of the convolutional neural network. Alternatively, or in addition, the region may be part of the channel output of the convolutional neural network, and the bitstream may further include information (position, shape, size) for the region definition. In other words, the region parameters may be defined in the bitstream. The artificial neural network does not necessarily have to be a CNN. Generally, the neural network may be a fully connected neural network.

[0287] The decoder may be further configured to read (extract from the bitstream) the feature map data by entropy decoding, such as Huffman (de)coding, range coding, arithmetic coding, asymmetric number system (ANS), or entropy of other types of variable length codes. The region presence flag may be encoded as a context-coded bin with corresponding probability updates.

[0288] Some exemplary implementations in hardware and software A corresponding system capable of deploying the above encoder-decoder processing chain is illustrated in FIG. 25. FIG. 25 is a schematic block diagram illustrating an exemplary coding system that can utilize the technology of the present application, for example, a video, image, audio, and / or other coding system (or short coding system). The video encoder 20 (or short encoder 20) and video decoder 30 (or short decoder 30) of the video coding system 10 represent examples of devices that can be configured to perform the techniques according to various examples described in the present application. For example, video encoding and decoding may employ a neural network as shown in FIGS. 1 to 7, which may be distributed, and the above-described bitstream analysis and / or bitstream generation may be applied to transmit feature maps between distributed computing nodes (two or more).

[0289] As shown in FIG. 25, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to a destination device 14, for example, for decoding encoded picture data 13.

[0290] The source device 12 includes an encoder 20 and, in addition, optionally, may include a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.

[0291] The picture source 16 may include, or may be, any kind of picture capture device, such as a camera for capturing real-world pictures, and / or any kind of picture generation device, such as a computer graphics processor for generating computer animation pictures, or any other device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures) and / or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may be any kind of memory or storage device that stores any of the aforementioned pictures.

[0292] Distinguished from the processing performed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17.

[0293] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optional component. Note that the preprocessing may also employ a neural network (such as in any of FIGS. 1 to 7) that uses presence indicator signaling.

[0294] The video encoder 20 is configured to receive the preprocessed picture data 19 and provide encoded picture data 21.

[0295] The communication interface 22 of the source device 12 receives the encoded picture data 21 and may be configured to transmit the encoded picture data 21 (or any further processed version thereof) on the communication channel 13 to other devices, such as the destination device 14 or any other device, for storage or direct reconstruction.

[0296] The destination device 14 includes a decoder 30 (e.g., a video decoder 30) and may further include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.

[0297] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof) from, for example, directly from the source device 12 or any other source, such as a storage device, e.g., an encoded picture data storage device, and provide the encoded picture data 21 to the decoder 30.

[0298] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network or any combination thereof, or any type of private and public network, or any combination thereof.

[0299] The communication interface 22 may be configured to, for example, package the encoded picture data 21 into an appropriate format, such as packets, and / or process the encoded picture data using any kind of transmission encoding or processing for transmission over a communication link or communication network.

[0300] The communication interface 28 that forms one half of the pair of communication interfaces 22 may be configured to, for example, receive the transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or unpacking to obtain the encoded picture data 21.

[0301] Both the communication interface 22 and the communication interface 28 may be configured as a unidirectional communication interface as indicated by the arrow of the communication channel 13 in FIG. 25 from the source device 12 to the destination device 14, or as a bidirectional communication interface, for example, to send and receive messages, for example, to set up a connection and perform acknowledgment responses and exchanges of any other information related to the communication link and / or data transmission, for example, encoded picture data transmission. The decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (for example, employing a neural network based on one or more of FIGS. 1 to 7).

[0302] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example, the decoded picture 31, to obtain post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing for preparing the decoded picture data 31 for display by, for example, the display device 34.

[0303] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33, for example, to display a picture to a user or viewer. The display device 34 may be any type of display for representing the reconstructed picture, for example, an integrated or external display or monitor, or may include them. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.

[0304] FIG. 25 illustrates the source device 12 and the destination device 14 as separate devices, but embodiments of the device may include both or both functionalities, i.e., the source device 12 or corresponding functionality, and the destination device 14 or corresponding functionality. In such embodiments, the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, separate hardware and / or software, or any combination thereof.

[0305] As will be apparent to those skilled in the art based on the description, the functionality of the different units and their (exact) partitioning, or the functionality of the source device 12 and / or destination device 14 as shown in FIG. 25, may vary depending on the actual devices and applications.

[0306] The encoder 20 (e.g., a video encoder 20) or the decoder 30 (e.g., a video decoder 30) or both the encoder 20 and the decoder 30 may be implemented via a processing circuit such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, video encoding dedicated devices, or any combination thereof. The encoder 20 may be implemented via the processing circuit 46 to embody various modules including a neural network such as those shown in any of FIGS. 1 to 7 or a portion thereof. The decoder 30 may be implemented via the processing circuit 46 to embody various modules as described with respect to FIGS. 1 to 7 and / or any other decoder system or subsystem described herein. The processing circuit may be configured to perform various operations as described hereinafter. When the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Either the video encoder 20 or the video decoder 30 may be integrated, for example, as part of a combined encoder / decoder (CODEC) in a single device as shown in FIG. 26.

[0307] Source device 12 and destination device 14 may include any of a wide range of devices, such as any type of handheld or stationary device, for example, a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiver device, a broadcast transmitter device, or the like, and may either not use an operating system at all or may use any type of operating system. In some cases, source device 12 and destination device 14 may be equipped for wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices.

[0308] In some cases, the video encoding system 10 illustrated in FIG. 25 is merely an example, and the technology of the present application may be applicable to video encoding settings (such as video encoding or video decoding) that do not necessarily include data communication between an encoding device and a decoding device. In other examples, the data is retrieved from local memory, streamed over a network, or otherwise processed. The video encoding device may encode data and store it in memory, and / or the video decoding device may retrieve the data from memory and decode it. In some examples, encoding and decoding do not communicate with each other, but are simply performed by devices that encode data and put it into memory and / or retrieve data from memory and decode it.

[0309] FIG. 27 is a schematic diagram of a video encoding device 1000 according to an embodiment of the present disclosure. The video encoding device 1000 is suitable for implementing the disclosed embodiments as described herein. In one embodiment, the video encoding device 1000 may be a decoder such as the video decoder 30 of FIG. 25, or an encoder such as the video encoder 20 of FIG. 25.

[0310] The video encoding device 1000 includes a receiving port 1010 (or input port 1010) and a receiver unit (Rx) 1020 for receiving data, a processor, logic unit, or central processing unit (CPU) 1030 for processing data, a transmitter unit (Tx) 1040 and a transmitting port 1050 (or output port 1050) for transmitting data, and a memory 1060 for storing data. The video encoding device 1000 may also include optoelectronic (OE) components and electro-optic (EO) components coupled to the receiving port 1010, the receiver unit 1020, the transmitter unit 1040, and the transmitting port 1050 for the transmission or reception of optical or electrical signals.

[0311] Processor 1030 is implemented by hardware and software. Processor 1030 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. Processor 1030 communicates with a receiving port 1010, a receiver unit 1020, a transmitter unit 1040, a transmitting port 1050, and a memory 1060. Processor 1030 includes an encoding module 1070. Encoding module 1070 implements the disclosed embodiments described above. For example, encoding module 1070 implements, processes, prepares, or provides various encoding operations. Thus, the inclusion of encoding module 1070 brings a substantial improvement to the functionality of video encoding device 1000 and results in the conversion of video encoding device 1000 to different states. Alternatively, encoding module 1070 is implemented as instructions stored in memory 1060 and executed by processor 1030.

[0312] Memory 1060 includes one or more disks, tape drives, and solid state drives, stores programs when such programs are selected for execution, and stores instructions and data read during program execution and can be used as an overflow data storage device. Memory 1060 may be, for example, volatile and / or non-volatile and may be read only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).

[0313] FIG. 28 is a simplified block diagram of an apparatus 800 that may be used as either or both of a source device 12 and a destination device 14 from FIG. 25 according to an exemplary embodiment.

[0314] The processor 1102 within the device 1100 may be a central processing unit. Alternatively, the processor 1102 may be any other type of device or devices capable of manipulating or processing information that currently exists or will be developed later. The disclosed implementations may be implemented by a single processor, such as processor 1102 as illustrated, but the advantages regarding speed and efficiency may be achieved by using multiple processors.

[0315] The memory 1104 within the device 1100 may be a read-only memory (ROM) device or a random access memory (RAM) in one implementation. Other types of storage device devices may also be used as the memory 1104 if suitable. The memory 1104 may include code and data 1106 accessed by the processor 1102 using the bus 1112. The memory 1104 may further include an operating system 1108 and an application program 1110, and the application program 1110 includes at least one program that enables the processor 1102 to perform the methods described herein. For example, the application program 1110 may be, for example, those including Applications 1 through N, which further include video encoding applications that perform the methods described herein.

[0316] The device 1100 may also include one or more output devices, such as a display 1118. The display 1118 may be, in one example, a touch sensor type display combined with a touch sensor type element operable to sense touch input to the display. The display 1118 may be coupled to the processor 1102 via the bus 1112.

[0317] Although depicted here as a single bus, the bus 1112 of the apparatus 1100 may be composed of multiple buses. Further, the secondary storage device may be directly coupled to other components of the apparatus 1100 or accessed via a network and may include a single built-in unit such as a memory card or multiple units such as multiple memory cards. Thus, the apparatus 1100 can be implemented in a wide variety of configurations.

[0318] Some mathematical operators and symbols The mathematical operators in the exemplary syntax description used in this application are similar to those used to describe syntax in existing codecs. The numbering and counting conventions generally start from 0. For example, "the first" corresponds to the 0th, "the second" corresponds to the 1st, and so on.

[0319] The following arithmetic operators are defined as follows. + Addition - Subtraction (as a 2-argument operator) or negation (as a unary prefix operator) * Multiplication, including matrix multiplication / Integer division that truncates the result towards zero. For example, 7 / 4 and -7 / -4 are truncated to 1, and -7 / 4 and 7 / -4 are truncated to -1. x%y Modulo. The remainder when x is divided by y, defined only for integer x and y, where x >= 0 and y > 0.

[0320] The following logical operators are defined as follows. x && y Boolean logical product "and" of x and y x || y Boolean logical sum "or" of x and y ! Boolean logical negation "not" x? y : z If x is TRUE or not equal to 0, evaluate the value of y; otherwise, evaluate the value of z.

[0321] The following relational operators are defined as follows. > greater than >= greater than or equal to < less than <= less than or equal to == equal to != not equal to

[0322] When a relational operator is applied to a syntactic element or variable to which the value "na" (not applicable) is assigned, the value "na" is treated as a distinct value for that syntactic element or variable. The value "na" is considered not equal to any other value.

[0323] The following bitwise operators are defined as follows. & bitwise "and". In the case of an operation with integer arguments, it is an operation in the two's complement representation of integer values. When an operation with a binary argument containing fewer bits than the other argument is performed, the short argument is extended by adding additional high-order bits equal to 0. | bitwise "or". In the case of an operation with integer arguments, it is an operation in the two's complement representation of integer values. When an operation with a binary argument containing fewer bits than the other argument is performed, the short argument is extended by adding additional high-order bits equal to 0. ^ bitwise "exclusive or". In the case of an operation with integer arguments, it is an operation in the two's complement representation of integer values. When an operation with a binary argument containing fewer bits than the other argument is performed, the short argument is extended by adding additional high-order bits equal to 0. x >> y Performs an arithmetic right shift of binary digit number y on the two's complement representation integer x. This function is defined only for non-negative integer values of y. For the result of the right shift, the bit shifted into the most significant bit (MSB) has the same value as the MSB of x before the shift operation. x << y Performs an arithmetic left shift of binary digit number y on the two's complement representation integer x. This function is defined only for non-negative integer values of y. For the result of the left shift, the bit shifted into the least significant bit (LSB) has a value equal to 0.

[0324] The following arithmetic operators are defined as follows. = Assignment operator ++ Increments, i.e., x++ is equivalent to x = x + 1. When used as an array index, the value of the variable before the increment operation is evaluated. -- Decrements, i.e., x-- is equivalent to x = x - 1. When used as an array index, the value of the variable before the decrement operation is evaluated. += Increments by the specified amount, i.e., x += 3 is equivalent to x = x + 3 and x += (-3) is equivalent to x = x + (-3). -= Decrements by the specified amount, i.e., x -= 3 is equivalent to x = x - 3 and x -= (-3) is equivalent to x = x - (-3).

[0325] The following notations are used to specify ranges of values. x = y..z x takes integer values from y to z, where x, y, and z are integers and z is greater than y.

[0326] When the precedence of an expression is not explicitly indicated by using parentheses, the following rules apply. - Higher precedence operations are evaluated before lower precedence operations. - Operations of the same precedence are evaluated sequentially from left to right.

[0327] The following table specifies the precedence of operations from highest to lowest, where the higher the position in the table, the higher the precedence.

[0328] For operators also used in the C programming language, the precedence used in this specification is the same as that used in the C programming language.

[0329] Table: Operator precedence from highest (top of the table) to lowest (bottom of the table)

[0330]

Table 14

[0331] In the text, a statement of a logical operation mathematically described in the following form if(condition 0) Statement 0 else if(condition 1) Statement 1 ... else / * Remarks for reference regarding the remaining conditions * / Statement n can be described in the following way. ... as follows / ... the following applies: - If condition 0, Statement 0 - Otherwise, if condition 1, Statement 1 -... - Otherwise (Remarks for reference regarding the remaining conditions), Statement n

[0332] Each "If... Otherwise, if... Otherwise,..." statement in the text is introduced immediately followed by "... as follows" or "... the following applies" where "If..." follows. The last condition of "If... Otherwise, if... Otherwise,..." is always "Otherwise,...". Interleaved "If... Otherwise, if... Otherwise,..." statements can be identified by matching "... as follows" or "... the following applies" with the final "Otherwise,...".

[0333] In the text, a statement of a logical operation mathematically described in the following form if( condition 0a && condition 0b ) Statement 0 else if (condition 1a || condition 1b) Statement 1 ... else Statement n can be described in the following manner. ... as follows / ... the following applies: - If all of the following conditions are true, Statement 0: - Condition 0a - Condition 0b - Otherwise, if one or more of the following conditions are true, Statement 1: - Condition 1a - Condition 1b -... - Otherwise, Statement n

[0334] In the text, a logical operation statement mathematically described in the following form if (condition 0) Statement 0 if (condition 1) Statement 1 can be described in the following manner. - When condition 0, Statement 0 - When condition 1, Statement 1

[0335] In summary, the present disclosure relates to efficient signaling of feature map information for systems employing neural networks. In particular, on the decoder side, an existence indicator is parsed from the bitstream or derived based on information parsed from the bitstream. Based on the value of the parsed existence indicator, further data related to the feature map region is parsed or the parsing is bypassed. The existence indicator can be, for example, a region existence indicator indicating whether the bitstream contains feature map data, or a side information existence indicator indicating whether the bitstream contains side information related to the feature map data. Similarly, an encoding method, and further an encoding and decoding device are also provided.

Explanation of Signs

[0336] 10 Video encoding system 12 Source device 13 Encoded picture data 14 Destination device 16 Picture source 17 Picture or picture data 18 Preprocessor, preprocessing unit 19 Preprocessed picture, preprocessed picture data 20 Video encoder, short encoder 21 Encoded picture data 22 Communication interface or communication unit 28 Communication interface or communication unit 30 Video decoder, short decoder, decoder 31 Decoded picture data 32 Postprocessor 32 Postprocessing unit 33 Postprocessed picture data, postprocessed picture 34 Display device 46 Processing circuit 800 Device 1000 Video Encoding Device 1010 Reception Port, Input Port 1020 Receiver Unit (Rx) 1030 Central Processing Unit (CPU), Processor 1040 Transmitter Unit (Tx) 1050 Transmission Port, Output Port 1060 Memory 1070 Encoding Module 1100 Device 1102 Processor 1104 Memory 1106 Code and Data 1108 Operating System 1110 Application Program 1112 Bus 1118 Display 2000 Block 2020 Block 2030 Block 2300 Decoding Device 2310 Region Existence Indicator Acquisition Module 2350 Device for Encoding 2360 Feature Map Region Existence Indicator Acquisition Module 2370 Encoding Control Module 2400 Decoding Device 2410 Side Information Indicator Acquisition Module 2420 Decoding Module 2450 Device for Encoding 2460 Feature Map Acquisition Module 2470 Encoding Control Module

Claims

1. A method for decoding a feature map input to a neural network for processing an image by the neural network based on a bitstream, comprising: obtaining an area presence indicator based on information from the bitstream for an area in the feature map (step S110), wherein the feature map includes a plurality of channels generated by detecting features in the image, and the area is part of the channels; (step S110); decoding the area (step S150), when the area presence indicator has a first value, analyzing data from the bitstream to decode the area (step S130), and when the area presence indicator has a second value, bypassing analyzing data from the bitstream to decode the area (step S140), the method comprising (step S150).

2. When the area presence indicator has a second value, the step (S150) of decoding the area further comprises setting the area according to a predetermined rule (step S230), the method according to claim 1.

3. The method according to claim 2, wherein the predetermined rule specifies setting the features of the area to a constant.

4. The method according to claim 3, wherein the constant is 0.

5. The method according to claim 3 or 4, further comprising decoding the constant from the bitstream (step S240).

6. The method according to any one of claims 1 to 5, wherein the bitstream includes the area presence indicator.

7. The method according to any one of claims 1 to 5, further comprising obtaining side information from the bitstream and obtaining the area presence indicator based on the side information.

8. obtaining a side information presence indicator from the bitstream (step S310), when the side information presence indicator has a third value, analyzing the side information from the bitstream (step S330), and when the side information presence indicator has a fourth value, bypassing analyzing the side information from the bitstream (step S340). The method according to claim 7, wherein the side information includes at least one of the region presence indicator and information regarding being processed by a neural network to obtain an estimated probability model for use in entropy decoding of the region.

9. The method according to claim 8, wherein when the side information presence indicator has the fourth value, the method includes a step (S580) of setting the side information to a predetermined side information value.

10. The method according to any one of claims 1 to 9, wherein the region presence indicator is a flag that can take only one of two values formed by the first value and the second value.

11. The method according to any one of claims 1 to 10, wherein the image is abstracted into a feature map including a plurality of channels corresponding to the number of filters in the neural network, and each filter generates channel data by detecting features in the image.

12. A step (S610) of obtaining a significance order indicating the significance of a plurality of channels of the feature map; A step (S630) of obtaining a last significant channel indicator; The method according to claim 11, further including steps (S640, S645) of obtaining the region presence indicator based on the last significant channel indicator.

13. The method according to claim 12, wherein the indication of the last significant channel corresponds to a quality indicator decoded from the bitstream and indicates the quality of the encoded feature map resulting from compression of the region in the feature map.

14. The method according to claim 12, wherein the indication of the last significant channel corresponds to the index of the last significant channel in the significance order.

15. The method according to any one of claims 12 to 14, wherein the step of obtaining the significance order includes a step of decoding an indication of the significance order from the bitstream.

16. The method according to any one of claims 12 to 15, wherein the step (S630) of obtaining the significance order includes a step of deriving the significance order based on previously decoded information regarding source data from which the feature map was generated.

17. The step (S630) of obtaining the significance order includes deriving the significance order based on previously decoded information regarding the type of source data from which the feature map was generated, according to any one of claims 12 to 16.

18. The method according to any one of claims 12 to 17 further includes a step (S680) of decoding channels sorted within the bitstream according to the significance order from the most significant channel to the last significant channel from the bitstream.

19. Decode region division information from the bitstream that instructs to divide the region in the feature map into units, The method according to any one of claims 1 to 18 further includes a step of obtaining a unit existence indication that indicates whether feature map data should be parsed from the bitstream according to the region division information or not for decoding the units of the region.

20. The region division information for the region includes a flag that indicates whether the bitstream includes unit information specifying the dimensions and / or positions of the units of the region, The method includes a step of obtaining the unit existence indication for each unit of the region based on information from the bitstream, The method according to claim 19 includes a step of parsing, or not parsing, the feature map data for the unit from the bitstream according to the value of the unit existence indication for the unit.

21. The bitstream includes the unit existence indication, according to the method of claim 19 or 20.

22. The unit information specifies a hierarchical division of the region (2000) including at least one of a quadtree, a binary tree, a ternary tree, or a triangular division, according to the method of claim 20.

23. Decoding the region from the bitstream is extracting a last significant coefficient indicator (S720) from the bitstream that indicates the position of the last coefficient among the coefficients of the region, decoding the significant coefficients of the region from the bitstream (S750, S760), setting the coefficients following the last significant coefficient indicator according to a predefined rule (S740); obtaining the feature data of the region by inverse-transforming the coefficients of the region (S770), the method according to any one of claims 1 to 22. **Claim 24** The inverse transformation (S770) is an inverse discrete cosine transform, an inverse discrete sine transform, or an inverse transform obtained by modifying an inverse discrete cosine transform or an inverse discrete cosine transform, or a convolutional neural network transform, the method according to claim 23. **Claim 25** further comprising decoding a side information presence flag from the bitstream that indicates whether the bitstream contains any side information for the feature map, the side information including information regarding being processed by a neural network to obtain an estimated probability model for use in entropy decoding of the feature map, the method according to any one of claims 1 to 24. **Claim 26** Decoding the region presence indicator includes decoding by a context adaptive entropy decoder, the method according to any one of claims 1 to 25. **Claim 27** A method for decoding a feature map input to a neural network for processing an image by a neural network from a bitstream, obtaining a side information indicator regarding the feature map from the bitstream (S310); decoding the feature map (S350), when the side information indicator has a fifth value, parsing side information for decoding the feature map from the bitstream (S330), and when the side information indicator has a sixth value, bypassing parsing the side information for decoding the feature map from the bitstream (S340), the side information corresponding to a region within the feature map, the feature map including a plurality of channels generated by detecting features within the image, the region being a part of the channels, step (S350). **Claim 28** The data included in the feature map is entropy-encoded, The method according to claim 27, further comprising the step (S590) of entropy-decoding the data included in the decoded feature map.

29. The method according to claim 28, wherein when the side information indicator has the sixth value, the method includes the step (S580) of setting the side information to a predetermined side information value.

30. The method according to claim 29, wherein the predetermined side information value is 0.

31. The method according to claim 29 or 30, comprising the step (S590) of decoding the predetermined side information value from the bitstream.

32. A method for decoding an image, A method according to any one of claims 1 to 31 for decoding, from a bitstream, a feature map input to a neural network for processing an image by the neural network, and Obtaining a decoded image including the step of processing the decoded feature map by the neural network.

33. The feature map is Encoded image data, and / or Represents encoded side information for decoding the image data, according to the method of claim 32.

34. A method for computer vision, A method according to any one of claims 1 to 31 for decoding, from a bitstream, a feature map input to a neural network for processing an image by the neural network, and Performing a computer vision task including processing the decoded feature map by the neural network.

35. The method according to claim 34, wherein the computer vision task is object detection, object classification, and / or object recognition.

36. A method for encoding, in a bitstream, a feature map that is an output of a neural network for processing an image by the neural network, A step (S160) of obtaining a region presence indicator for a region in the feature map, wherein the feature map includes a plurality of channels generated by detecting features in the image, and the region is a part of the channels, step (S160); Based on the obtained region presence indicator; When the region presence indicator has a first value, encoding the region in the feature map into the bitstream (S180), or When the region presence indicator has a second value, bypassing encoding the region in the feature map into the bitstream (S190) A method including a step (S170) of determining;

37. The method according to claim 36, wherein the region presence indicator is indicated in the bitstream.

38. The method for encoding according to claim 36, wherein the determining step (S170) includes a step of evaluating a value of a feature of the region.

39. The method for encoding according to any one of claims 36 to 38, wherein the determining step (S170) is based on an influence of the region on a quality of a result of processing by the neural network.

40. The determining step (S170) A step (S450) of gradually determining a total number of bits required for transmission of the feature map by starting from bits of the most significant region and continuing with bits of regions having lower significance until the total exceeds a pre-configured threshold; Encoding the region where the total does not exceed the pre-configured threshold and the region presence indicator having the first value for the encoded region (S270, S280); Encoding the region presence indicator having the second value for the non-encoded region (S470), a method for encoding according to any one of claims 36 to 39.

41. A method for encoding a feature map that is an output of a neural network for processing an image by the neural network into a bitstream, including: A step (S360) of obtaining the feature map; Determining whether to indicate side information regarding the feature map and in the bitstream; (S380) A third value and a side information indicator indicating the side information, or (S390) a step (S370) of indicating either a side information indicator that indicates a fourth value without the side information, wherein the side information corresponds to a region in the feature map, the feature map includes a plurality of channels generated by detecting features in the image, and the region is a part of the channels. The method includes step (S370).

42. A computer program stored in a non-transitory medium, including code for performing the steps of the method according to any one of claims 1 to 35 when executed by one or more processors.

43. A computer program stored in a non-transitory medium, including code for performing the steps of the method according to any one of claims 36 to 41 when executed by one or more processors.

44. A device (2300) for decoding a feature map input to a neural network for processing an image by the neural network based on a bitstream, a region existence indicator acquisition module (2310) configured to acquire a region existence indicator based on information from the bitstream for a region in the feature map, wherein the feature map includes a plurality of channels generated by detecting features in the image, and the region is a part of the channels. The region existence indicator acquisition module (2310), a decoding module (2320), when the region existence indicator has a first value, the step of analyzing data from the bitstream to decode the region, and when the region existence indicator has a second value, the step of bypassing the step of analyzing data from the bitstream to decode the region The device (2300) includes a decoding module (2320) configured to decode the region.

45. A device (2400) for decoding a feature map input to a neural network for processing an image by the neural network from a bitstream, A side information indicator acquisition module (2410) configured to acquire a side information indicator regarding the feature map from the bitstream; A decoding module (2420), When the side information indicator has a fifth value, the step of analyzing side information for decoding the feature map from the bitstream, and When the side information indicator has a sixth value, the step of bypassing the step of analyzing the side information for decoding the feature map from the bitstream A decoding module (2420) configured to decode the feature map, including the side information corresponding to a region in the feature map, the feature map including a plurality of channels generated by detecting features in the image, and the region being a part of the channels. A device (2400) including the decoding module (2420). **Claim 46** A device (2350) for encoding a feature map that is an output of a neural network for processing an image by the neural network in a bitstream, A feature map region presence indicator acquisition module (2360) configured to acquire a feature map region presence indicator; An encoding control module (2370), based on the acquired feature map region presence indicator, When the feature map region presence indicator has a first value, encoding a region in the feature map into the bitstream, When the feature map region presence indicator has a second value, bypassing encoding a region in the feature map into the bitstream An encoding control module (2370) configured to determine the above, the feature map including a plurality of channels generated by detecting features in the image, and the region being a part of the channels. A device (2350) including the encoding control module (2370). **Claim 47** A device (2450) for encoding a feature map that is an output of a neural network for processing an image by the neural network in a bitstream, A feature map acquisition module (2460) configured to acquire the feature map determines whether to indicate side information regarding the feature map, and in the bitstream, either a third value and a side information indicator indicating the side information, or a side information indicator indicating the fourth value without the side information An encoding control module (2470) configured to indicate, wherein the side information corresponds to a region in the feature map, the feature map includes a plurality of channels generated by detecting features in the image, and the region is a part of the channels. A device (2450) including the encoding control module (2470). **Claim 48** A device for decoding a feature map input to a neural network for processing an image by the neural network based on a bitstream, the device including a processing circuit configured to perform the steps of the method according to any one of claims 1 to 31. **Claim 49** A device for encoding a feature map that is an output of a neural network for processing an image by the neural network in a bitstream, the device including a processing circuit configured to perform the steps of the method according to any one of claims 36 to 41.

Citation Information

Patent Citations

  • Decoder-side region of interest: video processing

    JP2010515300A

  • Image coding method and apparatus and image decoding method and apparatus

    JP2020191077A

  • Neural network suppression

    US20160358069A1

  • Receptive-Field-Conforming Convolutional Models for Video Coding

    US20200092552A1

  • Image encoding / decoding method and apparatus for signaling image feature information, and method for transmitting bitstream

    US20230085554A1