Method, apparatus, and system for encoding and decoding tensors

The hierarchical encoding and decoding of CNN tensors using a bottleneck encoder and decoder efficiently compresses data for distributed processing, addressing computational and bandwidth challenges in CNNs, ensuring effective edge device operation.

JP2025520260APending Publication Date: 2025-07-03CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024565238
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-08
Filing Date
2023-06-13
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Convolutional neural networks (CNNs) require high computational complexity and power consumption, exceeding the capabilities of edge devices, necessitating distributed processing across edge devices and cloud servers, which demands efficient tensor data compression to reduce bandwidth and processing load.

Method used

A method for encoding and decoding tensors with a hierarchical representation of feature maps, using a bottleneck encoder and decoder to reduce spatial dimensions, and a quantization and packing module to convert floating-point tensors to integer precision for efficient transmission and decoding.

Benefits of technology

This approach reduces computational and bandwidth requirements, allowing CNNs to function effectively on edge devices while maintaining task performance, particularly in tasks like object detection and segmentation, by minimizing data loss and preserving spatial details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520260000001_ABST
    Figure 2025520260000001_ABST
Patent Text Reader

Abstract

The present invention relates to an apparatus and a method for decoding at least a plurality of tensors forming a hierarchical representation of a feature map for a single frame from a bit stream. The method includes decoding a first information unit from the bit stream, decoding a second information unit from the bit stream, and determining a first plurality of tensors from the first information unit, wherein a feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from a feature map of another tensor of the first plurality of tensors. The method also includes determining a second plurality of tensors from the second information unit, wherein a feature map of at least one tensor of the second plurality of tensors has a different spatial resolution from a feature map of another tensor of the second plurality of tensors. A feature map of each tensor of the first plurality of tensors has a different spatial resolution from a feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to a hierarchical representation of a feature map for a single frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Reference to Related Applications This application claims the benefit of the filing date under 35 U.S.C. § 119 of Australian Patent Application No. 2022204911, filed on July 8, 2022, and is incorporated herein by reference in its entirety as if fully set forth herein.

[0002] The present invention generally relates to digital video signal processing, and more particularly, to methods, apparatuses, and systems for encoding and decoding tensors from convolutional neural networks. The present invention also relates to a computer program product including a computer-readable medium having recorded thereon a computer program for encoding and decoding tensors from a convolutional neural network using video compression techniques.

Background Art

[0003] Convolutional neural networks (CNNs) are a new technology that can handle use cases related to machine vision, such as object detection, instance segmentation, object tracking, human pose estimation, and action recognition. For CNN applications, it is possible to use "edge devices" with sensors and a certain degree of processing power and combine them with application servers as part of the "cloud". CNNs may require relatively high computational complexity, exceeding the normal acceptable range in terms of both the computational power and power consumption of edge devices. Executing CNNs in a distributed manner has emerged as one solution for running state-of-the-art edge networks using edge devices with limited capabilities. In other words, through distributed processing, by dispersing the processing between an edge device and an external processing means such as a cloud server, it is possible to provide the functions of state-of-the-art CNNs even with legacy edge devices.

[0004] CNNs typically include many layers, such as convolutional layers and fully connected layers, and data is passed from one layer to the next in the form of "tensors". Such compression is sometimes called "feature compression", and the intermediate tensor data is often referred to as the "features" of the input, such as an image frame or a video frame. The International Organization for Standardization / International Electrotechnical Commission Joint Technical Committee 1 / Subcommittee 29 / Working Group 2-8 (ISO / IEC JTC1 / SC29 / WG2-8), also known as the "Moving Picture Experts Group" (MPEG), is tasked with researching compression techniques for video. The "MPEG Technical Requirements" of WG2 established an "Ad Hoc Group on Video Compression for Machines" (VCM) and mandated the research of video compression for machine consumption and feature compression. The directive for feature compression is in the exploratory stage where the issuance of a "Call for Evidence" (CfE) is expected, and it solicits technologies that can significantly outperform the feature compression results achieved using state-of-the-art standardization techniques.

[0005] In a CNN, it is necessary to determine the weights of each layer during the learning stage. During this stage, a very large amount of learning data passes through the CNN, and the determined results are compared with the ground truth related to the learning data. A process for updating the network weights, such as stochastic gradient descent, is applied, and the network weights are iteratively improved until the network operates with the desired accuracy. If the convolutional stage has a "stride" greater than 1, the output tensor from the convolution has a lower spatial resolution than the corresponding input tensor. The pooling operation results in an output tensor with dimensions smaller than the input tensor. An example of a pooling operation is "max pooling" (or "max pool"), which reduces the spatial size of the output tensor compared to the input tensor. Max pooling divides the input tensor into groups of data samples (e.g., groups of 2×2 data samples) and selects the maximum value as the output for the corresponding value of the output tensor from each group to generate the output tensor. The process of running a CNN on an input and gradually transforming the input into an output is generally called "inference".

[0006] Generally, a tensor has four dimensions, namely, batch, channel, height, and width. The first dimension, 'batch', has a size of '1' when inferring video data, indicating that one frame passes through the CNN at a time. When training the network, the value of the batch dimension can be increased so that multiple frames pass through the network before the network weights are updated according to a given "batch size". A video of multiple frames can be passed as a single tensor with an increased batch dimension according to the number of frames in the given video. However, due to practical considerations regarding memory consumption and access, inference of video data is usually performed frame by frame. The 'channel' dimension indicates the number of simultaneous "feature maps" of a given tensor, and the height and width dimensions indicate the size of the feature map at a specific stage of the CNN. The number of channels varies depending on the CNN network architecture. The size of the feature map also varies depending on the subsampling that occurs in a specific network layer.

[0007] The input to the first layer of the CNN is a batch of one or more images, such as one image or a video frame, and is usually resized to fit the dimensions of the tensor input to the first layer. It is also possible to supply images or video frames in a batch of a size larger than 1. The dimensions of the tensor depend on the CNN architecture and generally have dimensions related to the width and height of the input, as well as an additional "channel" dimension.

[0008] Slicing the tensor based on the channel dimension, i.e., reducing the tensor to a set of two-dimensional arrays, results in a set of two-dimensional "feature maps". Each slice of the tensor has some relationship with the corresponding input image and captures characteristics such as various edge types, and is thus called a so-called "feature map". In layers further away from the input to the network, the characteristics become more abstract. The "task performance" of the CNN is measured by comparing the results of the CNN that performed the task using a specific input with the provided ground truth (generally considered to be the "correct" result prepared by humans).

[0009] Once the network topology is determined, as more training data becomes available, the weights of the network are updated over time. The overall complexity of a CNN tends to be relatively high, with a relatively large number of multiplication operations being performed, and a large number of intermediate tensors being written to and read from memory. Depending on the application, since the CNN is implemented entirely in the "cloud", high-cost processing capabilities are required. In other applications, the CNN is implemented on edge devices such as cameras and mobile phones, where flexibility is reduced but the processing load is distributed. In the new architecture, the network is split, with one part being executed on an edge device and the other on the cloud. Such a distributed network architecture is called "collaborative intelligence" and has advantages such as the partial results obtained from the first part of the network being reusable in multiple different second parts. The collaborative intelligence architecture brings about the need to efficiently compress tensor data for transmission over a network such as a WAN.

[0010] Video compression standards can be used for feature compression as described below. Various methods can be used to narrow or reduce the data presented for compression. However, some of the methods used to narrow or reduce the data presented for compression can result in a reduction in accuracy that is not suitable for some of the tasks implemented by a CNN.

[0011] Feature compression may benefit from existing video compression standards such as Versatile Video Coding (VVC) developed by the Joint Video Expert Team (JVET). VVC is expected to meet the continuous demand for higher compression performance, especially with the increasing capabilities of video formats (such as higher resolution and higher frame rate), and the growing market need for service provision via WAN with relatively high bandwidth costs. VVC is implementable with modern silicon processes and provides an acceptable trade-off between the achieved performance and the implementation cost. The implementation cost can be considered in terms of, for example, one or more of silicon area, CPU processor load, memory utilization, and bandwidth. Part of the versatility of the VVC standard lies in the breadth of options for tools available for compressing video data and the breadth of applications for which VVC is suitable. Other video compression standards such as HEVC (High Efficiency Video Coding) and AV-1 can also be used for feature compression applications.

[0012] Video data includes a sequence of frames of image data, with each frame including one or more color channels. Generally, one primary color channel and two secondary color channels are required. The primary color channel is generally called the "luma" channel, and the secondary color channels are generally called the "chroma" channels. Video data is usually displayed in the RGB (red-green-blue) color space, but there is a high correlation between the three components of this color space. The representation of video data seen by encoders and decoders often uses a color space such as YCbCr. YCbCr concentrates the luminance mapped to "luma" according to a transfer function in the Y (primary) channel and the chroma in the Cb and Cr (secondary) channels. Because of the uncorrelated YCbCr signals, the statistics of the luma channel are significantly different from those of the chroma channels. The main difference is that after quantization, there are relatively fewer significant coefficients in a block in the chroma channels compared to the coefficients of the corresponding luma channel block. Additionally, the Cb and Cr channels may be spatially sampled (subsampled) at a lower rate than the luma channel. The 4:2:0 chroma format is commonly used in "consumer-oriented" applications such as Internet video streaming, broadcast television, and storage on Blu-Ray (trademark) discs. When only luma samples are present, the resulting monochrome frame is said to use the "4:0:0 chroma format".

[0013] The VVC standard defines a "block-based" architecture, where a frame is first divided into a square array of regions known as "Coding Tree Units" (CTUs). A CTU generally occupies a relatively large area, such as 128×128 luma samples. When using the VVC standard, other CTU sizes such as 32×32 and 64×64 are also considered. However, the CTUs at the right end and bottom end of each frame may have an implicit division where the CBs are surely left within the frame, and their area may become smaller. Each CTU is associated with an "encoding tree" ("shared tree") for both the luma channel and the chroma channel, or separate trees for each of the luma channel and the chroma channel. The encoding tree is defined to decompose the area of the CTU into a set of blocks also called "Coded Blocks" (CBs). When a shared tree is used, a single encoding tree specifies blocks for both the luma channel and the chroma channel. In this case, a collection of co-located coded blocks is called a "Coding Unit" (CU) (i.e., each CU has coded blocks for each color channel). The CBs are processed for encoding or decoding in a specific order. As a result of using the 4:2:0 chroma format, a CTU with a luma encoding tree for a 128×128 luma sample area has a corresponding chroma encoding tree for a 64×64 chroma sample area juxtaposed with the 128×128 luma sample area. When a single encoding tree is used for both the luma channel and the chroma channel, a collection of collocated blocks for a given area is generally called a "unit", similar to, for example, the above-mentioned CUs, as well as "Prediction Units" (PUs) and "Transformation Units" (TUs). A single tree for a CU spanning the color channels of video data in 4:2:0 chroma format results in chroma blocks that are half the width and height of the corresponding luma blocks. When separate encoding trees are used for a given area, in addition to the above-mentioned CBs, "Prediction Blocks" (PBs) and "Transformation Blocks" (TBs) are used.

[0014] Regardless of the above distinction between "units" and "blocks", the term "block" can be used as a general term for an area or region of a frame to which an operation is applied to all color channels.

[0015] For each CU, a prediction unit (PU) of the content (sample value) of the corresponding area of the frame data is generated ("prediction unit"). Further, an expression of the difference (or "spatial area" residual) between the predicted value and the content of the area as seen at the input to the coder is formed. The differences for each color channel are transformed, encoded as a sequence of residual coefficients, and can form one or more TUs for a given CU. The transformation applied may be a discrete cosine transform (DCT) or other transformation applied to a block of each residual value. The transformation is applied separately (i.e., the 2D transformation is performed in two passes). The block is first transformed by applying a 1D transformation to each row of samples within the block. Next, the partial results are transformed by applying a 1D transformation to each column of the partial results, generating a final block of transformation coefficients that substantially decorrelates the residual samples. The VVC standard supports various sizes of transformations, including the transformation of rectangular blocks where the dimension of each side is a power of 2. The transformation coefficients are quantized for entropy coding into the bitstream.

[0016] The features of VVC are intra-frame prediction and inter-frame prediction. In intra-frame prediction, previously processed samples within the frame are used to generate a prediction of the current block of data samples within the frame. Inter-frame prediction involves generating a prediction of the current block of samples within the frame using a block of samples obtained from a previously decoded frame. The block of samples obtained from the previously decoded frame is offset from the spatial position of the current block according to a motion vector, and filtering is often applied. Intra-frame prediction blocks can be (i) a uniform sample value (“DC intra prediction”), (ii) a plane with an offset and horizontal and vertical gradients (“plane intra prediction”), (iii) a group of blocks with adjacent samples applied in a specific direction (“angular intra prediction”), or (iv) the result of a matrix multiplication using adjacent samples and selected matrix coefficients. Further discrepancies between the predicted block and the corresponding input samples can be corrected to some extent by encoding the “residual” in the bitstream. The residual is generally transformed from the spatial domain to the frequency domain to form the residual coefficients in the “primary transform” domain. The residual coefficients can be further transformed by the application of a “secondary transform” to generate the residual coefficients in the “secondary transform domain”. The residual coefficients are quantized according to a quantization parameter, as a result of which the accuracy of the reconstruction of the samples generated at the decoder is reduced, but the bitrate of the bitstream is reduced. A sequence of pictures may be encoded according to a specified structure of pictures using intra prediction and pictures using intra or inter prediction, and a specified dependency on preceding pictures in the encoding order, which may be different from the display order or the delivery order. In the “random access” configuration, periodic intra pictures occur, forming an entry point where the decoder starts decoding the bitstream. Other pictures in the random access configuration generally use inter prediction to predict the content from pictures before and after the current picture in display order or delivery order according to a hierarchical structure of a specified depth.To predict the current picture, in order to use the pictures following the current picture in the display order, some picture buffering and delay are required between the decoding of a predetermined picture and the display (and removal from the buffer) of the predetermined picture.

Summary of the Invention

[0017] An object of the present invention is to substantially overcome or at least improve one or more drawbacks of existing devices.

[0018] One aspect of the present disclosure provides a method for decoding at least a plurality of tensors forming a hierarchical representation of a feature map for a single frame from a bitstream, the method comprising decoding a first information unit from the bitstream; decoding a second information unit from the bitstream; determining a first plurality of tensors from the first information unit, wherein a feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from a feature map of another tensor of the first plurality of tensors; determining a second plurality of tensors from the second information unit, wherein a feature map of at least one tensor of the second plurality of tensors has a different spatial resolution from a feature map of another tensor of the second plurality of tensors; and each feature map of each tensor of the first plurality of tensors has a different spatial resolution from each feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame.

[0019] Another aspect of the present disclosure provides a method for encoding at least a plurality of tensors into a bitstream, the plurality of tensors forming a hierarchical representation of feature maps for a single frame, the method comprising using a convolution operation to determine a first information unit from a first plurality of tensors, wherein a feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from feature maps of other tensors of the first plurality of tensors, said using; using a convolution operation to determine a second information unit from a second plurality of tensors, wherein a feature map of at least one tensor of the second plurality of tensors has a different spatial resolution from feature maps of other tensors of the second plurality of tensors, and a feature map of each tensor of the first plurality of tensors has a different spatial resolution from a feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of feature maps for the single frame, said using; encoding the first information unit into the bitstream; and encoding the second information unit into the bitstream.

[0020] Another aspect of the present disclosure provides a decoder that decodes at least a plurality of tensors forming a hierarchical representation of a feature map for a single frame from a bitstream, the decoder decoding a first information unit from the bitstream, decoding a second information unit from the bitstream, determining a first plurality of tensors from the first information unit, a feature map of at least one tensor of the first plurality of tensors having a different spatial resolution from a feature map of another tensor of the first plurality of tensors, determining a second plurality of tensors from the second information unit, a feature map of at least one tensor of the second plurality of tensors having a different spatial resolution from a feature map of another tensor of the second plurality of tensors, configured such that a feature map of each tensor of the first plurality of tensors has a different spatial resolution from a feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame.

[0021] Another aspect of the present disclosure provides an encoder that encodes at least a plurality of tensors into a bitstream, the plurality of tensors forming a hierarchical representation of feature maps for a single frame, the encoder using a convolution operation to determine a first information unit from a first plurality of tensors, a feature map of at least one tensor of the first plurality of tensors having a different spatial resolution from feature maps of other tensors of the first plurality of tensors, using a convolution operation to determine a second information unit from a second plurality of tensors, a feature map of at least one tensor of the second plurality of tensors having a different spatial resolution from feature maps of other tensors of the second plurality of tensors, a feature map of each tensor of the first plurality of tensors having a different spatial resolution from a feature map of each tensor of the second plurality of tensors, the plurality of tensors of the first plurality of tensors and the second plurality of tensors corresponding to the hierarchical representation of the feature maps for the single frame, encoding the first information unit into the bitstream, and encoding the second information unit into the bitstream, and is configured to.

[0022] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing a program for executing a method of decoding at least a plurality of tensors forming a hierarchical representation of a feature map for a single frame, the method comprising: decoding a first information unit from the bitstream; decoding a second information unit from the bitstream; determining a first plurality of tensors from the first information unit, wherein a feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from a feature map of another tensor of the first plurality of tensors; determining a second plurality of tensors from the second information unit, wherein a feature map of at least one tensor of the second plurality of tensors has a different spatial resolution from a feature map of another tensor of the second plurality of tensors; wherein a feature map of each tensor of the first plurality of tensors has a different spatial resolution from a feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame.

[0023] Another aspect of the present disclosure provides a system including a memory and a processor, the processor being configured to execute code stored on the memory to implement a method of decoding at least a plurality of tensors from a bitstream that form a hierarchical representation of a feature map for a single frame, the method comprising decoding a first information unit from the bitstream, decoding a second information unit from the bitstream, determining a first plurality of tensors from the first information unit, wherein a feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from a feature map of another tensor of the first plurality of tensors, determining a second plurality of tensors from the second information unit, wherein a feature map of at least one tensor of the second plurality of tensors has a different spatial resolution from a feature map of another tensor of the second plurality of tensors, each feature map of each tensor of the first plurality of tensors having a different spatial resolution from each feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors corresponding to the hierarchical representation of the feature map for the single frame.

[0024] Other aspects are also disclosed.

Brief Description of the Drawings

[0025] Next, at least one embodiment of the present invention will be described with reference to the following drawings and appendices:

[0026]

Figure 1

[0027]

Figure 2A

Figure 2B

[0028]

Figure 3A

[0029]

Figure 3B

[0030]

Figure 3C

[0031]

Figure 3D

[0032]

Figure 4

[0033]

Figure 5

[0034]

Figure 6

[0035]

Figure 7

[0036]

Figure 8

[0037]

Figure 9

[0038]

Figure 10A

[0039]

Figure 10B

[0040]

Figure 11

[0041]

Figure 12A

[0042]

Figure 12B

[0043]

Figure 12C

[0044]

Figure 13

[0045]

Figure 14

DETAILED DESCRIPTION OF THE INVENTION

[0046] If, in any one or more of the accompanying drawings, steps and / or features having the same reference numerals are referred to, those steps and / or features have the same function or operation in this specification, unless the contrary is intended.

[0047] The distributed machine task system can include edge devices such as network cameras and smartphones that generate intermediate compressed data. Also, the distributed machine task system can include end devices such as server farm-based (``cloud'') applications that operate on the intermediate compressed data to generate some task result. Further, the functionality of the edge device can be implemented in the cloud, and the intermediate compressed data can potentially be saved for later processing for a plurality of different tasks as needed. Examples of machine tasks include object detection and instance segmentation, both of which generate task results measured as the ``mean average precision'' (mAP) of detections that exceed an intersection-over-union (IoU) threshold such as 0.5. Another example machine task is object tracking, which has as a typical task result the mean object tracking accuracy (MOTA) score.

[0048] A convenient form of the intermediate compressed data is a compressed video bitstream because high-performance compression standards and their implementations are available. Video compression standards typically operate on integer samples of a given bit depth, such as 10 bits, arranged in a planar array. Color video has three planar arrays corresponding to color components Y, Cb, Cr, or R, G, B, depending on the application. CNNs typically operate on floating-point data in the form of tensors. Tensors generally have far fewer spatial dimensions than the input video data on which the CNN operates but have more channels than the typical 3 channels of color video data.

[0049] Tensors usually have the following dimensions: frame, channel, height, and width. For example, a tensor of dimension [1, 256, 76, 136] can be said to contain 256 feature maps of size 136×76 each. In the case of video data, inference is usually performed frame by frame rather than using a tensor that includes multiple frames at once.

[0050] VVC supports splitting a picture into multiple sub-pictures, and each sub-picture is encoded independently and decoded independently. In one approach, each sub-picture is encoded as one "slice", i.e., a continuous sequence of encoded CTUs. A "tile" mechanism that divides a picture into multiple independently decodable regions is also available. Sub-pictures can be specified in a somewhat flexible way, and various rectangular sets of CTUs are encoded as respective sub-pictures. By flexibly defining the dimensions of sub-pictures, it is possible to efficiently hold data of types that require different regions in one picture and avoid large "unused" regions, i.e., regions of frames that are not used for the reconstruction of tensor data.

[0051] Figure 1 is a schematic block diagram showing the functional modules of the distributed machine task system 100. The concept of distributing machine tasks across multiple systems is sometimes referred to as "Collaborative Intelligence" (CI). System 100 can be used to implement a method for decorrelating, packing, and quantizing a feature map into a planar frame in order to encode and decode the feature map from the encoded data. This method is implemented such that the associated overhead data is not overly burdensome, the task performance of the decoded feature map is robust to changes in the bitrate of the bitstream, and the quantized representation of the tensor does not consume unnecessary bits that do not provide a commensurate benefit in terms of task performance.

[0052] System 100 includes a source device 110 for generating encoded tensor data 115 in the form of an encoded video bitstream 121 from a CNN backbone 114. System 100 also includes a destination device 140 for decoding tensor data in the form of an encoded video bitstream 143. A communication channel 130 is used to communicate the encoded video bitstream 121 from the source device 110 to the destination device 140. In some arrangements, the source device 110 and the destination device 140 can each be composed of a respective mobile phone terminal (e.g., a "smartphone") or a network camera and cloud application. The communication channel 130 can be a wired connection such as Ethernet, or a wireless connection such as WiFi or 5G, and includes a connection that traverses a wide area network (WAN) or a connection that traverses an ad hoc connection. Further, the source device 110 and the destination device 140 can constitute an application in which the encoded video data is captured on some computer-readable storage medium such as a file server or a hard disk drive of a memory.

[0053] As shown in FIG. 1, the source device 110 includes a video source 112, a CNN backbone 114, a bottleneck encoder 116, a quantization and packing module 118, a feature map encoder 120, and a transmitter 122. The video source 112 typically consists of a source of captured video frame data (shown as 113) such as an image capture sensor, a previously captured video sequence stored on a non-transitory recording medium, or a video feed from a remote image capture sensor. The video source 112 may also be the output of a computer graphics card that displays the video output of an operating system and various applications running on a computing device (e.g., a tablet computer). Examples of source devices 110 that may include an image capture sensor as the video source 112 include smartphones, video camcorders, professional video cameras, and network video cameras. The system 100 uses additional network layers that "bottleneck," i.e., limit the dimensionality of the tensor on the encoding side and restore the dimensionality of the tensor on the decoder side, to reduce the dimensionality of the tensor at the interface between the first network portion and the second network portion. The multi-scale representations generated by the Feature Pyramid Network (FPN) are "fused" into a single tensor using an approach named "Multi-Scale Feature Compression" (MSFC). MSFC is typically used to integrate all FPN layers into a single tensor. Merging all FPN layers into one tensor is done at the expense of the spatial details of the less decomposed (larger) layers of the FPN. When the spatial details are lost, the accuracy of some of the tasks or operations implemented by the system 100 may be unacceptably reduced.

[0054] Instead of integrating all the layers of the FPN into one tensor, the tensors of the FPN are grouped and the MSFC technique is applied individually. By applying the MFSC technique individually, cross-layer fusion becomes possible to some extent without degrading the spatial detail too much. In tasks that require more retention of spatial detail, such as instance segmentation, using individual MSFC techniques results in a higher mAP than integrating all FPN layers into a single tensor at a low spatial resolution.

[0055] The CNN backbone 114 receives the video frame data 113, executes specific layers of the entire CNN, such as the layers corresponding to the "backbone" of the CNN, and outputs a tensor 115. The backbone layer of the CNN may, for example, generate a plurality of tensors corresponding to different spatial scales of the input image represented by the video frame data 113 as outputs, and may also be called a "Feature Pyramid Network" (FPN) architecture. The tensors obtained from the FPN backbone form a hierarchical representation of the frame data 113 that includes the data of the feature maps. Each successive layer of the hierarchical representation has a width and height that are half of the previous layer. The later layers generated deeper in the backbone network tend to contain features with a more abstract representation of the frame data 113. The less decomposed layers generated at an earlier stage of the backbone network tend to contain features that represent less abstract features of the frame data 113, such as various geometric characteristics like edges at various angles. When the "YOLOv3" network is executed by the system 100, the FPN may result in outputting three tensors corresponding to three layers as the tensor 115 from the backbone 114 while varying the spatial resolution and the number of channels. When the system 100 is executing a network such as "Faster RCNN X101 - FPN" or "Mask RCNN X101 - FPN", the tensor 115 includes tensors for four layers P2 to P5. The bottleneck encoder 116 receives the tensor 115. The bottleneck encoder 116 compresses one or more internal layers of the overall CNN. The internal layers of the overall CNN are trained to convert to a lower number of channels and a smaller spatial resolution than required by the tensor 115, and provide the output of the CNN backbone 114 compressed or shrunk by the bottleneck encoder 116 using a set of neural network layers. The bottleneck encoder 116 outputs a bottleneck tensor 117. The bottleneck tensor 117 is passed to the quantization and packing module 118.Each feature map of the bottleneck tensor 117 is quantized from floating - point to integer precision, packed into a monochrome frame by module 118, and a frame 119 is generated. The frame 119 is encoded by the feature map encoder 120 to generate a bitstream 121. The bitstream 121 is supplied to the transmitter 122 for transmission via the communication channel 130, or the bitstream 121 is written to the storage 132 for later use.

[0056] The source device 110 supports a specific network for the CNN backbone 114. However, the destination device 140 can use one of a plurality of networks for the head CNN 150. In this way, the partially processed data in the form of the packed feature maps can be saved for later use when performing various tasks without having to repeatedly execute the operations of the CNN backbone 114.

[0057] The bitstream 121 is transmitted by the transmitter 122 via the communication channel 130 as encoded video data (or "encoded video information"). In some implementations, the bitstream 121 can be stored in the storage 132 until it is later transmitted via the communication channel 130 (or instead of being transmitted via the communication channel 130), and the storage 132 is a non - volatile storage device such as a "flash" memory or a hard disk drive. For example, the encoded video data may be provided to a customer on demand via a wide - area network (WAN) for a video analysis application.

[0058] The destination device 140 includes a receiver 142, a feature map decoder 144, an inverse quantization and unpacking module 146, a bottleneck decoder 148, a CNN head 150, and a CNN task result buffer 152. The receiver 142 receives the encoded video data from the communication channel 130 and passes the video bitstream 143 to the feature map decoder 144. The feature map decoder 144 operates to decode the feature map and output the decoded frame 145. The decoded frame 145 is passed to the unpacking and inverse quantization module 146. The module 146 unpacks and inverse quantizes the tensor of the frame 145 to generate a dequantized tensor that is output as the decoded bottleneck tensor 147. The decoded bottleneck tensor 147 is supplied to the bottleneck decoder 148. The bottleneck decoder 148 performs the inverse operation of the bottleneck encoder 116 to generate the extracted tensor 149. The extracted tensor 149 is passed to the CNN head 150. The CNN head 150 executes the subsequent layers of the task started by the CNN backbone 114 to generate a task result 151, which is stored in the task result buffer 152. The content of the task result buffer 152 may be presented to the user, for example, via a graphical user interface, or provided to an analysis application where some action is determined based on the task result, which may include a summary-level presentation of the aggregated task results to the user. It is also possible to embody the respective functions of the source device 110 and the destination device 140 in a single device, examples of which include a mobile phone handset, a tablet computer, and a cloud application.

[0059] Regardless of the exemplary devices described above, each of the source device 110 and the destination device 140 may typically be configured within a general-purpose computing system through a combination of hardware and software components. FIG. 2A shows such a computer system 200, including an input device such as a computer module 201, a keyboard 202, a mouse pointer device 203, a scanner 226, a camera 227 which may be configured as a video source 112, and a microphone 280, and an output device including a printer 215, a display device 214 which may be configured as a display device for presenting the task result 151, and a speaker 217. The external modem transceiver device 216 may be used by the computer module 201 to communicate with the communication network 220 via the connection 221. The communication network 220, which may represent the communication channel 130, may be a wide area network (WAN) such as the Internet, a cellular telecommunications network, or a private WAN. If the connection 221 is a telephone line, the modem 216 may be a conventional "dial-up" modem. Alternatively, if the connection 221 is a high-capacity (e.g., cable or optical) connection, the modem 216 may be a broadband modem. A wireless modem may also be used for a wireless connection to the communication network 220. The transceiver device 216 can provide the functions of the transmitter 122 and the receiver 142, and the communication channel 130 can be embodied in the connection 221.

[0060] Computer module 201 typically includes at least one processor unit 205 and a memory unit 206. For example, the memory unit 206 can have a semiconductor random access memory (RAM) and a semiconductor read only memory (ROM). The computer module 201 also includes a number of input / output (I / O) interfaces, including an audio / video interface 207 coupled to a video display 214, a speaker 217, and a microphone 280, an I / O interface 213 coupled to a keyboard 202, a mouse 203, a scanner 226, a camera 227, and optionally a joystick or other human interface device (not shown), and an interface 208 for an external modem 216 and a printer 215. Signals from the audio / video interface 207 to the computer monitor 214 are generally outputs of a computer graphics card. In some implementations, the modem 216 can be incorporated within the computer module 201, for example within the interface 208. The computer module 201 also has a local network interface 211 that enables the coupling of the computer system 200 via a connection 223 to a local area communication network 222 known as a local area network (LAN). As shown in Figure 2A, the local communication network 222 can also be coupled to a wide network 220 via a connection 224, which typically includes a so-called "firewall" device or a device with similar functionality. The local network interface 211 can be configured with an Ethernet (registered trademark) circuit card, a Bluetooth (registered trademark) wireless arrangement, or an IEEE802.11 wireless arrangement, although a number of other types of interfaces can be implemented for the interface 211. The local network interface 211 can also provide the functions of a transmitter 122 and a receiver 142, and the communication channel 130 can also be embodied in the local communication network 222.

[0061] The I / O interfaces 208 and 213 can provide either or both of serial and parallel connections, the former typically being implemented in accordance with the Universal Serial Bus (USB) standard and having a corresponding USB connector (not shown). A storage device 209 is provided, typically including a hard disk drive (HDD) 210. Other storage devices such as a floppy disk drive or a magnetic tape drive (not shown) can also be used. The optical disk drive 212 is typically provided to function as a non-volatile source of data. For example, portable memory devices such as optical disks (e.g., CD-ROM, DVD, Blu-ray Disc (trademark)), USB-RAM, portable, external hard drives, and floppy disks can be used as appropriate data sources for the computer system 200. Typically, any of the HDD 210, optical drive 212, networks 220 and 222 can also be configured to operate as a video source 112 or as a destination for decoded video data stored for playback via the display 214. The source device 110 and the destination device 140 of the system 100 can be embodied in the computer system 200.

[0062] The components 205 to 213 of the computer module 201 typically communicate in a manner that provides a conventional mode of operation of the computer system 200 known to those skilled in the art via an interconnected bus 204. For example, the processor 205 is coupled to the system bus 204 using connection 218. Similarly, the memory 206 and the optical disk drive 212 are coupled to the system bus 204 by connection 219. Examples of computers that can implement the described arrangement include IBM-PCs and compatibles, Sun's SPARCstation, Apple's Mac (trademark), or similar computer systems.

[0063] When appropriate or desired, source device 110, destination device 140, and the methods described below can be implemented using computer system 200. In particular, source device 110, destination device 140, and the methods described can be implemented as one or more software application programs 233 executable within computer system 200. The steps of source device 110, destination device 140, and the methods described are realized by instructions 231 (see FIG. 2B) within software 233 executed within computer system 200. Software instructions 231 can each be formed as one or more code modules for performing one or more particular tasks. The software can also be split into two separate parts, in which case the first part and corresponding code modules perform the methods described, and the second part and corresponding code modules manage the user interface between the first part and the user.

[0064] The software can be stored on a computer-readable medium including, for example, the storage devices described below. The software is loaded from the computer-readable medium into computer system 200 and executed by computer system 200. A computer-readable medium having such software or a computer program recorded thereon is a computer program product. The use of the computer program product in computer system 200 preferably provides an advantageous apparatus for implementing source device 110, destination device 140, and the methods described.

[0065] Software 233 is typically stored in HDD 210 or memory 206. The software is loaded from the computer-readable medium into computer system 200 and executed by computer system 200. Thus, for example, software 233 can be stored on an optically readable disk storage medium (e.g., CD-ROM) 225 read by optical disk drive 212.

[0066] In one embodiment, the application program 233 may be encoded on one or more CD-ROMs 225 and supplied to the user, and may be read via the corresponding drive 212, or may be read by the user from the network 220 or 222. Further, the software can also be loaded into the computer system 200 from other computer-readable media. A computer-readable storage medium refers to any non-transitory tangible storage medium that provides instructions and / or data recorded in the computer system 200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray Discs (registered trademarks), hard disk drives, ROMs or integrated circuits, USB memories, magneto-optical disks, or computer-readable cards such as PCMCIA cards, regardless of whether such devices are internal or external to the computer module 201. Examples of transitory or non-tangible computer-readable transmission media that can also participate in providing software, application programs, instructions, and / or video data or encoded video data to the computer module 201 include wireless or infrared transmission channels, as well as network connections to another computer or networked device, and the Internet or intranet including information recorded on email transmissions and websites.

[0067] The second part of the application program 233 and the corresponding code module described above can be executed to implement one or more graphical user interfaces (GUIs) that are rendered on the display 214 or otherwise represented. Typically through the operation of the keyboard 202 and the mouse 203, the user of the computer system 200 and the application can operate the interface in a functionally adaptable manner to provide control commands and / or inputs to the application(s) associated with the GUI(s). Other forms of functionally adaptable user interfaces can also be implemented, such as audio prompts output via the loudspeaker 217 and an audio interface that utilizes the user's voice commands input via the microphone 280.

[0068] FIG. 2B is a detailed schematic block diagram of the processor 205 and the “memory” 234. The memory 234 represents the logical aggregation of all memory modules (including the storage device 209 and the semiconductor memory 206) accessible by the computer module 201 of FIG. 2A.

[0069] When the computer module 201 is first powered on, a power-on self-test (POST) program 250 is executed. The POST program 250 is typically stored in the ROM 249 of the semiconductor memory 206 in FIG. 2A. A hardware device such as the ROM 249 that stores software is sometimes called firmware. The POST program 250 inspects the hardware within the computer module 201 to ensure proper functionality, and typically inspects the processor 205, the memory 234 (209, 206), and the basic input / output system software (BIOS) module 251 (which is also typically stored in the ROM 249) to ensure correct operation. When the POST program 250 is executed successfully, the BIOS 251 activates the hard disk drive 210 in FIG. 2A. Activation of the hard disk drive 210 causes the bootstrap loader program 252 resident on the hard disk drive 210 to be executed via the processor 205. Thereby, the operating system 253 is loaded into the RAM memory 206 and the operating system 253 begins to operate. The operating system 253 is a system-level application executable by the processor 205 and performs various high-level functions including processor management, memory management, device management, storage management, software application interface, and general-purpose user interface.

[0070] The operating system 253 manages the memory 234 (209, 206) such that each process or application running on the computer module 201 has sufficient memory to execute without colliding with the memory allocated to other processes. Further, the different types of memory available in the computer system 200 of FIG. 2A need to be used appropriately so that each process can execute effectively. Thus, the aggregated memory 234 is not for explaining how specific segments of memory are allocated (unless otherwise specified), but rather for providing a general view of the memory accessible by the computer system 200 and how such memory is used.

[0071] As shown in FIG. 2B, the processor 205 includes a number of functional modules including a control unit 239, an arithmetic logic unit (ALU) 240, and a local or internal memory 248 sometimes referred to as a cache memory. The cache memory 248 typically includes a number of storage registers 244-246 within a register section. One or more internal buses 241 functionally interconnect these functional modules. The processor 205 also typically has one or more interfaces 242 for communicating with external devices via the system bus 204 using the connection 218. The memory 234 is coupled to the bus 204 using the connection 219.

[0072] The application program 233 includes a series of instructions 231 including conditional branch instructions and loop instructions. The program 233 can also include data 232 used for the execution of the program 233. The instructions 231 and the data 232 are stored in memory locations 228, 229, 230 and 235, 236, 237 respectively. Depending on the relative sizes of the instructions 231 and the memory locations 228 - 230, a particular instruction may be stored in a single memory location as depicted by the instruction shown at memory location 230. Alternatively, as shown by the instruction segments shown at memory locations 228 and 229, the instructions may be split into several parts, each stored in a separate memory location.

[0073] Generally, the processor 205 is provided with an instruction set to be executed therein. The processor 205 waits for subsequent inputs, and in response thereto, the processor 205 reacts by executing another instruction set. Each input may be provided from one or more of the input devices 202, 203, data generated by one or more of the input devices 202, 203, data received from an external source via one of the networks 220, 202, data retrieved from one of the storage devices 206, 209, or data retrieved from a storage medium 225 inserted into the corresponding reading device 212, all of which are depicted in FIG. 2A. The execution of a set of instructions may, in some cases, result in the output of data. The execution may also involve the storage of data or variables in the memory 234.

[0074] The bottleneck encoder 116, the bottleneck decoder 148, and the methods described can use input variables 254, which are stored in the memory 234 at the corresponding memory locations 255, 256, 257. The bottleneck encoder 116, the bottleneck decoder 148, and the methods described generate output variables 261, which are stored in the corresponding memory locations 262, 263, 264 within the memory 234. Intermediate variables 258 may be stored in the memory locations 259, 260, 266, 267.

[0075] Referring to the processor 205 in FIG. 2B, the registers 244, 245, 246, the arithmetic logic unit (ALU) 240, and the control unit 239 cooperate to execute a sequence of micro-operations necessary to perform a "fetch, decode, execute" cycle for each instruction in the instruction set that makes up the program 233. Each fetch, decode, execute cycle includes the following: A fetch operation to fetch or read the instruction 231 from the memory locations 228, 229, 230; A decode operation in which the control unit 239 determines which instruction was fetched; An execute operation in which the control unit 239 and / or the ALU 240 execute the instruction.

[0076] Thereafter, a fetch, decode, execute cycle for the next instruction is further executed. Similarly, a store cycle may be executed in which the control unit 239 stores or writes a value to the memory location 232.

[0077] Each step or sub-process in the methods of FIGS. 6 and 11, which will be described hereinafter, is associated with one or more segments of the program 233 and typically the register sections 244, 245, 247, the ALU 240, and the control section 239 within the processor 205 cooperate to execute the fetch, decode, and execute cycles for all instructions in the instruction set for the segment of the program 233 of interest.

[0078] FIG. 3A is a schematic block diagram showing the functional modules of the backbone portion 310 of a CNN that can function as the CNN backbone 114. The backbone portion 114 may also be referred to as "DarkNet-53" and forms the backbone of the "YOLOv3" object detection network. Different backbones are possible, and as a result, the number and dimensions of the layers of the tensor 115 for each frame are different.

[0079] As shown in FIG. 3A, video data 113 is passed to a resizer module 304. The resizer module 314 resizes the frame to a resolution suitable for processing by the CNN backbone 310 and generates resized frame data 312. If the resolution of the frame data 113 is already suitable for the CNN backbone 310, the operation of the resizer module 304 is not necessary. The resized frame data 312 is passed to a convolutional batch normalization leaky rectified linear (CBL) module 314, which generates a tensor 316. CBL 314 includes modules as described with reference to CBL module 360 as shown in FIG. 3D.

[0080] Referring to FIG. 3D, the CBL module 360 receives a tensor 361 as input. The tensor 361 is passed to a convolutional layer 362, which generates a tensor 363. When the convolutional layer 362 has a stride of 1, padding is set to k samples, and has a convolutional kernel of size 2k + 1, the tensor 363 has the same spatial dimensions as the tensor 361. When the convolutional layer 362 has a larger stride, such as 2, the tensor 363 has smaller spatial dimensions compared to the tensor 361, for example, half the size with a stride of 2. Regardless of the stride, the size of the channel dimension of the tensor 363 may be different compared to the channel dimension of the tensor 361 of a particular CBL block. The tensor 363 is passed to a batch normalization module 364 that outputs a tensor 365. The batch normalization module 364 normalizes the input tensor 363 and applies scaling coefficients and offset values to generate the output tensor 365. The scaling coefficients and offset values are derived from the learning process. The tensor 365 is passed to a leaky rectified linear activation (“LeakyReLU”) module 366, which generates a tensor 367. Module 366 provides a “leaky” activation function where positive values of the tensor pass through and negative values are significantly reduced in magnitude, for example, to 0.1 times the previous value.

[0081] Returning to FIG. 3A, the tensor 316 is passed from the CBL block 314 to the residual block 11 module 320. The module 320 includes a sequential connection of three residual blocks, each containing 1, 2, and 8 residual units internally.

[0082] Regarding the residual blocks as present in the module 320, it will be described with reference to the ResBlock340 as shown in FIG. 3B. The ResBlock340 receives the tensor 341. The tensor is zero-padded by the zero-padding module 342 to generate the tensor 343. The tensor 343 is passed to the CBL module 344 to generate the tensor 345. The tensor 345 is passed to the residual unit 346 where the residual block 340 includes a series of connected residual units. The last residual unit of the residual unit 346 outputs the tensor 347.

[0083] A residual unit such as the unit 346 is described with reference to the ResUnit350 as shown in FIG. 3C. The ResUnit350 receives the tensor 351 as input. The tensor 351 is passed to the CBL module 352 to generate the tensor 353. The tensor 353 is passed to the second CBL unit 354 to generate the tensor 355. The addition module 356 sums the tensor 355 with the tensor 351 to generate the tensor 357. Since the input tensor 351 substantially affects the output tensor 357, the addition module 356 is also called a "shortcut". In an untrained network, the ResUnit350 passes the tensor through. When training is executed, the CBL modules 352 and 354 deviate the tensor 357 from the tensor 351 according to the training data and the ground truth data.

[0084] The Res11 module 320 outputs a tensor 322. The tensor 322 is output as one of the layers from the backbone module 310 and is also provided to the Res8 module 324. The Res8 module 324 is a residual block (i.e., 340) and includes eight residual units (i.e., 350). The Res8 module 324 generates a tensor 326. The tensor 326 is passed to the Res4 module 328 and is output as one of the layers from the backbone module 310. The Res4 module is a residual block (i.e., 340) and includes four residual units (i.e., 350). The Res4 module 324 generates a tensor 329. The tensor 329 is output as one of the layers from the backbone module 310. The layer tensors 322, 326, and 329 are collectively output as tensor 115. The backbone CNN 310 takes a video frame with a resolution of 1088×608 as input and can generate three tensors corresponding to three layers in the following dimensions: [1,256,76,136], [1,512,38,68], [1,1024,19,34]. Another example of the three tensors corresponding to the three layers is [1,512,34,19], [1,256,68,38], [1,128,136,76], which are separated by the 75th network layer, 90th network layer, and 105th network layer of the CNN 310, respectively. Each tensor can have a different resolution from the next tensor. The resolution of each tensor may double in height and width between the respective tensors. When forming the output tensor 115, 322, 326, and 329 provide a hierarchical representation of the frame data including the data of the feature map for encoding into a bitstream. The separation point depends on the CNN 310.

[0085] FIG. 4 is a schematic block diagram showing functional modules of an alternative backbone part 400 of a CNN that can function as the CNN backbone 114. The backbone part 400 implements a residual network with a feature pyramid network (“ResNet FPN”) and serves as an alternative to the CNN backbone 114. Frame data 113 is input and passes through a stem network 408, a res2 module 412, a res3 module 416, a res4 module 420, a res5 module 424, and a max pooling module 428 via tensors 409, 413, 417, 425, and the max pooling module 428 generates a P6 tensor 429 as an output.

[0086] The stem network 408 includes a 7x7 convolution of stride 2(2) and a max pooling operation. The res2 module 412, res3 module 416, res4 module 420, and res5 module 424 perform a convolution operation and a LeakyReLU activation. Each module 412, 416, 420, 424 also halves the resolution of the processed tensor with a stride setting of 2. Tensors 413, 417, 421, 425 are passed to 1x1 horizontal convolution modules 446, 444, 442, 440 respectively. Modules 440, 442, 444, 446 generate tensors 441, 443, 445, 447 respectively. Tensor 441 is passed to a 3x3 output convolution module 470, which generates an output tensor P5 471. Tensor 441 is also passed to an upsample module 450, which generates an upsampled tensor 451. The sum module 460 sums tensors 443 and 451 to generate tensor 461. Tensor 461 is passed to an upsample module 452 and a 3x3 horizontal convolution module 472. Module 472 outputs a P4 tensor 473. The upsample module 452 generates an upsampled tensor 453. The sum module 462 sums tensors 445 and 453 to generate tensor 463. Tensor 463 is passed to a 3x3 horizontal convolution module 474 and an upsample module 454. Module 474 outputs a P3 tensor 475. The upsample module 454 outputs an upsampled tensor 455. The sum module 464 sums tensors 447 and 455 to generate tensor 465, which is passed to a 3x3 horizontal convolution module 476. Module 476 outputs a P2 tensor 477. The upsample modules 450, 452, 454 use nearest neighbor interpolation to reduce computational complexity. Tensors 429, 471, 473, 475, 477 form the output tensor 115 of the CNN backbone 400. When forming the output tensor 115, the FPNs of tensors 429, 471, 473, 475, 477 provide a hierarchical representation of the frame data containing the data of the feature maps for encoding into a bitstream.

[0087] FIG. 5 is a schematic block diagram showing one type of bottleneck encoder 500 that can function as the bottleneck encoder 116. FIG. 6 shows a method 600 for performing the first part of a CNN that compresses using the bottleneck encoder 500 and encodes the resulting compressed feature map. FIG. 7 shows the packing arrangement of the feature map from the compressed tensor to the monochrome video frame.

[0088] The bottleneck encoder 500 receives the FPN tensor 501 corresponding to the tensor 115, operates to limit the dimensions of the received tensor to fewer layers, and reduce the spatial size. By applying a bottleneck encoder and a decoder between the first part (backbone 114) and the second part (head 150) of the separate neural network, it becomes possible to reduce the spatial area within the frame of the packed tensor data. The reduction of the spatial area is achieved by using the interface between the bottleneck encoder and the bottleneck decoder as the split point between the first part and the second part (114 and 150) of the neural network. The bottleneck encoder 116 functions as an additional layer added to the first part of the neural network, and the bottleneck decoder 148 functions as an additional layer added to the second part of the neural network.

[0089] The sensitivity of the task results to bottlenecks depends on the nature of the task. In the case of object detection, since the spatial sensitivity is not very high, the large spatial downsampling of the FPN layers does not have much of an adverse effect on the resulting mAP. In contrast, segmentation and the resulting segmentation maps are more sensitive to the loss of spatial details, so they do not benefit much from significant spatial downsampling, especially in the spatially large tensors of FPN. In the described arrangement, the bottleneck encoder 500 operates at two separate scales with two different spatial resolutions, rather than applying a single scale across all FPN layers. The higher of the two resolutions is used for the larger FPN layers, and more detailed information is retained through the bottleneck encoder 500. The lower of the two resolution scales is used for the smaller FPN layers. The input FPN tensor 501 is composed of layers P2 502, P3 503, P4, and P5 505. The spatial resolution of layers P2 - P4 (502, 503, 504) is a power of two times the spatial resolution of P5 505. P5 505 has width and height (w, h), and P2 - P4 (502, 503, 504) have dimensions of (8w, 8h), (4w, 4h), and (2w, 2h) respectively. In other words, each tensor has a resolution that forms a geometric sequence where the width and height double between consecutive tensors. Layers P2 - P5 each have 256 channels.

[0090] In the above example, inputs P2 through P5 correspond to the hierarchical feature pyramid network outputs (P2 477, P3 475, P4 473, P5 471) generated by the CNN backbone 400 of FIG. 4. When the CNN backbone is implemented based on FIG. 3A and outputs tensors 329, 326, 322, tensors 329, 326, 322 are derived into two tensor sets: a first group having one FPN layer 329 and a second group having two FPN layers 322, 326. The first group having one FPN layer 329 has the smallest spatial resolution and does not require the operation of the MSFF module 510 within the bottleneck encoder 116, and tensor 329 is passed directly to the SSFC encoder 550 as tensor 529. The second group is processed by the MSFF 510 within the bottleneck encoder 116 with tensor 326 passed (input) as tensor 503 and tensor 322 passed as tensor 502.

[0091] Method 600 can be implemented using a device such as a configured FPGA, ASIC, or ASSP. Alternatively, as described below, method 600 may be implemented by source device 110 as one or more software code modules of application program 233 under the execution of processor 205. The software code modules of application program 233 implementing method 600 can be resident, for example, on hard disk drive 210 and / or memory 206. Method 600 is repeated for each frame of video data generated by video source 112. Method 600 may be stored in a computer-readable storage medium and / or memory 206. Method 600 is at step 610 of executing a first part of a neural network.

[0092] In step 610, the CNN backbone 114 executes a neural network layer corresponding to the first part of the neural network under the execution of the processor 205. Step 610 effectively executes the CNN backbone 114. For example, a CNN layer as described with reference to FIG. 4 can be executed to generate the P2-P5 tensor 501(115). The control of the processor 205 proceeds from step 610 to step 615 of selecting the first set of tensors.

[0093] The described arrangement effectively splits the tensor generated by the CNN backbone 114 into first and second sets of a plurality of tensors (also referred to as a plurality of tensors), where the tensors within each set (or at least one tensor) have feature maps with different spatial resolutions. In step 615, the bottleneck encoder 116, under the execution of the processor 205, selects a plurality of adjacent tensors among the tensors 501 as a first plurality of tensors, such as two tensors P4 504 and P5 505. The tensors 501 form a hierarchical representation of the frame data 113 resulting from the application of the FPN to the frame data 113. By using a stride equal to two convolutional stages of the FPN, that is, the stride in modules 412, 416, 420, and 424, the spatial dimensions of the tensors between the tensors 501 are halved in width and height for each tensor when arranged according to the decomposition levels from, for example, P2 to P5. There is a certain inter-layer correlation between the layers P2 to P5 of the tensors 501, despite the different spatial resolutions of the layers. By utilizing the inter-layer correlation, if the tensors are first spatially scaled to the same resolution, for example, the minimum resolution among the tensors to be combined, a reduction in the number of channels is possible compared to the concatenation of tensors across layers. When tensors with significantly different spatial resolutions are combined, the ratio of the downsampling operation becomes high, resulting in excessive loss of details of the high-resolution tensors. For example, to scale P2 502 to P5 505, it is necessary to reduce the width and height to one-eighth of the previous values and reduce the area to one-sixty-fourth of the area of P2 502. In tasks that rely on spatial details, such as instance segmentation, the mAP decreases. In the detection of small objects where the higher-resolution layers are trusted by the network head, a decrease in mAP occurs due to excessive downsampling of the larger layers. The control of the processor 205 proceeds from step 615 to step 620, which selects a second set of tensors.

[0094] In step 620, a second set of tensors is selected by the processor 205 and consists of tensors of adjacent spatial scales among the P2 - P5 layers, including the remaining tensors of P2 - P5 that were not selected in step 615 as the second plurality of tensors. For example, tensors P2 502 and P3 503 are selected in step 620. As a result of steps 615 and 520, the tensor 501 is assigned to two sets of plural tensors. The control of the processor 205 proceeds from step 620 to step 630 that combines the first plurality of tensors.

[0095] In step 630, under the execution of the processor 205, the MSFF module 510 (see FIG. 5) combines each tensor of the first set of tensors, i.e., 504 and 505, to generate a combined tensor 529. The downsampling module 522 operates on the tensor with a larger spatial scale, i.e., P4 504 at 2h, 2w, 256, and downsamples it to match the spatial scale of the tensor with a smaller spatial scale, i.e., P5 505 at h, w, 256, to generate a downscaled P5 tensor 523. The concatenation module 524 performs channel-wise concatenation of tensors 505 and 523 to generate a concatenated tensor 525 of dimension h, w, 512. The concatenated tensor 525 is passed to a squeeze-and-excitation (SE) module 526 to generate a tensor 527. The SE module 526 sequentially executes global pooling, a fully connected layer with reduced number of channels, a rectified linear unit activation, a second fully connected layer to restore the number of channels, and a sigmoid activation function to generate a scaling tensor. The tensor 525 is scaled according to the scaling tensor to generate an output as tensor 527. The SE block 526 is learnable to adaptively change the weighting of different channels of the passed tensor based on the output of the first fully connected layer. The output of the first fully connected layer reduces each feature map of each channel to a single value, which is then passed through a non-linear activation unit (ReLU) and restored to the total number of channels executed by the second fully connected layer, creating a conditional representation of the unit suitable for the weighting of other channels. Thus, the SE block 526 can extract non-linear inter-channel correlations to a greater extent than possible with purely convolutional (linear) layers when generating tensor 527 from tensor 525. Since tensors 525 and 527 contain 512 channels as a result of concatenating two FPN layers, the decorrelation achieved by the SE block 526 spans the two FPN layers P5 and P4.

[0096] The tensor 527 is passed to the convolutional layer 528. The convolutional layer 528 implements one or more convolutional layers to generate a first combined tensor 529 with the number of channels reduced to F channels, typically 256 channels. As a result of step 630, the tensors of the two FPN layers are reduced to a single tensor having the same number of channels as the tensors of the input FPN layer and the smaller spatial resolution of the tensors of the two FPN layers. This dimensionality reduction is achieved in multiple network layers and depends on the training of the layers (e.g., layers 526 and 528) rather than determining the correlations utilized in situ. Returning to FIG. 6, the control of the processor 205 proceeds from step 630 to step 640 of combining a second plurality of tensors.

[0097] In step 640, the MSFF module 510, under the execution of the processor 205, combines each of the second tensors, i.e., the tensors of 502 and 503, as described with reference to FIG. 5, to generate a combined tensor 519. The steps executed as part of step 640 correspond to the steps of step 630, except that the steps are applied to the tensor selected in step 620 rather than the tensor selected in step 615. The downsampling module 512 operates on the tensor with a larger spatial scale, i.e., P2 502 of 8h, 8w, 256, and downsamples it to match the spatial scale of the smaller tensor, i.e., P3 503 of 4h, 4w, 256, to generate a downscaled P2 tensor 513. The concatenation module 514 performs a channel-wise concatenation of tensors 503 and 513 to generate a concatenated tensor 515 of dimensions 4h, 4w, 512. The concatenated tensor 515 is passed to a squeeze-and-excitation (SE) module 516 to generate a tensor 517. The SE module 516 operates in the same manner as described with reference to the SE module 526. The tensor 517 is passed to a convolutional layer 518. The convolutional layer 518 operates in a similar manner to the convolutional layer 528 and generates a second combined tensor 519 with the number of channels reduced to F channels, typically 256 channels. As a result of step 640, the tensors of the two FPN layers are reduced to a single tensor having the same number of channels as the tensors of the input FPN layer and the smaller spatial resolution of the tensors of the two FPN layers. The dimensionality reduction is achieved in multiple network layers and depends on the training of the layers (e.g., the layers of 516 and 518) rather than determining the correlations utilized in situ. As shown in FIG. 6, the control within the processor 205 proceeds from step 640 to step 650 that singly-scale feature compression (SSFC) encodes the first tensor.

[0098] In step 650, as shown in FIG. 5, the SSFC encoder 550 is implemented under the execution of the processor 205. The operation of the SSFC encoder 550 reduces the dimensionality of the combined tensor 529 in order to generate the compressed tensor 557. The combined tensor 529 is passed to the convolutional layer 552 to generate the tensor 553. The tensor 553 has a reduced channel count from 256 to a smaller value C', such as 64. The value 96 can also be used for C', resulting in a larger area requirement for the packed frames as described with reference to FIG. 7. The tensor 553 is passed to the batch normalization module 554 to generate the tensor 555. The batch normalization tensor 555 has the same dimensionality as the tensor 553. The tensor 555 is passed to the tanh layer 556. The tanh layer 556 implements a hyperbolic tangent (tanh) layer similar to layer 536 and generates the compressed tensor 557. The compressed tensor 557 has the same dimensions as the tensor 553.

[0099] The control of the processor 205 proceeds from step 650 to step 660 where a second tensor is SSFC encoded, as shown in FIG. 6. In step 660, under the execution of the processor 205, the SSFC encoder 530 is implemented as described in relation to FIG. 5. The SSFC encoder 530 operates to further reduce the dimensions of the combined tensor 519 in order to generate the compressed tensor 537. The combined tensor 519 is passed to the convolutional layer 532 to generate the tensor 533 with a reduced channel count from 256 to a small value C', such as 64. The tensor 533 is passed to the batch normalization module 534 to generate the tensor 535. The batch normalization tensor 535 has the same dimensions as the tensor 533. The tensor 535 is passed to the tanh layer 536 to generate the compressed tensor 537. The compressed tensor 537 has the same dimensionality as the tensor 533. The use of the hyperbolic tangent (tanh) layer compresses the dynamic range of the values in the tensor 537 to [-1, 1] and removes outliers.

[0100] Each of the compressed tensors 557 and 537 provides a unit of information of the feature map of the frame data 113, such as obtained by a convolution operation of either (i) the MSFF 510 and the SSFC encoder 550 with respect to the tensors 505 and 504, or (ii) the MSFF 510 and the SSFC encoder 530 with respect to the tensors 503 and 502. The compressed tensors 557 and 537 provide a set of tensors 560 corresponding to the bottleneck encoded tensor 117. Upon completion of step 660, the control within the processor 205 proceeds, as shown in FIG. 6, to step 670 of packing the compressed tensors from step 660.

[0101] At step 670, under the execution of the processor 205, the quantization and packing module 118 quantizes the compressed tensors 537 and 557 and packs the quantized tensors into a single monochrome video frame. An example of a single monochrome video frame 700 is shown in FIG. 7. The frame 700 corresponds to the frame data 119. The ranges of the compressed tensors 537 and 557 are [-1, 1] because the tanh activation function is used with 536 and 556 respectively. The property of tanh to remove outliers results in a distribution compliant with linear quantization for the bit depth of the frame 700. The channels of the compressed tensor 537 are packed as a feature map of a specific size, such as the feature map 710 of the frame 700. The channels of the compressed tensor 557 are packed as a feature map of a different size, such as the feature map 712 in the frame 700. One channel of the compressed tensor 537 corresponds to one feature map shown as one rectangular region, such as region 710. One channel of the compressed tensor 557 corresponds to one feature map shown by one rectangular region, such as region 712. Returning to FIG. 6, the control within the processor 205 proceeds from step 670 to step 680 of compressing the frame.

[0102] In step 680, the feature map encoder 120 encodes the frame 700 under the execution of the processor 205 to generate a bitstream 121. The operation of the feature map encoder 120 will be described with reference to FIG. 8. By executing step 680, the method 600 ends with the dimensions of the FPN layer of the image frame 312 reduced and compressed in the video bitstream 121.

[0103] FIG. 8 is a schematic block diagram showing the functional modules of a video encoder 120, also referred to as a feature map encoder. The video encoder 120 encodes the packed frame 119 shown as frame 700 in the example of FIG. 7 to generate a bitstream 121. Generally, data passes through the functional modules within the video encoder 120 in groups of samples or coefficients, such as by dividing the blocks into fixed-size sub-blocks, or as an array. The video encoder 120 can be implemented using the general-purpose computer system 200 as shown in FIGS. 2A and 2B, and various functional modules can be implemented by software within the computer system 200, such as one or more software code modules of a software application program 233 that resides on the hard disk drive 205 and is controlled for execution by the processor 205, with the dedicated hardware within the computer system 200. Alternatively, the video encoder 120 can also be implemented by a combination of dedicated hardware and software executable within the computer system 200. The video encoder 120 and the described method can alternatively be implemented in dedicated hardware such as one or more integrated circuits that perform the functions or sub-functions of the described method. Such dedicated hardware can include a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific standard product (ASSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or one or more microprocessors and associated memory. In particular, the video encoder 120 is composed of modules 810 to 890, each of which can be implemented as one or more software code modules of a software application program 233.

[0104] The video encoder 120 of FIG. 8 is an example of a general-purpose video coding (VVC) video coding pipeline, but other video codecs can also be used to perform the processing stages described herein. The frame data 119 can be in any chroma format and bit depth supported by the profile in use, for example, 4:0:0, 4:2:0 for the "Main10" profile of the VVC standard, and the sample precision is from 8 (8) bits to 10 (10) bits.

[0105] The block partitioner 810 first divides the frame data 119 into CTUs, which are generally square-shaped and configured such that a specific size of CTU is used. The maximum valid size of the CTU can be, for example, 32×32, 64×64, or 128×128 luma samples, and is configured by the "sps_log2_ctu_size_minus5" syntax element present in the "sequence parameter set". The CTU size also provides the maximum CU size such that a CTU that is not further divided contains one CU. The block divider 810 further divides each CTU into one or more CUs according to the luma coding tree and the chroma coding tree. The luma channel may also be referred to as the primary color channel. Also, each chroma channel may also be referred to as the secondary color channel. The CUs have various sizes and can include both square and non-square aspect ratios. However, in the VVC standard, the CUs, CUs, PUs, and TUs always have side lengths that are powers of 2. Thus, the current CU represented as 812 proceeds according to iterations over one or more blocks of the CTU according to the luma coding tree and the chroma coding tree of the CTU and is output from the block partitioner 810.

[0106] The CTUs obtained from the first partition of frame data 119 may be scanned in raster scan order and grouped into one or more "slices". A slice may be an "intra" (or "I") slice. An intra slice (I slice) indicates that all CUs within the slice are intra predicted. Generally, the first picture of a coded layer video sequence (CLVS) contains only I slices and is called an "intra picture". The CLVS contains periodic intra pictures and may form "random access points" (i.e., intermediate frames within a video sequence from which decoding can start). Alternatively, a slice may be uni predicted or bi predicted (a 'P' or 'B' slice respectively), indicating the additional availability of uni prediction and bi prediction in the slice.

[0107] Video encoder 120 encodes a sequence of pictures according to a picture structure. One picture structure is "low latency", in which case pictures using inter prediction can only reference pictures that have occurred previously in the sequence. Low latency allows each picture to be output as soon as it is decoded, in addition to being stored for possible reference by subsequent pictures. Another picture structure is "random access", in which the encoding order of pictures is different from the display order. In random access, other pictures that have been decoded but not yet output can be referenced between predicted pictures. Some picture buffering is required, so future reference pictures from the perspective of display order are present in the decoded picture buffer, resulting in a latency of multiple frames.

[0108] When a chroma format other than 4:0:0 is used, in an I slice, the coding tree of each CTU may branch into two separate coding trees for luma and chroma at levels of 64×64 or below. By using separate trees, different block structures for luma and chroma can exist within the 64×64 luma area of a CTU. For example, large chroma CUs may be placed together with a number of small luma CUs, and vice versa. In a P or B slice, the single coding tree of a CTU defines a common block structure for luma and chroma. As a result, the blocks of the single tree are intra - predicted or inter - predicted.

[0109] In addition to dividing a picture into slices, a picture can also be divided into "tiles". A tile is a sequence of CUs that covers a rectangular area of a picture. The CTU scan is performed in a raster scan manner within each tile and proceeds from one tile to the next. A slice can be either an integer number of tiles or an integer number of consecutive rows of CUs within a given tile.

[0110] For each CTU, the video encoder 120 operates in two stages. In the first stage (referred to as the "search" stage), the block partitioner 810 tests various potential configurations of the coding tree. Each potential configuration of the coding tree has an associated "candidate" CB. The first stage involves testing various candidate CBs to select a CB that provides a relatively high compression efficiency with a relatively low distortion. The testing generally involves Lagrangian optimization, whereby candidate CBs are evaluated based on a weighted combination of rate (i.e., coding cost) and distortion (i.e., error with respect to the input frame data 119). The "best" candidate CB (i.e., the CB with the lowest evaluated rate / distortion) is selected for subsequent coding into the bitstream 121. Included in the evaluation of candidate CBs is the option of using a CB for a given region, or further dividing the region according to various partitioning options and coding each of the resulting smaller regions with a further CB, or the option of further dividing the region. As a result, both the coding tree and the CB itself are selected during the search stage.

[0111] The video encoder 120 generates a predicted block (PB) indicated by arrow 820 for each CB, e.g., CB812. The PB820 is a prediction of the content of the associated CB812. The subtractor module 822 generates a difference (or "residual", meaning the difference in the spatial domain) indicated as 824 between the PB820 and the CB812. The difference 824 is the block-size difference between corresponding samples of the PB820 and the CB812. The difference 824 is transformed, quantized, and represented as a transformed block (TB) indicated by arrow 836. The PB820 and the associated TB836 are typically selected from among a number of possible candidate CBs, e.g., based on the evaluated cost or distortion.

[0112] The candidate coded block (CB) is a CB resulting from one of the prediction modes available to video coder 120 for the associated PB and the resulting residual. When combined with the PB predicted in video coder 120, TB836 reduces the difference between the decoded CB and the original CB812 at the expense of additional signaling in the bitstream.

[0113] Each candidate coded block (CB), i.e., the predicted block (PB) combined with the transform block (TB), thus has an associated coding cost (or "rate") and an associated difference (or "distortion"). The distortion of a CB is typically estimated as the difference of sample values such as the sum of absolute differences (SAD), the sum of squared differences (SSD), or the Hadamard transform applied to the differences. The estimated values obtained from each candidate PB can be determined by mode selector 886 using difference 824 to determine prediction mode 887. Prediction mode 887 indicates a decision to use a particular prediction mode for the current CB, e.g., intra-frame prediction or inter-frame prediction. The estimation of the coding cost associated with the residual coding corresponding to each candidate prediction mode may be performed at a significantly lower cost than the entropy coding of the residual. Thus, even in a real-time video coder, a large number of candidate modes can be evaluated to determine the optimal mode in terms of rate-distortion.

[0114] To determine the optimal mode from a rate-distortion perspective, a variation of Lagrangian optimization is typically used.

[0115] Lagrange or similar optimization processing can be employed for both the selection of the optimal partitioning of CTUs into CUs (by block partitioner 810) and the selection of the best prediction mode from multiple possibilities. Through the application of the Lagrange optimization process for candidate modes in mode selector module 886, the intra prediction mode with the minimum cost measurement is selected as the "best" mode. The lowest cost mode includes the selected secondary transform index 888, which is also encoded into bitstream 121 by entropy encoder 838.

[0116] In the second stage of the operation of video encoder 120 (referred to as the "encoding" stage), in video encoder 120, an iteration for the determined encoding tree(s) of each CTU is performed. For a CTU using separate trees, for each 64×64 luma region of the CTU, first the luma encoding tree is encoded, and then the chroma encoding tree is encoded. Only luma CUs are encoded within the luma encoding tree, and only chroma CUs are encoded within the chroma encoding tree. For a CTU using a shared tree, a single tree describes the CUs (i.e., luma CUs and chroma CUs) according to the common block structure of the shared tree.

[0117] Entropy encoder 838 supports bit-level encoding of syntax elements using variable-length and fixed-length codewords, and an arithmetic coding mode for syntax elements. For example, a part of a bitstream such as a "parameter set" like a sequence parameter set (SPS) and a picture parameter set (PPS) uses a combination of fixed-length and variable-length codewords. A slice, also called a consecutive part, has slice data using arithmetic coding following a slice header using variable-length coding. The slice header defines parameters specific to the current slice, such as a slice-level quantization parameter offset. The slice data includes the syntax elements of each CTU within the slice. To use variable-length coding and arithmetic coding, sequential parsing is required within each part of the bitstream. Ports may be delimited by a start code to form a "network abstraction layer unit" or "NAL unit". Arithmetic coding is supported using a context-adaptive binary arithmetic coding process.

[0118] Arithmetically coded syntax elements are composed of a sequence of one or more "bins". A bin has a value of "0" or "1", similar to a bit. However, a bin is not encoded as discrete bits in the bitstream 121. A bin has an associated prediction value (or "likely" or "most likely") and an associated probability, known as a "context". If the actual bin to be encoded matches the prediction value, the "most probable symbol" (MPS) is encoded. Encoding the most probable symbol is relatively inexpensive in terms of consumed bits in the bitstream 121, including a cost of less than one discrete bit. If the actual bin to be encoded does not match the likely value, the "least probable symbol" (LPS) is encoded. Encoding the least probable symbol has a relatively high cost in terms of consumed bits. Bin coding techniques enable efficient coding of bins where the probability of "0" versus "1" is skewed. If there are two possible values for a syntax element (i.e., a "flag"), one bin is sufficient. For syntax elements with multiple possible values, a series of bins is required.

[0119] The presence of subsequent bins within a sequence may be determined based on the value of a previous bin within the sequence. Further, each bin may be associated with multiple contexts. The selection of a particular context may depend on previous bins within a syntax element, bin values of adjacent syntax elements (i.e., those from adjacent blocks), etc. Each time a context-encoded bin is encoded, the context (if any) selected for that bin is updated in a way that reflects the new bin value. Thus, the binary arithmetic coding scheme is said to be adaptive.

[0120] Also supported by the entropy encoder 838 are bins that lack a context, called "bypass bins". Bypass bins are encoded assuming an equiprobable distribution between "0" and "1". Thus, each bin has an encoding cost of 1 bit in the bitstream 121. Since there is no context, memory is saved and complexity is reduced. Thus, bypass bins are used when the distribution of the values of a particular bin is not skewed. An example of an entropy encoder that employs context and adaptation is known in the art as CABAC (Context Adaptive Binary Arithmetic Coder), and many variations of this coder have been adopted for video coding.

[0121] Entropy encoder 838 encodes quantization parameter 892 and, when used for the current CB, LFNST index 888, using a combination of context - encoded bins and bypass - encoded bins. The quantization parameter 892 is encoded using "delta QP" generated by QP controller module 890. Delta QP is signaled at most once in each region known as a "quantization group". The quantization parameter 892 is applied to the residual coefficients of the luma CB. The adjusted quantization parameter is applied to the residual coefficients of the collocated chroma CB. The adjusted quantization parameter may include a mapping from the luma quantization parameter 892 according to a CU - level offset selected from a list of mapping tables and offsets. The secondary transform index 888 is signaled when the residual associated with the transform block contains significant residual coefficients only at the coefficient positions targeted to be transformed into primary coefficients by the application of the secondary transform.

[0122] The residual coefficients of each TB related to the CB are encoded using residual syntax. The residual syntax mainly uses arithmetically - encoded bins to indicate the significance of the coefficients, secures bypass bins for higher - magnitude residual coefficients along with the magnitude of lower values, and is designed to efficiently encode low - magnitude coefficients. Thus, residual blocks consisting of very low - magnitude values and sparse arrangements of significant coefficients are efficiently compressed. Furthermore, there are two residual encoding schemes. The normal residual encoding scheme is optimized for TBs where significant coefficients are mainly located in the upper - left corner of the TB as seen when the transform is applied. The transform - skip residual encoding scheme is available for TBs where the transform is not executed and can efficiently encode the residual coefficients regardless of the distribution of the residual coefficients across the TB.

[0123] The multiplexer module 884 outputs PB820 from the intra prediction module 864 according to the determined best in-frame prediction mode selected from the tested prediction modes of each candidate CB. The candidate prediction modes do not have to include all possible prediction modes supported by the video encoder 120. Intra prediction is classified into three types. First, "DC intra prediction" involves inputting PB with a single value representing the average of neighboring reconstructed samples. Second, "plane intra prediction" involves inputting PB with samples following a plane where the DC offset and vertical and horizontal gradients are derived from neighboring reconstructed neighboring samples. The neighboring reconstructed samples typically include a column of reconstructed samples extending up to the range on the right side of the PB above the current PB and a column of reconstructed samples extending down beyond the PB on the left side of the current PB. Third, "angular intra prediction" involves inputting PB with reconstructed neighboring samples that are filtered and propagated across the PB in a specific direction (or "angle"). In VVC, 65 angles are supported, and additional angles that are available for rectangular blocks but not for square blocks can be utilized, resulting in a total of 87 angles that can be generated.

[0124] The fourth type of intra prediction is available for chroma PB, whereby the PB is generated from luma reconstruction samples rearranged according to the "Cross Component Linear Model" (CCLM) mode. Three different CCLM modes are available, and each mode uses a different model derived from adjacent luma and chroma samples. The derived model is used to generate a block of chroma PB samples from the arranged luma samples. The luma block can be intra predicted using matrix multiplication of reference samples using one matrix selected from a predefined set of matrices. This Matrix Intra Prediction (MIP) achieves gain by using a matrix learned on a large-scale video dataset that represents the relationship between reference samples and the prediction block that is not easily captured in the angular, planar, or DC intra prediction modes.

[0125] Module 864 can also generate a prediction unit by copying blocks from the vicinity of the current frame using the "Intra-Block Copy" (IBC) method. The position of the reference block is restricted to a region corresponding to one CTU, which is divided into 64×64 regions known as VPDUs. This region covers the processed VPDU of the current CTU and the VPDUs of the previous CTU(s) within each row or CTU and within each slice or tile, up to the region limit corresponding to one 128×128 luma sample, regardless of the configured CTU size of the bitstream. This region is known as the "IBC virtual buffer", and restricting the IBC reference region limits the required storage. Since the reconstructed samples 854 (i.e., before loop filtering) are input to the IBC buffer, a buffer separate from the frame buffer 872 is required. When the CTU size is 128×128, the virtual buffer contains only samples from the CTU adjacent to the current CTU and from the CTU to its left. When the CTU size is 32×32 or 64×64, the virtual buffer contains the CTUs up to 4 or 16 CTUs to the left of the current CTU. Regardless of the CTU size, access to adjacent CTUs to obtain samples for the IBC reference block is restricted by boundaries such as the edges of the picture, slice, tile, etc. In particular, in the feature maps of FPN layers with small dimensions, using CTU sizes such as 32×32 or 64×64 aligns the reference regions that cover the previous set of feature maps better. When the arrangement of the feature maps is ordered based on SAD, SSE, or other difference metrics, accessing similar feature maps for IBC prediction provides an advantage in coding efficiency.

[0126] The residual of the prediction block when encoding the feature map data is different from the residual seen in natural video. Such natural videos are generally captured by an imaging sensor or are screen content such as commonly seen in the user interface of an operating system. The feature map residual tends to contain many details and is more compliant with transform skip encoding than the dominance of the low-frequency coefficients of various transforms. According to experiments, the feature map residual has sufficient local similarity to benefit from transform encoding. However, the distribution of the feature map residual coefficients is not clustered towards the DC (top left) coefficient of the transform block. In other words, there is sufficient correlation for the transform to show a gain when encoding the feature map data, which also holds when block copy is used to generate the prediction block of the feature map data. Therefore, when evaluating the residual obtained from the candidate block vector for intra-block copy when encoding the feature map data, the Hadamard cost estimate can be used instead of relying only on the SAD or SSD cost estimate. The SAD or SSD cost estimate tends to select block vectors with residuals that are more compliant with transform skip encoding and may miss block vectors with residuals that are encoded compactly using a transform. The multiple transform selection (MTS) tool of the VVC standard can be used when encoding the feature map data so that in addition to the DCT-2 transform, a combination of the DST-7 transform and the DCT-8 transform can be utilized horizontally and vertically for residual encoding.

[0127] The intra-prediction luma encoding block can be divided into a set of prediction blocks of equal size either vertically or horizontally, and each block has a minimum area of 16 luma samples. This intra-subpartition (ISP) approach allows separate transform blocks to contribute to the generation of prediction blocks from one subpartition to the next within the luma encoding block, improving the compression efficiency.

[0128] When neighboring samples that have been previously reconstructed, such as the edges of the frame, are not available, a default halftone value of one-half of the sample range is used. For example, 512 is used for 10-bit video. Since there is no previous sample for the CB at the upper left position of the frame, the angular prediction mode and the in-plane prediction mode generate the same output as the DC prediction mode (i.e., a flat plane of samples with the halftone value as the magnitude).

[0129] In inter-frame prediction, the motion compensation module 880 generates a prediction block 882 using samples from one or two frames preceding the current frame in the coded order frames in the bitstream, and outputs it as PB820 by the multiplexer module 884. Further, for inter-frame prediction, usually, a single coding tree is used for both the luma channel and the chroma channel. The order of the coded frames in the bitstream may be different from the order of the frames when captured or displayed. When one frame is used for prediction, the block is said to be "uni-predicted" and has one associated motion vector. When two frames are used for prediction, the block is said to be "bi-predicted" and has two associated motion vectors. For P slices, each CU may be intra-predicted or uni-predicted. For B slices, each CU may be intra-predicted, uni-predicted, or bi-predicted.

[0130] Frames are typically encoded using a "Group of Pictures (GOP)" structure, which enables a temporal hierarchy of frames. A frame can be divided into multiple slices, and each slice encodes a part of the frame. The temporal hierarchy of frames allows a frame to reference preceding and subsequent pictures in the display order of the frames. Pictures are encoded in an order necessary to satisfy the dependencies for decoding each frame. The affine inter-prediction mode divides a prediction unit into multiple small blocks instead of selecting and filtering a reference sample block of the prediction unit using one or two motion vectors, and generates a motion field such that each small block has a different motion vector. The motion field uses the motion vectors of points close to the prediction unit as "control points". Affine prediction reduces the need to use a deeply divided coding tree and enables the coding of motions different from translations. The bi-prediction mode available in VVC geometrically blends two reference blocks along a selected axis. This geometric partitioning mode ("GPM") enables the use of larger coding units along the boundary between two objects, and the geometric shape of the boundary is encoded for the coding unit as an angle and a center offset. The motion vector difference can be encoded as a direction (up / down / left / right) and a distance instead of using Cartesian (x, y) offsets. The motion vector predictor is obtained from an adjacent block as if no offset were applied ("merge mode"). The current block shares the same motion vector as the selected adjacent block.

[0131] Samples are selected according to motion vector 878 and reference picture index. The motion vector 878 and reference picture index are applied to all color channels, and thus, inter prediction is described mainly from the perspective of operations on PUs rather than PBs. The decomposition of each CTU into more than one inter prediction block is described by a single coding tree. The inter prediction method may have different numbers and precisions of motion parameters. Motion parameters typically consist of a reference frame index indicating which reference frame(s) are used from a list of reference frames and a spatial transformation for each of the reference frames, but may include more frames, special frames, or complex affine parameters such as scaling and rotation. Further, a predetermined motion refinement process may be applied to generate a high-density motion estimate based on the referenced sample block.

[0132] PB820 is determined and selected, and after subtracting PB820 from the original sample block by subtractor 822, the residual with the lowest coding cost represented as 824 is obtained and subjected to irreversible compression. The irreversible compression process consists of steps of transformation, quantization, and entropy coding. The forward primary transformation module 826 applies a forward transformation to the difference 824, transforms the difference 824 from the spatial domain to the frequency domain, and generates primary transformation coefficients represented by arrow 828. The maximum primary transformation size in one dimension is either a 32-point DCT-2 transformation or a 64-point DCT-2 transformation, configured by "sps_max_luma_transform_size_64_flag" in the sequence parameter set. If the CB to be coded is larger than the maximum supported primary transformation size represented as the block size (e.g., 64×64 or 32×32), the primary transformation 826 is applied in a tiled manner to transform all samples of the difference 824. If a non-square CB is used, the tiling is also performed using the maximum transformation size available in each dimension of the CB. For example, if a maximum transformation size of 32 is used, a 64×16 CB uses two 32×16 primary transformations arranged in a tiled manner. If the size of the CB is larger than the maximum supported transformation size, the CB is filled with TBs in a tiled manner. For example, a 128×128 CB with a maximum size of 64pt transform is filled with four 64×64 TBs in a 2×2 arrangement. A 64×128 CB with a maximum transform size of 32pt is filled with eight 32×32 TBs in a 2×4 arrangement.

[0133] The application of transform 826 results in multiple TBs for a CB. When each application of the transform operates on a TB of a difference 824 larger than 32×32, e.g., 64×64, all resulting primary transform coefficients 828 outside the upper left 32×32 region of the TB are set to zero (i.e., discarded). The remaining primary transform coefficients 828 are passed to the quantizer module 834. The primary transform coefficients 828 are quantized according to quantization parameters 892 associated with the CB, generating primary transform coefficients 832. In addition to the quantization parameters 892, the quantizer module 834 can also apply a "scaling list" to enable non-uniform quantization within the TB by further scaling the residual coefficients according to the spatial position within the TB. The quantization parameters 892 may be different for each luma CB and each chroma CB. The primary transform coefficients 832 are passed to the forward secondary transform module 830, which generates the transform coefficients represented by arrow 836 by performing a non-separable secondary transform (NSST) operation or bypassing the secondary transform. The forward primary transform is typically separable and transforms a set of rows and then a set of columns for each TB. The forward primary transform module 826 uses either a type-II discrete cosine transform (DCT-2) in the horizontal and vertical directions, or bypassing of the transform in the horizontal and vertical directions, or a combination of a type-VII discrete sine transform (DST-7) and a type-VIII discrete cosine transform (DCT-8) in either the horizontal or vertical direction for luma TBs with a width and height not exceeding 16 samples. The use of the combination of DST-7 and DCT-8 is called "multi-transform selection set (MTS)" in the VVC standard.

[0134] The forward secondary transform of module 830 is generally a non-separable transform that is only applied to the residuals of the predicted CUs, although it may be bypassed. The forward secondary transform operates on either 16 samples (arranged as the top-left 4×4 sub-block of the primary transform coefficients 828) or 48 samples (arranged as three 4×4 sub-blocks of the top-left 8×8 coefficients of the primary transform coefficients 828) and generates a set of secondary transform coefficients. The set of secondary transform coefficients may be fewer in number than the set of primary transform coefficients from which they are derived. Since the secondary transform is only applied to sets of coefficients that are adjacent to each other and include the DC coefficient, the secondary transform is called a "low-frequency non-separable secondary transform" (LFNST). Such a secondary transform may be obtained through a training process and, due to its non-separable nature and trained origin, exploits additional redundancy within the residual signal that cannot be captured by separable transforms such as variants of the DCT or DST. Further, when the LFNST is applied, all of the remaining coefficients within the TB become zero in both the primary transform domain and the secondary transform domain.

[0135] The quantization parameter 892 is constant for a given TB, thus providing uniform scaling for the generation of residual coefficients in the primary transform region for the TB. The quantization parameter 892 can vary periodically with the signaled "delta quantization parameter". The delta quantization parameter (delta QP) is signaled once for CUs included within a given region called a "quantization group". If a CU is larger than the quantization group size, the delta QP is signaled once for one of the CUs TBs. That is, the delta QP is signaled once by the entropy encoder 838 for the first quantization group of the CU and not signaled for subsequent quantization groups of the CU. The scaling factor applied to each residual coefficient is derived from a combination of the quantization parameter 892 and the corresponding entry of the scaling matrix. The scaling matrix can have a size smaller than the size of the TB and when applied to the TB, a nearest neighbor approach is used to provide the scaling value for each residual coefficient from a scaling matrix of a size smaller than the size of the TB. The residual coefficients 836 are supplied to the entropy encoder 838 for encoding into the bitstream 121. Typically, the residual coefficients of each TB having at least one significant residual coefficient of a TU are scanned according to a scan pattern to generate an ordered list of values. The scan pattern generally scans the TB as a sequence of 4×4 "sub-blocks", providing a regular scan operation at the granularity of 4×4 sets of residual coefficients, and the arrangement of the sub-blocks depends on the size of the TB. The scan within each sub-block and the progression from one sub-block to the next typically follows a backward diagonal scan pattern. Further, the quantization parameter 892 is encoded into the bitstream 121 using the delta QP syntax element, and the slice QP for the initial value within a given slice or sub-picture and the secondary transform index 888 are encoded into the bitstream 121.

[0136] As described above, the video encoder 120 needs to access a frame representation corresponding to the decoded frame representation seen in the video decoder. Therefore, the residual coefficients 836 are passed to an inverse secondary transform module 844 that operates according to the secondary transform index 888 to generate the intermediate inverse transform coefficients represented by the arrow 842. The intermediate inverse transform coefficients 842 are inverse quantized by an inverse quantization module 840 according to the quantization parameter 892 to generate the inverse transform coefficients represented by the arrow 846. The inverse quantization module 840 can also perform inverse non-uniform scaling of the residual coefficients using a scaling list corresponding to the forward scaling performed by the quantization module 834. The inverse transform coefficients 846 are passed to an inverse primary transform module 848 to generate the residual samples represented by the arrow 850 of the TU. The inverse primary transform module 848 applies a DCT-2 transform in the horizontal and vertical directions while being constrained by the maximum transform size as described with reference to the forward primary transform module 826. The type of inverse transform performed by the inverse secondary transform module 844 corresponds to the type of forward transform performed by the forward secondary transform module 830. The type of inverse transform performed by the inverse primary transform module 848 corresponds to the type of primary transform performed by the primary transform module 826. An addition module 852 adds the residual samples 850 and the PU 820 to generate the reconstructed samples of the CU (indicated by the arrow 854).

[0137] The reconstructed sample 854 is passed to the reference sample cache 856 and the in-loop filter module 868. The reference sample cache 856 is typically implemented using static RAM on the ASIC, avoiding costly off-chip memory accesses and providing the minimum sample storage necessary to satisfy the dependencies for generating in-frame PBs for subsequent CUs within the frame. The minimum dependencies typically include a "line buffer" of samples along the bottom of the CTU row for use in the next row of the CTU, and column buffering over a range set by the CTU height. The reference sample cache 856 supplies reference samples (represented by arrow 858) to the reference sample filter 860. The sample filter 860 applies a smoothing operation to generate filtered reference samples (indicated by arrow 862). The filtered reference samples 862 are used by the in-frame prediction module 864 to generate an intra prediction block of samples represented by arrow 866. For each candidate intra prediction mode, the in-frame prediction module 864 generates a block of samples, which is 866. The block of samples 866 is generated by the module 864 using techniques such as DC, planar, or angular intra prediction. The block of samples 866 can also be generated using a matrix multiplication approach having as inputs adjacent reference samples and a matrix selected from a set of matrices by the video encoder 120, where the selected matrix is signaled within the bitstream 121 using an index to identify which matrix of the set of matrices is used by the video decoder 144.

[0138] The in-loop filter module 868 applies a plurality of filtering stages to the reconstructed samples 854. The filtering stages include a "deblocking filter" (DBF) that applies smoothing aligned to the CU boundary to reduce artifacts due to discontinuities. Another filtering stage present in the in-loop filter module 868 is the "adaptive loop filter" (ALF), which applies a Wiener-based adaptive filter to further reduce distortion. A further available filtering stage in the in-loop filter module 868 is the "sample adaptive offset" (SAO) filter. The SAO filter first classifies the reconstructed samples into one or more categories and operates by applying an offset at the sample level according to the assigned category.

[0139] The filtered samples represented by the arrow 870 are output from the in-loop filter module 868. The filtered samples 870 are stored in the frame buffer 872. The frame buffer 872 typically has the capacity to store a plurality (e.g., up to 16) of pictures and is thus stored in the memory 206. Since the frame buffer 872 requires a large amount of memory consumption, it is typically not stored using on-chip memory. Therefore, access to the frame buffer 872 is costly in terms of memory bandwidth. The frame buffer 872 provides a reference frame (represented by the arrow 874) to the motion estimation module 876 and the motion compensation module 880.

[0140] The motion estimation module 876 estimates a number of "motion vectors" (shown as 878), which are Cartesian space offsets from the position of the current CB, each of which references one block of the reference frame within the frame buffer 872. A filtered block of the reference samples (represented as 882) is generated for each motion vector. The filtered reference samples 882 form additional candidate modes available for potential selection by the mode selector 886. Further, for a given CU, the PU 820 is formed using one reference block ("uni-predicted") or two reference blocks ("bi-predicted"). For the selected motion vector, the motion compensation module 880 generates the PB 820 according to a filtering process that supports sub-pixel accuracy of the motion vector. Thus, the motion estimation module 876 (which operates on a number of candidate motion vectors) performs a simplified filtering process compared to the filtering process of the motion compensation module 880 (which operates only on the selected candidates), achieving a reduction in computational complexity. When the video encoder 120 selects inter-prediction for the CU, the motion vectors 878 are encoded into the bitstream 121.

[0141] Although video encoder 120 of FIG. 8 is described with reference to Versatile Video Coding (VVC), other video coding standards or implementations may also employ the processing stages of modules 810-890. Frame data 119 (and bitstream 121) may also be read from (or written to) memory 206, hard disk drive 210, CD-ROM, Blu-ray Disc (TM), or other computer-readable storage media. Further, frame data 119 (and bitstream 121) may be received from (or transmitted to) an external source such as a communication network 220 or a server connected to a high-frequency receiver. Communication network 220 may provide limited bandwidth, and rate control may be necessary in video encoder 120 to avoid network saturation when it is difficult to compress frame data 119. Further, bitstream 121 may be constructed from one or more slices representing spatial sections (a collection of CTUs) of frame data 119 generated by one or more instances of video encoder 120 operating in cooperation under the control of processor 205. Bitstream 121 may also include one slice corresponding to one subpicture output as a collection of subpictures forming one picture, and each slice is independently encodable and independently decodable with respect to any other slice or subpicture within the picture.

[0142] A video decoder 144, also referred to as a feature map decoder, is shown in FIG. 9. The video decoder 144 of FIG. 9 is an example of a general-purpose video coding (VVC) video decoding pipeline, but other video codecs can also be used to perform the processing steps described herein. As shown in FIG. 9, a bitstream 143 is input to the video decoder 144. The bitstream 143 can be read from a memory 206, a hard disk drive 210, a CD-ROM, a Blu-ray Disc (registered trademark), or other non-transitory computer-readable storage media. Alternatively, the bitstream 143 can also be received from an external source such as a server or a radio receiver connected to a communication network 220. The bitstream 143 includes encoded syntax elements representing the captured frame data to be decoded.

[0143] The bitstream 143 is input to the entropy decoder module 920. The entropy decoder module 920 extracts syntax elements from the bitstream 143 by decoding a sequence of "bins" and passes the values of the syntax elements to other modules within the video decoder 144. The entropy decoder module 920 uses variable length decoding and fixed length decoding to decode the SPS, PPS, or slice header and uses an arithmetic decoding engine to decode the syntax elements of the slice data as a sequence of one or more bins. Each bin can use one or more "contexts", where a context describes the probability levels used to encode the "1" and "0" values of the bin. When multiple contexts are available for a given bin, a "context modeling" or "context selection" step is performed to select one of the contexts available for decoding the bin. Since the process of decoding a bin forms a sequential feedback loop, each slice may be decoded in its entirety by a given instance of the entropy decoder 920. A single (or a few) high-performance instances of the entropy decoder 920 can decode all slices or sub-pictures for a frame or picture from the bitstream 143. Multiple low-performance instances of the entropy decoder 920 can decode slices simultaneously for a frame from the bitstream 143.

[0144] Entropy decoder module 920 applies an arithmetic coding algorithm, such as, for example, "Context Adaptive Binary Arithmetic Coding" (CABAC), to decode syntax elements from bitstream 143. The decoded syntax elements are used to reconstruct parameters within video decoder 144. The parameters include residual coefficients (represented by arrow 924), quantization parameter 974, secondary transform index 970, and mode selection information such as intra prediction mode (represented by arrow 958). The mode selection information also includes information such as motion vectors and the partitioning of each CTU into one or more CBs. The parameters are typically used, in combination with sample data from previously decoded CBs, to generate PBs.

[0145] Residual coefficient 924 is passed to inverse secondary transform module 936, where, according to the secondary transform index, a secondary transform is applied or the operation is bypassed. Inverse secondary transform module 936 generates reconstructed transform coefficient 932, i.e., primary transform domain coefficient, from the secondary transform domain coefficient. The reconstructed transform coefficient 932 is input to inverse quantization module 928. Inverse quantization module 928 performs inverse quantization (or "scaling") in the residual coefficient 932, i.e., the primary transform coefficient domain, and creates reconstructed intermediate transform coefficient represented by arrow 940 according to quantization parameter 974. Inverse quantization module 928 can also apply a scaling matrix to provide non-uniform inverse quantization within the TB corresponding to the operation of inverse quantization module 840. If the use of a non-uniform inverse quantization matrix is indicated in bitstream 143, video decoder 144 reads the quantization matrix from bitstream 143 as a sequence of scaling factors and arranges the scaling factors in a matrix. Inverse scaling uses the quantization matrix in combination with the quantization parameter to create reconstructed intermediate transform coefficient 940.

[0146] The reconfigured transform coefficient 940 is passed to the inverse primary transform module 944. Module 944 inverse-transforms the coefficient 940 from the frequency domain to the spatial domain. The inverse primary transform module 944 applies an inverse DCT-2 transform, constrained by the maximum transform size available, in the horizontal and vertical directions as described with reference to the forward primary transform module 726. The result of the operation of module 944 is a block of residual samples represented by arrow 948. The block of residual samples 948 is of the same size as the corresponding CB. The residual samples 948 are supplied to the sum module 950.

[0147] In the sum module 950, the residual samples 948 are added to the decoded PB (represented as 952), and a block of reconstructed samples represented by arrow 956 is generated. The reconstructed samples 956 are supplied to the reconstructed sample cache 960 and the in-loop filtering module 988. The in-loop filtering module 988 generates a reconstructed block of frame samples represented as 992. The frame samples 992 are written to the frame buffer 996.

[0148] The reconstructed sample cache 960 operates in the same manner as the reconstructed sample cache 856 of the video encoder 120. The reconstructed sample cache 960 provides storage for the reconstructed samples necessary for intra prediction of subsequent CBs without using the memory 206 (e.g., typically by using the data 232 which is on-chip memory instead). The reference samples indicated by the arrow 964 are obtained from the reconstructed sample cache 960 and supplied to the reference sample filter 968 to generate the filtered reference samples indicated by the arrow 972. The filtered reference samples 972 are supplied to the intra prediction module 976. The module 976 generates a block of intra prediction samples represented by the arrow 980 according to the intra prediction mode parameter 958 signaled in the bitstream 143 and decoded by the entropy decoder 920. The intra prediction module 976 supports the modes of the module 764 including IBC and MIP. The block of samples 980 is generated using modes such as DC, planar or angular intra prediction.

[0149] In the bitstream 143, when the prediction mode of the CB is indicated to use intra prediction, the intra predicted samples 980 form the decoded PB952 via the multiplexer module 984. Intra prediction generates a predicted block of samples (PB), which is a block in one color component and is derived using the "adjacent samples" in the same color component. The neighboring samples are samples adjacent to the current block and have already been reconstructed by preceding in the block decoding order. When the luma block and the chroma block are co-located, the luma block and the chroma block can use different intra prediction modes. However, the two chroma CBs share the same intra prediction mode.

[0150] If it is shown that the prediction mode of CB in the bitstream 143 is inter prediction, the motion compensation module 934 generates a block of inter prediction samples represented as 938. The block of inter prediction samples 938 is generated using the motion vectors decoded from the bitstream 143 by the entropy decoder 920 and the reference frame index for selecting and filtering a block of samples 998 from the frame buffer 996. The block of samples 998 is obtained from a previously decoded frame stored in the frame buffer 996. In the case of bi-prediction, two blocks of samples are generated and blended to generate the samples of the decoded PB952. Filtered block data 992 from the in-loop filtering module 988 is input to the frame buffer 996. Similar to the in-loop filtering module 868 of the video encoder 120, the in-loop filtering module 988 applies any of the DBF, ALF, and SAO filtering operations. Generally, the motion vectors are applied to both the luma channel and the chroma channel, but the filtering processes for sub-sample interpolation in the luma channel and the chroma channel are different. The frame from the frame buffer 996 is output as the decoded frame 145.

[0151] Although not shown in FIGS. 8 and 9, it is a module for preprocessing the video before encoding and postprocessing the video after decoding to shift the sample values so as to use the range of sample values in each chroma channel more uniformly. The multi-segment linear model is derived in the video encoder 120 and signaled in the bitstream used by the video decoder 804 to reverse the sample shift. This linear model chroma scaling (LMCS) tool provides compression advantages for certain color spaces and contents with some non-uniformities, especially the utilization of a limited range, in the utilization of the sample space where higher quality loss may occur due to the application of quantization.

[0152] FIG. 10A is a schematic block diagram showing a cross-layer tensor inverse bottleneck 148 for restoring the dimensionality of a tensor after compression. FIG. 11 shows a method 1100 for decoding a bitstream, reconstructing a decorrelated feature map, and executing a second part of a CNN.

[0153] Method 1100 can be implemented using a device such as a configured FPGA, ASIC, or ASSP. Alternatively, as described below, method 1100 may be implemented by destination device 140 as one or more software code modules of application program 233 under the execution of processor 205. The software code modules of application program 233 implementing method 1100 can be resident, for example, on hard disk drive 210 and / or memory 206. Method 1100 is repeated for each frame of compressed data in bitstream 143. Method 1100 can be stored in a computer-readable storage medium and / or memory 206. Method 1100 begins with step 1110 of decoding the bitstream.

[0154]

[0155] In step 1110, video decoder 144 decodes one frame 145 from bitstream 143 under the execution of processor 205 as described with reference to FIG. 9. Frame 145 contains the feature map packed as described with reference to FIG. 7. Control of processor 205 proceeds from step 1110 to step 1120 of extracting a first combined tensor.In step 1120, the unpack and inverse quantization module 146 extracts a first combined tensor, for example, a feature map of 557, from the decoded frame 145. The extracted feature map is inverse quantized from the integer domain to the floating-point domain in step 1120. The inverse quantized feature map generated by the operation of step 1120 is shown as tensor 1011 in FIG. 10A. The control of the processor 205 proceeds to step 1130 that extracts a second combined tensor from step 1120.

[0156] In step 1130, the unpack and inverse quantization module 146 extracts a second combined tensor, for example, a feature map of 537, from the frame 145. The extracted feature map is inverse quantized from the integer domain to the floating-point domain in step 1130. The inverse quantized feature map generated by the operation of step 1130 is shown as tensor 1021 in FIG. 10A. The inverse quantized feature quantities 1011 and 1021 correspond to the decoded bottleneck (compressed) tensor 147 as shown in FIG. 10A. The control in the processor 205 proceeds to step 1140 that SSFC decodes the first combined tensor from step 1130 as shown in FIG. 11.

[0157] Steps 1110, 1120, and 1130 operate to decode the frames of the bitstream and obtain two information units of the frame. Tensor 1011 provides the first information unit, and tensor 1021 provides the second information unit. Each information unit corresponds to the feature map of the frame encoded in the bitstream.

[0158] In step 1140, the SSFC decoder 1010 is implemented under the execution of the processor 205. The SSFC decoder executes a neural network layer to decode the first decoded compressed tensor 1011 and generate the first decoded combined tensor 1017. The convolutional layer 1012 receives the tensor 1011 with C’ = 64 channels and outputs the tensor 1013 with F = 256 channels. The tensor 1013 is passed to the batch normalization layer 1014. The batch normalization layer outputs the tensor 1015. The tensor 1015 is passed to the parameterized leaky rectified linear (PReLU) layer 1016. The PReLU layer 1016 outputs the tensor 1017. Returning to FIG. 11, the control of the processor 205 proceeds from step 1140 to step 1150 where the second combined tensor is SSFC decoded.

[0159] In step 1150, the SSFC decoder 1020 is implemented under the execution of the processor 205. The SSFC executes a neural network layer to decode the second decoded compressed tensor 1021 and generate the second decoded combined tensor 1027. The convolutional layer 1022 receives the tensor 1021 with C’ = 64 channels and outputs the tensor 1023 with F = 256 channels. The tensor 1023 is passed to the batch normalization layer 1024. The batch normalization layer 1024 outputs the tensor 1025. The tensor 1025 is passed to the PReLU layer 1026. The PReLU layer 1026 outputs the tensor 1027. Returning to FIG. 11, the control of the processor 205 proceeds from step 1150 to step 1160 where the first plurality of tensors are reconstructed.

[0160] In step 1160, under the execution of the processor 205, the multi-scale feature reconstruction (MSFR) module 1030 generates a first tensor as described in connection with FIG. 10A. Step 1160 receives the tensor 1017 generated by the operation of step 1140. The tensor 1017 is passed to the MSFR module 1030 that generates the decoded tensors 1055 and 1033. When implementing step 1160, the MSFR module 1030 uses an upsampling module 1032, a downsampling module 1034, a convolutional layer 1036, and a summation module 1038. The upsampling module 1032 receives the tensor 1017 and performs interpolation to generate the tensor 1033. The tensor 1033 has a width and height that are twice that of the tensor 1017. The tensor 1033 is output as P'4 and passed to the downsample module 1034. The downsample module 1034 downsamples the tensor 1033 to generate a tensor 1035 having the same number of dimensions as the tensor 1017. The tensor 1035 is provided to the convolutional layer 1036 that outputs the tensor 1037. The summation module 1038 adds the tensors 1037 and 1017 to generate the tensor 1055. The tensor 1055 is output as P'5. The tensor P'5 is the smallest tensor among the set of tensors P'4 and P'5, and is generated, for example, by adding another tensor 1033 generated by the convolution operation in the block 1036 and the first information unit (1011). Returning to FIG. 11, the control of the processor 205 proceeds from step 1160 to step 1170 that reconstructs a second plurality of tensors.

[0161] In step 1170, under the execution of the processor 205, the MSFR module 1030 generates a second tensor from the tensor 1027. The MSFR module 1030 that generates the decoded tensors 1053 and 1043 uses the upsampling module 1042, the downsampling module 1044, the convolutional layer 1046, and the sum module 1048 when executing step 1170. The tensor 1027 is passed to the upsampling module 1042. The upsampling module 1042 performs interpolation and generates a tensor 1043 having a width and height that are twice that of the tensor 1027. The tensor 1043 is output as P'2 and passed to the downsample module 1044. The downsample module 1044 downsamples the tensor 1043 to generate a tensor 1045 having the same dimensions as the tensor 1027. The tensor 1045 is supplied to the convolutional layer 1046 that outputs the tensor 1047. The summation module 1048 adds the tensors 1047 and 1027 to generate a tensor 1053 that is output as P'3. The tensor P'3 is the smallest tensor in the set of tensors P'2 and P'3, and is generated, for example, by adding another tensor 1053 generated by the convolution operation in the block 1046 and the first information unit (1021). The decoded tensors P'2 to P'5 corresponding to the tensors 1043, 1033, 1055, that is, the tensors P2 to P5, form the tensor 149. Referring to FIG. 11, the control in the processor 205 proceeds from step 1170 to step 1180 that executes the second part of the neural network. As shown in FIG. 10A, each tensor of each set of a plurality of tensors (1055, 1033) and (1053, 1043) has a resolution that forms an exponential sequence in which the width and height are doubled between consecutive tensors. For example, the width and height of 1033 are twice the width and height of 1055. The largest tensor in each set of tensors (1033 and 1043) is determined based on the upsampling operation (1032 or 1042 respectively) applied to the feature map of the information unit.

[0162] In step 1180, the CNN head 150 receives the tensor 149 as an input for executing the remaining part of the neural network implemented by the system 100 under the execution of the processor 205. The method 1100 ends by executing step 1180 after processing the tensor related to one frame of the video data. The method 1100 is recalled for each frame of the video data encoded in the bitstream 143.

[0163] FIG. 12A is a schematic block diagram showing the head portion 150 of the CNN for object detection. Different networks can be substituted for the CNN head portion 150 according to the tasks executed by the destination device 140. The input tensor 149 is separated into tensors of each layer (i.e., tensors 1210, 1220, 1234). The tensor 1210 is passed to the CBL module 1212 to generate the tensor 1214. The tensor 1214 is passed to the detection module 1216 and the upscaler module 1222. The detection module 1216 operates to detect the bounding box 1218. The bounding box 1218 is in the form of a detection tensor. The bounding box 1218 is passed to the non-maximum suppression (NMS) module 1248. The NMS module 1248 selects one of the multiple inputs generated by the detection module to generate the detection result 151. Scaling by the width and height of the original video is performed to generate a bounding box that addresses the coordinates in the original video data 113 before resizing for the backbone portion of the network 114. The upscaler module 1222 generates the upscaled tensor 1224 scaled by the width and height of the original video. The upscaled tensor 1224 is passed to the CBL module 1226. The CBL module 1226 generates the tensor 1228 as an output. The tensor 1228 is passed to the detection module 1230 and the upscaler module 1236. The detection module 1230 operates in the same manner as the detection module 1216 to generate the detection tensor 1232. The detection tensor 1232 is supplied to the NMS module 1248.

[0164] The upscaler module 1236 operates in the same manner as the module 1260 and outputs an upscaled tensor 1238. The upscaled tensor 1238 is passed to the CBL module 1240. The CBL module 1240 operates in the same way as the modules 1212 and 1226 and outputs a tensor 1242 to the detection module 1244. The detection module 1244 operates in a similar way to the detection module 1216 and generates a detection tensor 1246. The detection tensor 1246 is supplied to the NMS module 1248.

[0165] The CBL modules 1212, 1226, and 1240 each include a concatenation of five CBL modules. The upscaler modules 1222 and 1236 are each instances of the upscaler module 1260 as shown in FIG. 12B.

[0166] The upscaler module 1260 receives the tensor 1262 and the tensor 1264 as inputs. The tensor 1262 is passed to the CBL module 1266 to generate a tensor 1268. The tensor 1268 is passed to the upsampler 1270 to generate an upsampled tensor 1272 using nearest neighbor interpolation or various other methods. The concatenation module 1274 generates a tensor 1276 by concatenating the upsampled tensor 1272 with the input tensor 1264.

[0167] The detection modules 1216, 1230, and 1244 are instances of the detection module 1280 as shown in FIG. 12C. The detection module 1260 receives the tensor 1282, which is passed to the CBL module 1284 to generate the tensor 1286. The tensor 1286 is passed to the convolutional module 1288 that implements the detection kernel. The detection kernel is a 1×1 kernel applied to generate the output of the three-layer feature map. The detection kernel is 1×1×(B×(5 + C)), where B is the number of bounding boxes that a particular cell can predict, typically 3, and C is the number of classes, which may be 80, resulting in a kernel size of 255 detection attributes. The module 1288 outputs the tensor 1290. The constant "5" represents four bounding box attributes (box center x, y and size scale x, y) and one object confidence ("objectness"). The result of the detection kernel has the same spatial dimensions as the input feature map, but the output depth corresponds to the detection attributes. The detection kernel is applied to each layer, typically 3 layers, resulting in a large number of bounding box candidates. The non-maximum suppression process is applied to the obtained bounding boxes by the NMS module 1048, and redundant boxes such as overlapping prediction values at similar scales are discarded, resulting in the final set of bounding boxes as the output for object detection.

[0168] Figure 13 is a schematic block diagram showing an alternative head portion 1300 of a CNN. The head portion 1300 forms part of an overall network known as 'Faster RCNN' and includes a feature network (i.e., backbone portion 400), a region proposal network, and a detection network. The input to the head portion 1300 is tensor 149. Tensor 149 includes P2 - P6 layer tensors 1310, 1312, 1314, 1316, 1318. The P2 - P6 tensors 1310, 1312, 1314, 1316, and 1318 are input to a region proposal network (RPN) head module 1320. The RPN head module 1320 performs convolution on the input tensor to generate an intermediate tensor. The intermediate tensor is input to two subsequent sibling layers of module 1320, one for classification and the other for bounding box, i.e., 'region of interest' (ROI) regression, as classification and bounding box 1322. Classification and bounding box 1322 are passed to an NMS module 1324. The NMS module prunes redundant bounding boxes by removing overlapping boxes with lower scores to generate a pruned bounding box 1326. The bounding box 1326 is passed to a region of interest (ROI) pooler 1328. The ROI pooler 1328 also receives tensors P2 through P6 and generates a fixed - size feature map from various input - size maps using a max - pooling operation. In the operation performed by 1328, subsampling takes the maximum value of each group of input values to generate one output value for the output tensor.

[0169] In the arrangement of the CNN backbone 400 and the CNN head 1300, the 'P6' layer tensor 429 is omitted from the output tensor 115, and in the CNN head 1300, the P6 input tensor 1318 is generated by performing a 'Maxpool' operation with a stride equal to 2 on the P5 tensor 1316. Since the P6 layer can be reconstructed from the P5 layer, there is no need to separately encode and decode the P6 layer as an explicit FPN layer within the first set of tensors or the second set of tensors.

[0170] The inputs to the ROI Pooler 1328 are the P2 - P5 feature maps 1310, 1312, 1314, and 1316 (corresponding to 1043, 1053, 1033, and 1055 in Figure 10A respectively), and the region of interest proposals 1326. Each proposal (ROI) from 1326 is associated with a part of the feature maps (1310 - 1316) to generate a fixed - size map. The fixed - size map is of a size independent of the underlying part of the feature maps 1310 - 1316. One of the feature maps 1310 - 1316 is selected, for example, according to the following rule: floor(4 + log2(sqrt(box_area) / 224)) (where 224 is the canonical box size), so that the resulting cropped map has sufficient detail. In this way, the ROI Pooler 1328 crops the input feature map according to the proposal 1326 to generate the tensor 1330. The tensor 1330 is fed to the fully - connected (FC) neural network head 1332. The FC head 1332 executes two fully - connected layers to generate the class scores and the bounding box predictor delta tensor 1334. The class scores are generally tensors of 80 elements, and each element corresponds to the predicted score of the corresponding object category. The bounding box prediction delta tensor is a tensor of 80×4 = 320 elements and contains the bounding boxes of the corresponding object categories. The final processing is executed by the output layer module 1336, which receives the tensor 1334 and executes filtering operations to generate the filtered tensor 1338. Objects with low scores (low classifications) are removed from further consideration. The non - maximum suppression module 1340 receives the filtered tensor 1334 and removes overlapping bounding boxes by removing overlapping boxes with low classification scores, resulting in the inference output tensor 151.

[0171] FIG. 14 is a schematic block diagram showing a top-down multi-scale feature reconstruction (MSFR) stage 1400 of a destination device 140. The top-down MSFR stage 1400 is an alternative embodiment to the MSFR 1030 of FIG. 10A. Top-down feature reconstruction includes upsampling from a lower-resolution layer (at the "top" of the FPN) to generate a higher-resolution version of the layer. When multiple spatial representations of the FPN layers are used, the top-down operation is performed to reconstruct each tensor within each given spatial representation. The first combined tensor 1017 and the second combined tensor 1027 are received by the top-down MSFR 1400 and are directly output as tensors P'5 1055a and P'3 1053a, respectively. Tensor 1017 is passed to an upsampling module 1410. Module 1410 performs 2x upsampling using an interpolation filter to generate tensor 1411. Tensor 1017 is also passed to an upsampling module 1412. Module 1412 performs nearest-neighbor upsampling to generate a tensor 1413 having twice the width and height of tensor 1017. Tensor 1413 is passed to a convolutional layer 1414 to generate a tensor 1415 having the same dimensionality as tensor 1413. An addition module 1416 adds tensors 1415 and 1411 to generate tensor 1417. Tensor 1417 is passed to a convolutional stage 1418 to generate a tensor P'4 1033a. Tensor 1033a has the same dimensionality as tensor 1417.

[0172] Tensor 1027 is passed to the upsampling module 1420, and the upsampling module 1420 performs 2x upsampling using an interpolation filter to generate tensor 1421. Tensor 1027 is also passed to the upsampling module 1422. Module 1422 performs nearest neighbor upsampling to generate tensor 1423 which has twice the width and height of tensor 1027. Tensor 1423 is passed to the convolutional layer 1424 to generate tensor 1425 which has the same dimensionality as tensor 1423. The summation module 1426 sums tensors 1425 and 1421 to generate tensor 1427. Tensor 1427 is passed to the convolutional stage 1428 to generate tensor P’2 1043a which has the same dimension as tensor 1427.

[0173] When the top-down MSFR module 1400 is in use, the output tensors 1055a, 1033a, 1053a, and 1043a are supplied to the rest of the network as tensors 1055, 1033, 1053, and 1043 respectively. Various approaches are possible to reconstruct each FPN layer, but in any case the FPN layers can be grouped into multiple (e.g., two) layer sets, each layer is individually fused by the bottleneck encoder 116 and extracted by the bottleneck decoder 148.

[0174] Regardless of whether the bottleneck decoder 148 implements the MSFR decoder as decoder 1030 or the MSFR decoder as decoder 1400, it operates to generate the first and second tensor sets respectively using the first and second information units decoded by the operations of steps 1110 to 1130. Each set of tensors can be determined independently such that steps 1120 and 1130 can be executed independently, simultaneously, or in a different order as required. The first set or sets of tensors are provided by tensors P’4 and P’5 (1055 and 1033 in FIG. 10A, or 1055a and 1033a in FIG. 14). The second set or sets of tensors are provided by tensors P’2 and P’3 (1053 and 1043 in FIG. 10A, or 1053a and 1043a in FIG. 14). Since the first set of tensors is related to P4 and P5, the associated feature maps have a different spatial resolution from the second set related to P3 and P2. Combining both sets of tensors corresponds to a single hierarchical representation of the image frame. For example, the hierarchical representation can be a feature pyramid network generated by the operation of the CNN backbone 400 in FIG. 4, or a set of tensors generated by the operation of the CNN backbone 300 in FIG. 3A.

[0175] Regardless of the implementation used for the MSFR decoder, the described arrangement enables each set of the first and second plurality of tensors to be determined using neural network layers. This layer is typically a convolutional layer, a batch normalization layer, a PReLu layer. Other layers can be used for the SSFC decoders 1010 and 1020 provided that decoding from a single packed frame such as frame 700 can be achieved.

[0176] The first and second sets of tensors can be provided to the CNN head 150 to perform tasks such as object detection, frame segmentation. For example, P’5 to P’2 can be processed as shown in relation to FIG. 13.

[0177] Alternatively, the CNN 150 of FIG. 12 can be supplied with tensors 1210, 1220, 1234 from the bottleneck decoder 148. In an implementation using the CNN head 150 of FIG. 12, the bottleneck decoder 148 operates to output the tensor 1017 from the SSFC decoder 1010 and processes the first group. The tensor 1017 is convolved with a 256-channel input and a 1024-channel output, and provides the tensor 1210 with modules 1032, 1034, 1036, 1038 omitted to the CNN head. FIG. 10B shows a convolution arrangement 1090 for use with the bottleneck decoder 148 of FIG. 10A when the CNN is implemented in a manner similar to the embodiment of FIG. 12. This arrangement 1090 includes convolutional layers 1060, 1070, 1080. The tensor 1017 of FIG. 10A is input to the convolutional layer 1060, and the convolutional layer 1060 outputs the tensor 1210 having 1024 channels. For the second group, tensors 1053 and 1043 are output from the bottleneck decoder 148 each having 256 channels. The tensor 1053 is input to the convolutional layer 1070 to generate a tensor 1220 having 512 channels. The tensor 1043 is input to the convolutional layer 1080 that generates a tensor 1234 having 256 channels. Referring to FIG. 3, the tensor 329 has 1024 channels and is reduced to 64 channels by the SSFC encoder 550. Tensors 326 and 322 have 512 channels and 256 channels respectively, which are first reduced to 256 channels by the MSFF module 510 and further reduced to 64 channels by the SSFC encoder 530.

[0178] In the exemplary implementations described in connection with FIGS. 5 and 10, the first and second tensors (e.g., 537 and 557) have different dimensions but the same number of channels. In another arrangement of the bottleneck encoder 116 and the bottleneck decoder 148, the number of channels of the second compressed tensors 557, 1011 is different from the number of channels of the first compressed tensors 537, 1021. For example, the second compressed tensor can have a lower or smaller number of channels than the first compressed tensor. Preferably, the second compressed tensor for the higher spatial resolution FPN layers (P’2 and P’3) has a smaller number of channels than the first compressed tensor for the lower spatial resolution FPN layers (P’4 and P’5). The number of channels of the second compressed tensor 537 can be 32 channels or 48 channels. The second compressed tensor 537 has a width and height four times that of the first compressed tensor 557, and the reduction in the number of channels results in a smaller area required for frame 700. The second compressed tensor 537 represents fewer decomposed features of frame 113, and since most of the semantically meaningful information is included in the first compressed tensor 557, showing a smaller number of channels is appropriate for achieving high task performance.

[0179] In an arrangement of the source device 1,100 and the destination device 140 with a three-layer FPN, such as when the “YOLOv3” network is used, the tensors can be grouped into two sets, divided into a first group with two FPN layers and a second group with one FPN layer. The group with one FPN layer requires a bottleneck layer associated with the SSFC encoder (530 or 550) and the SSFC decoder (1010 or 1020), but does not require a layer associated with the MSFF510 or the MSFR1030. When the number of channels varies between the layers of the FPN, the connection of the MSFF module 510 (e.g., 514 or 522) connects such tensors. Additional convolutional layers are inserted into the MSFR (e.g., 1030 or 1400) to restore the number of channels of each FPN layer.

[0180] Industrial Applicability The described arrangement is applicable to the computer and data processing industries, particularly digital signal processing for encoding and decoding signals such as video and image signals, and achieves high compression efficiency.

[0181] This agreement separately describes a method of splitting the MFSC into two operations and compressing (and expanding correspondingly after decoding) the hierarchical structure of the data frame (such as a feature pyramid network) before encoding. By applying the MFSC technology separately, a certain degree of cross-layer fusion can be achieved without significant degradation of spatial details generated by a single MFSC operation. In tasks that require retaining larger spatial details, such as instance segmentation, using separate MSFC technologies results in a higher mAP than integrating all FPN (hierarchical) layers into a single tensor. Furthermore, different numbers of channels can be used in different instances of MFSC, so the increase in area can be suppressed as needed. MFSC has conventionally been used to "fuse" or "constrain" all data within the FPN. In contrast to fusing or converging all data within the FPN, the MFSC can be implemented while suppressing the unfavorable effects on mAP by using multiple MFSC instances in a single hierarchical structure or FPN.

[0182] What has been described above are only some embodiments of the present invention, and modifications and / or changes can be made without departing from the scope and spirit of the present invention, and the embodiments are illustrative and not restrictive.

[0183] As used herein, the term "comprising" means "mainly including, but not necessarily alone" or "having" or "including", and does not mean "consisting only of". Variations of the word "comprising" such as "comprises" and "comprises" have corresponding various meanings.

Claims

Claim 1 A method for decoding at least a plurality of tensors forming a hierarchical representation of a feature map for a single frame from a bitstream, comprising: decoding a first information unit from the bitstream; decoding a second information unit from the bitstream; determining a first plurality of tensors from the first information unit, wherein at least one tensor of the first plurality of tensors has a spatial resolution different from that of other tensors of the first plurality of tensors; determining a second plurality of tensors from the second information unit, wherein at least one tensor of the second plurality of tensors has a spatial resolution different from that of other tensors of the second plurality of tensors; wherein each tensor of the first plurality of tensors has a spatial resolution different from that of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame. A method. Claim 2 Each tensor of the first plurality of tensors and the second plurality of tensors has a resolution forming an exponential sequence in which the width and height are doubled between consecutive tensors. The method according to claim 1. Claim 3 The first plurality of tensors and the second plurality of tensors have different numbers of channels. The method according to claim 1. Claim 4 Among the first plurality of tensors and the second plurality of tensors, the plurality of tensors with higher spatial resolution have fewer channels than the other plurality of tensors. The method according to claim 1. Claim 5 Each maximum tensor of the first plurality of tensors and the second plurality of tensors is determined based on an upsampling operation applied to a corresponding one of the feature maps of the first information unit and the second information unit. The method according to claim 1. Claim 6 The determination of the first plurality of tensors and the determination of the second plurality of tensors are independent of each other. The method according to claim 1. Claim 7 The first plurality of tensors and the second plurality of tensors are determined using neural network layers. The method according to claim 1. Claim 8 The first information unit is used to determine the smallest tensor among the first plurality of tensors The method according to claim 1

9. The second information unit is used to determine the smallest tensor among the second plurality of tensors The method according to claim 1

10. A method for encoding at least a plurality of tensors into a bitstream, the plurality of tensors forming a hierarchical representation of a feature map for a single frame, the method comprising using a convolution operation to determine a first information unit from a first plurality of tensors, wherein a feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from a feature map of other tensors of the first plurality of tensors, said using using a convolution operation to determine a second information unit from a second plurality of tensors, wherein a feature map of at least one tensor of the second plurality of tensors has a different spatial resolution from a feature map of other tensors of the second plurality of tensors, and a feature map of each tensor of the first plurality of tensors has a different spatial resolution from a feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame, said using encoding the first information unit into the bitstream encoding the second information unit into the bitstream A method comprising

11. A decoder for decoding at least a plurality of tensors forming a hierarchical representation of a feature map for a single frame from a bitstream, the decoder comprising decoding a first information unit from the bitstream decoding a second information unit from the bitstream determining a first plurality of tensors from the first information unit, wherein a feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from a feature map of other tensors of the first plurality of tensors determining a second plurality of tensors from the second information unit, wherein a feature map of at least one tensor of the second plurality of tensors has a different spatial resolution from a feature map of other tensors of the second plurality of tensors configured as such The feature map of each tensor of the first plurality of tensors has a different spatial resolution from the feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame Decoder Claim 12 An encoder for encoding at least a plurality of tensors into a bitstream, the plurality of tensors forming a hierarchical representation of a feature map for a single frame, the encoder comprising: using a convolution operation to determine a first information unit from a first plurality of tensors, wherein the feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from the feature maps of other tensors of the first plurality of tensors; using a convolution operation to determine a second information unit from a second plurality of tensors, wherein the feature map of at least one tensor of the second plurality of tensors has a different spatial resolution from the feature maps of other tensors of the second plurality of tensors, and the feature map of each tensor of the first plurality of tensors has a different spatial resolution from the feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame; encoding the first information unit into the bitstream; encoding the second information unit into the bitstream; An encoder configured as described above Claim 13 A non-transitory computer-readable storage medium storing a program for executing a method of decoding at least a plurality of tensors forming a hierarchical representation of a feature map for a single frame from a bitstream, the method comprising: decoding a first information unit from the bitstream; decoding a second information unit from the bitstream; determining a first plurality of tensors from the first information unit, wherein the feature map of at least one tensor of the first plurality of tensors has a different spatial resolution from the feature maps of other tensors of the first plurality of tensors; Determining a second plurality of tensors from the second information unit, wherein a feature map of at least one tensor of the second plurality of tensors has a spatial resolution different from that of feature maps of other tensors of the second plurality of tensors, said determining; comprising; A feature map of each tensor of the first plurality of tensors has a spatial resolution different from that of a feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame A non-transitory computer-readable storage medium. [

14. ] A system including a memory and a processor, The processor is configured to execute code stored on the memory to implement a method of decoding at least a plurality of tensors that form a hierarchical representation of a feature map for a single frame, The method includes, decoding a first information unit from the bitstream; decoding a second information unit from the bitstream; determining a first plurality of tensors from the first information unit, wherein a feature map of at least one tensor of the first plurality of tensors has a spatial resolution different from that of feature maps of other tensors of the first plurality of tensors, said determining; determining a second plurality of tensors from the second information unit, wherein a feature map of at least one tensor of the second plurality of tensors has a spatial resolution different from that of feature maps of other tensors of the second plurality of tensors, said determining; comprising; A feature map of each tensor of the first plurality of tensors has a spatial resolution different from that of a feature map of each tensor of the second plurality of tensors, and the plurality of tensors of the first plurality of tensors and the second plurality of tensors correspond to the hierarchical representation of the feature map for the single frame System