Method, apparatus and system for encoding and decoding tensor

By introducing bottleneck encoder and decoder into the convolutional neural network, the intermediate tensor data is compressed, and the problem of excessive computational complexity in the prior art is solved, thereby achieving efficient feature compression and optimization of computing resources.

CN120019662APending Publication Date: 2025-05-16CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071969.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-07-28
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress intermediate tensor data in convolutional neural networks, resulting in excessive computational complexity in running leading-edge CNNs in edge devices.

Method used

By segmenting the neural network between the main body and head of the CNN and introducing an additional neural network layer at this interface, the feature map is compressed and decoded using bottleneck encoder and decoder to reduce the space area.

Benefits of technology

It realizes efficient compression of intermediate tensor data, reduces the computational complexity and power consumption of CNN running in edge devices, and improves the system's collaborative intelligence capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019662A_ABST
    Figure CN120019662A_ABST
Patent Text Reader

Abstract

Systems and methods for encoding at least a plurality of tensors in a bitstream, the plurality of tensors forming a hierarchical representation for a single frame. The method includes deriving a first information unit from a plurality of tensors forming a hierarchical representation, the plurality of tensors including at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a first mode, at least a first information unit is encoded in a bitstream. In the second mode, the method further comprises deriving a second information unit from at least the first tensor; and encoding the second information unit and the first information unit in a bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit under 35 U.S.C. §119 of the filing date of Australian patent application 2022252784, filed on October 13, 2022, which is incorporated herein by reference in its entirety as if fully set forth herein. Technical Field

[0003] The present invention generally relates to digital video signal processing, and in particular, to methods, devices and systems for encoding and decoding tensors from convolutional neural networks. The present invention also relates to a computer program product comprising a computer-readable medium having recorded thereon a computer program for encoding and decoding tensors from convolutional neural networks using video compression techniques. Background Art

[0004] Convolutional neural networks (CNNs) are emerging technologies for use cases involving machine vision, such as object detection, instance segmentation, object tracking, human pose estimation, and action recognition. Applications of CNNs may involve the use of "edge devices" with sensors and some processing capabilities that are coupled to application servers as part of the "cloud". CNNs may require relatively high computational complexity, which exceeds the computational complexity that can typically be provided in terms of utilizing the computing capacity or power consumption of edge devices. Executing CNNs in a distributed manner has become a solution for running cutting-edge networks using limited-capacity edge devices. In other words, distributed processing enables traditional edge devices to still provide the capabilities of cutting-edge CNNs by distributing processing between edge devices and external processing components such as cloud servers. This distributed network architecture can be called "collaborative intelligence" and provides benefits such as reusing partial results from the first part of the network for several different second parts (possibly each part is optimized for different tasks). Collaborative intelligence architectures introduce the need for efficient compression of tensor data for transmission over networks such as WANs.

[0005] CNNs typically include many layers such as convolutional layers and fully connected layers, where data is passed from one layer to the next in the form of "tensors". Splitting the network across different devices introduces the need to compress the intermediate tensor data passed from one layer to the next within the CNN, such compression can be called "feature compression" because the intermediate tensor data is often called "features" or "feature maps" and represents a partially processed form of an input such as an image frame or video frame. The International Organization for Standardization / International Electrotechnical Commission Joint Technical Committee 1 / Subcommittee 29 / Working Group 2-8 (ISO / IEC JTC1 / SC29 / WG2-8) (also known as the "Moving Picture Experts Group" (MPEG)) has been assigned the task of studying compression techniques in various contexts and often related to video. WG2 "MPEG Technical Requirements" has established the "Machine Video Compression" (VCM) ad-hoc group, which is entrusted to study compression used for machine consumption, as well as feature compression. The feature compression commission is in the exploratory phase of publishing a Call for Evidence (CfE) for techniques that can significantly outperform feature compression results achieved using state-of-the-art standardized techniques.

[0006] CNNs typically require pre-determination of weights for various layers during a training phase, in which a very large amount of training data is passed through the CNN and the results determined by the trained network are compared with the ground truth associated with the training data. The difference between the obtained result and the expected result is represented as a "loss" and measured using a "loss function". Using the determined loss, a process (such as stochastic gradient descent (SGD)) is performed to update the network weights. Network weight updates typically involve backpropagation of a "gradient" that indicates the delta to be applied to the network weights, starting from the output layer of the network and ending at the input layer to the network, and covering all intermediate or "hidden" layers of the network. The rate of weight updates is scaled by a "learning rate" hyperparameter, which is typically set to facilitate the training process to find a global minimum in terms of loss (i.e., the highest possible task performance for the network architecture and training data), while avoiding the training process from being "stuck" in a local minimum. Being stuck in a local minimum corresponds to obtaining suboptimal task performance for the network architecture and being unable to find new weight values ​​that can lead to higher task performance. The network weights are repeatedly updated by supplying input data and ground truth data organized into "batches" to iteratively refine the network performance until no further improvement in accuracy can be achieved. Iterations of the entire training data set form an "epoch" of training, and training typically requires multiple epochs to achieve a high level of performance for the task. The trained network can then be used for deployment, operating in a mode where the weights are fixed and gradients for weight updates are omitted. The process of executing a pre-trained CNN with an input and gradually transforming the input into an output according to the topology of the CNN is generally referred to as "inference."

[0007] Typically, a tensor has four dimensions, namely: batch, channel, height, and width. The first dimension, "batch", is typically of size 1 when performing inference on video data, and indicates that one frame is passed through the CNN as a batch. When training the network, the value of the batch dimension can increase so that multiple frames are passed through the network in each batch before updating the network weights according to a predetermined "batch size". Multiple frames of video can be passed through as a single tensor with a batch dimension whose size increases according to the number of frames of a given video. However, for practical considerations related to memory consumption and access, inference on video data is typically performed frame by frame. The "channel" dimension indicates the number of concurrent "feature maps" for a given tensor, and the height and width dimensions indicate the size of the feature map of a particular stage of the CNN. The number of channels through the CNN varies depending on the network architecture. The size of the feature map also varies depending on the subsampling that occurs in a particular network layer.

[0008] The overall complexity of CNNs tends to be relatively high with a relatively large number of multiply-accumulate (MAC) operations performed and a large number of intermediate tensors written and read from memory, as well as weights read for the performance of each layer of the CNN. Therefore, partitioning the neural network into parts enables more complex networks to be implemented even in less capable edge devices.

[0009] Feature compression can benefit from existing video compression standards, such as the universal video coding (VVC) developed by the Joint Video Experts Group (JVET). It is expected that VVC will address the continuous demand for even higher compression performance, and address the growing market demand for service delivery over WAN (wherein the bandwidth cost is relatively high), especially with the increase in the capabilities of video formats (e.g., with higher resolution and higher frame rate). VVC can be implemented in contemporary silicon processes and provides an acceptable compromise between the achieved performance and the implementation cost. The implementation cost can be considered to be one or more than one aspect in terms of silicon area, CPU processor load, memory utilization, and bandwidth. Part of the versatility of the VVC standard is the wide selection of tools that can be used to compress video data, and the wide range of applications that VVC is suitable for. Other video compression standards (such as high-efficiency video coding (HEVC) and AV-1, etc.) can also be used for feature compression applications.

[0010] The video data comprises a sequence of frames of image data, each frame comprising one or more color channels. Where feature map data is to be represented in packed frames, a monochrome frame having only luma and no color channels is generally sufficient. When only luma samples are present, the resulting monochrome frame is said to use a "4:0:0 chroma format".

[0011] The VVC standard specifies a "block-based" architecture in which a frame is first partitioned into an array of square regions called "coding tree units" (CTUs). In VVC, a CTU generally occupies 128×128 luminance samples. Other possible CTU sizes when using the VVC standard are 32×32 and 64×64. However, the CTUs at the right and lower edges of each frame may be smaller in area, where implicit splitting occurs to ensure that the coding blocks remain in the frame. Associated with each CTU is a "coding tree" (also called a "coding unit" (CU)) that defines the decomposition of the area of ​​the CTU into a set of blocks. Blocks applicable to only luminance channels or only chrominance channels are called "coding blocks" (CBs). The prediction of the content of the coding block is maintained in a "prediction block" (PB) or a "prediction unit" (PU), and the residual block defining the array of sample values ​​to be added and combined with the PB or PU is called a "transform block" (TB) or a "transform unit" (TU), due to the typical use of transform processing in the generation of TBs or TUs.

[0012] Despite the above distinction between "unit" and "block", the term "block" may be used as a general term for an area or region of a frame for which operations are applied to all color channels.

[0013] For each CU, a prediction unit (PU) ("prediction unit") is generated for the content (sample values) of the corresponding region of the frame data. In addition, a representation of the difference (or "spatial domain" residual) between the prediction and the content of the region seen at the input of the encoder is formed. The differences in each color channel can be transformed and encoded into a sequence of residual coefficients, thereby forming one or more TUs for a given CU. The applied transform can be a discrete cosine transform (DCT) or other transform applied to each block of residual values. The transform is applied separately (i.e., a two-dimensional transform is performed twice (once horizontally and once vertically)). The block is first transformed by applying a one-dimensional transform to each row of samples in the block. The partial result is then transformed by applying a one-dimensional transform to each column of the partial result to produce a final block of transform coefficients that substantially decorrelate the residual samples. The VVC standard supports transforms of various sizes, including transforms of rectangular blocks whose dimensions on each side are powers of 2. The transform coefficients are quantized for entropy encoding into the bitstream.

[0014] A PB or PU in a VVC may be generated using either an intra prediction or an inter prediction process. Intra prediction involves using previously processed samples in the frame being used to generate a prediction of the current block of data samples in that frame. Inter prediction involves using a block of samples obtained from a previously decoded frame to generate a prediction of the current block of samples in a frame. The block of samples obtained from the previously decoded frame is offset relative to the spatial position of the current block according to a motion vector, which is typically filtered. The intra prediction block can be: (i) uniform sample values ​​("DC intra prediction"), (ii) a plane with an offset and horizontal and vertical gradients ("planar intra prediction"), (iii) a population of blocks of neighboring samples applied in a particular direction ("angular intra prediction"), or (iv) the result of a matrix multiplication using neighboring samples and selected matrix coefficients.

[0015] VVC can be used to compress intermediate feature maps from the first part (the "backbone") of a neural network that is separated into two parts. In compression, feature maps from the backbone are arranged into frames and quantized from a floating point domain to a sample domain suitable for compression as video data. In order to reduce the spatial area of ​​the feature maps, additional neural network layers can be implemented at the interface between the VVC encoder and decoder and the intermediate point in the CNN where the split occurs. Training is performed for such additional network layers that may not be suitable for the varying and unpredictable feature map data encountered. Training may not result in a CNN having adaptability to operating points of various qualities in terms of task performance. The operating points of the encoder and decoder may also change during operation, where it is necessary to support varying quality levels of the reconstructed tensors to be supplied to the rest of the network on the decoder side. Summary of the invention

[0016] It is an object of the present invention to substantially overcome or at least ameliorate one or more disadvantages of existing arrangements.

[0017] An aspect of the present disclosure provides a method for encoding at least multiple tensors into a bitstream, wherein the multiple tensors form a hierarchical representation for a single frame, the method comprising: deriving a first information unit from the multiple tensors forming the hierarchical representation, the multiple tensors including at least a first tensor and a second tensor, a feature map of the first tensor having a larger spatial resolution than a feature map of the second tensor; in a first mode, encoding at least the first information unit into the bitstream; in a second mode, deriving a second information unit from at least the first tensor; and in the second mode, encoding the second information unit and the first information unit into the bitstream.

[0018] Another aspect of the present disclosure provides a method for decoding at least multiple tensors from a bitstream, the multiple tensors forming a hierarchical representation for a single frame, the method comprising: decoding a bitstream including at least a first information unit from the bitstream; in a first mode, deriving multiple tensors forming the hierarchical representation from the first information unit, the multiple tensors including at least a first tensor and a second tensor, and a feature map of the first tensor has a larger spatial resolution than the feature map of the second tensor; in a second mode, decoding a second information unit from the bitstream; and in the second mode, deriving multiple tensors forming at least a part of the hierarchical representation from at least a part of the first information unit and the second information unit, the second information unit corresponding to the first tensor.

[0019] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing a program for executing a method for encoding at least multiple tensors into a bitstream, wherein the multiple tensors form a hierarchical representation for a single frame, the method comprising: deriving a first information unit from multiple tensors forming the hierarchical representation, the multiple tensors including at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a first mode, encoding at least the first information unit into the bitstream; in a second mode, deriving a second information unit from the first tensor; and in the second mode, encoding the second information unit and the first information unit into the bitstream.

[0020] Another aspect of the present disclosure provides an encoder configured to encode at least multiple tensors into a bitstream by the following operations, wherein the multiple tensors form a hierarchical representation for a single frame: deriving a first information unit from the multiple tensors forming the hierarchical representation, the multiple tensors including at least a first tensor and a second tensor, wherein a feature map of the first tensor has a larger spatial resolution than a feature map of the second tensor; in a first mode, encoding at least the first information unit into the bitstream; in a second mode, deriving a second information unit from at least the first tensor; and in the second mode, encoding the second information unit and the first information unit into the bitstream.

[0021] Another aspect of the present disclosure provides a system, comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory, the code being used to implement a method for encoding at least multiple tensors into a bitstream, the multiple tensors forming a hierarchical representation for a single frame, the method comprising: deriving a first information unit from multiple tensors forming the hierarchical representation, the multiple tensors including at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a first mode, encoding at least the first information unit into the bitstream; in a second mode, deriving a second information unit from the first tensor; and in the second mode, encoding the second information unit and the first information unit into the bitstream.

[0022] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing a program for executing a method for decoding at least multiple tensors from a bitstream, the multiple tensors forming a hierarchical representation for a single frame, the method comprising: decoding a bitstream including at least a first information unit from the bitstream; in a first mode, deriving multiple tensors forming the hierarchical representation from the first information unit, the multiple tensors including at least a first tensor and a second tensor, the feature map of the first tensor having a greater spatial resolution than the feature map of the second tensor; in a second mode, decoding a second information unit from the bitstream; and in the second mode, deriving multiple tensors forming at least a part of the hierarchical representation from at least a part of the first information unit and the second information unit, the second information unit corresponding to at least the first tensor.

[0023] Another aspect of the present disclosure provides a decoder configured to decode at least multiple tensors from a bitstream, the multiple tensors forming a hierarchical representation for a single frame by: decoding a bitstream including at least a first information unit from the bitstream; in a first mode, deriving multiple tensors forming the hierarchical representation from the first information unit, the multiple tensors including at least a first tensor and a second tensor, wherein a feature map of the first tensor has a greater spatial resolution than a feature map of the second tensor; in a second mode, decoding a second information unit from the bitstream; and in the second mode, deriving multiple tensors forming at least a portion of the hierarchical representation from at least a portion of the first information unit and the second information unit, the second information unit corresponding to at least the first tensor.

[0024] Another aspect of the present disclosure provides a system, comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory, the code being used to implement a method for decoding at least multiple tensors from a bitstream, the multiple tensors forming a hierarchical representation for a single frame, the method comprising: decoding a bitstream including at least a first information unit from the bitstream; in a first mode, deriving multiple tensors forming the hierarchical representation from the first information unit, the multiple tensors including at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a second mode, decoding a second information unit from the bitstream; and in the second mode, deriving multiple tensors forming at least a part of the hierarchical representation from at least a part of the first information unit and the second information unit, the second information unit corresponding to at least the first tensor.

[0025] Other aspects are also disclosed. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] At least one embodiment of the present invention will now be described with reference to the following drawings, in which:

[0027] Figure 1 is a schematic block diagram illustrating a distributed machine task system;

[0028] Figure 2A and Figure 2B Form a practical Figure 1 A schematic block diagram of a general computer system for a distributed machine task system;

[0029] Figure 3A is a schematic block diagram showing the functional modules of the main body of CNN;

[0030] Figure 3B It is shown Figure 3A A schematic block diagram of a residual block;

[0031] Figure 3C It is shown Figure 3A A schematic block diagram of a residual unit of ;

[0032] Figure 3D It is shown Figure 3A Schematic block diagram of the CBL module;

[0033] Figure 4 is a schematic block diagram showing the functional modules of an alternative backbone of a CNN;

[0034] Figure 5 is a schematic block diagram illustrating a cross-layer tensor bottleneck encoder for reducing tensor dimensionality prior to compression;

[0035] Figure 6 Show a method for performing the first part of a CNN, shrinking using a bottleneck encoder, and encoding the resulting shrunken feature map;

[0036] Fig. 7A is a schematic block diagram illustrating a packing arrangement for a plurality of compressed tensors;

[0037] Figure 7B is a schematic block diagram illustrating a bit stream for a plurality of compressed tensors;

[0038] Figure 8 is a schematic block diagram showing the functional modules of a video encoder;

[0039] Fig. 9 is a schematic block diagram showing the functional modules of a video decoder;

[0040] Fig.10 is a schematic block diagram illustrating a cross-layer tensor bottleneck decoder for restoring tensor dimensions after compression;

[0041] Fig.11 A method for decoding a bitstream, reconstructing a decorrelated feature map, and performing the second part of a CNN is shown;

[0042] Fig. 12A is a schematic block diagram showing the head of a CNN;

[0043] Fig. 12B It is shown Fig. 12A A schematic block diagram of an upgrader module;

[0044] Fig. 12C It is shown Fig. 12A A schematic block diagram of a detection module;

[0045] Fig.13 is a schematic block diagram showing an alternative head of a CNN;

[0046] Fig.14 is a method for encoding a tensor into a reduced-dimensional form; and

[0047] Fig.15 is a method for decoding a tensor from a reduced-dimensional form into a tensor of original dimension. DETAILED DESCRIPTION

[0048] Where steps and / or features having the same reference numerals are referenced in any one or more of the accompanying drawings, these steps and / or features have the same function(s) or operation(s) for the purposes of this specification, unless otherwise intended.

[0049] A distributed machine task system may include edge devices, such as a web camera or smart phone that generates intermediate compressed data. The distributed machine task system may also include end devices, such as server farm ("cloud") based applications that operate on the intermediate compressed data to generate some task result. In addition, edge device functionality may be embodied in the cloud, and the intermediate compressed data may be stored for later processing (potentially for multiple different tasks as needed). Examples of machine tasks include object detection and instance segmentation, both of which generate task results measured as "mean average precision" (mAP) for detections that exceed a threshold (such as 0.5) of intersection over union (IoU). Another example machine task is object tracking, with a mean object tracking accuracy (MOTA) score as a typical task result.

[0050] A convenient form of intermediate compressed data is a compressed video bitstream, due to the availability of high performance compression standards and their implementations.Video compression standards typically operate on integer samples of some given bit depth (such as 8 or 10 bits per sample, etc.) arranged in a planar array.

[0051] Tensors typically have the following dimensions: batch size, number of channels, height, and width. For example, a tensor of dimension [1,256,76,136] would be considered to contain a batch of tensors containing two hundred and fifty-six (256) feature maps, each of size 136 × 76. For video data, inference is typically performed on one frame at a time, rather than using tensors containing multiple frames, resulting in a batch size of 1.

[0052] Figure 1 1 is a schematic block diagram showing the functional modules of a distributed machine task system 100 that implements a neural network that is split into two parts (e.g., one part can be in an edge device and the other part can be in a cloud server). The system 100 can be used to implement methods for decorrelating, packing, and quantizing feature maps into a plane frame to encode feature maps and decode feature maps from encoded data. These methods can be implemented so that compressed data is encoded to reduce bit rate while adapting to changing statistics encountered in the input data. In this way, the system 100 provides the ability to perform "real-time" training (or "refined training") on input tensor data to generate weight updates for active networks in encoders and decoders. Although training the task network requires ground truth for the input data, the refined training is only applied to the part of the network that forms the bottleneck encoder and decoder. The goal of the refined training is to preserve the data passed with minimal degradation. Therefore, the tensor at the input to the bottleneck encoder forms the ground truth for the output of the bottleneck decoder. Refined training alleviates the need to fully train the predetermined network weights to anticipate all conceivable input data.

[0053] refer to Figure 6 (which illustrates method 600 for performing the first part of a CNN) and Fig.11 (which shows a method 1100 for performing the second part of CNN) to illustrate Figure 1 . Also refer to Fig. 7A (which shows the packed arrangement of feature maps from compressed tensors into monochrome video frames) and Figure 7B (which shows the bitstream format used when encoding and decoding tensors).

[0054] System 100 implements a "FasterRCNN" network, which is used for object detection and is split into a trunk and a head at a midpoint typically described as a "P layer" in the described example. Other networks such as "MaskRCNN" and the like may be implemented in system 100. Notably, the trunks of FasterRCNN and MaskRCNN have the same topology and dimensions of convolutions, batch normalization, activation functions, etc. The head of FasterRCNN is a subset of the head of MaskRCNN, where MaskRCNN includes a "mask head" for generating instance segmentation maps in addition to the bounding box outputs present in both FasterRCNN and MaskRCN. The mask head includes two convolutional layers and is for (refer to Fig.13 Each “region of interest” produced by the RoIAlign stage (as illustrated) produces a segmentation map. Therefore, MaskRCNN can be used to perform both object detection and instance segmentation, with additional complexity in the network head due to the use of mask headers.

[0055] The system 100 includes a source device 110 for generating encoded tensor data from a video source 112 in the form of an encoded video bitstream 123. The system 100 also includes a destination device 140 for decoding the tensor data in the form of the encoded video bitstream 123 to produce a task result 153. A communication channel 130 is used to communicate the encoded video bitstream 123 from the source device 110 to the destination device 140. In some arrangements, one or both of the source device 110 and the destination device 140 may include a respective mobile phone handset (e.g., a "smart phone") or a web camera and a cloud application. The communication channel 130 may be a wired connection such as Ethernet or a wireless connection such as WiFi or 5G, including a connection across a wide area network (WAN) or across an ad hoc connection. In addition, the source device 110 and the destination device 140 may include an application that captures the encoded video data on some computer-readable storage media (such as a hard drive or memory in a file server, etc.).

[0056] The source device 140 generates a signal according to the signal stored in the memory 206 and in the processor 205 (see Figure 2A ) is executed by the source device 110. The method 600 may be implemented using a device such as a configured FPGA, ASIC, or ASSP. Alternatively, as described below, the method 600 may be implemented by the source device 110 as an application 233 (see Figure 2A ). The software code module of the application 233 implementing the method 600 may reside, for example, in ( Figure 2AThe method 600 encodes a tensor for a frame of video data and includes functionality for updating weights used to encode the tensor. The updated weights are used to encode subsequent frames of video data based on the performance of the currently used weights compared to an internal model with weights that were updated when the source device 110 received the image frame.

[0057] Video source 112 provides a source of captured video frame data (denoted as 113), such as a camera sensor, a previously captured video sequence stored on a non-transitory recording medium, or a video feed from a remote camera sensor, etc. Video source 112 may also be the output of a computer graphics card, for example, displaying the video output of an operating system and various applications executed on a computing device (e.g., a tablet computer). Examples of source devices 110 that may include a camera sensor as video source 112 include smartphones, video camcorders, professional video cameras, and web video cameras.

[0058] The source device 140 begins the machine task by performing the first part of the CNN (referred to as the backbone network 114) to produce an intermediate tensor. The intermediate tensor is shown as 115. To facilitate bit rate reduction of the bitstream 123, the source device 140 uses a compression network called a "bottleneck encoder" (shown as encoder 116) to reduce the dimensionality of the tensor 115. The encoder 116 operates to reduce the dimensionality of the tensor 115. In the destination device 140, the "bottleneck decoder" (shown as 150) restores the tensor dimensions so that the output tensor 151 corresponds to the tensor dimensions of the tensor 115. Since the bottleneck encoder 116 and the bottleneck decoder 150 include trainable network layers, there is a need to specify weights.

[0059] For the task network (including the backbone network 114 and the head network 152), offline training (or "pre-training") is a possible mechanism for specifying weights. One disadvantage of using offline trained weights is that the training process needs to anticipate a very wide range of various input data. Typical video compression standards adapt to widely varying inputs by providing a certain degree of data adaptability (such as by using "context adaptive binary arithmetic coding" (CABAC) to model the changing input data statistics, etc.). Neural networks typically use predetermined (fixed) weights and are therefore not adaptable to the input data. Tensor 115 can form a multi-scale representation generated by the feature pyramid network (FPN) in the backbone 114, which is "fused" together into a single tensor, for example, at the encoder 116 using a path called "multi-scale feature compression" (MSFC). MSFC is typically used to merge all FPN layers into a single tensor and reduce the spatial dimensions of a single tensor to the spatial dimensions of the minimum spatial resolution tensor in the FPN layer tensor. Merging all FPN layers into a single tensor is achieved at the expense of spatial detail in the less decomposed (larger) layers of FPN.

[0060] refer to Figure 6 The operation of the source device 110 is described with reference to method 600. The method 600 begins with step 610 of performing a first portion of a neural network.

[0061] In step 610, the CNN backbone 114 receives a frame of video frame data 113 from the video source 112, and performs a specific early layer of the overall CNN, such as a layer corresponding to the "backbone" of the CNN. Step 610 outputs a tensor 115. The backbone layer of the CNN 114 can generate multiple tensors as outputs for each frame, corresponding to, for example, different spatial scales of applying the backbone including FPN to the input image. The tensor 115 generated by the backbone 114 with FPN forms a hierarchical representation of the frame data 113 including data of the feature map. Each successive layer of the hierarchical representation has half the width and half the height of the previous layer. When the system 100 executes the "YOLOv3" network, the FPN can obtain three tensors in the tensor 115 corresponding to the three layers output from the backbone 114, where the tensor 115 has different spatial resolutions and numbers of channels. When the system 100 is performing a network such as "Faster RCNN X101-FPN" or "Mask RCNN X101-FPN", tensor 115 includes four layers of tensors P2 to P5. Although the layers of the first part performed in the backbone module 114 can be referred to as the "backbone" of the overall task network, the specific division of the layers between the source device 110 and the destination device 140 does not need to correspond to the boundary results from the layers typically defined as the "backbone" in the task network. The terms "backbone" and "head" are used herein to refer to any division of the network into a first part and a second part, and such division can be selected based on considerations other than the machine task network architecture (such as available computing resources in the source device 110 or the destination device 140, etc.).

[0062] The method 600 continues from step 610 to a bottleneck encoding step 615 under the control of the processor 205. In step 615, the source device 110 uses the bottleneck encoder 116 to reduce the number and dimension of the tensors 115 by performing a plurality of network layers corresponding to the compression operation using the weight set also associated in the destination device 140. Step 615 receives as input at least one tensor in which a neural network for a machine task (that is, a backbone network has been implemented) has been partially performed, and generates a tensor 117. For a given frame of the frame data 113, the tensor 115 includes a plurality of tensors with different spatial resolutions between FPN layers by using the FPN in the backbone 114. This "multi-scale" means being converted into fewer tensors in the bottleneck encoder 116, for example, into one ("base layer") or two ("base layer" and "enhancement layer") tensors (which also have reduced spatial dimensions) within the reduced tensor 117 for each frame of the frame data 113. The cross-layer fusion and spatial reduction implemented at encoder 116 is an example of a method for deep feature compression of features for fusing neural networks. In the described example, cross-layer fusion is implemented using a technique called "multi-scale feature compression" (MSFC). Figure 5 To illustrate the operation of MFSC. Figure 5 Variations of the described MSFC are possible, including applying single-scale feature fusion (SSFC) to each FPN layer without cross-layer fusion. Applying a separate SSFC stage to each FPN layer enables channel count reduction for each layer, but does not provide any means for spatial reduction of layers having resolutions larger than the resolution of the FPN layer, nor does it exploit any cross-layer redundancy. Step 615 operates to perform data compression on data associated with the image frame (in the described example processed by the backbone module 114). Data compression is performed by the encoder's neural network using the currently associated or applied set of weights. Reference Fig.14 The method 1400 of FIG. 1400 is used to illustrate the method of cross-layer fusion and spatial reduction implemented at step 615. From step 615, control in the processor 205 enters the encoding compressed tensor step 620.

[0063] At step 620, the quantization and packing module 184 quantizes the tensor 117 from the floating point domain into integer samples (such as 10-bit samples, etc.). The module 184 also uses the reference Fig. 7A The packing format packs the feature maps of each channel of the tensor 117 into a frame to generate a packed frame 185. The video encoder 120 encodes the packed frame 185 into the bitstream 123 as a compressed tensor 121. From step 620, control in the processor 205 enters step 625 for bottleneck decoding.

[0064] At step 625, tensor 117 is fed to bottleneck decoder 118. Bottleneck decoder 118 operates with a set of weights known to destination device 140 to produce a recovered tensor 119 from tensor 117. Recovered tensor 119 has the same tensor dimensions as tensor 115. Tensor 119 represents a degraded version of tensor 115 with a loss due to the contracted dimensions of tensor 117 and the optimality of weights of modules 116 and 118 with respect to input tensor statistics. Bottleneck decoder 118 is initialized with the same weights as used in bottleneck decoder 150. Reference Fig.15 The method 1500 depicted illustrates one approach for generating the restored tensor 119 from the tensor 117. From step 625, control in the processor 205 passes to a measure bottleneck performance step 630.

[0065] At step 630, a value for evaluating data compression using the weights determined at step 640 is determined or obtained. For example, mean square error (MSE) module 178 compares tensor 115 to tensor 119, thereby generating MSE value 179, which indicates the loss caused by the conversion to and from the reduced dimensionality of tensor 117. Over time, the statistics of tensor 115 can change, resulting in a decrease in performance of bottleneck encoder 116 and bottleneck decoder 118, which is seen as a decrease in MSE 179, although short-term changes in MSE can also be expected. MSE 179 provides an indication of the expected signal quality of tensor 151 in destination device 140, except that the lossy video encoding process of video encoder 120 is not included in the indication. Therefore, MSE 179 gives an indication of the performance of the bottleneck encoder and decoder (i.e., 116, 118, 150) in preserving tensor data passed from backbone 114 to head 152. The value 179 is obtained or acquired using the tensor 115 input to the encoder 116 and the tensor 119 after data compression using the weights associated with the neural network at steps 615 and 630. In other implementations, mechanisms other than MSE may be used to evaluate data compression. From step 630, control in the processor 205 enters a training state step 635.

[0066] At step 635, the processor 205 updates a training state variable based on the MSE value 179. The training state variable is stored in the memory 206 and indicates whether the source device 110 is currently training the set of modules. If the MSE value 179 drops below a threshold, the training state variable is set to a "TRAINING" state. The threshold can be determined in a number of ways (e.g., using a moving average of previously generated MSE values ​​179, based on an expected MSE value compared to a moving average of the MSE values, or compared to a predetermined or configurable threshold, etc.). If no weight updates are made after a period of training, the training state variable can be reset to stop further training. If the training state variable indicates that training is in progress, control in the processor 205 passes from step 635 to a perform trainable bottleneck encoding / decoding step 640 ( Figure 6 Independent of metrics such as the MSE value 179, step 635 may be operable to enter the TRAINING state periodically (e.g., once a day or once a week) to provide a continuous means of determining whether an improved model can be generated even in the absence of clear indications of degraded performance (such as a decrease in the measured MSE value 179). If the training state variable does not indicate that training is in progress, control in the processor 205 passes from step 635 to the encoding weight update flag step 660 ( Figure 6 in the “NOT TRAIN” section).

[0067] At step 640, the trainable bottleneck encoder 170 and the trainable bottleneck decoder 174 operate to compress and decompress the tensor 115 to produce tensor 175. The trainable bottleneck encoder 170 has the same structure as the encoder 116. The trainable bottleneck decoder 174 has the same structure as the decoder 150. The structural compatibility between modules 170 and modules 116 and between modules 174 and modules 150 enables the weights generated in modules 170 and 174 to be transmitted to modules 116 and 150, respectively, for subsequent processing. The compressed tensor 171 is output from the trainable bottleneck encoder 170 and passed to the trainable bottleneck decoder 174. The forward pass through modules 170 and 174 is performed using the weights currently present in modules 170 and 174. The network topology and layer dimensions of modules 170 and 174 correspond to the network topology and layer dimensions of modules 116 and 118, respectively. The operation of modules 170 and 174 conforms to the reference Fig.14 and Fig.15 The methods 1400 and 1500 are described. From step 640, control in processor 205 passes to a measure trainable bottleneck performance step 645.

[0068] At step 645, a value for evaluating data compression using the weights determined at step 640 is determined or obtained. Step 645 uses the same mechanism as step 630 to evaluate data compression performance. For example, mean square error (MSE) module 180 generates measurement loss 181 by performing a mean square error calculation on tensors 115 and 175. The measurement loss 181 value provides a measure of the ability of modules 170 and 174 to accurately restore the tensor after shrinking in dimension by using reduced dimension tensor 171. From step 645, control in processor 205 enters back propagation step 650.

[0069] At step 650, weight updates are made in modules 170 and 174 using a process of "back-propagation," whereby weights are updated based on a process (such as stochastic gradient descent (SGD) or the like) that attempts to minimize the measured loss 181. Trainable bottleneck encoder 170 and trainable bottleneck decoder 174 are able to adapt to such changes in the statistics of tensor 115 by means of ongoing weight updates due to back-propagation, which results in the potential for achieving a lower MSE 181 compared to MSE 179. The rate of weight updates is scaled by a "learning rate." Higher learning rates generally result in minimizing the measured loss with fewer back-propagation operations (i.e., a faster training process), but suffer from the risk of instability in training caused by over-adjusting weights, which prevents the discovery of local minima of the loss function. Smaller learning rates may take longer to train the network, but smaller learning rate values ​​are also less likely to over-adjust the rate. When using a small batch size such as one image, a smaller learning rate is desired to reduce the impact of individual frames that may be statistically outliers. A typical learning rate may be a value such as 0.01 or 0.001, and may be scaled by the inverse of the batch size or by the inverse of the square root of the batch size to arrive at a final learning rate for a given batch. The learning rate may vary over time, with a larger value initially used when the network weights are away from the final value of the weights, and a smaller value later used when the network approaches an optimal or acceptable state in terms of MSE. In the context of refinement training of bottleneck encoders and decoders, a smaller learning rate (such as between 0.001 and 0.0001, etc.) is appropriate. Although modules 170 and 174 receive one tensor at a time, the input tensors may be grouped into batches of size greater than 1. Increasing the batch size may improve training because each weight update step is affected by a variety of inputs. Increasing the batch size may increase the memory requirements for modules 170 and 174. In addition, for video data, statistical changes in consecutive frames or consecutive tensors from the backbone 114 are less pronounced, which reduces the benefit of using a larger batch size. Other processes for updating weights may also be used, such as an "AdamW" optimizer that utilizes momentum and scaling and decouples weight decay from gradient updates. Reducing tensor 171 forms the result of bottleneck encoders and decoders (i.e., 170 and 174) training or adapting to actual input data (i.e., frame data 113 converted to tensor 115) as they encounter it. Control in processor 205 passes from step 650 to a weight update determination step 655.

[0070] At step 655, trigger module 182 compares MSE 179 to MSE 181. If MSE 179 is observed to be lower than MSE 181 for a period of time and the difference exceeds a threshold, a weight update process is initiated. The period of time may correspond to a number of frames, a moving average, or a combination thereof, which indicates substandard performance of modules 116 and 118 compared to achievable performance as indicated by modules 170 and 174. The heuristics used to initiate weight updates are adapted to capture the point at which the weights currently in use for tensor reduction and recovery (i.e., conversion of tensor 115 to tensor 117 and ultimately to tensor 119 or 151) are no longer well suited to the statistics of the received input data. The weights used in modules 170 and 174 are expensive to encode in the bitstream 123 and are therefore not sent out to the destination device 140 on each update operation. Less frequent transmission of updated weights resulting from the determination of step 655 from the source device 110 to the destination device 140, for example based on the detection of performance degradation while using the currently active weights, is sufficient for adequate system performance. The result of step 655 is a decision whether to perform a weight update in the form of a weight update flag 750. Steps 625 to 655 operate to determine whether to change the set of weights used for data compression applied in the next round of data compression. A determination is made based on a comparison of the MSE values ​​181 and 179 whether to implement the change. If a weight update or change is to be performed, the training state variables are also reset to stop further training for subsequent calls to method 600. Control in the processor 205 passes from step 655 to an encoding weight update flag step 660.

[0071] In step 660, reference is made to Figure 8 The illustrated entropy encoder 838 encodes the weight update flag into the bitstream 123 as a weight update flag 750. The indication to update the weights may be included in a supplemental enhancement information (SEI) message associated with the current packed frame 185. The SEI message 744 contains the weights and may be associated with the current packed frame 185 that includes the weight update flag 750 or has a separate SEI message that contains the weight update flag 750. Alternatively, the presence of an SEI message containing the weights may be an indication from the source device 140 to the destination device 140 that a weight update is to be performed. Control in the processor 205 passes from step 660 to a weight update flag test step 665.

[0072] At step 665, if the weight update flag 750 indicates that the determination in trigger module 182 is to change or update the weight (e.g. Figure 6If the weight update flag is not indicated, "UPDATE", then control in processor 205 proceeds to load updated weights step 667. Otherwise, if the weight update flag does not indicate a determination in trigger module 182 to perform an update of the weights, then control in processor 205 proceeds from step 665 to process the next frame of video data at step 675 of the first portion of the neural network (e.g., Figure 6 as shown, "NO UPDATE").

[0073] At step 667, the bottleneck decoder weights 176 are passed from the bottleneck decoder 174 to the bottleneck decoder 118 and to the weight encoder 186. The bottleneck encoder weights 172 from the trainable bottleneck encoder 170 are loaded or applied to the bottleneck encoder 116. The change in weights can point to a specific stored set of values ​​or replace previously stored values. At step 667, the association of the encoder 116 is changed from the set of weights used at step 615 to the weights 172. Associating the weights 172 with the encoder 116 can involve loading the weights 172 into a region of the memory 206 referenced by the encoder 116, or changing a pointer to select the weights 172 for subsequent use in the encoder 116. Control in the processor 205 enters the encoding updated weights step 670 from step 667.

[0074] At step 670, weight encoder 186 encodes bottleneck decoder weights 176 to produce encoded weights 187. Encoded weights 187 are stored in bitstream 123 (as Figure 7B752). Encoding the weights 187 may involve encoding the weights 187 directly using variable length codewords to compress the weight values. The weights 187 may be encoded using an arithmetic coding scheme such as context adaptive binary arithmetic coding (CABAC). Alternatively, the encoded information may represent a difference between weights, such as a difference relative to a weight previously used by the bottleneck decoder 118 and also known by the bottleneck decoder 146 by means of a previous weight update and synchronization initial state. An example syntax that may be used to represent weights is standard ISO / IEC 15938-17 (sometimes referred to as "neural network coding" or "neural network representation" or "MPEG-NNR"), although other means for efficiently encoding neural network weights into a bitstream may be used, such as "Open Neural Network Exchange Intermediate Representation" and the like. Upon completion of step 670, encoding of the partial task result (i.e., tensor 115) is completed for a frame from the video source 112. Advance to a subsequent frame (such as the next frame, etc.) occurs. The remaining steps in method 600 relate to subsequent frames from video source 112 and illustrate the use or non-use of updated weights as determined at step 655. From step 670, control in processor 205 passes to step 675 where a second neural network first portion is performed.

[0075] At step 675, similar to step 610, a subsequent frame of frame data from video source 112 (e.g., frame 113a) is subjected to the first portion of the neural network to generate an updated tensor for frame 113a, such as tensor 115a, etc. From step 675, control in processor 205 passes to step 680 where a second bottleneck encoding is performed.

[0076] At step 680, similar to step 615, the bottleneck encoder 116 performs encoding of the updated tensor 115a to produce the updated tensor 117a. As determined at step 655, if a weight update is performed at step 667, step 680 performs encoding using the bottleneck encoder weights 172 received by the bottleneck encoder 116 from the trainable bottleneck encoder 172. In other words, data compression is performed using the associated updated weight set 176. From step 680, control in the processor 205 proceeds to an encode second compressed tensor step 685.

[0077] At step 685, similar to step 620, the updated tensor 117a represented as a packed frame according to the format of the frame 700 is encoded into the video bitstream 121 as a compressed frame N+1 (see Figure 7B746). The video bitstream 121 is multiplexed into the bitstream 123 using the multiplexer 122. The source device 140, under the execution of the processor 205, continues to encode feature maps of consecutive frames of the video data 113 from the backbone 114, wherein the trainable layers such as the trainable layers of modules 116 and 118 are updated from time to time as determined by the trigger module 182. As a result of the operation of the method 600, the bitstream 123 includes the encoded tensor from the first part of the neural network, and in some instances includes dynamic updates of weights for reducing the dimensionality of the tensor. The updates depend on the performance of the bottleneck encoder 116 relative to the achievable performance using the bottleneck encoder 170 and the decoder 174. Training on the received data in modules 170 and 174 takes advantage of the following property: for the compression task, the goal is to recover the input data with minimal loss. In other words, for the compression task, the ground truth is the input data. The training performed by modules 170 and 174 occurs simultaneously with the use of the "deployed" network weights present in modules 116, 118, and 150, and uses the same input data. Therefore, the training performed by modules 170 and 174 can be considered as "overfitting" to the current input data. However, since the network is capable of dynamic weight updates, this overfitting behavior can be considered as data-driven adaptation. Support for refined training alleviates the need to increase the complexity of the training of the bottleneck encoder and decoder and / or the complexity of the bottleneck encoder and decoder themselves to accommodate a wider range of statistical diversity input data. Considering the data path from the input frame 113 to the task result 153 provided in the system 100 (i.e., modules 114, 116, 150, and 152), the trainable portion of the path corresponds to modules 116 and 150, thereby forming a subset of the total CNN layers in the data path.

[0078] refer to Fig.11 The operation of the destination device 140 portion of the system 100 is described with reference to method 1100. The method 1100 involves data encoded at a source device, where neural network processing for machine tasks has been partially performed (i.e., utilizing the backbone network 114). The method 1100 may be implemented by a software code module of an application 233 stored on the memory 206 and controlled by execution of the processor 205. The method 1100 begins with a decoded packed frame step 1110.

[0079] At step 1110, the demultiplexer 142 receives the bitstream 123. The demultiplexer 142 extracts a video bitstream 143 corresponding to the bitstream 121 from the bitstream 123. The video bitstream 143 is supplied to the video decoder 146 to produce a decoded packed frame 162. The decoded packed frame 162 is passed to the unpacking and dequantization module 160. The module 160 extracts and inversely quantizes each feature map from the packed frame 162 from the integer sample domain to the floating point domain, thereby arranging the feature maps into tensors, such as the decoded tensor 147. From step 1110, control in the processor 205 passes to a bottleneck decoding step 1120.

[0080] At step 1120, the bottleneck decoder 150, which contains weights for layers such as convolutional layers, converts the tensor 147 into a tensor 151 having an increased dimension compared to the tensor 147. The bottleneck decoder 150 operates to decode data associated with an image (including in the case of an image frame of a video) that has been compressed at the source device 110. The decoding is performed by the neural network of the bottleneck decoder 150 using the currently associated set of weights. The bottleneck decoder 150 uses a decoding method (e.g., MFSC) associated with the encoding or compression implemented by the bottleneck encoder 116. The operation of the bottleneck decoder 150 is consistent with the reference Fig.15 The method 1500 is described. From step 1120, control in the processor 205 proceeds to step 1130 where a second portion of the neural network is performed.

[0081] At step 1130, the head module 152 performs a second portion of the overall neural network that can be implemented in the system 100 to produce a task result 153. The task result 153 is stored in a task result buffer 154, which is generally implemented in the memory 206. The method 1100 accordingly implements the remainder of the neural network machine task at step 1130. From step 1130, control in the processor 205 passes to a decode weight update flag step 1140.

[0082] In step 1140, reference is made to Fig. 9 The entropy decoder 920 decodes the weight update flag 750 from the bitstream 123. The weight update flag 750 provides information indicating whether the set of weights used for bottleneck decoding is to be changed (that is, whether the weights in the bottleneck decoder 150 are to be updated), as determined by the trigger module 182. Control in the processor 205 passes from step 1140 to a weight update indication test step 1150.

[0083] At step 1150, the application 233 determines whether to update the weights based on the decoded weight update flag. If an update or change in the weights is indicated by the weight update flag 750 ("Update" at step 1150), control in the processor 205 proceeds to a decode updated weight step 1160. As described with respect to step 670, the information may be related to the weight value or the difference between the weight values. Otherwise, control in the processor 205 proceeds to a decode packed second frame step 1180 ("No Update" at step 1150).

[0084] At step 1160, information indicative of the encoded weights 145 is extracted from the bitstream 123 using the demultiplexer 142. The weight decoder 148 converts or decodes the encoded weights 145 into decoded weights 149. From step 1160, control in the processor 205 passes to an apply updated weights step 1170.

[0085] At step 1170, the bottleneck decoder 150 is updated to use or apply the decoded weights 149, thereby maintaining synchronization of the weights with the weights present in the bottleneck decoder 118 in the source device 140. In other words, the neural network association of the bottleneck decoder 150 is updated to use the weights 149 from the weights used in step 1120. The update can be directed to a specific stored set of values ​​or replace previously stored values. At the completion of step 1170, the CNN tasks implemented in the backbone 114 and the head 152 have been performed for one frame, and advancement to a subsequent frame (such as the next frame, etc.) occurs. Control in the processor 205 passes from step 1170 to step 1180.

[0086] At step 1180, video decoder 146 decodes second encoded frame 746 from bitstream 123 to produce a second packed frame (eg, frame 147a), similar to step 1110. From step 1180, control in processor 205 passes to perform second bottleneck decoding step 1190.

[0087] At step 1190, the bottleneck decoder 150 decodes the frame 147a using the weights 149 (where indicated by the weight update flag 750) to produce a second decoded tensor 151a. The destination device 140 continues to run the CNN head 152 using the tensor 151a to produce another task result (such as result 153a), and continues to decode subsequent frames using the weights used in the bottleneck decoder that are updated from time to time (as indicated by the decoded weight update flag 750).

[0088] The contents of task result buffer 154 may be presented to a user, for example, via a graphical user interface, or provided to an analysis application that decides some action based on the task results, which may include a summary-level presentation of the aggregated task results to the user. The functionality of each of source device 110 and destination device 140 may be embodied in a single device in some implementations, examples of which include mobile phones and tablet computers, as well as cloud applications.

[0089] Although example devices are described above, source device 110 and destination device 140 may each typically be configured within a general purpose computer system via a combination of hardware and software components. Figure 2A Such a computer system 200 is illustrated, and includes: a computer module 201; input devices such as a keyboard 202, a mouse pointer device 203, a scanner 226, a camera 227 that can be configured as a video source 112, and a microphone 280; and output devices including a printer 215, a display device 214 (which can be configured as a display device for presenting task results 151), and a speaker 217. The computer module 201 can use an external modulator-demodulator (modem) transceiver device 216 to communicate with a communication network 220 via a connection 221. The communication network 220, which can represent a communication channel 130, can be a (WAN), such as the Internet, a cellular telecommunications network, or a private WAN. In the case where the connection 221 is a telephone line, the modem 216 can be a traditional "dial-up" modem. Alternatively, in the case where the connection 221 is a high-capacity (e.g., cable or optical) connection, the modem 216 can be a broadband modem. A wireless modem may also be used to make a wireless connection to the communication network 220. The transceiver device 216 may provide the functionality of the transmitter 122 and the receiver 142, and the communication channel 130 may be embodied in the connection 221.

[0090] The computer module 201 typically includes at least one processor unit 205 and a memory unit 206. For example, the memory unit 206 may have a semiconductor random access memory (RAM) and a semiconductor read-only memory (ROM). The computer module 201 also includes a plurality of input / output (I / O) interfaces, wherein the plurality of input / output (I / O) interfaces include: an audio-video interface 207 coupled to a video display 214, a speaker 217, and a microphone 280; an I / O interface 213 coupled to a keyboard 202, a mouse 203, a scanner 226, a camera 227, and an optional joystick or other human interface device (not illustrated); and an interface 208 for an external modem 216 and a printer 215. The signal from the audio-video interface 207 to the computer monitor 214 is typically the output of a computer graphics card. In some implementations, the modem 216 may be built into the computer module 201, such as built into the interface 208. The computer module 201 also has a local network interface 211, which allows the computer system 200 to be coupled to a local area communication network 222, known as a local area network (LAN), via a connection 223. Figure 2A As shown, the local area communication network 222 may also be connected to the wide area network 220 via a connection 224, wherein the local area communication network 222 will typically include a so-called "firewall" device or a device with similar functionality. The local network interface 211 may include an Ethernet TM )Circuit card, Bluetooth TM ) wireless arrangement or IEEE 802.11 wireless arrangement; however, for the interface 211, a variety of other types of interfaces can be implemented. The local network interface 211 can also provide the functionality of the transmitter 122 and the receiver 142, and the communication channel 130 can also be embodied in the local communication network 222.

[0091] I / O interfaces 208 and 213 may provide either or both serial connectivity and parallel connectivity, wherein the former is typically implemented according to the Universal Serial Bus (USB) standard and has a corresponding USB connector (not illustrated). A storage device 209 is provided, and the storage device 209 typically includes a hard disk drive (HDD) 210. Other storage devices (not illustrated) such as floppy disk drives and tape drives may also be used. An optical disk drive 212 is typically provided to serve as a non-volatile source of data. For example, optical disks (e.g., CD-ROM, DVD, Blu-ray Disc, etc.) may be used. TM)), portable memory devices such as USB-RAM, portable external hard drives and floppy disks as suitable sources of data for computer system 200. Typically, any of HDD 210, optical drive 212, networks 220 and 222 may also be configured to operate as video source 112, or as a destination for decoded video data to be stored for reproduction via display 214. Source device 110 and destination device 140 of system 100 may be embodied in computer system 200.

[0092] The components 205 to 213 of the computer module 201 typically communicate via an interconnect bus 204 and in a manner that results in conventional modes of operation of the computer system 200 known to those skilled in the relevant art. For example, the processor 205 is coupled to the system bus 204 using a connection 218. Likewise, the memory 206 and the optical drive 212 are coupled to the system bus 204 via a connection 219. Examples of computers on which the described arrangement may be practiced include IBM-PC and compatibles, Sun SPARCstation, Apple Mac TM or similar computer system.

[0093] Where appropriate or desired, the source device 110 and the destination device 140 and the methods described below may be implemented using the computer system 200. In particular, the source device 110, the destination device 140 and the methods to be described may be implemented as one or more software applications 233 executable within the computer system 200. The instructions 231 (see FIG. 233 ) executed within the computer system 200 may be used to execute the source device 110, the destination device 140 and the methods to be described. Figure 2B ) to implement the source device 110, the destination device 140 and the steps of the method. The software instructions 231 may be formed into one or more code modules, each for performing one or more specific tasks. The software may also be split into two separate parts, with a first part and corresponding code modules performing the method, and a second part and corresponding code modules managing a user interface between the first part and a user.

[0094] For example, the software may be stored in a computer-readable medium including a storage device described below. The software is loaded from the computer-readable medium into the computer system 200 and then executed by the computer system 200. A computer-readable medium having such software or a computer program recorded on the computer-readable medium is a computer program product. The use of the computer program product in the computer system 200 preferably implements an advantageous device for implementing the source device 110 and the destination device 140 and the method.

[0095] The software 233 is typically stored in the HDD 210 or the memory 206. The software is loaded into the computer system 200 from a computer readable medium and executed by the computer system 200. Thus, for example, the software 233 may be stored on an optically readable disk storage medium (e.g., a CD-ROM) 225 read by the optical drive 212.

[0096] In some examples, the application 233 is provided to the user in a manner encoded on one or more CD-ROMs 225 and read via corresponding drives 212, or alternatively, the application 233 can be read by the user from the network 220 or 222. Furthermore, the software can also be loaded into the computer system 200 from other computer-readable media. Computer-readable storage media refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to the computer system 200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray discs, hard disk drives, ROMs or integrated circuits, USB memories, magneto-optical disks, or computer-readable cards such as PCMCIA cards, etc., regardless of whether these devices are internal or external to the computer module 201. Examples of transient or non-tangible computer-readable transmission media that may also participate in the provision of software, applications, instructions and / or video data or encoded video data to the computer module 201 include: radio or infrared transmission channels and network connections to another computer or networked device, and the Internet or Intranet including e-mail transmissions and information recorded on websites.

[0097] The second part of the above-mentioned application program 233 and the corresponding code module can be executed to implement one or more graphical user interfaces (GUIs) to be rendered or otherwise represented on the display 214. By typically manipulating the keyboard 202 and the mouse 203, users and applications of the computer system 200 can manipulate the interface in a functionally applicable manner to provide control commands and / or input to the applications associated with these (one or more) GUIs. Other forms of user interfaces that are functionally applicable can also be implemented, such as audio interfaces that utilize voice prompts output via the speaker 217 and user voice commands input via the microphone 280, etc.

[0098] Figure 2B is a detailed schematic block diagram of the processor 205 and the "memory" 234. The memory 234 represents Figure 2A A logical aggregation of all memory modules (including storage device 209 and semiconductor memory 206) that can be accessed by computer module 201 in.

[0099] When the computer module 201 is initially powered on, a power-on self-test (POST) program 250 is executed. The POST program 250 is typically stored in Figure 2A 249 of the semiconductor memory 206. Hardware devices such as ROM 249 storing software are sometimes referred to as firmware. POST program 250 checks the hardware within computer module 201 to ensure proper operation, and typically checks processor 205, memory 234 (209, 206) and basic input-output system software (BIOS) module 251, which is also typically stored in ROM 249, for correct operation. Once POST program 250 runs successfully, BIOS 251 activates Figure 2A Activating the hard disk drive 210 causes the bootstrap loader program 252 residing on the hard disk drive 210 to be executed via the processor 205. This loads the operating system 253 into the RAM memory 206, where the operating system 253 begins to operate on the RAM memory 206. The operating system 253 is a system-level application executable by the processor 205 to implement various high-level functions including processor management, memory management, device management, storage management, software application interface, and general user interface.

[0100] The operating system 253 manages the memory 234 (209, 206) to ensure that each process or application running on the computer module 201 has sufficient memory to execute without conflicting with the memory allocated to another process. Figure 2A The different types of memory available in the computer system 200 are described so that each process can run efficiently. Therefore, the aggregate memory 234 is not intended to illustrate how to allocate specific segments of memory (unless otherwise specified), but rather to provide an overview of the memory accessible by the computer system 200 and how to use such memory.

[0101] like Figure 2B As shown, the processor 205 includes a plurality of functional modules, wherein the plurality of functional modules include a control unit 239, an arithmetic logic unit (ALU) 240, and a local or internal memory 248 sometimes referred to as a cache memory. The cache memory 248 typically includes a plurality of storage registers 244 to 246 in a register section. One or more internal buses 241 functionally connect these functional modules to each other. The processor 205 typically also has one or more interfaces 242 for communicating with external devices via the system bus 204 using a connection 218. The memory 234 is coupled to the bus 204 using a connection 219.

[0102] The application program 233 includes an instruction sequence 231 that may include conditional branch instructions and loop instructions. The program 233 may also include data 232 used when executing the program 233. The instructions 231 and data 232 are stored in memory locations 228, 229, 230 and 235, 236, 237, respectively. Depending on the relative size of the instructions 231 and the memory locations 228 to 230, a particular instruction may be stored in a single memory location, as described by the instruction shown in memory location 230. Alternatively, the instruction may be segmented into multiple portions, each stored in a separate memory location, as described by the instruction segments shown in memory locations 228 and 229.

[0103] Typically, a processor 205 is given a set of instructions, which are executed within the processor 205. The processor 205 awaits a subsequent input, which the processor 205 reacts to by executing another set of instructions. Each input may be provided from one or more of a plurality of sources, including data generated by one or more of the input devices 202, 203, data received from an external source across one of the networks 220, 202, data retrieved from one of the storage devices 206, 209, or data retrieved from a storage medium 225 inserted into a corresponding reader 212 (all of which are described in detail in the accompanying drawings). Figure 2A Execution of the instruction set may result in outputting data in some cases. Execution may also involve storing data or variables to memory 234.

[0104] The bottleneck encoder 116, the bottleneck decoder 148, and the method may use input variables 254 stored in corresponding memory locations 255, 256, 257 within the memory 234. The bottleneck encoder 116, the bottleneck decoder 148, and the method produce output variables 261 stored in corresponding memory locations 262, 263, 264 within the memory 234. The intermediate variables 258 may be stored in memory locations 259, 260, 266, and 267.

[0105] refer to Figure 2B The processor 205, registers 244, 245, 246, arithmetic logic unit (ALU) 240 and control unit 239 work together to perform the micro-operation sequence required to perform the "fetch, decode and execute" cycle for each instruction in the instruction set that constitutes the program 233. Each fetch, decode and execute cycle includes:

[0106] A fetch operation for fetching or reading an instruction 231 from a memory location 228 , 229 , 230 ;

[0107] a decode operation in which the control unit 239 determines which instruction was fetched; and

[0108] An execution operation, wherein in the execution operation, the control unit 239 and / or the ALU 240 executes the instruction.

[0109] Thereafter, a further fetch, decode and execute cycle for the next instruction may be performed. Similarly, a store cycle may be performed, whereby the control unit 239 stores or writes a value to the memory location 232 .

[0110] Figures 4 to 15 Each step or sub-process in the method is associated with one or more segments of the program 233, and is typically performed by the register segments 244, 245, 247, ALU 240 and control unit 239 in the processor 205 working together to perform fetch, decode and execute cycles for each instruction in the instruction set of the segment of the program 233.

[0111] Figure 3A is a schematic block diagram showing the functional modules of a backbone 310 of a CNN that can be used as a CNN backbone 114. Backbone 114 is sometimes referred to as "DarkNet-53" and forms the backbone of a "YOLOv3" object detection network. Different backbones are also possible, resulting in different numbers and dimensions of layers of tensor 115 for each frame.

[0112] like Figure 3A As shown in FIG. 1 , the video data 113 is passed to a resizer module 304. The resizer module 304 resizes the frame to a resolution suitable for processing by the CNN backbone 310, thereby generating resized frame data 312. If the resolution of the frame data 113 is already suitable for the CNN backbone 310, then the operation of the resizer module 304 is not required. The resized frame data 312 is passed to a convolutional batch normalisation leaky rectified linear (CBL) module 314 to generate a tensor 316. The CBL 314 contains reference Figure 3D The modules described by the CBL module 360 ​​are shown.

[0113] refer to Figure 3D, the CBL module 360 ​​takes tensor 361 as input. The tensor 361 is passed to the convolution layer 362 to produce a tensor 363. When the convolution layer 362 has a stride of 1 and the padding is set to k samples (where the size of the convolution kernel is 2k+1), the tensor 363 has the same spatial dimensions as the tensor 361. When the convolution layer 362 has a larger stride (such as 2, etc.), the tensor 363 has a smaller spatial dimension than the tensor 361, for example, for a stride of 2, the size of the tensor 363 is halved. Regardless of the stride, for a particular CBL block, the size of the channel dimension of the tensor 363 may vary compared to the channel dimension of the tensor 361. The tensor 363 is passed to the batch normalization module 364, which outputs a tensor 365. The batch normalization module 364 normalizes the input tensor 363 and applies a scaling factor and an offset value to produce an output tensor 365. The scaling factors and offset values ​​are derived from the training process. Tensor 365 is passed to a leaky rectified linear activation ("LeakyReLU") module 366 to produce tensor 367. Module 366 provides a "leaky" activation function whereby positive values ​​in the tensor are passed through and negative values ​​are severely reduced in magnitude, e.g., to 0.1 times the previous value.

[0114] Return to Figure 3A , tensor 316 is passed from CBL block 314 to residual block 11 (Res11) module 320. Module 320 comprises a sequential cascade of three residual blocks internally containing 1, 2 and 8 residual units respectively.

[0115] References Figure 3B ResBlock 340 is shown describing a residual block such as that present in module 320. ResBlock 340 receives tensor 341. Tensor 341 is zero-filled by zero-filling module 342 to produce tensor 343. Tensor 343 is passed to CBL module 344 to produce tensor 345. Tensor 345 is passed to residual unit 346, where residual block 340 contains a series of cascaded residual units. The last residual unit in residual unit 346 outputs tensor 347.

[0116] References Figure 3C350 is used to describe a residual unit such as unit 346. ResUnit 350 takes tensor 351 as input. Tensor 351 is passed to CBL module 352 to produce tensor 353. Tensor 353 is passed to a second CBL unit 354 to produce tensor 355. Addition module 356 sums tensor 355 with tensor 351 to produce tensor 357. Addition module 356 may also be referred to as a "shortcut" because input tensor 351 substantially affects output tensor 357. For an untrained network, ResUnit 350 acts to pass through tensors. When training, CBL modules 352 and 354 act to deviate tensor 357 from tensor 351 based on training data and ground truth data.

[0117] Return to Figure 3A , the Res11 module 320 outputs a tensor 322. The tensor 322 is output from the backbone module 310 as one of the layers and is also provided to the Res8 module 324. The Res8 module 324 is a residual block (i.e., 340) including eight residual units (i.e., 350). The Res8 module 324 generates a tensor 326. The tensor 326 is passed to the Res4 module 328 and is output from the backbone module 310 as one of the layers. The Res4 module 328 is a residual block (i.e., 340) including four residual units (i.e., 350). The Res4 module 328 generates a tensor 329. The tensor 329 is output from the backbone module 310 as one of the layers.

[0118] In general, layer tensors 322, 326, and 329 are output as tensor 115. Backbone CNN 310 may take a video frame with a resolution of 1088×608 as input and produce three tensors corresponding to three layers, which have the following dimensions: [1,256,76,136], [1,512,38,68], [1,1024,19,34]. Another example of three tensors 115 corresponding to three layers may be [1,512,34,19], [1,256,68,38], [1,128,136,76], which are separated at the 75th network layer, the 90th network layer, and the 105th network layer in CNN 310, respectively. Each tensor may have a different resolution from the next tensor. The resolution of each tensor may be doubled in height and width between corresponding tensors. In forming the output tensor 115 , the layer tensors 322 , 326 , and 329 provide a hierarchical representation of the frame data including data for feature maps encoded into a bitstream. The separation point depends on the CNN 310 .

[0119] Figure 4is a schematic block diagram showing the functional modules of an alternative backbone 400 of a CNN that can be used as a CNN backbone 114. Backbone 400 implements a residual network with a feature pyramid network ("ResNet FPN") and is an alternative to CNN backbone 114. Frame data 113 is input and passed through stem network 408, res2 module 412, res3 module 416, res4 module 420, res5 module 424 and max pooling module 428 via tensors 409, 413, 417, 425, where max pooling module 428 produces P6 tensor 429 as output.

[0120] The stem network 408 includes 7×7 convolutions and max pooling operations with a stride of two (2). The res2 module 412, the res3 module 416, the res4 module 420, and the res5 module 424 perform convolution operations (LeakyReLU activation). Each module 412, 416, 420, and 424 also halves the resolution of the processed tensor via a stride setting of 2. Tensors 413, 417, 421, and 425 are passed to 1×1 lateral convolution modules 446, 444, 442, and 440, respectively. Modules 440, 442, 444, and 446 produce tensors 441, 443, 445, and 447, respectively. Tensor 441 is passed to a 3×3 output convolution module 470, which produces an output tensor P5 471. Tensor 441 is also passed to an upsampler module 450 to produce an upsampled tensor 451. Summing module 460 sums tensors 443 and 451 to produce tensor 461. Tensor 461 is passed to upsampler module 452 and 3×3 horizontal convolution module 472. Module 472 outputs P4 tensor 473. Upsampler module 452 produces upsampled tensor 453. Summing module 462 sums tensors 445 and 453 to produce tensor 463. Tensor 463 is passed to 3×3 horizontal convolution module 474 and upsampler module 454. Module 474 outputs P3 tensor 475. Upsampler module 454 outputs upsampled tensor 455. Summing module 464 sums tensors 447 and 455 to produce tensor 465, which is passed to 3×3 horizontal convolution module 476. Module 476 outputs P2 tensor 477. Upsampler modules 450, 452, and 454 use nearest neighbor interpolation to reduce computational complexity. Tensors 429, 471, 473, 475, and 477 form the output tensor 115 of CNN backbone 400. In forming output tensor 115, the FPN of tensors 429, 471, 473, 475, and 477 provides a hierarchical representation of frame data including data for feature maps encoded into a bitstream.

[0121] Figure 5is a schematic block diagram illustrating one type of bottleneck encoder 500 that may be used as the bottleneck encoder 116 or, when implemented with support for back-propagation of gradients with weight updates, as the trainable bottleneck encoder 170 . Fig. 7A Showing the packing arrangement from compressed tensors to feature maps in monochrome video frames.

[0122] Fig.14 Shown for use Figure 5 The method 1400 is a method for reducing tensor dimensionality by using a bottleneck encoder 500 of the source device 110. The method 1400 can be implemented using a device such as a configured FPGA, ASIC, or ASSP. Alternatively, as described below, the method 1400 can be implemented by the source device 110 under execution of the processor 205 as one or more software code modules of the application 233. The software code modules of the application 233 implementing the method 1400 can reside, for example, in the hard drive 210 and / or the memory 206. The method 1400 is repeated for each frame of compressed data in the bitstream 123. The method 1400 can be stored on a computer-readable storage medium and / or in the memory 206.

[0123] Method 1400 begins with step 1410 of selecting a first FPN tensor. Bottleneck encoder 500 receives FPN tensor 501 corresponding to tensor 115 and operates according to method 1400 to restrict the dimension of the received tensor to fewer layers and with a reduced spatial size. Applying a bottleneck encoder and decoder between the first part (trunk 114) and the second part (head 152) of the separated neural network enables the spatial area in the packed tensor data frame to be reduced. The reduction in spatial area is achieved by using the interface between the bottleneck encoder and the bottleneck decoder as the split point of the first part and the second part (114 and 152) of the neural network. Bottleneck encoder 116 acts as an additional layer attached to the first part of the neural network, and bottleneck decoder 150 acts as an additional layer pre-set to the second part of the neural network. Tensor 115 can be considered to include a first set of tensors (e.g., P2, P3, P4, and P5) and a second set of tensors (e.g., P2 and P3), wherein the feature map of the second set of tensors is a subset of the first set of tensors. The feature map of the second tensor set includes tensors with a larger spatial resolution in the spatial resolution of the feature map of the first tensor set. A first tensor (e.g., P2 or P3) belonging to both the first tensor set and the second tensor set is represented in the base layer (first information unit) and the enhancement layer (second information unit). A second tensor (e.g., P4 or P5) belonging to the first tensor set but not belonging to the second tensor set is represented in the base layer (first information unit) but not in the enhancement layer (second information unit). The first information unit typically encodes all tensors within tensor 115 (i.e., P2 to P5), and the second information unit typically encodes a subset of tensors (e.g., P2 and P3) encoded by the first information unit, which contains tensors with a larger feature map resolution in the feature map resolution of the P2 to P5 tensors. Other combinations of tensor sets are also possible. The second information unit may, for example, encode only tensor P2 or may encode tensors P2 to P4. The first information unit may include only tensors P2 to P4, wherein tensor P5 is separately packed into frame 700 for encoding into bitstream 123. When the first information unit includes tensors P2 to P4, the second information unit may include tensor P2 or tensors P2 and P3.

[0124] The sensitivity of the task result to the bottleneck depends on the nature of the task. For object detection, there is less spatial sensitivity, and therefore the spatial downsampling of the larger layer of FPN is less harmful to the resulting mAP. In particular, for the larger spatial tensor of FPN, the instance segmentation and the resulting segmentation map are more sensitive to the loss of spatial details, and therefore benefit from less severe spatial downsampling. In the described arrangement, the bottleneck encoder 500 operates at two scales. The first scale covers all FPN layers, and the second scale covers a subset of the FPN layer. The second scale can include larger layers (e.g., P2 and P3 (502 and 503)), thereby providing additional fidelity for these higher resolution feature maps. The input FPN tensor 501 includes layers P2502, P3 503, P4 504, and P5 505. The spatial resolution of layers P2 to P4 (502, 503, 504) is a power of 2 of the spatial resolution of P5 505. Where P5 505 has a width and height (w, h), P2 to P4 (502, 503, 504) have dimensions (8w, 8h), (4w, 4h), (2w, 2h), respectively. In other words, the individual tensors have a resolution that forms an exponential sequence that doubles in width and height between consecutive tensors. Layers P2 to P5 each have 256 channels. Although the base layer is described as mandatory and the enhancement layer is described as optional (depending on circumstances such as fidelity), an arrangement including a second enhancement layer is possible. The second enhancement layer can only be included if the enhancement layer is included, and the tensors of the second enhancement layer are a subset of the tensors of the first enhancement layer (as described above). For example, the tensors of the second subset can be used for P2 only if the (first) enhancement layer is related to P2 and P3. In other words, the second enhancement layer provides further fidelity for a specific tensor that has been enhanced by the first enhancement layer. Cascading enhancement layers provides further flexibility in providing incremental quality improvements by including tensors for each progressive enhancement layer in a packed frame.

[0125] In the above example, inputs P2 to P5 are Figure 4 The hierarchical feature pyramid network output (P2 477, P3 475, P4 473 and P5 471) generated by the CNN backbone 400 of Figure 3AThe tensors 329, 326, and 322 are implemented and output, and the tensors 329, 326, and 322 are derived into two sets of tensors: the first group has one FPN layer 329, and the second group has two FPN layers 322 and 326. The first group with one FPN layer 329 has the smallest spatial resolution and does not require the operation of the MSFF module 510 in the bottleneck encoder 116, where the tensor 329 is directly passed to the SSFC encoder 550 as tensor 529. The second group is processed by the MSFF 510 in the bottleneck encoder 116, where the tensor 326 is passed in (input) as tensor 503, and the tensor 322 is passed in as tensor 502. At step 1410, the bottleneck encoder 500, under the execution of the processor 205, selects a plurality of adjacent tensors in the tensor 501 as the first plurality of tensors (such as four tensors P2 502, P3 503, P4 504, and P5 505, etc.). Tensor 501 forms a hierarchical representation of frame data 113, which is generated by applying FPN to frame data 113. Using a stride equal to two convolution stages in FPN (i.e., at modules 412, 416, 420, and 424) results in the spatial dimensions of the tensors in tensor 501 being halved in width and height as each corresponding tensor is sorted according to the decomposition level, for example, from P2 to P5. Although the layers have different spatial resolutions, there is a certain degree of inter-layer correlation between layers P2 to P5 of tensor 501. As long as the tensors are first spatially scaled to the same resolution (e.g., the minimum resolution in the tensors to be combined), the use of inter-layer correlation permits a reduction in the number of channels relative to the cascade of tensors across layers. Due to the higher rate of downsampling operations, combining tensors with greatly different spatial resolutions typically results in a relatively high loss of details in the higher resolution tensor. For example, scaling P2 502 to P5 505 requires reducing the width and height to one-eighth of their previous values ​​to reduce the area to one-sixty-fourth the area of ​​P2 502. For tasks that rely on spatial detail, such as instance segmentation, the mAP degrades. For detection of small objects where the network head relies on higher resolution layers, the reduction in mAP due to this high downsampling of the larger layers occurs. From step 1410, control in processor 205 passes to generate first bottleneck tensor step 1420.

[0126] At step 1420, the MSFF module 510 (see Figure 5) combines the tensors of the first set of tensors (i.e., 502, 503, 504, 505) under execution by processor 205 to produce a combined tensor 529. The combined tensor 529 is encoded as a compressed tensor 557. The combined tensor 529 forms a "base layer" representation of the FPN layer tensor. Downsampling modules 522, 522a, 522b operate on tensors with larger spatial scales (i.e., P4 504 at 2h, 2w, 256 and P3 503 at 4h, 4w, 256 and P2 502 at 8h, 8w, 256), respectively. Modules 522, 522a, and 522b downsample to match the spatial scale of the smallest tensor (i.e., P5 505 at h, w, 256), thereby producing downscaled P5 tensors 523, 523a, 523b, respectively. The concatenation module 524 performs channel-by-channel concatenation of tensors 505, 523, 523a, and 523b to produce a concatenated tensor 525 of dimensions h, w, 512. The concatenated tensor 525 is passed to a squeeze and excitation (SE) module 526 to produce a tensor 527. The SE module 526 sequentially performs global pooling, a fully connected layer with reduced channel number, a rectified linear unit activation, a second fully connected layer with restored channel number, and a sigmoid activation function to produce a scaled tensor. Tensor 525 is scaled according to the scaled tensor to produce an output as tensor 527. The SE block 526 can be trained to adaptively change the weighting of different channels in the tensor passed based on the first fully connected layer output. The first fully connected layer output reduces each feature map for each channel to a single value. Each single value is passed through a nonlinear activation unit (ReLU) to create a conditional representation of the unit suitable for the weighting of other channels, where the second fully connected layer performs restoration to the full channel number. Thus, SE block 526 is able to extract nonlinear inter-channel correlations to a greater extent than is possible using purely convolutional (linear) layers when generating tensor 527 from tensor 525. Since tensors 525 and 527 contain 512 channels (the result of the concatenation of two FPN layers), the decorrelation achieved by SE block 526 spans two FPN layers P5 and P4.

[0127] The tensor 527 is passed to the convolutional layer 528. The convolutional layer 528 implements one or more convolutional layers to produce a first combined tensor 529, in which the number of channels is reduced to F channels (typically 256 channels) (ie, F=256).

[0128] The operation of the SSFC encoder 550 reduces the dimensionality of the combined tensor 529 to produce a compressed tensor 557. The combined tensor 529 is passed to the convolution layer 552 of the encoder 550. The encoder 550 produces a tensor 553. The tensor 553 has a channel number reduced from 256 to a smaller value C' (such as 64, etc.). The value 96 can also be used for C', which results in a reference to Fig. 7A501. The larger area requirement for the packed frame is illustrated. Tensor 553 is passed to batch normalization module 554 to produce tensor 555. Batch normalized tensor 555 has the same dimensions as tensor 553. Tensor 555 is passed to TanH layer 556. TanH layer 556 implements a hyperbolic tangent (TanH) layer as layer 536 to produce compressed tensor 557. Compressed tensor 557 has the same dimensions as tensor 553. Step 1420 operates to derive or decode a first information unit from tensor 501. Control in processor 205 passes from step 1420 to determine a second bottleneck tensor exists step 1430.

[0129] At step 1430, the processor 205 determines whether to generate and encode a second set of tensors, the second set of encoded tensors being indicated as the second bottleneck tensors 537. The determination at step 1430 may depend on at least one of the configuration of the system 100 and the machine task to be completed at the head network 152. For example, if the system 100 is configured to perform tasks that require a high degree of spatial acuity (such as instance segmentation, etc.), a determination may be made to include an enhancement layer, thereby permitting the destination device 140 to achieve a higher mAP. When the system 100 is configured to perform tasks that require a lower degree of spatial acuity (or "normal quality") (such as object detection, etc.), a determination may be made to omit the enhancement layer, thereby saving the bit rate overhead of the additional layer. If the system 100 is configured for "normal quality" operation, the second set of tensors is not generated, and the system 100 is considered to be operating in a "first mode". If the system 100 is configured for "high quality" operation, a second set of tensors is generated and the system 100 is set to operate in a "second mode". If the machine task being performed by the destination device 140 requires relatively high preservation of spatial detail (e.g., instance segmentation), a determination is made to include a second set of tensors. Control in the processor 205 passes from step 1420 to an encode second bottleneck tensor presence indication step 1440. In an arrangement of the system 100, a consumer of the task results 153 (such as a human operator or an algorithm that aggregates tasks from many networks, etc.) can determine the need for higher quality and signal an indication to include (or omit) enhancement layers to the source device 110 via an out-of-band communication channel. One example arrangement would involve the neural network head 152 of the destination device 140 performing general person detection and, when a person is detected, signaling to the source device 110 to include enhancement layers before proceeding to a more capable alternative network head as the network head 152. The more capable alternative network head can perform object detection tasks with greater specificity, such as identifying a person of interest.

[0130] At step 1440, the entropy encoder 838, under execution by the processor 205, encodes a flag into the bitstream 123 indicating the decision made at step 1420 to operate in the first mode (base layer only) or the second mode (base layer and enhancement layer). This flag may be included in the SEI message 744 as flag 751. From step 1430, control in the processor 205 passes to a second bottleneck tensor existence test step 1450.

[0131] At step 1450, the software 233 determines whether to generate a second bottleneck tensor 537. If there is a flag indicating that a second bottleneck tensor is to be generated ("PRESENT" at step 150), control in the processor 205 enters a second FPN tensor selection step 1460 from step 1450. Otherwise, if the flag is not present ("ABSENT" at step 1450), the method 1400 terminates. Terminating the method 400 when implementing step 1450 ("ABSENT" at step 1450) can be considered as a first operating mode of the encoder 500. Entering step 1460 and the following steps from step 1450 can be considered as a second operating mode of the encoder 500. Therefore, step 1450 determines whether to operate in the first mode or the second mode based on at least one of the quality configuration and the machine task to be completed. According to the above example, if the machine task to be completed is instance segmentation, operation in the second mode is determined.

[0132] At step 1460, a second set of FPN tensors is selected. The second set of FPN tensors is a subset of the tensors selected as part of step 1410, and generally includes tensors with larger spatial resolution and adjacent spatial scales (e.g., P2 502 and P3 503). Switch 570 is activated so that the second set of FPN tensors is assigned to a subsequent processing stage. In particular, when switch 270 is closed, tensor 502 is provided as 502a, and tensor 503 is provided as 503a. Control in processor 205 proceeds from step 1460 to generate a second bottleneck tensor step 1470.

[0133] At step 1470, the MSFF module 510, under execution of the processor 205, combines the tensors in the second tensor (i.e., 502a, 503a) to generate the reference Figure 5The combined tensor 519 described. The combined tensor 519 provides an "enhancement layer" representation in the form of a subset of the FPN layer tensor. The downsampling module 512 operates on the tensor with a larger spatial scale (i.e., P2502 at 8h, 8w, 256), thereby downsampling to match the spatial scale of the smaller tensor (i.e., P3 503 at 4h, 4w, 256), thereby producing a downscaled P2 tensor 513. The cascade module 514 performs a channel-by-channel cascade of tensors 503a and 513 to produce a cascaded tensor 515. Tensor 515 has dimensions 4h, 4w, 512. The cascaded tensor 515 is passed to a squeeze and excitation (SE) module 516 to produce a tensor 517. The SE module 516 operates in the same manner as described with reference to the SE module 526. Tensor 517 is passed to the convolutional layer 518. Convolutional layer 518 operates in a similar manner to convolutional layer 528 to produce a second combined tensor 519. The second combined tensor 519 has a channel number reduced to F channels (typically 256 channels) (i.e., F = 256). As a result of modules 512, 514, 516, and 518, the tensors of the two FPN layers are reduced to a single tensor having the same number of channels as the input FPN layer tensor and the spatial resolution of the smaller of the two FPN layer tensors. Dimensionality reduction is achieved using several network layers and relies on training layers (e.g., layers 516 and 518) rather than on-the-fly determination of correlations to be exploited.

[0134] SSFC encoder 530 operates to further reduce the dimension of combined tensor 519 to produce compressed tensor 537. Combined tensor 519 is passed to convolution layer 532 to produce tensor 533. Tensor 533 has a channel number reduced from 256 to a smaller value C' (such as 64, etc.). Tensor 533 is passed to batch normalization module 534 to produce tensor 535. Batch normalized tensor 535 has the same dimension as tensor 533. Tensor 535 is passed to TanH layer 536 to produce compressed tensor 537. Compressed tensor 537 has the same dimension as tensor 533. The use of hyperbolic tangent (TanH) layer compresses the dynamic range of values ​​within tensor 537 to [-1,1], thereby removing outliers. Layers 532, 534, and 536 operate in a similar manner to layers 552, 554, and 556, respectively.

[0135] The compressed tensors 557 and 537 provide information units of feature maps of the frame data 113 obtained using the following convolution operations: (i) MSFF 510 and SSFC encoder 550 for tensors 505, 504, 503 and 502, and (if present) (ii) MSFF 510 and SSFC encoder 530 for tensors 503 and 502. The compressed tensors 557 and 537 provide a set of tensors 560 corresponding to the bottleneck encoded tensor 117. If the encoder 500 is operating in the first mode ("absent" at 1450 and switch 570 is closed), the tensor 557 provides the tensor 560. The tensor 560 is provided to the quantization and packing module 184 as the tensor 117 for encoding into the bitstream by the video encoder 120.

[0136] In the arrangement of the bottleneck encoder 500, the TanH modules 536 and 556 are omitted, resulting in tensors 535 and 555 being passed as tensors 537 and 557, respectively. In other words, the output 537 in the arrangement omitting 536 involves applying the combined tensor 519 to the convolution layer 532 and the batch normalization layer. The output 557 in the arrangement omitting 556 involves applying the combined tensor 529 to the convolution layer 532 and the batch normalization layer. Omitting the TanH module results in retaining outliers or large magnitude values, which experiments have found to make a disproportionate contribution to the final task performance in the head 152.

[0137] exist Fig. 7AAn example single monochrome video frame (frame 700) is shown in FIG. Frame 700 corresponds to packed and quantized feature map data 185. The properties of TanH in removing outliers result in a distribution suitable for linear quantization of the bit depth of frame 700. The channels of compressed tensor 557 are packed into feature maps of a particular size, such as feature map 712 in region 714 in frame 700, etc. The channels of compressed tensor 537, if present (as determined at step 1430), are packed into feature maps of different sizes in region 715 of frame 700, such as feature map 710, etc. One channel of compressed tensor 537 corresponds to one feature map indicated by one rectangular region, such as region 710, etc. One channel of compressed tensor 557 corresponds to one feature map indicated by one rectangular region, such as region 712, etc. Region 714 and region 716 (if present) form a packed representation of tensors 557 and 537, where tensors 557 and 537 once compressed by the video encoder 120 form a first information unit and a second information unit, respectively. The first information unit and the second information unit may be stored in a manner that permits independent encoding and decoding, such as by using separate slices, tiles, or sub-pictures of the frame 700, or completely separate pictures. The video decoder 148 may decode only region 714 (i.e., the first information unit) and discard region 716 (i.e., the second information unit), and still provide tensor 151 to the CNN head 152 to produce a task result, albeit with a lower fidelity than would be achievable if region 716 or the second information unit had been decoded.

[0138] Figure 7B 7 is a schematic block diagram showing a bitstream 723 (which may be a portion of the bitstream 123) encoding tensor data. The compressed frame n 742 contains the following information: Fig. 7A The SEI message 744 includes a weight update flag 750 and, if indicated by the weight update flag 750, includes a neural network weight 752. Compressed frame N+1 746 contains a compressed tensor using weights as derived from the neural network weight 752. The presence of a second bottleneck tensor (such as 537, etc.) can be encoded in the SEI message 744 as a presence flag 751.

[0139] Figure 8 1 is a schematic block diagram showing the functional modules of the video encoder 120 (also referred to as the feature map encoder). The video encoder 120 processes the packed feature map frames 185 (in Fig. 7A 700) to produce a video bitstream 121. Typically, data is passed between functional modules within the video encoder 120 in groups of samples or coefficients (such as a partition of a block into fixed-size sub-blocks, etc.) or as an array. Figure 2A and Figure 2BAs shown, the video encoder 120 can be implemented using a general-purpose computer system 200, wherein various functional modules can be implemented using dedicated hardware within the computer system 200, using software executable within the computer system 200 (such as one or more software code modules of a software application 233 residing on the hard disk drive 205 and controlled by the processor 205 for execution, etc.). Alternatively, the video encoder 120 can be implemented using a combination of dedicated hardware and software executable within the computer system 200. The video encoder 120 and the method can alternatively be implemented in dedicated hardware such as one or more integrated circuits that perform the functions or sub-functions of the method. Such dedicated hardware may include a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific standard product (ASSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or one or more microprocessors and associated memory. In particular, the video encoder 120 includes modules 810 to 890, wherein each of these modules can be implemented as one or more software code modules of the software application 233.

[0140] although Figure 8 The video encoder 120 is an example of a Versatile Video Coding (VVC) video encoding pipeline, but other video encoding standards or implementations may also use the processing stages of modules 810 to 890. The video encoder 120 may also be from the memory 206, the hard drive 210, the CD-ROM, the Blu-ray Disc TM ) or other computer-readable storage medium (or write to memory 206, hard drive 210, CD-ROM, Blu-ray disc or other computer-readable storage medium). In addition, frame data 185 (and bitstream 121) can be received from (or sent to) an external source (such as a server or radio frequency receiver connected to communication network 220). Communication network 220 may provide limited bandwidth, thereby requiring the use of rate control in video encoder 120 to avoid saturating the network when frame data 185 is difficult to compress. Frame data 185 can be in any chroma format and bit depth supported by the profile used, for example, 4:0:0, 4:2:0 of ​​the "Main 10" profile of the VVC standard, with a sample precision of eight (8) to ten (10) bits.

[0141] The block partitioner 810 first partitions the frame data 185 into CTUs, which are typically square in shape and configured so that a specific size of the CTU is used. The maximum effective size of the CTU can be, for example, 32×32, 64×64, or 128×128 luminance samples, configured by the "sps_log2_ctu_size_minus5" syntax element present in the "sequence parameter set". The CTU size also provides the maximum CU size, because a CTU that is not further split will contain one CU. The block partitioner 810 also partitions each CTU into one or more CBs based on the luminance coding tree and the chrominance coding tree. The luminance channel may also be referred to as the primary color channel. Each chrominance channel may also be referred to as a secondary color channel. CBs have various sizes and may include both square and non-square aspect ratios. However, in the VVC standard, CBs, CUs, PUs, and TUs always have side lengths that are powers of two. Thus, based on the luma coding tree and chroma coding tree of the CTU, a current CB denoted as 812 is output from the block partitioner 810 (advancing based on iteration over one or more blocks of the CTU).

[0142] The CTUs resulting from the first partitioning of the frame data 185 may be scanned in a raster scan order, and may be grouped into one or more "slices". A slice may be an "intra" (or "I") slice. An intra slice (I slice) indicates that each CU in the slice is intra predicted. Typically, the first picture in a coding layer video sequence (CLVS) contains only I slices and is referred to as an "intra picture". A CLVS may contain periodic intra pictures that form "random access points" (i.e., intermediate frames in a video sequence from which decoding may begin). Alternatively, a slice may be uni-predicted or bi-predicted ("P" or "B" slices, respectively), indicating the additional availability of uni-prediction and bi-prediction in a slice, respectively.

[0143] The video encoder 120 encodes a sequence of pictures according to a picture structure. One picture structure is "low delay", in which a picture using inter-frame prediction may only reference pictures that appeared previously in the sequence. Low delay allows each picture to be output as soon as it is decoded, and is also stored for possible reference by subsequent pictures. Another picture structure is "random access", in which the encoding order of the pictures is different from the display order. Random access allows inter-frame predicted pictures to reference other pictures that have been decoded but have not yet been output. A certain degree of picture buffering is required so that future reference pictures in display order are present in the decoded picture buffer, resulting in multiple frame delays.

[0144] When using a chroma format other than 4:0:0, in an I slice, the coding tree for each CTU can diverge below the 64×64 level into two separate coding trees, one for luma and the other for chroma. The use of separate trees allows different block structures to exist between luma and chroma within the luma 64×64 region of the CTU. For example, a large chroma CB can be co-located with many smaller luma CBs, and vice versa. In a P or B slice, a single coding tree for a CTU defines a block structure common to luma and chroma. The resulting blocks of a single tree can be intra-predicted or inter-predicted.

[0145] In addition to partitioning a picture into slices, a picture can also be partitioned into "tiles". A tile is a sequence of CTUs covering a rectangular area of ​​the picture. CTU scanning is performed within each tile in a raster scan, advancing from one tile to the next. A slice can be an integer number of tiles, or an integer number of consecutive CTU rows within a given tile.

[0146] For each CTU, the video encoder 120 operates in two stages. In the first stage (referred to as the "search" stage), the block partitioner 810 tests various potential configurations of the coding tree. Each potential configuration of the coding tree has an associated "candidate" CB. The first stage involves testing various candidate CBs to select a CB that provides relatively high compression efficiency and relatively low distortion. The testing stage typically involves Lagrangian optimization, whereby the candidate CBs are evaluated based on a weighted combination of rate (i.e., coding cost) and distortion (i.e., error relative to the input frame data 185). The "best" candidate CB (i.e., the CB with the lowest evaluated rate / distortion) is selected for subsequent encoding into the bitstream 121. Included in the evaluation of the candidate CBs is the option of using a CB for a given region, or further splitting the region according to various splitting options and encoding each smaller resulting region with a further CB, or even further splitting the region. Therefore, both the coding tree and the CB itself are selected in the search stage.

[0147] The video encoder 120 generates a prediction block (PB) indicated by arrow 820 for each CB (e.g., CB 812). PB 820 is a prediction of the content of the associated CB 812. Subtractor module 822 generates a difference (or "residual", meaning the difference is in the spatial domain) between PB 820 and CB 812, represented as 824. Difference 824 is the block size difference between corresponding samples in PB 820 and CB 812. Difference 824 is transformed, quantized, and represented as a transform block (TB) indicated by arrow 836. PB 820 and associated TB 836 are typically selected from one of a plurality of possible candidate CBs, for example, based on an assessed cost or distortion.

[0148] A candidate coding block (CB) is a CB obtained from one of the prediction modes available to the video encoder 120 for the associated PB and the resulting residual. When combined with a predicted PB in the video encoder 120, the TB 836 reduces the difference between the decoded CB and the original CB 812 at the expense of additional signaling in the bitstream.

[0149] Thus, each candidate coding block (CB) (i.e., a prediction block (PB) in combination with a transform block (TB)) has an associated coding cost (or "rate") and an associated difference (or "distortion"). The distortion of the CB is typically estimated as a difference in sample values, such as a sum of absolute differences (SAD), a sum of squared differences (SSD), or a Hadamard transform applied to the difference, etc. The mode selector 886 uses the difference 824 to determine the estimates obtained from each candidate PB to determine a prediction mode 887. The prediction mode 887 indicates a decision to use a particular prediction mode (e.g., intra-frame prediction or inter-frame prediction) for the current CB. The estimation of the coding cost associated with each candidate prediction mode and the corresponding residual encoding can be performed at a significantly lower cost than entropy encoding of the residual. Therefore, even in a real-time video encoder, multiple candidate modes can be evaluated to determine the best mode in terms of rate-distortion.

[0150] Determining the best mode based on rate-distortion is typically implemented using a variation of Lagrangian optimization.

[0151] A Lagrangian or similar optimization process may be employed to select both the best partitioning of a CTU into CBs (using the block partitioner 810) and the selection of the best prediction mode from among multiple possibilities. The intra prediction mode with the lowest cost measure is selected as the "best" mode by applying a Lagrangian optimization process of the candidate modes in the mode selector module 886. The lowest cost mode includes the selected secondary transform index 888, which is also encoded into the bitstream 121 by the entropy encoder 838.

[0152] In the second stage of the operation of the video encoder 120 (referred to as the "encoding" stage), the determined coding tree(s) for each CTU are iterated in the video encoder 120. For CTUs using separate trees, the luma coding tree is encoded first, followed by the chroma coding tree, for each 64x64 luma region of the CTU. Within the luma coding tree, only the luma CBs are encoded, and within the chroma coding tree, only the chroma CBs are encoded. For CTUs using shared trees, a single tree describes the CU (i.e., luma CBs and chroma CBs) according to the common block structure of the shared tree.

[0153] The entropy encoder 838 supports bitwise encoding of syntax elements using variable length and fixed length codewords, as well as arithmetic coding modes of syntax elements. Portions of the bitstream such as "parameter sets" (e.g., sequence parameter sets (SPS) and picture parameter sets (PPS)) use a combination of fixed length codewords and variable length codewords. Slices (also called continuous portions) have a slice header using variable length encoding, followed by slice data using arithmetic encoding. The slice header defines parameters specific to the current slice, such as slice-level quantization parameter offsets, etc. For a given slice, the slice data includes syntax elements for each CTU in the slice. The use of variable length encoding and arithmetic coding requires sequential parsing within each portion of the bitstream. These portions can be described with start codes to form "network abstraction layer units" or "NAL units". Arithmetic coding is supported using context-adaptive binary arithmetic coding processing.

[0154] The arithmetically coded syntactic elements consist of a sequence of one or more "bins" (binary files). Like bits, bin values ​​are "0" or "1". However, bins are not encoded in the bitstream 121 as discrete bits. Bins have associated predicted (or "likely" or "maximum probability") values ​​and associated probabilities (known as "context"). When the actual bin to be encoded matches the predicted value, the "maximum probability symbol" (MPS) is encoded. Encoding the maximum probability symbol is relatively cheap in terms of consumed bits in the bitstream 121, including a total cost of less than one discrete bit. When the actual bin to be encoded does not match the possible value, the "minimum probability symbol" (LPS) is encoded. Encoding the minimum probability symbol has a relatively high cost in terms of consumed bits. The bin encoding technique enables efficient encoding of bins that skew the probability of "0" vs "1". For syntactic elements with two possible values ​​(i.e., "flag"), a single bin is sufficient. For syntactic elements with many possible values, a sequence of bins is required.

[0155] The decomposition of the value of a syntactic element into a sequence of one or more bins is called "binarization" of the syntactic element. Binarization can include the conditional presence of a later bin on the value of an earlier bin, making variable bin length binarization possible. In addition, each bin can be associated with more than one context. The selection of a context for a bin is called "context modeling". Context modeling can depend on the bin values ​​of earlier bins in the syntactic element, neighboring syntactic elements (i.e., neighboring syntactic elements from neighboring blocks), etc. Each time a context coding bin is encoded, the context selected for that bin (if any) is updated to reflect the new bin value. Therefore, the binary arithmetic coding scheme is called adaptive.

[0156] The entropy encoder 838 also supports bins lacking context (referred to as "bypass bins"). The bypass bins are encoded using an equal probability distribution between "0" and "1". Therefore, each bin has a coding cost of one bit in the bitstream 121, and is generally used when there is no (not easily exploitable) statistical offset in the probability distribution of the bin value. The absence of context saves memory and reduces complexity, so bypass bins where the distribution of the value of a particular bin is not skewed are used. An example of an entropy encoder that uses context and adaptation is known in the art as CABAC (Context Adaptive Binary Arithmetic Coder), and many variations of this encoder have been used in video coding.

[0157] The entropy encoder 838 encodes a quantization parameter 892 using a combination of context-encoded and bypass-encoded bins, and, if used for the current CB, a secondary transform index 888. The quantization parameter 892 is encoded using a "delta QP" generated by the QP controller module 890. The delta QP is signaled at most once in each region known as a "quantization group". The quantization parameter 892 is applied to the residual coefficients of the luma CB. The adjusted quantization parameter is applied to the residual coefficients of the collocated chroma CB. The adjusted quantization parameter may include mapping from the luma quantization parameter 892 according to a mapping table and a CU-level offset selected from an offset list. The secondary transform index 888 is signaled when the residual associated with the transform block includes valid residual coefficients only in those coefficient positions that are transformed into primary coefficients by applying the secondary transform.

[0158] The residual coefficients of each TB associated with a CB are encoded using a residual syntax. The residual syntax is designed to efficiently encode coefficients with low amplitudes, primarily using arithmetic coded bins to indicate the significance of the coefficients and the amplitude of lower values, and reserving bypass bins for residual coefficients of higher amplitudes. Therefore, residual blocks consisting of very low amplitude values ​​and sparse placement of significant coefficients are efficiently compressed. In addition, there are two residual coding schemes. As seen when the transform is applied, the conventional residual coding scheme is optimized for TBs where significant coefficients are primarily located in the upper left corner of the TB. The transform skip residual coding scheme can be used for TBs that are not transformed, and is able to efficiently encode residual coefficients regardless of their distribution throughout the TB.

[0159] The multiplexer module 884 outputs the PB 820 from the intra prediction module 864 according to the determined best intra prediction mode selected from the test prediction modes of each candidate CB. The candidate prediction modes need not include every conceivable prediction mode supported by the video encoder 120. Intra prediction is divided into three types: first, "DC intra prediction", which involves filling the PB with a single value representing the average of nearby reconstructed samples; second, "plane intra prediction", which involves filling the PB with samples according to a plane, using a DC offset and vertical and horizontal gradients derived from nearby reconstructed neighboring samples. The nearby reconstructed samples typically include a row of reconstructed samples above the current PB (extending to a certain extent to the right of the PB) and a column of reconstructed samples to the left of the current PB (extending to a certain extent downward beyond the PB); and third, "angle intra prediction", which involves filling the PB with reconstructed neighboring samples filtered and propagated across the PB in a particular direction (or "angle"). In VVC, sixty-five (65) angles are supported, with rectangular blocks being able to take advantage of additional angles not available to square blocks, yielding a total of eighty-seven (87) angles.

[0160] A fourth type of intra prediction can be used for chroma PBs, whereby the PBs are generated from collocated luma reconstructed samples according to a "cross component linear model" (CCLM) mode. Three different CCLM modes are available, each using a different model derived from adjacent luma and chroma samples. The derived model is used to generate a block of samples of the chroma PB from collocated luma samples. Matrix multiplication of reference samples can be used to intra predict luma blocks using a matrix selected from a predefined set of matrices. This matrix intra prediction (MIP) achieves gains by using matrices trained on a large set of video data, where the matrices represent the relationship between reference samples and prediction blocks that is not easily captured in angular, planar, or DC intra prediction modes.

[0161] Module 864 can also generate prediction units by copying blocks from nearby in the current frame using an "intra-block copy" (IBC) method. The location of the reference block is constrained to be an area equivalent to one CTU. For 128×128 CTUs, segmentation into 64×64 quadrants (sometimes called "virtual pipeline data units" (VPDUs)) occurs. The referenceable area includes the VPDUs for which all CUs in the current CTU have been decoded, and the VPDUs in the previous CTU (except when the current CTU is the first in a slice, tile, or sub-picture), up to a total area of ​​128×128 luma samples. This area is called the "IBC virtual buffer" and limits the IBC reference area, thereby limiting the required storage. The IBC buffer is filled with reconstructed samples 854 (i.e., before loop filtering), so a buffer separate from the frame buffer 872 is required. When the CTU size is 128×128, the virtual buffer includes samples only from the CTU adjacent to the left of the current CTU. When the CTU size is 32×32 or 64×64, the virtual buffer includes CTUs from up to four or sixteen CTUs to the left of the current CTU. Regardless of the CTU size, access to neighboring CTUs to obtain samples of the IBC reference block is constrained by boundaries (such as the edges of pictures, slices, or tiles). In particular, for feature maps of FPN layers with smaller sizes, using CTU sizes such as 32×32 or 64×64 results in more aligned reference areas to cover the set of previous feature maps. Accessing similar feature maps for IBC prediction provides the advantage of efficient coding when feature map placement is sorted based on SAD, SSE, or other difference metrics.

[0162] The residual of the prediction block when encoding feature map data is different from the residual seen for natural video. Such natural video is typically captured by a camera sensor or screen content, as commonly seen in operating system user interfaces, etc. Feature map residuals tend to contain many details, which are more suitable for transform skip coding than the significant low-frequency coefficients of various transforms. Experiments show that feature map residuals have sufficient local similarity to benefit from transform coding. However, the distribution of feature map residual coefficients does not cluster toward the DC (upper left) coefficient of the transform block. In other words, when encoding feature map data, there is enough correlation to make the transform show gain, and this also applies when using intra-block replication to generate prediction blocks of feature map data. Therefore, when encoding feature map data, when evaluating the residuals generated by candidate block vectors for intra-block replication, Hadamard cost estimates can be used instead of relying solely on SAD or SSD cost estimates. SAD or SSD cost estimates tend to select block vectors with residuals that are more suitable for transform skip coding, and may miss block vectors with residuals that will be compactly encoded using transforms. When encoding feature map data, the Multiple Transform Selection (MTS) tool of the VVC standard may be used so that, in addition to the DCT-2 transform, a combination of DCT-7 and DCT-8 transforms may also be used for residual coding horizontally and vertically.

[0163] An intra-predicted luma coding block can be partitioned vertically or horizontally into a set of equally sized prediction blocks, each with a minimum area of ​​sixteen (16) luma samples. This intra sub-partitioning (ISP) approach enables a single transform block to contribute to the prediction block generation from one sub-partition to the next in the luma coding block, thereby improving compression efficiency.

[0164] In cases where previously reconstructed neighboring samples are not available, such as at the edge of a frame, a default halftone value of half the sample range is used. For example, for 10-bit video, a value of five hundred twelve (512) is used. Since no previous samples are available for the CB located at the top left position of the frame, the angular and planar intra prediction modes produce the same output as the output of the DC prediction mode (i.e., a flat plane of samples with halftone values ​​as amplitudes).

[0165] For inter-frame prediction, a prediction block 882 is generated by a motion compensation module 880 using samples from one or two frames before the current frame in the coding order frame in the bitstream, and the prediction block 882 is output by a multiplexer module 884 as a PB820. In addition, for inter-frame prediction, a single coding tree is typically used for both the luminance channel and the chrominance channel. The order of the coded frames in the bitstream may be different from the order of the frames when captured or displayed. When one frame is used for prediction, the block is called "single prediction" and has one associated motion vector. When two frames are used for prediction, the block is called "double prediction" and has two associated motion vectors. For P slices, each CU can be intra-predicted or uni-predicted. For B slices, each CU can be intra-predicted, uni-predicted, or bi-predicted.

[0166] Frames are typically encoded using a "group of pictures" (GOP) structure to achieve a temporal hierarchy of frames. Frames can be divided into multiple slices, each of which encodes a portion of a frame. The temporal hierarchy of frames allows frames to reference previous and next pictures in the order in which the frames are displayed. Images are encoded in the necessary order to ensure that the correlation for decoding each frame is met. Affine inter-frame prediction mode is available, where instead of using one or two motion vectors to select and filter the reference sample blocks of a prediction unit, the prediction unit is divided into multiple smaller blocks and a motion field is generated so that each smaller block has a different motion vector. The motion field uses the motion vectors of nearby points of the prediction unit as "control points". Affine prediction allows encoding of motions other than translation, with less need for coding trees that use depth splitting. The dual prediction mode available with VVC performs geometric blending of two reference blocks along the selected axis, and signals the angle and offset relative to the center of the block. This geometric partitioning mode ("GPM") allows the use of larger coding units along the boundary between two objects, and the geometry of the boundary for the coding of the coding unit is used as the angle and center offset. Motion vector differences can be encoded as direction (up / down / left / right) and distance (a set of power-of-2 distances are supported) instead of using Cartesian (x,y) offsets. Motion vector predictors are obtained from neighboring blocks ("merge mode") as if no offset was applied. The current block will share the same motion vector as the selected neighboring block.

[0167] Samples are selected based on the motion vector 878 and the reference picture index. The motion vector 878 and the reference picture index are applied to all color channels, so inter-frame prediction is described mainly in terms of operations on PUs rather than PBs. The decomposition of each CTU into one or more inter-frame prediction blocks is described with a single coding tree. The inter-frame prediction method can vary in the number of motion parameters and their precision. The motion parameters typically include a reference frame index that indicates which reference frame(s) in the reference frame list will be used plus the spatial translation of each reference frame, but may include more frames, specific frames, or complex affine parameters (such as scaling and rotation, etc.). In addition, a predetermined motion refinement process can be applied to generate a dense motion estimate based on the reference sample block.

[0168] The PB 820 has been determined and selected, and is subtracted from the original sample block at a subtractor 822, obtaining a residual (denoted as 824) with the lowest coding cost, and the residual is lossily compressed. The lossy compression results from the quantization process of the coefficients produced by the forward transform to the residual coefficients, where these coefficients are ready to be entropy encoded into the bitstream. The forward main transform module 826 applies a forward transform to the difference 824, converting the difference 824 from the spatial domain to the frequency domain, and producing the main transform coefficients represented by arrow 828. The maximum main transform size in one dimension is a 32-point DCT-2 or 64-point DCT-2 transform configured by the "sps_max_luma_transform_size_64_flag" in the sequence parameter set. If the CB being encoded is larger than the maximum supported main transform size (e.g., 64×64 or 32×32) represented as a block size, the main transform 826 is applied in a tiled manner to transform all samples of the difference 824. In case of non-square CBs, tiling is also performed using the maximum available transform size in each dimension of the CB. For example, a 64×16 CB uses two 32×16 primary transforms arranged in a tiled manner when a maximum transform size of thirty-two (32) is used. When the size of the CB is larger than the maximum supported transform size, the CB is filled with TBs in a tiled manner. For example, a 128×128 CB with a 64-pt transform maximum size is filled with four 64×64 TBs in a 2×2 arrangement. A 64×128 CB with a 32-pt maximum size is filled with eight 32×32 TB transforms in a 2×4 arrangement.

[0169] The application of the transform 826 results in multiple TBs for the CBs. When each application of the transform operates on a TB of difference 824 greater than 32×32 (e.g., 64×64), all resulting main transform coefficients 828 outside the upper left 32×32 region of the TB are set to zero (i.e., discarded). The remaining main transform coefficients 828 are passed to a quantizer module 834. The main transform coefficients 828 are quantized according to a quantization parameter 892 associated with the CB to produce main transform coefficients 832. In addition to the quantization parameter 892, the quantizer module 834 may also apply a "scaling list" to allow non-uniform quantization within a TB by further scaling the residual coefficients according to their spatial position within the TB. The quantization parameter 892 may be different for the luma CB and each chroma CB. The main transform coefficients 832 are passed to a forward secondary transform module 830 to produce transform coefficients represented by arrow 836 by performing a non-separable secondary transform (NSST) operation or bypassing the secondary transform. The forward main transform is typically separable, transforming a set of rows and then a set of columns of each TB. For luma TBs with a width and height not exceeding 16 samples, the forward main transform module 826 uses a type II discrete cosine transform (DCT-2) in the horizontal and vertical directions, or a bypass of the transform in the horizontal and vertical directions, or a combination of a type VII discrete sine transform (DST-7) and a type VIII discrete cosine transform (DCT-8) in the horizontal or vertical directions. In the VVC standard, the combination of using DST-7 and DCT-8 is called a "multi-transform selection set" (MTS).

[0170] The forward secondary transform of module 830 is typically a non-separable transform that is applied only to the residual of the intra-predicted CU and can nevertheless be bypassed. The forward secondary transform operates on sixteen (16) samples (arranged as the upper left 4×4 sub-block of the main transform coefficients 828) or forty-eight (48) samples (arranged as three 4×4 sub-blocks of the upper left 8×8 coefficients of the main transform coefficients 828) to produce a set of secondary transform coefficients. The set of secondary transform coefficients can be fewer in number than the set of main transform coefficients from which they are derived. Since the secondary transform is only applied to a set of coefficients that are adjacent to each other and include the DC coefficient, the secondary transform is called a "low-frequency non-separable secondary transform" (LFNST). This secondary transform can be obtained through a training process and, due to its non-separable nature and trained origin, can exploit additional redundancy in the residual signal that cannot be captured by a separable transform (such as a variant of DCT and DST applied horizontally and vertically). In addition, when LFNST is applied, all remaining coefficients in the TB are zero in both the main transform domain and the secondary transform domain.

[0171] The quantization parameter 892 is constant for a given TB, and thus results in uniform scaling of the generation of residual coefficients in the main transform domain for the TB. The quantization parameter 892 may vary periodically with a signaled "delta quantization parameter". The delta quantization parameter (delta QP) is signaled once for a CU contained in a given region (referred to as a "quantization group"). If the CU is larger than the quantization group size, the delta QP is signaled once using one of the TBs of the CU. That is, the delta QP is signaled once for the first quantization group of the CU by the entropy encoder 838, and the delta QP is not signaled for any subsequent quantization groups of the CU. Non-uniform scaling may also be achieved by applying a "quantization matrix", whereby the scaling factor applied for each residual coefficient is derived from a combination of the quantization parameter 892 and the corresponding entry in the scaling matrix. The scaling matrix may have a size less than the size of the TB, and when applied to a TB, a nearest neighbor approach is used to provide scaling values ​​for each residual coefficient from a scaling matrix of a size less than the size of the TB. The residual coefficients 836 are supplied to the entropy encoder 838 for encoding in the bitstream 121. Typically, the residual coefficients of each TB of a TU having at least one valid residual coefficient are scanned according to a scan pattern to produce an ordered list of values. The scan pattern typically scans the TB as a sequence of 4×4 "sub-blocks", thereby providing a conventional scanning operation with a granularity of 4×4 sets of residual coefficients, where the arrangement of the sub-blocks depends on the size of the TB. The scanning within each sub-block and the progression from one sub-block to the next typically follows a reverse diagonal scan pattern. In addition, the quantization parameter 892 is encoded into the bitstream 121 using a differential QP syntax element, and the slice QP and secondary transform index 888 of the initial value in a given slice or sub-picture are encoded into the bitstream 121.

[0172] As described above, the video encoder 120 needs to access a frame representation that corresponds to the decoded frame representation seen in the video decoder. Therefore, the residual coefficients 836 are passed through an inverse secondary transform module 844, which operates according to the secondary transform index 888 to produce intermediate inverse transform coefficients represented by arrows 842. The intermediate inverse transform coefficients 842 are inversely quantized by a dequantizer module 840 according to a quantization parameter 892 to produce inverse transform coefficients represented by arrows 846. The dequantizer module 840 can also use a scaling list to perform inverse non-uniform scaling of the residual coefficients, which corresponds to the forward scaling performed in the quantizer module 834. The inverse transform coefficients 846 are passed to an inverse main transform module 848 to produce residual samples for the TU (represented by arrows 850). The inverse main transform module 848 applies a DCT-2 transform horizontally and vertically, which is constrained by the maximum available transform size described with reference to the forward main transform module 826. The type of inverse transform performed by inverse secondary transform module 844 corresponds to the type of forward transform performed by forward secondary transform module 830. The type of inverse transform performed by inverse main transform module 848 corresponds to the type of main transform performed by main transform module 826. Summation module 852 adds residual samples 850 and PU 820 to produce reconstructed samples of the CU (indicated by arrow 854).

[0173] The reconstructed samples 854 are passed to the reference sample cache 856 and the in-loop filter module 868. The reference sample cache 856, which is typically implemented using static RAM on an ASIC to avoid expensive off-chip memory accesses, provides the minimum sample storage required to satisfy the dependencies for generating intra PBs for subsequent CUs in a frame. The minimum dependencies typically include a "line buffer" of samples along the bottom of a row of CTUs for use by the next row of CTUs and a column buffer (whose range is set by the height of the CTU). The reference sample cache 856 supplies reference samples (represented by arrow 858) to the reference sample filter 860. The sample filter 860 applies a smoothing operation to produce filtered reference samples (indicated by arrow 862). The filtered reference samples 862 are used by the intra prediction module 864 to produce an intra prediction block of samples represented by arrow 866. For each candidate intra prediction mode, the intra prediction module 864 generates a sample block (i.e., 866). The sample block 866 is generated by the module 864 using techniques such as DC, planar or angular intra prediction. A matrix multiplication approach may also be used to generate sample block 866, with neighboring reference samples as input and a matrix selected by video encoder 120 from a set of matrices, with the selected matrix signaled in bitstream 121 using an index to identify which matrix in the set of matrices is to be used by video decoder 144.

[0174] The in-loop filter module 868 applies several filtering stages to the reconstructed samples 854. The filtering stages include a "deblocking filter" (DBF), which applies smoothing aligned with CU boundaries to reduce artifacts caused by discontinuities. Another filtering stage present in the in-loop filter module 868 is an "adaptive loop filter" (ALF), which applies a Wiener-based adaptive filter to further reduce distortion. Another available filtering stage in the in-loop filter module 868 is a "sample adaptive offset" (SAO) filter. The SAO filter operates by first classifying the reconstructed samples into one or more categories and applying an offset at the sample level according to the assigned category.

[0175] Filtered samples, represented by arrow 870, are output from the in-loop filter module 868. The filtered samples 870 are stored in a frame buffer 872. The frame buffer 872 typically has the capacity to store several (e.g., up to sixteen (16)) pictures and is therefore stored in the memory 206. Due to the large memory consumption required, on-chip memory is typically not used to store the frame buffer 872. As such, access to the frame buffer 872 is expensive in terms of memory bandwidth. The frame buffer 872 provides a reference frame (represented by arrow 874) to the motion estimation module 876 and the motion compensation module 880.

[0176] The motion estimation module 876 estimates multiple "motion vectors" (represented as 878), each of which is a Cartesian spatial offset relative to the position of the current CB, thereby referencing a block in one of the reference frames in the frame buffer 872. A filtered block of reference samples (represented as 882) is generated for each motion vector. The filtered reference samples 882 form a further candidate mode for potential selection by the mode selector 886. In addition, for a given CU, the PB 820 can be formed using one reference block ("single prediction"), or can be formed using two reference blocks ("double prediction"). For the selected motion vector, the motion compensation module 880 generates the PB 820 based on a filtering process that supports sub-pixel precision in the motion vector. In this way, the motion estimation module 876 (which operates on many candidate motion vectors) can perform a simplified filtering process compared to the motion compensation module 880 (which operates only on the selected candidates) to achieve reduced computational complexity. When the video encoder 120 selects inter-frame prediction for the CU, the motion vector 878 is encoded into the bitstream 121. In Fig. 9 The video decoder 146 (also called a feature map decoder) is shown in FIG. Fig. 9 The video decoder 146 of is an example of a Versatile Video Coding (VVC) video decoding pipeline, but other video codecs may also be used to perform the processing stages described herein. Fig. 9As shown, the bitstream 143 is input to the video decoder 146. The bitstream 143 can be read from the memory 206, the hard drive 210, the CD-ROM, the Blu-ray disc or other non-transitory computer-readable storage medium. Alternatively, the bitstream 143 can be received from an external source (such as a server or a radio frequency receiver connected to the communication network 220). The bitstream 143 contains encoded syntax elements representing the captured frame data to be decoded.

[0177] The bitstream 143 is input to the entropy decoder module 920. The entropy decoder module 920 extracts the syntax elements from the bitstream 143 by decoding the "bin" sequence, and passes the values ​​of the syntax elements to other modules in the video decoder 146. The entropy decoder module 920 uses variable length and fixed length decoding to decode SPS, PPS or slice headers, and uses an arithmetic decoding engine to decode the syntax elements of the slice data into a sequence of one or more bins. Each bin can use one or more "contexts", where the context describes the probability level of the "one" and "zero" values ​​to be used for encoding the bin. In the case where multiple contexts are available for a given bin, a "context modeling" or "context selection" step is performed to select one of the available contexts to decode the bin.

[0178] The entropy decoder module 920 applies an arithmetic coding algorithm, such as "context adaptive binary arithmetic coding" (CABAC), to decode syntax elements from the bitstream 143. The decoded syntax elements are used to reconstruct parameters within the video decoder 146. The parameters include residual coefficients (represented by arrow 924), quantization parameters 974, secondary transform indexes 970, and mode selection information (represented by arrow 958) such as intra-frame prediction modes. The mode selection information also includes information such as motion vectors and partitioning of each CTU into one or more than one CB. The parameters are used to generate PBs, typically in combination with sample data from previously decoded CBs.

[0179] The residual coefficients 924 are passed to the inverse secondary transform module 936, where the secondary transform is applied or not performed (bypassed) according to the secondary transform index. The inverse secondary transform module 936 generates reconstructed transform coefficients 932, i.e., main transform domain coefficients, from the secondary transform domain coefficients. The reconstructed transform coefficients 932 are input to the dequantizer module 928. The dequantizer module 928 inverse quantizes (or "scales") the residual coefficients 932 (i.e., in the main transform coefficient domain) to create a reconstructed intermediate transform coefficient represented by arrow 940 according to the quantization parameter 974. The dequantizer module 928 can also apply a scaling matrix to provide non-uniform dequantization within the TB, which corresponds to the operation of the dequantizer module 840. If the use of a non-uniform inverse quantization matrix is ​​indicated in the bitstream 143, the video decoder 144 reads the quantization matrix from the bitstream 143 as a sequence of scaling factors and arranges the scaling factors into a matrix. Inverse scaling uses the quantization matrix in combination with the quantization parameter to create the reconstructed intermediate transform coefficients 940.

[0180] The reconstructed transform coefficients 940 are passed to an inverse main transform module 944. Module 944 transforms the coefficients 940 from the frequency domain back to the spatial domain. The inverse main transform module 944 applies an inverse DCT-2 transform horizontally and vertically, subject to the constraints of the maximum available transform size as described with reference to the forward main transform module 826. The result of the operation of module 944 is a block of residual samples represented by arrow 948. The residual sample block 948 is equal in size to the corresponding CB. The residual samples 948 are supplied to a summation module 950.

[0181] At summation module 950, residual samples 948 are added to the decoded PB (denoted as 952) to produce a block of reconstructed samples represented by arrow 956. Reconstructed samples 956 are supplied to a reconstructed sample cache 960 and an in-loop filtering module 988. The in-loop filtering module 988 produces a reconstructed block of frame samples denoted as 992. Frame samples 992 are written to a frame buffer 996.

[0182] The reconstructed sample cache 960 operates in a manner similar to the reconstructed sample cache 856 of the video encoder 120. The reconstructed sample cache 960 provides storage for the reconstructed samples required for intra prediction of subsequent CBs without the memory 206 (e.g., by using data 232, which is typically on-chip memory, instead). Reference samples, represented by arrows 964, are obtained from the reconstructed sample cache 960 and are supplied to a reference sample filter 968 to produce filtered reference samples, represented by arrows 972. The filtered reference samples 972 are supplied to an intra prediction module 976. The module 976 produces a block of intra prediction samples, represented as 980, based on the intra prediction mode parameters 958 signaled in the bitstream 143 and decoded by the entropy decoder 920. The intra prediction module 976 supports the modes of module 864, including IBC and MIP. The sample block 980 is generated using a mode such as DC, planar or angular intra prediction.

[0183] When the prediction mode of the CB is indicated in the bitstream 143 to use intra prediction, the intra prediction samples 980 form the decoded PB 952 via the multiplexer module 984. Intra prediction produces a prediction block (PB) of samples, which is a block in one color component that is derived using "neighboring samples" in the same color component. Neighboring samples are samples that are adjacent to the current block and have been reconstructed because they are at the front in the block decoding order. In the case where the luma block and the chroma block are collocated, the luma block and the chroma block can use different intra prediction modes. However, the two chroma CBs share the same intra prediction mode.

[0184] When the prediction mode of the CB is indicated as inter-prediction in the bitstream 143, the motion compensation module 934 generates a block of inter-prediction samples, denoted as 938. The inter-prediction sample block 938 is generated by selecting and filtering the sample block 998 from the frame buffer 996 using the motion vector and reference frame index decoded from the bitstream 143 by the entropy decoder 920. The sample block 998 is obtained from a previously decoded frame stored in the frame buffer 996. For bi-prediction, two sample blocks are generated and mixed to generate samples for the decoded PB 952. The frame buffer 996 is filled with filter block data 992 from the in-loop filtering module 988. Like the in-loop filtering module 868 of the video encoder 120, the in-loop filtering module 988 applies any of the DBF, ALF, and SAO filtering operations. Generally, the motion vector is applied to both the luma and chroma channels, although the filtering process for sub-sample interpolation in the luma and chroma channels is different. Frames from frame buffer 996 are output as decoded frames 162 .

[0185] Fig.10is a schematic block diagram illustrating a cross-layer tensor inverse bottleneck decoder 1000 , which corresponds to the decoder 150 (and similarly, the decoders 118 and 174 ) for restoring tensor dimensions after compression. Fig.15 Shown for use Fig.10 Method 1500 for recovering tensor dimensions by bottleneck decoder 150 of the present invention. Method 1500 may be implemented using a device such as a configured FPGA, ASIC, or ASSP. Alternatively, as described below, method 1500 may be implemented by destination device 140 as one or more software code modules of application 233 under execution of processor 205. Software code modules of application 233 implementing method 1500 may reside, for example, in hard drive 210 and / or memory 206. Method 1500 is repeated for each compressed data frame in bitstream 123. Method 1500 may be stored on a computer-readable storage medium and / or in memory 206. Method 1500 provides a "switchable" component to decode a bitstream containing a compressed representation of the entire FPN and optionally an additional compressed representation of a portion of the FPN, and combines the two representations on the FPN (the entire and the additional portion) into a final decoded FPN tensor for providing to CNN head 152.

[0186] The decoder 150 receives a tensor 147 as generated by the operation of the bitstream 143 by the video decoder 146 and the module 160. The tensor 147 includes tensors 1011 and 1021 as first information units and second information units. The tensor 1021 corresponds to a decoded version of the tensor 537. Similarly, the tensor 1011 corresponds to a decoded version of the tensor 557. If the bitstream 143 is generated by the operation of the first mode of the encoder 500, at least the tensor 1011 is decoded from the bitstream by the video decoder 146. A flag related to the mode operation may also be decoded. If the bitstream 143 is generated by the operation of the first mode of the encoder 500, the tensor 1021 is also decoded by the video decoder 146. The method 1500 starts with a decoding first bottleneck tensor step 1510.

[0187] At step 1510, the SSFC decoder 1010 is implemented under execution of the processor 205. The SSFC decoder performs a neural network layer to decompress the first decoded compressed tensor 1011 of the tensor 147 to produce a first decoded combined tensor 1017. The convolution layer 1012 receives the tensor 1011 with C'=64 channels and outputs a tensor 1013 with F=256 channels. The tensor 1013 is passed to the batch normalization layer 1014. The batch normalization layer outputs a tensor 1015. The tensor 1015 is passed to the parameterized leaky rectified linear (PReLU) layer 1016. The PReLU layer 1016 outputs a tensor 1017. The tensor 1017 is passed to the MSFR module 1030, which generates tensors 1051, 1053, 1055, and 1037, forming the "base layer" of the decoded FPN layer. The base layer provides a lower degree of fidelity than that present when an additional "enhancement layer" decoded FPN layer is included. Upsampling modules 1032, 1034, and 1036 receive tensor 1017 and perform interpolation at 2x, 4x, and 8x scales to produce tensors 1033, 1035, and 1037, respectively. For example, the width and height of tensor 1033 are twice the width and height of tensor 1017.

[0188] Tensor 1037 forms one output from the MSFR module 1030 and is passed to a downsampling module 1042. The downsampling module 1042 downsamples tensor 1037 by a factor of 2, horizontally and vertically, to produce a tensor 1043 having the same dimensions as tensor 1035. Tensor 1043 is provided to a convolutional layer 1048, which outputs a tensor 1049. A summation module 1054 adds tensors 1035 and 1049 to produce tensor 1055 as an output of the MSFR module 1030. The downsampling module 1040 downsamples tensor 1035 by a factor of 2, horizontally and vertically, to produce a tensor 1041 having the same dimensions as tensor 1033. Tensor 1041 is provided to a convolutional layer 1046, which outputs a tensor 1047. The summing module 1052 adds the tensors 1033 and 1047 to produce the tensor 1053 as the output of the MSFR module 1030. The downsampling module 1038 downsamples the tensor 1033 horizontally and vertically by a factor of 2 to produce a tensor 1039 having the same dimensions as the tensor 1017. The tensor 1039 is provided to the convolution layer 1044, which outputs the tensor 1045. The summing module 1050 adds the tensors 1017 and 1045 to produce the tensor 1051 as the output of the MSFR module 1030.

[0189] Step 1510 generates tensors 1051, 1053, 1055, and 1037. Tensors 1051, 1053, 1055, and 1037 form a hierarchical representation of the image frame and can be considered to include a first tensor (e.g., 1037P'2 or 1055P'3) and a second tensor (e.g., 1053P'4 or 1051P'5), the feature map in the first tensor having a greater spatial resolution than the feature map of the second tensor. Control in processor 205 enters decoding second bottleneck tensor existence indication step 1520 from step 1160.

[0190] At step 1520, the entropy decoder 920, under execution by the processor 205, decodes an indication from the bitstream 123 indicating whether the bitstream 123 includes a second bottleneck tensor. The presence of the second bottleneck tensor may be determined based on the decoded presence flag 751 obtained from the SEI message 744. Control in the processor 205 passes from step 1520 to a second bottleneck tensor presence test step 1530.

[0191] At step 1530, the application 233 executes to determine whether the decoding of the bitstream 123 at step 1520 indicates that a second bottleneck tensor is included. If the presence flag 751 indicates that the second bottleneck tensor is encoded ("present" at step 1530), control in the processor 205 passes from step 1530 to a decode second bottleneck tensor step 1540. Otherwise, if the presence of the second bottleneck tensor is not determined ("not present" at step 1530), control in the processor 205 passes to a combine first tensor and second tensor step 1550. When the second bottleneck tensor is not determined to be present ("not present" at 1530), the decoder 1000 can be considered to be operating in a first (basic) operating mode. When the second bottleneck tensor is determined to be present ("present" at 1530), the decoder 1000 can be considered to be operating in a second (enhanced) operating mode.

[0192] At step 1540, the SSFC decoder 1020 is implemented under execution of the processor 205. The SSFC decoder 1020 performs a neural network layer to decompress the second decoded compressed tensor 1021 of the tensor 147 to generate a second decoded combined tensor 1027.

[0193] Convolutional layer 1022 receives tensor 1021 with C'=64 channels and outputs tensor 1023 with F=256 channels. Tensor 1023 is passed to batch normalization layer 1024. Batch normalization layer 1024 outputs tensor 1025. Tensor 1025 is passed to PReLU layer 1026. PReLU layer 1026 outputs tensor 1027.

[0194] The MSFR module 1030 generates decoded tensors 1061 and 1067 using an upsampling module 1060, a downsampling module 1062, a convolutional layer 1064, and a summing module 1066. Tensor 1027 is passed to the upsampling module 1060. The upsampling module 1060 interpolates to generate a tensor 1061 having a width and height that are twice the width and height of tensor 1027. Tensor 1061 is output from the MSFR module 1030 and passed to the downsampling module 1062. The downsampling module 1062 downsamples tensor 1061 to generate a tensor 1063 having the same dimensions as tensor 1027. Tensor 1063 is provided to a convolutional layer 1064 having a stride of 1, which outputs a tensor 1065. Summation module 1066 adds tensors 1065 and 1027 to produce tensor 1067, which is output from MSFR module 1030. From step 1540, control in processor 205 proceeds to step 1550.

[0195] At step 1550, decoded FPN tensors P'2 to P'5, i.e., 1051, 1053, 1073, and 1077, are generated. Tensors 1051 and 1053 from the base layer portion of the MSFR module 1030 are ready to be passed to the CNN head 150. Tensors 1073 and 1077 are determined at step 1550 based on tensors 1055 and 1037 and optionally based on tensors 1071 and 1078. If it is determined that the enhancement layer is not used (i.e., the second bottleneck tensor is omitted), multiplexers 1074 and 1076 output tensors 1055 and 1037 as tensors 1073 and 1077, respectively. When the second bottleneck tensor is omitted, the CNN head is provided with tensors for all FPN layers that allow the task to be performed, however with reduced task performance due to the lower spatial fidelity in the higher resolution layers (e.g., P2 and P3). The reduced task performance due to using only the base layer tends to limit the maximum achievable mapping for instance segments where near-lossless compression is achieved in the video encoder 120 and losses within the bottleneck encoder and decoder are minimized. The presence of enhancement layers (in addition to the base layer) can increase the maximum achievable mAP to almost the mAP achieved when the neural network is run as a single operation (i.e., without separation into two parts).

[0196] If it is determined that an enhancement layer is included ("present" at step 1030), convolutions 1070 and 1072 are performed. Convolution 1070 takes tensors 1055 and 1067 concatenated along the channel dimension as input and produces an output tensor 1071. Convolution 1072 takes tensors 1037 and 1061 concatenated along the channel dimension as input and produces an output tensor 1078. Multiplexers 1074 and 1076 pass tensors 1071 and 1078 as tensors 1073 and 1077. When the enhancement layer is included, the output tensors 1071 and 1073 for P'2 and P'3 have increased spatial fidelity, which is beneficial for tasks such as instance segmentation.

[0197] Thus, in the second mode, a plurality of tensors forming at least a portion of a hierarchical representation of image data are derived. These tensors are derived from tensor 1021 and at least a portion of the tensors (e.g., 1051 and 1053) decoded at step 1510 and related to at least the first tensor.

[0198] In the arrangement of method 1500, multiplexers 1074 and 1076 are omitted, and convolutions 1072 and 1074 are initialized based on a determination of whether an enhancement layer is used. When an enhancement layer is used, pre-trained weights are used to initialize convolutions 1072 and 1074. Thus, in the second (enhancement) mode, the convolution layer receives tensors for P'2 and P'3 tensors (each of which is an example of a first tensor) derived from each of the first information unit and the second information unit. When the enhancement layer is not in use (in the first mode), convolutions 1072 and 1074 are initialized so that the convolution weights corresponding to input tensors 1055 and 1037 form an identity matrix, and the weights corresponding to input tensors 1061 and 1067 are zeroed. Applying the identity matrix to the input tensors 1055 and 1037 using convolution modules 1072 and 1074 results in outputting tensors 1055 and 1037 as 1071 and 1078, respectively, where input tensors 1061 and 1067 do not contribute to the output since their corresponding weights are zeroed.

[0199] The method 1500 terminates upon implementation of step 1550, having produced decoded FPN ready for processing using the CNN head 150. The method 1500 is re-invoked for each frame of video data encoded in the bitstream 123.

[0200] The operation of the bottleneck encoder 116 and the bottleneck decoder 150 utilizing the base layer and enhancement layer representations of the FPN layer tensors provides a form of quality scalability. In order to ensure the expected operation of the enhancement layer as a "delta" to improve the fidelity of the decoded FPN tensor 119 or 151 relative to the FPN tensor 115 emitted from the backbone 114, the trainable layers in the bottleneck encoder and decoder associated with the base layer must be initially trained with the enhancement layer inactive. SE modules 526, convolutions 528, SSFC encoder 550, SSFC decoder 1010, convolutions 1044, 1046, and 1048 are trained to provide the base layer capabilities of the bottleneck encoder and decoder. In order to train the enhancement layer, the modules associated with the base layer in the bottleneck encoder and decoder are fixed, and the modules associated with the enhancement layer are set to be trainable. The enhancement layer module (SE module 516, convolution 518, SSFC encoder 530, SSFC decoder 1020, convolution 1064) and the two convolutions 1070 and 1072 for merging the base layer and enhancement layer tensors together then learn to provide a "delta" improvement in performance over that achieved using only the base layer.

[0201] Fig. 12A 1 is a schematic block diagram showing a head 152 of a CNN for object detection. Depending on the task to be performed in the destination device 140, the CNN head 152 may be replaced with a different network. The input tensor 151 is separated into tensors for each layer (i.e., tensors 1210, 1220, and 1234). Tensor 1210 is passed to a CBL module 1212 to produce a tensor 1214. Tensor 1214 is passed to a detection module 1216 and an upscaling module 1222. The detection module 1216 operates to detect a bounding box 1218. The bounding box 1218 is in the form of a detection tensor. The bounding box 1218 is passed to a non-maximum suppression (NMS) module 1248. The NMS module 1248 selects one of the multiple inputs generated by the detection module to produce a detection result 153. In order to generate a bounding box addressing a coordinate in the original video data 113, scaling by the original video width and height is performed before resizing the trunk of the network 114. The upscaling module 1222 produces an upscaled tensor 1224 that is scaled to the original video width and height. The upscaled tensor 1224 is passed to the CBL module 1226. The CBL module 1226 produces a tensor 1228 as an output. The tensor 1228 is passed to the detection module 1230 and the upscaling module 1236. The detection module 1230 operates in a similar manner to the detection module 1216 and produces a detection tensor 1232. The detection tensor 1232 is fed to the NMS module 1248.

[0202] The upgrader module 1236 operates in the same manner as the module 1260 and outputs an upgraded tensor 1238. The upgraded tensor 1238 is passed to the CBL module 1240. The CBL module 1240 operates in the same manner as the modules 1212 and 1226 to output a tensor 1242 to the detection module 1244. The detection module 1244 operates in a similar manner as the detection module 1216 and produces a detection tensor 1246. The detection tensor 1246 is fed to the NMS module 1248.

[0203] CBL modules 1212, 1226, and 1240 each include a cascade of five CBL modules, each CBL module as shown in FIG. Figure 3D The upgrader modules 1222 and 1236 are as follows Fig. 12B Various instances of the upgrader module 1260 are shown.

[0204] Upscaling module 1260 accepts tensor 1262 and tensor 1264 as input. Tensor 1262 is passed to CBL module 1266 to produce tensor 1268. Using nearest neighbor interpolation or other various methods, tensor 1268 is passed to upsampler 1270 to produce upsampled tensor 1272. Concatenation module 1274 produces tensor 1276 by concatenating upsampled tensor 1272 with input tensor 1264.

[0205] Detection modules 1216, 1230, and 1244 are as follows Fig. 12C 12 is an example of a detection module 1280 shown. The detection module 1260 receives tensor 1282, which is passed to a CBL module 1284 to produce tensor 1286. Tensor 1286 is passed to a convolution module 1288, which implements a detection kernel. The detection kernel applies a 1×1 kernel to produce outputs on feature maps at three layers. The detection kernel is 1×1×(B×(5+C)), where B is the number of bounding boxes that a particular cell can predict, typically three (3), and C is the number of classes, which can be eighty (80), making the kernel size two hundred and fifty-five (255) detection attributes. Module 1288 outputs tensor 1290. The constant "5" represents four bounding box attributes (box center x, y and size scale x, y) and an object confidence level ("objectness"). The result of the detection kernel has the same spatial dimensions as the input feature map, but the depth of the output corresponds to the detection attributes. The detection kernel is applied at each layer, typically three layers, which results in a large number of candidate bounding boxes. The NMS module 1048 applies non-maximum suppression to the obtained bounding boxes to discard redundant boxes, such as overlapping predictions of similar scales, etc., thereby obtaining the final set of bounding boxes as the output of object detection.

[0206] Fig.131 is a schematic block diagram showing an alternative head 1300 of a CNN as may be implemented for module 152. Head 1300 forms part of an overall network known as "Faster RCNN" and includes a feature network (i.e., backbone portion 400), a region proposal network, and a detection network. The input to head 1300 is tensor 151. Tensor 151 includes P2 to P6 layer tensors 1310, 1312, 1314, 1316, and 1318, respectively. P2 to P6 layer tensors 1310, 1312, 1314, 1316, and 1318 are input to a region proposal network (RPN) head module 1320. RPN head module 1320 convolves the input tensor, producing an intermediate tensor. The intermediate tensor is fed into two subsequent sibling layers in module 1320, one for classification and one for bounding box or "region of interest" (ROI) regression, thereby generating an output of classification and bounding box 1322. The classification and bounding box 1322 are passed to the NMS module 1324. The NMS module prunes the redundant bounding boxes by removing overlapping boxes with lower scores to produce pruned bounding boxes 1326.

[0207] The bounding box 1326 is passed to a region of interest (ROI) alignment module 1328 (or "RoIAlign" stage). The ROI alignment module 1328 also receives tensors P2 to P6 and generates fixed-size feature maps from various input size maps using bilinear interpolation operations. In the operation performed by the ROI alignment module 1328, subsampling is generated by bilinear interpolation of multiple sub-regions in the received region of interest as 3×3 sub-regions to generate an output region of interest as an output value in the output tensor.

[0208] In the arrangement of CNN backbone 400 and CNN head 1300, "P6" layer tensor 429 is omitted from output tensor 115, and in CNN head 1300, P6 input tensor 1318 is generated by performing a "Maxpool" operation with stride equal to 2 on P5 tensor 1316. Since the P6 layer can be reconstructed from the P5 layer, there is no need to separately encode and decode the P6 layer as an explicit FPN layer in the first set of tensors or the second set of tensors.

[0209] The input to the ROI alignment module 1328 is the P2 to P5 feature maps 1310, 1312, 1314 and 1316 (corresponding to Fig.101077, 1073, 1053 and 1051) and region of interest proposals 1326. Each proposal (ROI) from 1326 is associated with a portion of the feature map (1310-1316) to produce a fixed-size map. The size of the fixed-size map is independent of the underlying portion of the feature maps 1310-1316. For example, one of the feature maps 1310-1316 is selected so that the resulting cropped map has sufficient detail according to the following rule: floor(4+log2(sqrt(box_area) / 224)), where 224 is a typical box size. The ROI alignment module 1328 thus crops the incoming feature map according to the proposal 1326, thereby producing a tensor 1330. The tensor 1330 is fed to a fully connected (FC) neural network head 1332. The FC head 1332 performs two fully connected layers to produce a class score and bounding box prediction sub-difference tensor 1334. The class score is typically a tensor of 80 elements, each element corresponding to the predicted score of the corresponding object class. The bounding box prediction difference tensor is a tensor of 80×4=320 elements, containing the bounding boxes of the corresponding object class. Final processing is performed by the output layer module 1336, which receives the tensor 1334 and performs a filtering operation to produce a filtered tensor 1338. The final processing uses an indication related to the object classification and the confidence ("objectness" value) that the bounding box does correspond to the object, and encodes one or more bounding boxes for each position in the tensor of each FPN layer. Low-scoring (low-classification) objects are no longer considered further. The non-maximum suppression module 1340 receives the filtered tensor 1334 and removes the overlapping bounding boxes encoded in the received tensor by removing overlapping boxes with lower classification scores, thereby obtaining the inferred output tensor 151.

[0210] In the arrangement of source device 110 and destination device 140, backbone 114 and head 152 are omitted, and bottleneck encoders and decoders (i.e., 170, 174, 116, 118, and 150) may be operated as end-to-end learned image compression and decompression neural networks, taking image frames as input to the encoding stage and outputting decoded image frames from the decoding stage. The end-to-end learned image compression network is trained as needed during operation to adapt to changing input frame data. This enables the use of potentially smaller networks that are dynamically updated to match the specific video or image being compressed, rather than relying on pre-trained networks that need to have been trained on a variety of source materials to achieve consistent performance. Examples of different types of source materials that may need to be adapted include screen content, camera captured content (under various lighting and other conditions), and rendered content, etc. The metric used to measure the performance of the bottleneck encoder and decoder may be different from MSE, for example, MS-SSIM may be generated at steps 630 and 645. In the case where system 100 is operable to provide trainable end-to-end learned image compression, the performance metric of MS-SSIM may be useful.

[0211] The initial weights of the bottleneck encoder and decoder (i.e., 170, 174, 116, 118, and 150) can be derived by training using a data set and ground truth suitable for the original task network (i.e., suitable for the network formed by the backbone 114 and the head 152). Training using such a data set can be performed with the network weights of the backbone 114 and the head 152 fixed and the network weights of the inserted bottleneck encoder and decoder allowed to be updated. Such initial weights can be used in the source device 110 and the destination device 140 before any refinement training. Such initial weights can show the following trend: during training, when measured on a channel-by-channel basis, the MSE varies between channels. The MSE variance between channels under the loss function corresponding to the final task performance indicates the relative contribution of each channel to the final task result. In the arrangement of the source device 110, modules 178 and 180 generate per-channel weight MSEs so that channels that contribute less to the final task performance are "derated" or scaled down in terms of their contribution to the final MSE. Per-channel (or "channel-by-channel") scaling based on a predetermined weighted MSE enables adaptation of the refined training without over-allocating importance to retain channels that make relatively small contributions to the final task score.

[0212] In the bottleneck encoder and decoder arrangement, modules 116, 120, 170, 174, and 150 operate to merge the tensors of all FPN layers into a single tensor having dimensions set according to the smallest resolution tensor among the FPN tensors for one image. In other words, Figure 5 and Fig.10, one MSFF module 510 fuses together the tensors for all FPN layer tensors (e.g., 502, 503, 504, and 505) into a single tensor, and the MSFR module 1030 reconstructs the tensors for all FPN layers (e.g., 1077, 1073, 1053, and 1051). In addition, if desired, switch 570 can be closed so that additional data for the P2 and P3 tensors can be as described with respect to Figure 5 The are encoded and as for Fig.10 The SSFC decoder 1020 associated with the MSFR 1030 is decoded.

[0213] In the arrangement of source device 110, multiple sets of trained weights are available, e.g., weights optimized for screen content, camera captured content, rendered content. Source device 110 is operable to select an optimal set of weights among the available predetermined weights and signal the selected weights to destination device 140. Source device 110 may "test" each set of weights in modules 170 and 174 to determine which set should be used. Changes in the content type of frame data 113 may result in a reduction in performance as measured by module 182, prompting a re-evaluation of which set of predetermined weights should be used.

[0214] In the arrangement of source device 110, modules 170 and 174 are operable to train weights associated with the enhancement layer (second information unit) and convolutions 1070 and 1072, but not to train weights associated with the base layer (first information unit). Signaling support associated with weight updates in bitstream 123 indicates that only the enhancement layer is updated when a determination is made to update the weights.

[0215] In another arrangement of the source device 110, the modules 170 and 174 are operable to train the base layer and the enhancement layer as separate training phases. While the base layer is being trained, the enhancement layer is disabled, allowing optimal base layer weights to be derived. Once updated weights for the base layer are determined in the source device 110 and communicated to the destination device 140, the enhancement layer needs to be retrained before the enhancement layer can be enabled. The retraining is required because the enhancement layer operates in combination with the base layer that has been retrained. Once the enhancement layer weights have been trained for operation on the new base layer, the enhancement layer weights must be communicated to the destination device 140 before the enhancement layer can be re-enabled for compression of the tensor 115.

[0216] In arrangements where the base layer and enhancement layers are trained separately, additional weight update flags (ie, flags in addition to weight update flags 750) are used to indicate which layer is to be weight updated. Weights 752 include weights for the indicated layer.

[0217] In the arrangement of system 100, each feature map of enhancement layer 537 is represented as a set of coefficients applicable to a set of basis vectors. The basis vectors are derived from the enhancement layer using a principal component analysis (PCA) method, such as singular value decomposition (SVD), etc. When the PCA method is in use, region 716 of frame 700 includes basis vectors and coefficients, with one coefficient per basis vector per feature map of enhancement layer 537. In source device 110, the transformation from enhancement layer 537 to coefficients is performed using a dot product with the basis vectors. In destination device 140, the transformation from coefficients back to reconstructing enhancement layer 1021 is performed using a dot product of the coefficients and the basis vectors. The PCA method can be applied to both the base layer and the enhancement layer, resulting in two sets of basis vectors and two sets of coefficients. The operation of the PCA encoder and PCA decoder is described in detail in the document "[VCMTrack 1] Tensor compression using VVC" (ISO / IEC JTC 1 / SC 29 / WG 2 document m59591). The PCA method can be applied to both the base layer and the enhancement layer, where separate basis vectors and coefficients are generated for each layer. The PCA method can be applied only to the enhancement layer, where basis vectors and coefficients are generated only for that layer, while all feature maps of the base layer are packed directly into frame 700.

[0218] In the case where separate bottleneck encoders are applied to different non-overlapping sets of tensors of the FPN layer, the PCA method can be independently applied to any compressed tensor in all the resulting compressed tensors. When the PCA encoder is in use, the input tensors to the PCA encoder are received from the output of the SSFC encoders 550 and 530 (i.e., from tensors 557 and 537 (which are the outputs of TanH modules 556 and 536 (if present) or the outputs of batch normalization modules 554 and 534)). The output basis vectors, coefficients, and average feature maps are forwarded for quantization and packing into frame 700. When the PCA decoder is in use, the input basis vectors, coefficients, and average feature maps are obtained from the packed frame 700, and the input basis vectors, coefficients, and average feature maps are supplied to each PCA decoder, and each PCA decoder outputs tensors 1011 and 1021 for use by SSFC decoders 1010 and 1020, respectively. In the arrangement of system 100 , where a PCA encoder and decoder are used, batch normalization 554 and 534 are postponed until after the PCA decoder, i.e., modules 554 and 534 are omitted from the SSFC encoders 550 and 530 and performed after the PCA decoder in the SSFC decoders 1010 and 1020 .

[0219] In the arrangement of system 100, regardless of whether the task to be performed is object detection or instance segmentation, the source device 140 implements the MaskRCNN backbone and is initialized with pre-trained weights for the MaskRCNN network. When performing object detection, the destination device 140 implements the FasterRCNN head at 152 and initializes it with pre-trained weights for the FasterRCNN network. If the FasterRCNN head is implemented, the bottleneck decoder 150 is required to act as an interface between feature maps generated from the MaskRCNN backbone but supplied to the FasterRCNN head. The bottleneck encoder 116 remains optimized in terms of training for the MaskRCNN backbone and head, because it may not be known what the head network will do when encoding. In order to prepare the initial weights for the bottleneck decoder 150, a "hybrid" training process can be performed. The hybrid training process includes instantiating a FasterRCNN network, in which the backbone is initialized with MaskRCNN weights, and the head is initialized with FasterRCNN weights. The MSFC encoder 116 is initialized with weights corresponding to the MaskRCNN training of the MSFC (bottleneck encoder and decoder). Then, only the bottleneck decoder 150 is set to be trainable, and all other network layers are set to be fixed. A training operation is performed, which trains the bottleneck decoder to not only decode the compressed FPN tensor, but also adapt the resulting feature map to match the expected input to the FasterRCNN head, resulting in minimizing the loss. As a result of the training, the bottleneck decoder 150 provides an adaptation between the MaskRCNN backbone and the FasterRCNN head, where the bottleneck encoder 16 remains optimized for the more capable network (i.e., MaskRCNN). When the bottleneck decoder 150 is initialized with weights trained for a "hybrid" operating system (MaskRCNN backbone with FasterRCNN head), the resulting bitstream from the source device 110 is suitable for both object detection and instance segmentation when it is generated, that is, the same bitstream can be used for both tasks later without any transcoding or other operations. When a single bitstream generated by the source device 110 can be used to provide tensors to different neural network heads (assuming that different neural networks corresponding to each neural network head share the same backbone topology and dimensions), the source device 110 is said to support a "shared backbone" operating mode.

[0220] In another arrangement of the system 100, the number of C' channels for the SSFC encoder 530 and the SSFC decoder 1020 (i.e., the enhancement layer) is reduced compared to the number of C' channels for the SSFC encoder 550 and the SSFC decoder 1010 (i.e., the base layer), respectively. The enhancement layer may use a C' value of 32. Such an arrangement packs fewer feature maps of larger size (i.e., 710) into the frame 700. The ability to retrain the bottleneck encoder and the bottleneck decoder applied to the enhancement layer permits the use of fewer channels to encode the enhancement details present in the P2 and P3 layers, thereby accommodating the changing statistics encountered across the full number of channels of the applicable tensor (i.e., 256 channels across the P2 and P3 FPN layers).

[0221] Industrial Applicability

[0222] The described arrangement is suitable for use in the computer and data processing industry, and in particular in digital signal processing for encoding and decoding signals such as video and image signals, thereby achieving high compression efficiency.

[0223] about Figure 1 The arrangement described in the weight encoding provides a system that can adapt to dynamically changing statistics of input video data by undergoing a refinement training process from time to time (such as is deemed necessary by continuous monitoring of the performance of the bottleneck encoder and decoder during use). When a refinement weight is determined that provides improved performance, the weights actively used in the encoder and decoder are updated to use the refinement weights, thereby maintaining operation during the training process. The continuous monitoring of performance allows training of the MFSC unit 116 and the MFSC decoding unit 150 to be implemented based on changes in the data input or based on specific data type input during inference operations. Therefore, the feature compression operation can be adjusted or updated without the need for separate off-system training for specific image types (such as natural or computer-generated images) or scenarios.

[0224] As about Fig.14 and Fig.15 As stated, about Figure 5 and Fig.10 The arrangement described allows flexible operation between high performance and low performance requirements. Different architectures are not required for each system, but flexibility is provided without loading different networks.

[0225] The foregoing describes only some embodiments of the present invention, and modifications and / or changes may be made thereto without departing from the scope and spirit of the present invention, the embodiments being illustrative rather than restrictive.

Claims

1. A method for encoding at least a plurality of tensors into a bitstream, the plurality of tensors forming a hierarchical representation for a single frame, the method comprising: deriving a first information unit from a plurality of tensors forming the hierarchical representation, the plurality of tensors comprising at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a first mode, encoding at least the first information unit into the bitstream; in a second mode, deriving a second information unit from at least said first tensor; as well as In the second mode, the second information unit and the first information unit are encoded into the bitstream.

2. The method according to claim 1, further comprising: Whether to operate in the first mode or the second mode is determined based on at least one of a quality configuration for encoding and a machine task to be completed.

3. The method according to claim 2, further comprising: In case the machine task to be completed is instance segmentation, it is determined to operate in the second mode.

4. The method according to claim 1, wherein: Deriving the first information unit includes combining at least the first tensor and the second tensor into a first combined tensor, and applying a convolution layer to the first combined tensor, followed by a batch normalization layer.

5. The method according to claim 4, wherein: Deriving the first information unit also includes: providing the output of the batch normalization layer to a TanH layer.

6. The method according to claim 1, wherein: Deriving the first information unit includes: combining at least the first tensor and the second tensor into a first combined tensor, and applying a first convolution layer and a first batch normalization layer to the first combined tensor, and Deriving the second information unit includes combining at least the first tensor and another tensor into a second combined tensor, and applying a second convolutional layer and a second batch normalization layer to the second combined tensor.

7. The method according to claim 6, wherein: Deriving the first information unit further includes: providing the output of the first batch normalization layer to a TanH layer, and Deriving the second information unit also includes: providing the output of the second batch normalization layer to a TanH layer.

8. A method for decoding at least a plurality of tensors from a bitstream, the plurality of tensors forming a hierarchical representation for a single frame, the method comprising: decoding a bitstream including at least a first information unit from the bitstream; In a first mode, a plurality of tensors forming the hierarchical representation are derived from the first information unit, the plurality of tensors comprising at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a second mode, decoding a second information unit from the bitstream; as well as In the second mode, a plurality of tensors forming at least a portion of the hierarchical representation are derived from at least a portion of the first information unit and the second information unit, the second information unit corresponding to the first tensor.

9. The method according to claim 8, further comprising: An indication of whether the second mode is used is decoded from the bitstream.

10. The method according to claim 8, wherein: In the second mode, a multiplexer is used to select a plurality of tensors from the second information unit.

11. The method according to claim 8, wherein: A convolutional layer is used to select a tensor corresponding to at least the first tensor.

12. The method according to claim 11, wherein: In the second mode, the convolutional layer receives a tensor for at least the first tensor derived from each of the first information unit and the second information unit.

13. The method according to claim 11, wherein: In the first mode, the convolutional layer receives (i) at least the first tensor from the first information unit, and (ii) an identity matrix representing a tensor derived from the second information unit.

14. A non-transitory computer-readable storage medium storing a program for executing a method for encoding at least a plurality of tensors into a bitstream, the plurality of tensors forming a hierarchical representation for a single frame, the method comprising: deriving a first information unit from a plurality of tensors forming the hierarchical representation, the plurality of tensors comprising at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a first mode, encoding at least the first information unit into the bitstream; in a second mode, deriving a second information unit from said first tensor; as well as In the second mode, the second information unit and the first information unit are encoded into the bitstream.

15. An encoder configured to encode at least a plurality of tensors into a bitstream, the plurality of tensors forming a hierarchical representation for a single frame, by: deriving a first information unit from a plurality of tensors forming the hierarchical representation, the plurality of tensors comprising at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a first mode, encoding at least the first information unit into the bitstream; in a second mode, deriving a second information unit from at least said first tensor; as well as In the second mode, the second information unit and the first information unit are encoded into the bitstream.

16. A system comprising: Memory; as well as A processor, wherein the processor is configured to execute code stored on the memory, the code for implementing a method for encoding at least a plurality of tensors into a bitstream, the plurality of tensors forming a hierarchical representation for a single frame, the method comprising: deriving a first information unit from a plurality of tensors forming the hierarchical representation, the plurality of tensors comprising at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a first mode, encoding at least the first information unit into the bitstream; In a second mode, deriving a second information unit from said first tensor; and In the second mode, the second information unit and the first information unit are encoded into the bitstream.

17. A non-transitory computer readable storage medium storing a program for executing a method for decoding at least a plurality of tensors from a bitstream, the plurality of tensors forming a hierarchical representation for a single frame, the method comprising: decoding a bitstream including at least a first information unit from the bitstream; In a first mode, a plurality of tensors forming the hierarchical representation are derived from the first information unit, the plurality of tensors comprising at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a second mode, decoding a second information unit from the bitstream; as well as In the second mode, a plurality of tensors forming at least a part of the hierarchical representation are derived from at least a part of the first information unit and the second information unit, the second information unit corresponding to at least the first tensor.

18. A decoder configured to decode at least a plurality of tensors from a bitstream, the plurality of tensors forming a hierarchical representation for a single frame by: decoding a bitstream including at least a first information unit from the bitstream; In a first mode, a plurality of tensors forming the hierarchical representation are derived from the first information unit, the plurality of tensors comprising at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; in a second mode, decoding a second information unit from the bitstream; as well as In the second mode, a plurality of tensors forming at least a part of the hierarchical representation are derived from at least a part of the first information unit and the second information unit, the second information unit corresponding to at least the first tensor.

19. A system comprising: Memory; as well as A processor, wherein the processor is configured to execute code stored on the memory, the code for implementing a method for decoding at least a plurality of tensors from a bitstream, the plurality of tensors forming a hierarchical representation for a single frame, the method comprising: decoding a bitstream including at least a first information unit from the bitstream; In a first mode, a plurality of tensors forming the hierarchical representation are derived from the first information unit, the plurality of tensors comprising at least a first tensor and a second tensor, a feature map of the first tensor having a greater spatial resolution than a feature map of the second tensor; In a second mode, decoding a second information unit from the bitstream; and In the second mode, a plurality of tensors forming at least a part of the hierarchical representation are derived from at least a part of the first information unit and the second information unit, the second information unit corresponding to at least the first tensor.