Method, apparatus and system for encoding and decoding a plurality of tensors

AU2024202416B2Pending Publication Date: 2026-08-06CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
AU · AU
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-04-12
Publication Date
2026-08-06

AI Technical Summary

Technical Problem

Existing split network solutions for convolutional neural networks (CNNs) are inflexible, require separate instantiations for each task, and often result in a trade-off between accuracy and efficiency, especially when implemented across edge devices and cloud servers for tasks like object detection and instance segmentation.

Method used

A method and system for encoding and decoding tensors from a CNN that includes encoding additional information about machine tasks and split points, allowing for flexible task performance across distributed networks by separating tensor data into groups and applying multi-scale feature compression techniques.

Benefits of technology

Enhances accuracy and flexibility in distributed CNN processing by preserving spatial detail and enabling efficient compression and decompression of tensors, suitable for tasks like object detection and instance segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000094_0000
    Figure 00000094_0000
  • Figure 00000095_0000
    Figure 00000095_0000
Patent Text Reader

Abstract

49313826_1 Abstract METHOD, APPARATUS AND SYSTEM FOR ENCODING AND DECODING A PLURALITY OF TENSORS An apparatus and method for method of encoding tensors. The method comprises encoding one or more tensors produced by a portion of a neural network into a bitstream; and encoding additional data into the bitstream. The additional data comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network. Abstract METHOD, APPARATUS AND SYSTEM FOR ENCODING AND DECODING A PLURALITY OF TENSORS An apparatus and method for method of encoding tensors. The method comprises encoding one or more tensors produced by a portion of a neural network into a bitstream; and encoding additional data into the bitstream. The additional data comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network. 49313826_1 20 24 20 24 16 12 A pr 2 02 4 A b s t r a c t 2 0 2 4 2 0 2 4 1 6 1 2 2 0 2 4 A p r M E T H O D , A P P A R A T U S A N D S Y S T E M F O R E N C O D I N G A N D D E C O D I N G A P L U R A L I T Y O F T E N S O R S A n a p p a r a t u s a n d m e t h o d f o r m e t h o d o f e n c o d i n g t e n s o r s . T h e m e t h o d c o m p r i s e s e n c o d i n g o n e o r m o r e t e n s o r s p r o d u c e d b y a p o r t i o n o f a n e u r a l n e t w o r k i n t o a b i t s t r e a m ; a n d e n c o d i n g a d d i t i o n a l d a t a i n t o t h e b i t s t r e a m . T h e a d d i t i o n a l d a t a c o m p r i s i n g ( a ) a p l u r a l i t y o f p i e c e s o f t a s k i n f o r m a t i o n e a c h o f w h i c h i n d i c a t e s m a c h i n e t a s k p r o c e s s w h i c h i s c a p a b l e o f b e i n g p e r f o r m e d f o r t h e o n e o r m o r e t e n s o r s u p o n b e i n g d e c o d e d f r o m t h e b i t s t r e a m a n d ( b ) a p l u r a l i t y o f p i e c e s o f s p l i t p o i n t i n f o r m a t i o n f o r t h e p o r t i o n o f t h e n e u r a l n e t w o r k . 4 9 3 1 3 8 2 6 _ 1 2 0 2 4 2 0 2 4 1 6 1 2 2 0 2 4 A p r P L U R A L I T Y O F T E N S O R S 19 / 19 23467558_1 Start 1600 Fig. 16 Stop Decode additional data 1610 1660 (1180) Perform neural network second portion Perform MSFC decoder Select MSFC decoder 1630 (1110) 1640 (1120, 1140, 1160) Select task information and split point info 1620 Decode tensors 1624 19 / 19 Start 1600 1610 Decode additional data 1620 Select task information and split point info 1624 Select MSFC decoder 1630 (1110) Decode tensors 1640 (1120, 1140, 1160) Perform MSFC decoder 1660 (1180) Perform neural network second portion Stop Fig. 16 23467558_1 20 24 20 24 16 12 A pr 2 02 4 1 9 / 1 9 S t a r t 1 6 0 0 2 0 2 4 2 0 2 4 1 6 1 2 A p r 2 0 2 4 1 6 1 0 D e c o d e a d d i t i o n a l d a t a 1 6 2 0 S e l e c t t a s k i n f o r m a t i o n a n d s p l i t p o i n t i n f o 1 6 2 4 S e l e c t M S F C d e c o d e r 1 6 3 0 ( 1 1 1 0 ) D e c o d e t e n s o r s 1 6 4 0 ( 1 1 2 0 , 1 1 4 0 , 1 1 6 0 ) P e r f o r m M S F C d e c o d e r 1 6 6 0 ( 1 1 8 0 ) P e r f o r m n e u r a l n e t w o r k s e c o n d p o r t i o n S t o p F i g . 1 6 2 3 4 6 7 5 5 8 _ 1 1 9 / 1 9 2 0 2 4 2 0 2 4 1 6 1 2 A p r 2 0 2 4 1 6 1 0 1 6 2 4 S e l e c t M S F C d e c o d e r D e c o d e t e n s o r s P e r f o r m M S F C d e c o d e r
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates generally to digital video signal processing and, in particular, to a method, apparatus and system for encoding and decoding tensors from a convolutional neural network. The present invention also relates to a computer program product including a computer readable medium having recorded thereon a computer program for encoding and decoding tensors from a convolutional neural network using video compression technology. BACKGROUND

[0002] Convolution neural networks (CNNs) are an emerging technology addressing, among other things, use cases involving machine vision such as object detection, instance segmentation, object tracking, human pose estimation and action recognition. Applications for CNNs can involve use of ‘edge devices’ with sensors and some processing capability, coupled to application servers as part of a ‘cloud’. CNNs can require relatively high computational complexity, more than can typically be afforded either in computing capacity or power consumption by an edge device. Executing a CNN in a distributed manner has emerged as one solution to running leading edge networks using limited capability edge devices. In other words, distributed processing allows legacy edge devices to still provide the capability of leading edge CNNs by distributing processing between the edge device and external processing means, such as cloud servers.

[0003] CNNs typically include many layers, such as convolution layers and fully connected layers, with data passing from one layer to the next in the form of ‘tensors’. Splitting a network across different devices introduces a need to compress the intermediate tensor data that passes from one layer to the next within a CNN, such compression may be referred to as ‘feature compression’, as the intermediate tensor data is often termed ‘features’ of input such as an image frame or video frame. International Organisation for Standardisation / International Electrotechnical Commission Joint Technical Committee 1 / Subcommittee 29 / Working Groups 2-8 (ISO / IEC JTC1 / SC29 / WG2-8), also known as the “Moving Picture Experts Group” (MPEG) are tasked with studying compression technology relating to video. WG2 ‘MPEG 2024202416   12 Apr 2024 Technical Requirements’ has established a ‘Video Compression for Machines’ (VCM) ad-hoc group, mandated to study video compression for machine consumption and feature compression. The feature compression mandate is in an exploratory phase with a ‘Call for Evidence’ (CfE) anticipated to be issued, to solicit technology that can significantly outperform feature compression results achieved using state-of-the-art standardised technology.

[0004] CNNs require weights for each of the layers to be determined in a training stage, where a very large amount of training data is passed through the CNN and a determined result is compared to ground truth associated with the training data. A process for updating network weights, such as stochastic gradient descent, is applied to iteratively refine the network weights until the network performs at a desired level of accuracy. Where a convolution stage has a ‘stride’ greater than one, an output tensor from the convolution has a lower spatial resolution than a corresponding input tensor. Pooling operations result in an output tensor having smaller dimensions than the input tensor. One example of a pooling operation is ‘max pooling’ (or ‘Maxpool’), which reduces the spatial size of the output tensor compared to the input tensor. Max pooling produces an output tensor by dividing the input tensor into groups of data samples (e.g., a 2*2 group of data samples), and from each group selecting a maximum value as output for a corresponding value in the output tensor. The process of executing a CNN with an input and progressively transforming the input into an output is commonly referred to as ‘inferencing’.

[0005] Generally, a tensor has four dimensions, namely: batch, channels, height, and width. The first dimension, ‘batch’, of size ‘one’ when inferencing on video data indicates that one frame is passed through a CNN at a time. When training a network, the value of the batch dimension may be increased so that multiple frames are passed through the network before the network weights are updated, according to a predetermined ‘batch size’. A multi-frame video may be passed through as a single tensor with the batch dimension increased in size according to the number of frames of a given video. However, for practical considerations relating to memory consumption and access, inferencing on video data is typically performed on a framewise basis. The ‘channels’ dimension indicates the number of concurrent ‘feature maps’ for a given tensor and the height and width dimensions indicate the size of the feature maps at the particular stage of the CNN. Channel count varies through a CNN according to the network architecture. Feature map size also varies, depending on subsampling occurring in specific network layers. 2024202416   12 Apr 2024

[0006] Input to the first layer of a CNN is a batch of one or more images, for example, a single image or video frame, typically resized for compatibility with the dimensionality of the tensor input to the first layer. It is also possible to supply images or video frames in batches of size larger than one. The dimensionality of tensors is dependent on the CNN architecture, generally having some dimensions relating to input width and height and a further ‘channel’ dimension.

[0007] Slicing, or reducing a tensor to a collection of two-dimensional arrays, a tensor based on the channel dimension results in a set of two-dimensional ‘feature maps’, so-called because each slice of the tensor has some relationship to the corresponding input image, capturing properties such as various edge types. At layers further from the input to the network, the property can be more abstract. The ‘task performance’ of a CNN is measured by comparing the result of the CNN in performing a task using specific input with a provided ground truth, generally prepared by humans and deemed to indicate a ‘correct’ result.

[0008] Once a network topology is decided, the network weights may be updated over time as more training data becomes available. The overall complexity of the CNN tends to be relatively high, with relatively large numbers of multiply-accumulate operations being performed and numerous intermediate tensors being written to and read from memory. In some applications, the CNN is implemented entirely in the ‘cloud’, resulting in a need for high and costly processing power. In other applications, the CNN is implemented in an edge device, such as a camera or mobile phone, resulting in less flexibility but a more distributed processing load. An emerging architecture involves splitting a network into portions, one of the portions run in an edge device and another portion run in the cloud. Such a distributed network architecture may be referred to as ‘collaborative intelligence’ and offers benefits such as re-using a partial result from a first portion of the network with several different second portions, perhaps each portion being optimised for a different task. Collaborative intelligence architectures introduce a need for efficient compression of tensor data, for transmission over a network such as a WAN.

[0009] Some existing split network solutions have low performance in an end to end network. For example, a relatively high level of accuracy may be achieved at a cost of producing a relatively large encoded feature map. Some variants may be required to train multiple stages to achieve a sufficiently high performance. Training multiple stages can be complex and timeconsuming. 2024202416   12 Apr 2024

[00010] Existing split network solutions have generally been developed and trained for one task at one split point only. However, as one CNN backbone and feature compression encoder can be used for multiple CNN heads, solutions that are trained for one task at one split point only can be inflexible and require separate instantiations for each CNN backbone.

[00011] Video compression standards can be used for feature compression, as described below. Various methods can be used to constrict or reduce the data being presented for compression. However, some methods used to constrict or reduce the data being presented for compression can result in a decrease in accuracy unsuitable for some tasks implemented by CNNs.

[00012] Feature compression may benefit from existing video compression standards, such as Versatile Video Coding (VVC), developed by the Joint Video Experts Team (JVET). VVC is anticipated to address ongoing demand for ever-higher compression performance, especially as video formats increase in capability (e.g., with higher resolution and higher frame rate) and to address increasing market demand for service delivery over WANs, where bandwidth costs are relatively high. VVC is implementable in contemporary silicon processes and offers an acceptable trade-off between achieved performance versus implementation cost. The implementation cost may be considered for example, in terms of one or more of silicon area, CPU processor load, memory utilisation and bandwidth. Part of the versatility of the VVC standard is in the wide selection of tools available for compressing video data, as well as the wide range of applications for which VVC is suitable. Other video compression standards, such as High Efficiency Video Coding (HEVC) and AV-1, may also be used for feature compression applications.

[00013] Video data includes a sequence of frames of image data, each frame including one or more colour channels. Generally, one primary colour channel and two secondary colour channels are needed. The primary colour channel is generally referred to as the ‘luma’ channel and the secondary colour channel(s) are generally referred to as the ‘chroma’ channels. Although video data is typically displayed in an RGB (red-green-blue) colour space, this colour space has a high degree of correlation between the three respective components. The video data representation seen by an encoder or a decoder is often using a colour space such as YCbCr. YCbCr concentrates luminance, mapped to ‘luma’ according to a transfer function, in a Y (primary) channel and chroma in Cb and Cr (secondary) channels. Due to the use of a decorrelated YCbCr signal, the statistics of the luma channel differ markedly from those of the chroma channels. A primary difference is that after quantisation, the chroma channels contain 2024202416   12 Apr 2024 relatively few significant coefficients for a given block compared to the coefficients for a corresponding luma channel block. Moreover, the Cb and Cr channels may be sampled spatially at a lower rate (subsampled) compared to the luma channel, for example half horizontally and half vertically - known as a ‘4:2:0 chroma format’. The 4:2:0 chroma format is commonly used in ‘consumer’ applications, such as internet video streaming, broadcast television, and storage on Blu-RayTM disks. When only luma samples are present, the resulting monochrome frames are said to use a “4:0:0 chroma format”.

[00014] The VVC standard specifies a ‘block based’ architecture, in which frames are firstly divided into a square array of regions known as ‘coding tree units’ (CTUs). CTUs generally occupy a relatively large area, such as 128x128 luma samples. Other possible CTU sizes when using the VVC standard are 32x32 and 64x64. However, CTUs at the right and bottom edge of each frame may be smaller in area, with implicit splitting occurring the ensure the CBs remain in the frame. Associated with each CTU is a ‘coding tree’ either for both the luma channel and the chroma channels (a ‘shared tree’) or a separate tree each for the luma channel and the chroma channels. A coding tree defines a decomposition of the area of the CTU into a set of blocks, also referred to as ‘coding blocks’ (CBs). When a shared tree is in use a single coding tree specifies blocks both for the luma channel and the chroma channels, in which case the collections of collocated coding blocks are referred to as ‘coding units’ (CUs) (i.e., each CU having a coding block for each colour channel). The CBs are processed for encoding or decoding in a particular order. As a consequence of the use of the 4:2:0 chroma format, a CTU with a luma coding tree for a 128x128 luma sample area has a corresponding chroma coding tree for a 64x64 chroma sample area, collocated with the 128x128 luma sample area. When a single coding tree is in use for the luma channel and the chroma channels, the collections of collocated blocks for a given area are generally referred to as ‘units’, for example, the abovementioned CUs, as well as ‘prediction units’ (PUs), and ‘transform units’ (TUs). A single tree with CUs spanning the colour channels of 4:2:0 chroma format video data result in chroma blocks half the width and height of the corresponding luma blocks. When separate coding trees are used for a given area, the above-mentioned CBs, as well as ‘prediction blocks’ (PBs), and ‘transform blocks’ (TBs) are used.

[00015] Notwithstanding the above distinction between ‘units’ and ‘blocks’, the term ‘block’ may be used as a general term for areas or regions of a frame for which operations are applied to all colour channels. 2024202416   12 Apr 2024

[00016] For each CU, a prediction unit (PU) of the contents (sample values) of the corresponding area of frame data is generated (a ‘prediction unit’). Further, a representation of the difference (or ‘spatial domain’ residual) between the prediction and the contents of the area as seen at input to the encoder is formed. The difference in each colour channel may be transformed and coded as a sequence of residual coefficients, forming one or more TUs for a given CU. The applied transform may be a Discrete Cosine Transform (DCT) or other transform, applied to each block of residual values. The transform is applied separably, (i.e., the two-dimensional transform is performed in two passes). The block is firstly transformed by applying a one-dimensional transform to each row of samples in the block. Then, the partial result is transformed by applying a one-dimensional transform to each column of the partial result to produce a final block of transform coefficients that substantially decorrelates the residual samples. Transforms of various sizes are supported by the VVC standard, including transforms of rectangular-shaped blocks, with each side dimension being a power of two. Transform coefficients are quantised for entropy encoding into a bitstream.

[00017] VVC features intra-frame prediction and inter-frame prediction. Intra-frame prediction involves the use of previously processed samples in a frame being used to generate a prediction of a current block of data samples in the frame. Inter-frame prediction involves generating a prediction of a current block of samples in a frame using a block of samples obtained from a previously decoded frame. The block of samples obtained from a previously decoded frame is offset from the spatial location of the current block according to a motion vector, which often has filtering applied. Intra-frame prediction blocks can be (i) a uniform sample value (“DC intra prediction”), (ii) a plane having an offset and horizontal and vertical gradient (“planar intra prediction”), (iii) a population of the block with neighbouring samples applied in a particular direction (“angular intra prediction”) or (iv) the result of a matrix multiplication using neighbouring samples and selected matrix coefficients. Further discrepancy between a predicted block and the corresponding input samples may be corrected to an extent by encoding a ‘residual’ into the bitstream. The residual is generally transformed from the spatial domain to the frequency domain to form residual coefficients in a ‘primary transform’ domain. The residual coefficients may be further transformed by application of a ‘secondary transform’ to produce residual coefficients in a ‘secondary transform domain’. Residual coefficients are quantised according to a quantisation parameter, resulting in a loss of accuracy of the reconstruction of the samples produced at the decoder but with a reduction in bitrate in the bitstream. Sequences of pictures may be encoded according to a specified structure of pictures 2024202416   12 Apr 2024 using intra-prediction and pictures using intra- or inter-prediction, and specified dependencies on preceding pictures in coding order, which may differ from display or delivery order. A ‘random access’ configuration results in periodic intra-pictures, forming entry points at which a decoder and commence decoding a bitstream. Other pictures in a random-access configuration generally use inter-prediction to predict content from pictures preceding and following a current picture in display or delivery order, according to a hierarchical structure of specified depth. The use of pictures after a current picture in display order for predicting a current picture requires a degree of picture buffering and delay between the decoding of a given picture and the display (and removal from the buffer) of the given picture. SUMMARY

[00018] It is an object of the present invention to substantially overcome, or at least ameliorate, one or more disadvantages of existing arrangements.

[00019] One aspect of the present disclosure provides a method of encoding tensors, the method comprising: encoding one or more tensors produced by a portion of a neural network into a bitstream; and encoding additional information into the bitstream, the additional information comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network.

[00020] Another aspect of the present disclosure provides a method of decoding one or more tensors from a bitstream, the method comprising: decoding one or more initial tensors from the bitstream; decoding additional information from the bitstream, the additional data comprising (a) a plurality of pieces of task information each of which indicates machine task which is capable of being performed for the one or more decoded tensors and (b) a plurality of pieces of split point information; selecting one of the plurality of pieces of task information and one of the plurality of pieces of split point information; and producing the decoded one or more tensors from the initial one or more tensors based on the selected task information and the selected split point information.

[00021] Another aspect of the present disclosure provides an encoder for encoding tensors, the encoder configured to: encode a one or more tensors produced by a portion of a neural network into a bitstream; and encode additional information into the bitstream, the additional 2024202416   12 Apr 2024 information comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network.

[00022] Another aspect of the present disclosure provides a computer-implemented medium non-transitory computer-readable storage medium which stores a program for executing a method of encoding tensors, the method comprising: encoding one or more tensors produced by a portion of a neural network into a bitstream; and encoding additional information into the bitstream, the additional information comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network.

[00023] Another aspect of the present disclosure provides a system comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory for implementing a method of encoding tensors, the method comprising: encoding one or more tensors produced by a portion of a neural network into a bitstream; and encoding additional information into the bitstream, the additional information comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network.

[00024] Another aspect of the present disclosure provides a decoder for decoding one or more tensors from a bitstream, the decoder configured to: decode one or more initial tensors from the bitstream; decode additional information from the bitstream, the additional data comprising (a) a plurality of pieces of task information each of which indicates machine task which is capable of being performed for the one or more decoded tensors and (b) a plurality of pieces of split point information; select one of the plurality of pieces of task information and one of the plurality of pieces of split point information; and produce the decoded one or more tensors from the initial one or more tensors based on the selected task information and the selected split point information.

[00025] Another aspect of the present disclosure provides a computer-implemented medium non-transitory computer-readable storage medium which stores a program for executing a method of 2024202416   12 Apr 2024 decoding one or more tensors from a bitstream, the method comprising: decoding one or more initial tensors from the bitstream; decoding additional information from the bitstream, the additional data comprising (a) a plurality of pieces of task information each of which indicates machine task which is capable of being performed for the one or more decoded tensors and (b) a plurality of pieces of split point information; selecting one of the plurality of pieces of task information and one of the plurality of pieces of split point information; and producing the decoded one or more tensors from the initial one or more tensors based on the selected task information and the selected split point information.

[00026] Another aspect of the present disclosure provides a system comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory for implementing a method of decoding one or more tensors from a bitstream, the method comprising: decoding one or more initial tensors from the bitstream; decoding additional information from the bitstream, the additional data comprising (a) a plurality of pieces of task information each of which indicates machine task which is capable of being performed for the one or more decoded tensors and (b) a plurality of pieces of split point information; selecting one of the plurality of pieces of task information and one of the plurality of pieces of split point information; and producing the decoded one or more tensors from the initial one or more tensors based on the selected task information and the selected split point information.

[00027] Other aspects are also disclosed. BRIEF DESCRIPTION OF THE DRAWINGS

[00028] At least one embodiment of the present invention will now be described with reference to the following drawings, in which:

[00029] Fig. 1 is a schematic block diagram showing a distributed machine task system;

[00030] Figs. 2A and 2B form a schematic block diagram of a general-purpose computer system upon which the distributed machine task system of Fig. 1 may be practiced;

[00031] Fig. 3A is a schematic block diagram showing functional modules of a backbone portion of a CNN;

[00032] Fig. 3B is a schematic block diagram showing a residual block of Fig. 3A;

[00033] Fig. 3C is a schematic block diagram showing a residual unit of Fig. 3A; 2024202416   12 Apr 2024

[00034] Fig. 3D is a schematic block diagram showing a convolutional batch normalisation leaky rectified linear (CBL)module of Fig. 3A;

[00035] Fig. 3E shows an example structure of a convolutional neural network;

[00036] Fig. 4 is a schematic block diagram showing functional modules of an alternative backbone portion of a CNN;

[00037] Fig. 5A is a schematic block diagram showing a cross-layer tensor bottleneck for reducing tensor dimensionality prior to compression;

[00038] Fig. 5B is a schematic block diagram showing a functional block used in Fig. 5A;

[00039] Fig. 6 shows a method for performing a first portion of a CNN, constricting using a bottleneck encoder, and encoding resulting constricted feature maps;

[00040] Fig. 7 is a schematic block diagram showing a packing arrangement for a plurality of compressed tensors;

[00041] Fig. 8A is a schematic block diagram showing functional modules of a video encoder;

[00042] Fig. 8B is a schematic block diagram showing a bitstream holding encoded interchannel decorrelated feature maps and associated metadata;

[00043] Fig. 9 is a schematic block diagram showing functional modules of a video decoder;

[00044] Fig. 10A is a schematic block diagram showing a cross-layer tensor inverse bottleneck for restoring tensor dimensionality after compression;

[00045] Fig. 10B shows a schematic block if a functional block used in Fig. 10A;

[00046] Fig. 10C shows a schematic block diagram showing a reconstruction block used in Fig. 10A;

[00047] Fig. 11 shows a method for decoding a bitstream, reconstructing decorrelated feature maps, and performing a second portion of the CNN;

[00048] Fig. 12A is a schematic block diagrams showing a head portion of a CNN; 2024202416   12 Apr 2024

[00049] Fig. 12B is a schematic block diagram showing an upscaler module of Fig. 12A;

[00050] Fig. 12C is a schematic block diagram showing a detection module of Fig. 12A;

[00051] Fig. 13 is a schematic block diagram showing an alternative head portion of a CNN;

[00052] Fig. 14A is a schematic block diagram showing a cross-layer tensor inverse bottleneck for restoring tensor dimensionality after compression;

[00053] Fig. 14B is a schematic block diagram showing another implementation of a crosslayer tensor inverse bottleneck for restoring tensor dimensionality after compression;

[00054] Fig. 15 is a schematic block diagram showing a cross-layer tensor from bottleneck encoder to bottleneck decoder with multiple choices of bottleneck decoders; and

[00055] Fig. 16 shows a method for decoding and selecting task information, split point information from additional data, selecting MSFC decoder, and decoding tensors, reconstructing decorrelated feature maps, and performing a second portion of the CNN. DETAILED DESCRIPTION INCLUDING BEST MODE

[00056] Where reference is made in any one or more of the accompanying drawings to steps and / or features, which have the same reference numerals, those steps and / or features have for the purposes of this description the same function(s) or operation(s), unless the contrary intention appears.

[00057] A distributed machine task system may include an edge device, such as a network camera or smartphone producing intermediate compressed data. The distributed machine task system may also include a final device, such as a server-farm based (‘cloud’) application, operating on the intermediate compressed data to produce (generate) some task result. Additionally, the edge device functionality may be embodied in the cloud and the intermediate compressed data may be stored for later processing, potentially for multiple different tasks depending on need. Examples of machine task include object detection and instance segmentation, both of which produce a task result measured as ‘mean average precision’ (mAP) for detection over a threshold value of intersection-over-union (IoU), such as 0.5. Another example machine task is object tracking, with mean object tracking accuracy (MOTA) score as a typical task result. 2024202416   12 Apr 2024

[00058] A convenient form of intermediate compressed data is a compressed video bitstream, owing to the availability of high-performing compression standards and implementations thereof. Video compression standards typically operate on integer samples of some given bit depth, such as 10 bits, arranged in planar arrays. Colour video has three planar arrays, corresponding, for example, to colour components Y, Cb, Cr, or R, G, B, depending on application. CNNs typically operate on floating point data in the form of tensors. Tensors generally have a much smaller spatial dimensionality compared to incoming video data upon which the CNN operates but have many more channels than the three channels typical of colour video data.

[00059] Tensors typically have the following dimensions: Frames, channels, height, and width. For example, a tensor of dimensions [1, 256, 76, 136] would be said to contain two-hundred and fifty-six (256) feature maps, each of size 136*76. For video data, inferencing is typically performed one frame at a time, rather than using tensors containing multiple frames.

[00060] VVC supports a division of a picture into multiple subpictures, each of which may be independently encoded and independently decoded. In one approach, each subpicture is coded as one ‘slice’, or contiguous sequence of coded CTUs. A ‘tile’ mechanism is also available to divide a picture into a number of independently decodeable regions. Subpictures may be specified in a somewhat flexible manner, with various rectangular sets of CTUs coded as respective subpictures. Flexible definition of subpicture dimensions allows efficiently holding types of data requiring different areas in one picture, avoiding large ‘unused’ areas, i.e., areas of a frame that are not used for reconstruction of tensor data.

[00061] Fig. 1 is a schematic block diagram showing functional modules of a distributed machine task system 100. The notion of distributing a machine task across multiple systems is sometimes referred to as ‘collaborative intelligence’ (CI). The system 100 may be used for implementing methods for decorrelating, packing and quantising feature maps into planar frames for encoding and decoding feature maps from encoded data. The methods may be implemented such that associated overhead data is not too burdensome and task performance on the decoded feature maps is resilient to changing bitrate of the bitstream and the quantised representation of the tensors does not needlessly consume bits where the bits do not provide a commensurate benefit in terms of task performance. 2024202416   12 Apr 2024

[00062] The system 100 includes a source device 110 for generating encoded tensor data 115 from a CNN backbone 114 in the form of encoded video bitstream 121. The system 100 also includes a destination device 140 for decoding tensor data in the form of an encoded video bitstream 143. A communication channel 130 is used to communicate the encoded video bitstream 121 from the source device 110 to the destination device 140. In some arrangements, the source device 110 and destination device 140 may either or both comprise respective mobile telephone handsets (e.g., “smartphones”) or network cameras and cloud applications. The communication channel 130 may be a wired connection, such as Ethernet, or a wireless connection, such as WiFi or 5G, including connections across a Wide Area Network (WAN) or across ad-hoc connections. Moreover, the source device 110 and the destination device 140 may comprise applications where encoded video data is captured on some computer-readable storage medium, such as a hard disk drive in a file server or memory.

[00063] As shown in Fig. 1, the source device 110 includes a video source 112, the CNN backbone 114, a bottleneck encoder 116, a quantise and pack module 118, a feature map encoder 120, and a transmitter 122. The video source 112 typically comprises a source of captured video frame data (shown as 113), such as an image capture sensor, a previously captured video sequence stored on a non-transitory recording medium, or a video feed from a remote image capture sensor. The video source 112 may also be an output of a computer graphics card, for example, displaying the video output of an operating system and various applications executing upon a computing device (e.g., a tablet computer). Examples of source devices 110 that may include an image capture sensor as the video source 112 include smartphones, video camcorders, professional video cameras, and network video cameras. The system 100 reduces dimensionality of tensors at the interface between the first network portion and the second network portion using a ‘bottleneck’, i.e., additional network layers that restrict tensor dimensionality on the encoding side and restore tensor dimensionality at the decoder side. The multi-scale representation produced by a feature pyramid network (FPN) is ‘fused’ together into a single tensor using an approach named ‘multi-scale feature compression’ (MSFC). MSFC is ordinarily used to merge all FPN layers into a single tensor. Merging all FPN layers into a single tensor is implemented at the expense of spatial detail for the less decomposed (larger) layers of the FPN. The loss of spatial detail can result in an unacceptable decrease in accuracy for some tasks or operations implemented by the system 100.

[00064] The arrangements described separate tensors of the FPN into groups and separately apply MSFC techniques rather than merging all FPN layers into a single tensor. Separately 2024202416   12 Apr 2024 applying MFSC techniques permits a degree of cross-layer fusion without such severe degradation of spatial detail. For tasks requiring preservation of greater spatial detail, such as instance segmentation, the resulting mAP is higher using separate MSFC techniques than if all FPN layers are merged into a single tensor with low spatial resolution.

[00065] The CNN backbone 114 receives the video frame data 113 and performs specific layers of an overall CNN, such as layers corresponding to the ‘backbone’ of the CNN, outputting the tensors 115. The backbone layers of the CNN may produce (generate) multiple tensors as output, for example, corresponding to different spatial scales of an input image represented by the video frame data 113, sometimes referred to as a ‘feature pyramid network’ (FPN) architecture. The tensors resulting from an FPN backbone form a hierarchical representation of the frame data 113 including data of feature maps. Each successive layer of the hierarchical representation has half the width and height of the preceding layer. Later layers that are produced further into in the backbone network tend to contain feature having a more abstract representation of the frame data 113. Less decomposed layers, produced earlier in the backbone network, tend to contain features representing less abstract features of the frame data 113, such as various geometric properties such as edges of various angles. An FPN may result in three tensors, corresponding to three layers, output from the backbone 114 as the tensors 115 when a ‘YOLOv3’ network is performed by the system 100, with varying spatial resolution and channel count. When the system 100 is performing networks such as ‘Faster RCNN X101-FPN” or “Mask RCNN X101-FPN” the tensors 115 include tensors for four layers P2-P5. The bottleneck encoder 116 receives the tensors 115. The bottleneck encoder 116 acts to compress one or more internal layers of the overall CNN. The internal layers of the overall CNN provide the output of the CNN backbone 114, compressed or constricted by the bottleneck encoder 116 using a set of neural network layers trained to convert to a lower channel count and smaller spatial resolution than required by the tensors 115. The bottleneck encoder 116 outputs bottleneck tensors 117. The bottleneck tensors 117 are passed to the quantise and pack module 118. Each feature map of the bottleneck tensors 117 is quantised from floating point to integer precision and packed into a monochrome frame by the module 118 to produce a frame 119. The frame 119 is encoded by the feature map encoder 120 to produce the bitstream 121. The bitstream 121 is supplied to the transmitter 122 for transmission over the communications channel 130 or the bitstream 121 is written to storage 132 for later use.

[00066] The source device 110 supports a particular network for the CNN backbone 114. However, the destination device 140 may use one of several networks for a corresponding head 2024202416   12 Apr 2024 CNN 150. In this way, partially processed data in the form of packed feature maps may be stored for later use in performing various tasks without needing to repeatedly perform the operation of the CNN backbone 114. In some implementations, the CNN backbone 114 may also be referred to as a neural network part 1 or “NN part 1” and the CNN head 150 may also be referred to as a neural network part 2 or “NN part 2”.

[00067] The bitstream 121 is transmitted by the transmitter 122 over the communication channel 130 as encoded video data (or “encoded video information”). The bitstream 121 can in some implementations be stored in the storage 132, where the storage 132 is a non-transitory storage device such as a “Flash” memory or a hard disk drive, until later being transmitted over the communication channel 130 (or in-lieu of transmission over the communication channel 130). For example, encoded video data may be served upon demand to customers over a wide area network (WAN) for a video analytics application.

[00068] The destination device 140 includes a receiver 142, a feature map decoder 144, an unpack and inverse quantise module 146, a bottle neck decoder 148, the CNN head 150, and a CNN task result buffer 152. The receiver 142 receives encoded video data from the communication channel 130 and passes the video bitstream 143 to the feature map decoder 144. The feature map decoder 144 operates to decode the feature maps and output a decoded frame 145. The decoded frame 145 is passed to the unpack and inverse quantise module 146. The module 146 unpacks and inverse quantises the tensors of the frame 145 to generate dequantized tensors, output as decoded bottleneck tensors 147. The decoded bottleneck tensors 147 are supplied to the bottleneck decoder 148. The bottleneck decoder 148 performs the inverse operation of the bottleneck encoder 116, to produce extracted tensors 149. The extracted tensors 149 are passed to the CNN head 150. The CNN head 150 performs the later layers of the task that began with the CNN backbone 114 to produce (generate) a task result 151, which is stored in a task result buffer 152. The contents of the task result buffer 152 may be presented to the user, e.g., via a graphical user interface, or provided to an analytics application where some action is decided based on the task result, which may include summary level presentation of aggregated task results to a user. It is also possible for the functionality of each of the source device 110 and the destination device 140 to be embodied in a single device, examples of which include mobile telephone handsets and tablet computers and cloud applications.

[00069] Notwithstanding the example devices mentioned above, each of the source device 110 and destination device 140 may be configured within a general-purpose computing system, 2024202416   12 Apr 2024 typically through a combination of hardware and software components. Fig. 2A illustrates such a computer system 200, which includes: a computer module 201; input devices such as a keyboard 202, a mouse pointer device 203, a scanner 226, a camera 227, which may be configured as the video source 112, and a microphone 280; and output devices including a printer 215, a display device 214, which may be configured as a display device presenting the task result 151, and loudspeakers 217. An external Modulator-Demodulator (Modem) transceiver device 216 may be used by the computer module 201 for communicating to and from a communications network 220 via a connection 221. The communications network 220, which may represent the communication channel 130, may be a (WAN), such as the Internet, a cellular telecommunications network, or a private WAN. Where the connection 221 is a telephone line, the modem 216 may be a traditional “dial-up” modem. Alternatively, where the connection 221 is a high capacity (e.g., cable or optical) connection, the modem 216 may be a broadband modem. A wireless modem may also be used for wireless connection to the communications network 220. The transceiver device 216 may provide the functionality of the transmitter 122 and the receiver 142 and the communication channel 130 may be embodied in the connection 221.

[00070] The computer module 201 typically includes at least one processor unit 205, and a memory unit 206. For example, the memory unit 206 may have semiconductor random access memory (RAM) and semiconductor read only memory (ROM). The computer module 201 also includes a number of input / output (I / O) interfaces including: an audio-video interface 207 that couples to the video display 214, loudspeakers 217 and microphone 280; an I / O interface 213 that couples to the keyboard 202, mouse 203, scanner 226, camera 227 and optionally a joystick or other human interface device (not illustrated); and an interface 208 for the external modem 216 and printer 215. The signal from the audio-video interface 207 to the computer monitor 214 is generally the output of a computer graphics card. In some implementations, the modem 216 may be incorporated within the computer module 201, for example within the interface 208. The computer module 201 also has a local network interface 211, which permits coupling of the computer system 200 via a connection 223 to a local-area communications network 222, known as a Local Area Network (LAN). As illustrated in Fig. 2A, the local communications network 222 may also couple to the wide network 220 via a connection 224, which would typically include a so-called “firewall” device or device of similar functionality. The local network interface 211 may comprise an EthernetTM circuit card, a BluetoothTM wireless arrangement or an IEEE 802.11 wireless arrangement; however, numerous other types 2024202416   12 Apr 2024 of interfaces may be practiced for the interface 211. The local network interface 211 may also provide the functionality of the transmitter 122 and the receiver 142 and communication channel 130 may also be embodied in the local communications network 222.

[00071] The I / O interfaces 208 and 213 may afford either or both of serial and parallel connectivity, the former typically being implemented according to the Universal Serial Bus (USB) standards and having corresponding USB connectors (not illustrated). Storage devices 209 are provided and typically include a hard disk drive (HDD) 210. Other storage devices such as a floppy disk drive and a magnetic tape drive (not illustrated) may also be used. An optical disk drive 212 is typically provided to act as a non-volatile source of data. Portable memory devices, such optical disks (e.g., CD-ROM, DVD, Blu ray DiscTM), USB-RAM, portable, external hard drives, and floppy disks, for example, may be used as appropriate sources of data to the computer system 200. Typically, any of the HDD 210, optical drive 212, networks 220 and 222 may also be configured to operate as the video source 112, or as a destination for decoded video data to be stored for reproduction via the display 214. The source device 110 and the destination device 140 of the system 100 may be embodied in the computer system 200.

[00072] The components 205 to 213 of the computer module 201 typically communicate via an interconnected bus 204 and in a manner that results in a conventional mode of operation of the computer system 200 known to those in the relevant art. For example, the processor 205 is coupled to the system bus 204 using a connection 218. Likewise, the memory 206 and optical disk drive 212 are coupled to the system bus 204 by connections 219. Examples of computers on which the described arrangements can be practised include IBM-PC’s and compatibles, Sun SPARCstations, Apple MacTM or alike computer systems.

[00073] Where appropriate or desired, the source device 110 and the destination device 140, as well as methods described below, may be implemented using the computer system 200. In particular, the source device 110, the destination device 140 and methods to be described, may be implemented as one or more software application programs 233 executable within the computer system 200. The source device 110, the destination device 140 and the steps of the described methods are affected by instructions 231 (see Fig. 2B) in the software 233 that are carried out within the computer system 200. The software instructions 231 may be formed as one or more code modules, each for performing one or more particular tasks. The software may also be divided into two separate parts, in which a first part and the corresponding code 2024202416   12 Apr 2024 modules performs the described methods, and a second part and the corresponding code modules manage a user interface between the first part and the user.

[00074] The software may be stored in a computer readable medium, including the storage devices described below, for example. The software is loaded into the computer system 200 from the computer readable medium, and then executed by the computer system 200. A computer readable medium having such software or computer program recorded on the computer readable medium is a computer program product. The use of the computer program product in the computer system 200 preferably effects an advantageous apparatus for implementing the source device 110 and the destination device 140 and the described methods.

[00075] The software 233 is typically stored in the HDD 210 or the memory 206. The software is loaded into the computer system 200 from a computer readable medium and executed by the computer system 200. Thus, for example, the software 233 may be stored on an optically readable disk storage medium (e.g., CD-ROM) 225 that is read by the optical disk drive 212.

[00076] In some instances, the application programs 233 may be supplied to the user encoded on one or more CD-ROMs 225 and read via the corresponding drive 212, or alternatively may be read by the user from the networks 220 or 222. Still further, the software can also be loaded into the computer system 200 from other computer readable media. Computer readable storage media refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to the computer system 200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tape, CD-ROM, DVD, Blu-ray DiscTM, a hard disk drive, a ROM or integrated circuit, USB memory, a magneto-optical disk, or a computer readable card such as a PCMCIA card and the like, whether or not such devices are internal or external of the computer module 201. Examples of transitory or non-tangible computer readable transmission media that may also participate in the provision of the software, application programs, instructions and / or video data or encoded video data to the computer module 201 include radio or infra-red transmission channels, as well as a network connection to another computer or networked device, and the Internet or Intranets including e-mail transmissions and information recorded on Websites and the like.

[00077] The second part of the application program 233 and the corresponding code modules mentioned above may be executed to implement one or more graphical user interfaces (GUIs) to be rendered or otherwise represented upon the display 214. Through manipulation of 2024202416   12 Apr 2024 typically the keyboard 202 and the mouse 203, a user of the computer system 200 and the application may manipulate the interface in a functionally adaptable manner to provide controlling commands and / or input to the applications associated with the GUI(s). Other forms of functionally adaptable user interfaces may also be implemented, such as an audio interface utilizing speech prompts output via the loudspeakers 217 and user voice commands input via the microphone 280.

[00078] Fig. 2B is a detailed schematic block diagram of the processor 205 and a “memory” 234. The memory 234 represents a logical aggregation of all the memory modules (including the storage devices 209 and semiconductor memory 206) that can be accessed by the computer module 201 in Fig. 2A.

[00079] When the computer module 201 is initially powered up, a power-on self-test (POST) program 250 executes. The POST program 250 is typically stored in a ROM 249 of the semiconductor memory 206 of Fig. 2A. A hardware device such as the ROM 249 storing software is sometimes referred to as firmware. The POST program 250 examines hardware within the computer module 201 to ensure proper functioning and typically checks the processor 205, the memory 234 (209, 206), and a basic input-output systems software (BIOS) module 251, also typically stored in the ROM 249, for correct operation. Once the POST program 250 has run successfully, the BIOS 251 activates the hard disk drive 210 of Fig. 2A. Activation of the hard disk drive 210 causes a bootstrap loader program 252 that is resident on the hard disk drive 210 to execute via the processor 205. This loads an operating system 253 into the RAM memory 206, upon which the operating system 253 commences operation. The operating system 253 is a system level application, executable by the processor 205, to fulfil various high-level functions, including processor management, memory management, device management, storage management, software application interface, and generic user interface.

[00080] The operating system 253 manages the memory 234 (209, 206) to ensure that each process or application running on the computer module 201 has sufficient memory in which to execute without colliding with memory allocated to another process. Furthermore, the different types of memory available in the computer system 200 of Fig. 2A need to be used properly so that each process can run effectively. Accordingly, the aggregated memory 234 is not intended to illustrate how particular segments of memory are allocated (unless otherwise stated), but rather to provide a general view of the memory accessible by the computer system 200 and how such memory is used. 2024202416   12 Apr 2024

[00081] As shown in Fig. 2B, the processor 205 includes a number of functional modules including a control unit 239, an arithmetic logic unit (ALU) 240, and a local or internal memory 248, sometimes called a cache memory. The cache memory 248 typically includes a number of storage registers 244-246 in a register section. One or more internal busses 241 functionally interconnect these functional modules. The processor 205 typically also has one or more interfaces 242 for communicating with external devices via the system bus 204, using a connection 218. The memory 234 is coupled to the bus 204 using a connection 219.

[00082] The application program 233 includes a sequence of instructions 231 that may include conditional branch and loop instructions. The program 233 may also include data 232 which is used in execution of the program 233. The instructions 231 and the data 232 are stored in memory locations 228, 229, 230 and 235, 236, 237, respectively. Depending upon the relative size of the instructions 231 and the memory locations 228-230, a particular instruction may be stored in a single memory location as depicted by the instruction shown in the memory location 230. Alternately, an instruction may be segmented into a number of parts each of which is stored in a separate memory location, as depicted by the instruction segments shown in the memory locations 228 and 229.

[00083] In general, the processor 205 is given a set of instructions which are executed therein. The processor 205 waits for a subsequent input, to which the processor 205 reacts to by executing another set of instructions. Each input may be provided from one or more of a number of sources, including data generated by one or more of the input devices 202, 203, data received from an external source across one of the networks 220, 202, data retrieved from one of the storage devices 206, 209 or data retrieved from a storage medium 225 inserted into the corresponding reader 212, all depicted in Fig. 2A. The execution of a set of the instructions may in some cases result in output of data. Execution may also involve storing data or variables to the memory 234.

[00084] The bottleneck encoder 116, the bottleneck decoder 148 and the described methods may use input variables 254, which are stored in the memory 234 in corresponding memory locations 255, 256, 257. The bottleneck encoder 116, the bottleneck decoder 148 and the described methods produce output variables 261, which are stored in the memory 234 in corresponding memory locations 262, 263, 264. Intermediate variables 258 may be stored in memory locations 259, 260, 266 and 267. 2024202416   12 Apr 2024

[00085] Referring to the processor 205 of Fig. 2B, the registers 244, 245, 246, the arithmetic logic unit (ALU) 240, and the control unit 239 work together to perform sequences of microoperations needed to perform “fetch, decode, and execute” cycles for every instruction in the instruction set making up the program 233. Each fetch, decode, and execute cycle comprises: a fetch operation, which fetches or reads an instruction 231 from a memory location 228, 229, 230; a decode operation in which the control unit 239 determines which instruction has been fetched; and an execute operation in which the control unit 239 and / or the ALU 240 execute the instruction.

[00086] Thereafter, a further fetch, decode, and execute cycle for the next instruction may be executed. Similarly, a store cycle may be performed by which the control unit 239 stores or writes a value to a memory location 232.

[00087] Each step or sub-process in the methods of Figs. 6, 11 and 16, to be described, is associated with one or more segments of the program 233 and is typically performed by the register section 244, 245, 247, the ALU 240, and the control unit 239 in the processor 205 working together to perform the fetch, decode, and execute cycles for every instruction in the instruction set for the noted segments of the program 233.

[00088] Fig. 3A is a schematic block diagram 300 showing functional modules of a backbone portion 310 of a CNN, which may serve as the CNN backbone 114. The backbone portion 114 is sometimes referred to as ‘DarkNet-53’ and forms the backbone of a ‘YOLOv3’ object detection network. Different backbones are also possible, resulting in a different number of and dimensionality of layers of the tensors 115 for each frame.

[00089] As shown in Fig. 3A, the video data 113 is passed to a resizer module 304. The resizer module 304 resizes the frame to a resolution suitable for processing by the CNN backbone 310, producing resized frame data 312. If the resolution of the frame data 113 is already suitable for the CNN backbone 310, operation of the resizer module 304 is not needed. The resized frame data 312 is passed to a convolutional batch normalisation leaky rectified linear (CBL) module 314 to produce tensors 316. The CBL 314 contains modules as described with reference to a CBL module 360, as shown in Fig 3D. 2024202416   12 Apr 2024

[00090] Referring to Fig. 3D, the CBL module 360 takes as input a tensor 361. The tensor 361 is passed to a convolutional layer 362 to produce tensor 363. When the convolutional layer 362 has a stride of one and padding is set to k samples, with a convolutional kernel of size 2k+1, the tensor 363 has the same spatial dimensions as the tensor 361. When the convolution layer 362 has a larger stride, such as two, the tensor 363 has smaller spatial dimensions (size) compared to the tensor 361, for example, halved in size for the stride of two. Regardless of the stride, the size of channel dimension of the tensor 363 may vary compared to the channel dimension of the tensor 361 for a particular CBL block. The tensor 363 is passed to a batch normalisation module 364 which outputs a tensor 365. The batch normalisation module 364 normalises the input tensor 363, applies a scaling factor and offset value to produce the output tensor 365. The scaling factor and offset value are derived from a training process. The tensor 365 is passed to a leaky rectified linear activation (“LeakyReLU”) module 366 to produce a tensor 367. The module 366 provides a ‘leaky’ activation function whereby positive values in the tensor are passed through and negative values are severely reduced in magnitude, for example, to 0.1X their former value.

[00091] Returning to Fig. 3A, the tensor 316 is passed from the CBL block 314 to a residual block 11 module 320. The module 320 contains a sequential concatenation of three residual blocks, containing 1, 2, and 8 residual units internally, respectively.

[00092] A residual block, such as present in the module 320, is described with reference to a ResBlock 340 as shown in Fig. 3B. The ResBlock 340 receives a tensor 341. The tensor is zero-padded by a zero-padding module 342 to produce a tensor 343. The tensor 343 is passed to a CBL module 344 to produce (generate) a tensor 345. The tensor 345 is passed to a residual unit 346, of which the residual block 340 includes a series of concatenated residual units. The last residual unit of the residual units 346 outputs a tensor 347.

[00093] A residual unit, such as the unit 346, is described with reference to a ResUnit 350 as shown in Fig. 3C. The ResUnit 350 takes a tensor 351 as input. The tensor 351 is passed to a CBL module 352 to produce a tensor 353. The tensor 353 is passed to a second CBL unit 354 to produce a tensor 355. An add module 356 sums the tensor 355 with the tensor 351 to produce a tensor 357. The add module 356 may also be referred to as a ‘shortcut’ as the input tensor 351 substantially influences the output tensor 357. For an untrained network, ResUnit 350 acts to pass-through tensors. As training is performed, the CBL modules 352 and 354 act to deviate the tensor 357 away from the tensor 351 in accordance with training data and ground truth data. 2024202416   12 Apr 2024

[00094] The Res11 module 320 outputs a tensor 322. The tensor 322 is output from the backbone module 310 as one of the layers 115 and also provided to a Res8 module 324. The Res8 module 324 is a residual block (i.e., 340), which includes eight residual units (i.e., 350). The Res8 module 324 produces a tensor 326. The tensor 326 is passed to a Res4 module 328 and output from the backbone module 310 as one of the layers 115. The Res4 module 328 contains a sequence of four residual blocks (i.e., 340), which includes residual unit (i.e., 350). The Res4 module 328 produces a tensor 329. The tensor 329 is output from the backbone module 310 as one of the layers 115. Collectively, the layer tensors 322, 326, and 329 are output as the tensors 115. The backbone CNN 310 may take as input a video frame of resolution 1088x608 and produce three tensors, corresponding to three layers, with the following dimensions: [1, 256, 76, 136], [1, 512, 38, 68], [1, 1024, 19, 34]. Another example of the three tensors corresponding to three layers may be [1, 512, 34, 19], [1, 256, 68, 38], [1, 128, 136, 76] which are respectively separated at 75th network layer, 90th network layer, and 105th network layer in the CNN 310. Each tensor can have a different resolution to the next tensor. The resolution of each tensor can double in height and width between respective tensors. In forming the output tenors 115, the 322, 326, and 329 provide a hierarchical representation of the frame data including data of feature maps for encoding to the bitstream. The separating points depend on the CNN310.

[00095] Fig. 4 is a schematic block diagram showing functional modules of an alternative backbone portion 400 of a CNN, which may serve as the CNN backbone 114. The backbone portion 400 implements a residual network with feature pyramid network (‘ResNet FPN’) and is an alternative to the CNN backbone 114. Frame data 113 is input and passes through a stem network 408, a res2 module 412, a res3 module 416, a res4 module 420, and a res5 module 424 via tensors 409, 413, 417, 421 and 425.

[00096] The stem network 408 includes a 7x7 convolution with a stride of two (2) and a max pooling operation. The res2 module 412, the res3 module 416, the res4 module 420, and the res5 module 424 perform convolution operations, LeakyReLU activations. Each module 412, 416, 420 and 424 also performs one halving of the resolution of the processed tensors via a stride setting of two. The tensors 413, 417, 421, and 425 are passed to size 1x1 pixel lateral convolution modules 446, 444, 442, and 440 respectively. The modules 440, 442, 444, and 446 produce tensors 441, 443, 445 and 447 respectively. The tensor 441 is passed to a 3x3 output convolution module 470, which produces an output tensor P5 471. The tensor 441 is also passed to upsampler module 450 to produce an upsampled tensor 451. 2024202416   12 Apr 2024

[00097] A summation module 460 sums the tensors 443 and 451 to produce a tensor 461. The tensor 461 is passed to an upsampler module 452 and a 3x3 lateral convolution module 472. The module 472 outputs a P4 tensor 473. The upsampler module 452 produces an upsampled tensor 453. A summation module 462 sums tensors 445 and 453 to produce a tensor 463. The tensor 463 is passed to a 3x3 lateral convolution module 474 and an upsampler module 454. The module 474 outputs a P3 tensor 475. The upsampler module 454 outputs an upsampled tensor 455. A summation module 464 sums the tensors 447 and 455 to produce tensor 465, which is passed to a 3x3 lateral convolution module 476. The module 476 outputs a P2 tensor 477.

[00098] The upsampler modules 450, 452, and 454 use nearest neighbour interpolation for low computational complexity. A max pool module 428 produces P6 tensor 429 as output. The tensors 429, 471, 473, 475, and 477 form the output tensor 115 of the CNN backbone 400. In forming the output tenors 115, the FPN of tensors 429, 471, 473, 475, and 477 provide a hierarchical representation of the frame data including data of feature maps for encoding to the bitstream.

[00099] Fig. 5A is a schematic block diagram showing one type of bottleneck encoder 500, which may serve as the bottleneck encoder 116. Fig. 6 shows a method 600 for encoding frame data, the encoding including performing a first portion of a CNN, constricting using the bottleneck encoder 500, and encoding resulting constricted feature maps. Fig. 7 shows a packing arrangement of feature maps from compressed tensors into a monochrome video frame. [000100] The bottleneck encoder 500 receives FPN tensors 501, corresponding to the tensors 115, and operates to restrict the dimensionality of the received tensors to fewer layers. Applying a bottleneck encoder and decoder between the first portion (backbone 114) and the second portion (head 150) of a separated neural network enables a reduction in the spatial area in a frame of packed tensor data. The reduction in spatial area is achieved by using the interface between the bottleneck encoder and the bottleneck decoder as the split point of the first and second portions (114 and 150) of the neural network. The bottleneck encoder 116 acts as additional layers appended to the neural network first portion and the bottleneck decoder 148 acts as additional layers prepended to the neural network second portion. [000101] The sensitivity of the task result to the bottleneck depends on the nature of the task. For object detection, there is less spatial sensitivity and so spatial downsampling of larger layers of the FPN is less detrimental to the resulting mAP. In contrast, segmentation and the resulting 2024202416   12 Apr 2024 segmentation maps are more sensitive to a loss of spatial detail and so benefit from less severe spatial downsampling, especially for spatially larger tensors of the FPN. In the present disclosure, smaller tensors are upsampled spatially before being inputted to the bottleneck encoder. [000102] In the example arrangements described, the input FPN tensors 501 comprise layers P2 502, P3 503, P4 504, and P5 505. With P5 505 having width and height (w,h), P2-P4 (502, 503, 504) have dimensions (8w,8h), (4w,4h), (2w,2h), respectively. In other words, respective tensors (each of tensor in P2-P5) have resolutions forming an exponential sequence with a doubling in width and height between successive tensors. The layers P2-P5 each have 256 channels. The bottleneck encoder 500 inputs four feature maps P2-P5. The larger feature maps in P2 contain more object details and the smaller feature maps in P3, P4, and P5 contain more abstract features as P5, P4, P3 are produced from P2, for example via a deeper network of Res3 416, Res4 420, and Res5 424. [000103] In the example described above, the inputs P2 to P5 correspond with the hierarchical feature pyramid network outputs (P2 477, P3 475, P4 473 and P5 471) generated by the CNN backbone 400 of Fig. 4. If the CNN backbone is implemented based on Fig. 3A, the tensors 329, 326 and 322, are derived into a set of tensors 115. The set of tensors 115 is processed by a multi-scale feature fusion (MSFF) module 510 in the bottleneck encoder 116. [000104] Referring to Fig. 6, the method 600 may be implemented using apparatus such as a configured FPGA, an ASIC, or an ASSP. Alternatively, as described below, the method 600 may be implemented by the source device 110, as one or more software code modules of the application programs 233, under execution of the processor 205. The software code modules of the application programs 233 implementing the method 600 may be resident, for example, in the hard disk drive 210 and / or the memory 206. The method 600 is repeated for each frame of video data produced by the video source 112. The method 600 may be stored on computer-readable storage medium and / or in the memory 206. The method 600 begins at a perform neural network first portion step 610. [000105] At the step 610 the CNN backbone 114, under execution of the processor 205, performs neural network layers corresponding to the first portion of the neural network. The step 610 effectively implements the CNN backbone 114. For example, the CNN layers as described with reference to Fig. 4 may be performed to produce P2-P5 tensors 501 (115). Control in the processor 205 progresses from the step 610 to a select a set of tensors step 615. 2024202416   12 Apr 2024 [000106] The arrangements described effectively divide the tensors produced by CNN backbone 114 into a set of multiple tensors (also referred to as a plurality of tensors), the tensors in the set having different spatial resolution feature maps to one another. At the step 615 bottleneck encoder 116, under execution of the processor 205, selects multiple tensors adjacent among the tensors 501 as the plurality of tensors. In selecting the set of tensors, the step 615 operates to implement three sub-steps. In executing the step 615, the method 600 firstly implements a select largest spatial tensor step 617. At the step 617, one tensor having a largest spatial resolution is selected, such as tensor P2 502 for example. The largest spatial resolution tensor is kept unchanged before being inputted to the combine tensors 630. Control in the processor 205 continues from sub-step 617 to a select other smaller spatial tensors step 618. Tensors having a smaller spatial resolution are selected at the step 618, such as tensor P3 503, P4 504, and P5 505. In some arrangements, the order or 617 and 618 may be reversed, or 617 and 618 may be executed concurrently. Each chosen tensor from the step 618 input a convolutional layer in step 619 and is upsampled to the spatial size of the largest tensor. Control in the processor 205 progresses from the step 615 to a combine tensors step 630. [000107] The tensors 501 form a hierarchical representation of the frame data 113 that results from application of a FPN to the frame data 113. Use of convolution stages with stride equal to two in the FPN, i.e., at CBL block within the modules 412, 416, 420, and 424, results in the spatial dimensions of tensors among the tensors 501 halving in width and height with each respective (or ‘successive’) tensor, when ordered according to decompositional level, for example, from P2 to P5. A degree of inter-layer correlation exists among the layers P2 to P5 of the tensors 501 despite the layers having different spatial resolution. Exploiting inter-layer correlation permits a channel count reduction relative to a concatenation of tensors across layers, provided the tensors are firstly spatially scaled to the same resolution, e.g., the largest resolution among the tensors to be combined. When tensors of greatly differing spatial resolution are combined by downsampling, unacceptable loss of detail in the higher-resolution tensor can occur due to the higher ratio of the downsampling operation. For example, downsampling P2 502 to the spatial size of P5 505 requires reducing width and height to one eighth of their former values, for an area reduction to one sixty-fourth of the P2 502 area. For tasks dependent on spatial detail, such as instance segmentation, mAP is degraded. Reductions in mAP due to excessive downsampling of larger layers occurs for detection of small objects where the higher resolution layers are relied upon by the network head. 2024202416   12 Apr 2024 [000108] In contrast, in the arrangements described, convolutional modules 506, 507, and 508 (see Fig. 5A) upsample (at the step 619) the smaller tensors selected at step 618. For example at step 619 the tensors P5 505, P4 504, and P3 503 are upsampled to the width and height of the largest tensor selected at step 617, such as the tensor P2 502. The convolutional layers 506, 507, and 508 operate to produce upsampled tensors 510, 511, and 512 from the tensors P5, P4, and P3 respectively. The upsampled abstract features from the P5, P4, and P3 feature maps 510, 511, 512 are combined with the larger feature map P2 502 using a concatenate function 514. The concatenate function 514 produces a tensor 515 by concatenating the tensor P2 502 and the upsampled P5 tensor 510, the upsampled P4 511, and the upsampled P3 512 along the channel count dimension. As a result of concatenation, the channel count of the output 515 of the concatenate function 514 has the channel count equal to the total of channel count of all input tensors, e.g., 1024 channels. [000109] Referring to Fig. 6, at the step 630 the concatenate function 514 of the MSFF module 510 (see Fig. 5A), under execution of the processor 205, combines each tensor of the selected tensors 501, i.e., the largest tensor 502, other smaller tensors 503, 504, and 505, to produce the combined tensor 515. [000110] When receiving the tensors 115 from the backbone 400, the concatenate module 514 produces the concatenation of the tensor 510, 511, 512 and the tensor 502 to produce the tensor 515, of dimensions 8h, 8w, 256. [000111] The tensor 515 is passed to a squeeze and excitation (SE) module 516 in execution of step 630 to produce a tensor 517. The SE module 516 sequentially performs a global pooling, a fully-connected layer with a reduction in channel count, a rectified linear unit activation, a second fully-connected layer restoring the channel count, and a sigmoid activation function to produce a scaling tensor. The tensor 517 has a spatial dimension is the same as the largest spatial tensor selected at the step 617 (for example, tensor P2 502). The SE block 516 is capable of being trained to adaptively alter the weighting of different channels in the tensor passed through, based on the first fully-connected layer output. The first fully-connected layer output reduces each feature map for each channel to a single value, which is then passed through the non-linear activation unit (ReLU) to create a conditional representation of the feature map, suitable for weighting of other channels, with restoration to the full channel count performed by the second fully-connected layer. The SE block 516 is thus capable of extracting non-linear inter-channel correlation in producing the tensor 517 from the tensor 515, to a greater extent 2024202416   12 Apr 2024 than is possible purely with convolutional (linear) layers. As the tensors 515 and 517 contain 1024 (256 x 4) channels in the example described, a result of the concatenation of four FPN layers, the decorrelation achieved by the SE block 516 spans the four FPN layers P2, P3, P4, and P5. The tensor 517 is passed to a convolutional layer 518 in execution of the step 630. The convolutional layer 518 implements one or more convolutional layers to produce a combined tensor 519, with channel count reducing from the channel count of the tensor 517, typically 1024, to F channels, typically 256 channels. In other words, the channel count of the combined tensor 519 produced by the convolutional layer 518 is smaller than the channel count of the tensor 517 input to the convolutional layer 518. The convolutional layer 518 can use a kernel size having at least one size (width or height) more than 1 pixel, such as 3x3, 5x5, or 7x7 to extract more abstract features in the input data. In the described arrangements, kernel size 1x1 is used in the convolutional layer 518 to focus more on reduction the channel count, and learning abstract features is put to later layers. As a result of the step 630, a plurality of tensors (tensors of four FPN layers) are reduced to a single tensor. The single tensor has the same channel count as the input FPN layer tensors and the spatial resolution of the largest of the four FPN layer tensors selected at the sub step 617 in the step 615. This dimensionality decrease is achieved with several network functional blocks and layers (such as 516 and 518) and relies upon training the layers (for example layers of 516 and 518) rather than on-the-fly determination of correlation between tensors. The tensor 519 can be produced from the plurality of tensors 501 using one or more convolutional layers such as the layers 506-508. Alternatively, one or more upsampling layers may be used instead of the layers 506-508. Each of the tensors 515 and 517 can be considered to provide a concatenated tensor produced from the plurality of tensors 501. The tensor 519 is produced using the convolutional layer 518 with a kernel having a size of 1x1 pixels for the concatenated tensor. [000112] Returning to Fig. 6, control in the processor 205 progresses from the step 630 to a single-scale feature compression (SSFC) encode first tensor step 650. [000113] At the step 650, an SSFC encoder 530, shown in Fig. 5A, is implemented under execution of the processor 205. The SSFC encoder 530 operates to further reduce the dimensionality of the combined tensor 519 in both spatial dimension and channel count. In performing SSFC encoding, the step 650 can be considered to implement three sub-steps in the example described. In executing the step 650, the method 600 firstly executes a perform first convolutional layer step 652. At the step 652, a functional block 532, also referred to as a CBP block 532, performs a convolutional function on the tensor 519 with a stride is equal 2 to reduce 2024202416   12 Apr 2024 the spatial dimension by a factor of two in height and width and reduce to half of channel count of the input tensor. As a result, an output tensor 533 has a spatial size of 4h, 4w, and having channel count reduced from 256 (F) to a smaller value C2, such as 128 than the tensor 519. [000114] Fig. 5B shows an example CBP (Convolutional, Batch, PReLu) block 550, as used to implement the functional block 532. Implementations of the CBP block 550 are also used at functional blocks 534 and 536 of Fig. 5A. The CBP block 550 receives as input a tensor 551, for example corresponding to the tensor 519 for the block 532. The tensor 551 is input to a convolutional layer 552, which outputs a tensor 553 having a smaller spatial size and / or reduced channel count. The tensor 553 is passed to a batch normalisation module 554 to produce a tensor 555. The batch normalised tensor 555 has the same dimensionality as the tensor 553. The tensor 555 is passed to an activation layer 556, typically a Parametric Rectified Linear Unit (PreLU) layer, to produce a compressed tensor 558. Examples of other activation functions that can be used in the activation layer 566 include Rectified Linear Unit (ReLU) and hyperbolic tangent function (TanH). The compressed tensor 558 has the same dimensionality as the tensor 555. The output tensor 558 of the CBP block is changed in one or in both dimensions of channel count and spatial size. [000115] In the arrangements described, the SSFC encoder 530 reduces the dimensionality of the input tensor 519 in both spatial resolution and channel count at three different stages, such as three CBP blocks. Each stage contains a convolutional layer having a stride greater than one to reduce the spatial resolution. Two of the stages can have output channels being one half the number of input channels to reduce the number of channels. The first reduction, i.e., performed at the step 652 by implementation of a convolution at the CBP block 532, produces the tensor 533 having a size reduced to half the height, half the width and half the channel count compared to the input tensor, i.e., the tensor 519. After implementing the step 652, the method 600 continues to a perform second convolutional layer step 654. The step 654 implements the block 534. The CBP block 534 implements a second stage of dimensionality reduction. In the second reduction, performed at the step 654, the CBP block 534 contains a convolutional layer with stride parameter equal to 2 is performed on the tensor 533. The block 534 operates to produce a tensor 535. The tensor 535 has a spatial width and height of one-quarter of the tensor 519 and having reduced channel count of C1, typically 64. [000116] The method 600 continues under control of the processor 205 from step 654 to a perform third convolutional layer step 656. The step 656 performs a third reduction in height 2024202416   12 Apr 2024 and width by implementation of the block 536. The block 536 can be implemented as an instance of the CBP block 550. In the third reduction, performed at the step 656, the CBP block 536 containing a convolutional layer with the stride parameter equal to 2 is performed on the tensor 535. The block 536 produces a tensor 557. The tensor 557 has a spatial width and height of one-eighth of the tensor 519 and having reduced channel count of C, typically 64. The step 650 operates to perform at least one convolutional layer on the tensor 519. The steps 652, 654, and 656 in total operate to perform the first block of CBP 532 (convolutional layer, Batchnormalization layer, Activation layer) on the tensor 519 to generate an intermediate tensor 533, perform the second block of CBP 534 using the tensor 533 to generate the tensor 535, and perform the third block of CBP3 536 using the intermediate tensor 535 to generate the tensor 557. In other arrangements, a different number of convolutional layers may be applied. Examples include implementing a single convolutional layer, implementing additional convolutional layers with a stride of one another single convolutional layer with a stride greater than or equal to one. [000117] The SSFC encoder 530 reduces the channel count and spatial resolution gradually by performing the three trainable stages (532, 534 and 536) each containing at least one convolutional layer. Performing downsampling in the SSFC encoder 530 using convolutions with a stride of two results in the encoded tensor 557 having a smaller channel count and smaller spatial resolution (size) than the tensor 519 while retaining the most important features of the tensor. In the described arrangement, the two first CBP blocks (CBP 532, CBP 534) decrease their input tensors in both spatial size and channel count, the last CBP block (CBP 536) decreases the input tensor in spatial size and does not change the channel count. While different number of convolutional layers may be used, use of at least three convolutional layers can provide a particular benefit of improved end-to-end network performance by allowing sufficiently high performance to be achieved through without an overly large encoded feature map being required though dimensionality reduction. The three convolutional layers reduce feature size gradually during training, therefore improve effectiveness of the feature compression training without using a complex architecture and costly time. [000118] The step 650 produces the tensor 557 from the tensor 519 by using at least three blocks, each of the three blocks comprising a convolutional layer. As shown in Fig. 5A, the blocks 532, 534 and 536 are implemented in series, such that the tensor 557 is produced based on at least three convolutional layers being progressively applied to the tensor 519. As described above, two of the blocks 532, 534 and 536 reduce both channel count and spatial size 2024202416   12 Apr 2024 and one block reduces at least one of channel count and spatial size. The block 534 is located after the block 532, and the block 536 is located after the block 534 in the example shown. The order or sequential location of the blocks reducing both or least one of channel count and spatial size can be varied. [000119] Table 1 is a summary of size of tensors that input and output each CBP block in the SSFC encoder 530. The tensor size is in a format of height, width, channel count, where h, w are height and width of the smallest tensor, e.g., P5. CBP1 CBP2 CBP3 Input tensor size 8h, 8w, 256 4h, 4w, 128 2h, 2w, 64 Output tensor size 4h, 44, 128 2h, 2w, 64 h, w, 64 Table 1. Summary of tensor size in SSFC encoder [000120] The compressed tensor 557 provides feature maps of the frame data 113 as obtained using convolutional operations of the MSFF 510 and the SSFC encoder 530 on the tensors 505, 504, 503, and 502 (or in other implementations on 329, 326 and 322). The compressed tensor 557 provide a set of tensors corresponding to the bottleneck encoded tensors 117. On completion of the step 630, control in the processor 205 progresses from the step 650 to a pack compressed tensors step 670 as shown in Fig. 6. [000121] At the step 670 the quantise and pack module 118, under execution of the processor 205, quantises the compressed tensor 557 and packs the quantised tensors into a single monochrome video frame. An example single monochrome video frame 700 is shown in Fig. 7. The frame 700 corresponds to the frame data 119. Channels of the compressed tensor 557 (117) are packed as feature maps of a particular size, such as a feature map 710 in the frame 700. One channel of the compressed tensor 557 corresponds to one feature map indicated by one rectangular area, such as the area 710 or an area 720. The area 720 relates to a feature map or may also be referred to as a picture or a sub-picture. Returning to Fig. 6, control in the processor 205 progresses from the step 670 to an encode frame step 680. [000122] At the step 680 the feature map encoder 120, under execution of the processor 205, encodes the frame 700 to produce the bitstream 121. Operation of the feature map encoder 120 is described with reference to Fig. 8A. The step 680 also operates to encode construct the 2024202416   12 Apr 2024 bitstream to include information additional to the frame data, in accordance with the compression protocol. For example, the step 680 operates to encode NAL unit header(s), a sequence parameter set and an SEI, as described in relation to Fig. 8B. In particular, in the arrangements described, the step 680 encodes additional data including task information and split point information to the bitstream 121 for a picture. [000123] In encoding the packed tensors, the step 680 can be considered to encode one or more tensors produced by the neural network portion 114 into the bitstream. Further, the step 680 can be considered to encode additional data into the bitstream 121, the additional data (8097) comprising (a) a plurality of pieces of task information (8110) each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information (8120) for the portion of the neural network. [000124] In some implementations, the step 680 further encodes weight information into the bitstream, being the weight information 8093. The weight information identifies weight parameters used for the corresponding machine task. For example, 8093 may indicate weight parameters used for one or more of index values 1-4 of Table 2. The weight information may be encoded directly into the bitstream 121. In other implementations, information may be encoded into the bitstream by which the weight parameters may be determined, e.g. a pointer to a memory location for stored weights or a flag indicating a corresponding set of weights. The weigh parameters are generated at training for a particular neural network. [000125] The method 600 terminates on execution of step 680, with the FPN layers of an image frame 312 reduced in dimensionality and compressed into the video bitstream 121. [000126] Fig. 8A is a schematic block diagram showing functional modules of the video encoder 120, also referred to as a feature map encoder. The video encoder 120 encodes the packed frame 119, for example corresponding to the frame 700 in the example of Fig. 7, to produce the bitstream 121. Generally, data passes between functional modules within the video encoder 120 in groups of samples or coefficients, such as divisions of blocks into sub-blocks of a fixed size, or as arrays. The video encoder 120 may be implemented using a general-purpose computer system 200, as shown in Figs. 2A and 2B, where the various functional modules may be implemented by dedicated hardware within the computer system 200, by software executable within the computer system 200 such as one or more software code modules of the software application program 233 resident on the hard disk drive 205 and being controlled in its 2024202416   12 Apr 2024 execution by the processor 205. Alternatively, the video encoder 120 may be implemented by a combination of dedicated hardware and software executable within the computer system 200. The video encoder 120 and the described methods may alternatively be implemented in dedicated hardware, such as one or more integrated circuits performing the functions or sub functions of the described methods. Such dedicated hardware may include graphic processing units (GPUs), digital signal processors (DSPs), application-specific standard products (ASSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or one or more microprocessors and associated memories. In particular, the video encoder 120 comprises modules 810-890 which may each be implemented as one or more software code modules of the software application program 233. [000127] Although the video encoder 120 of Fig. 8A is an example of a versatile video coding (VVC) video encoding pipeline, other video codecs may also be used to perform the processing stages described herein. The frame data 119 may be in any chroma format and bit depth supported by the profile in use, for example 4:0:0, 4:2:0 for the “Main 10” profile of the VVC standard, at eight (8) to ten (10) bits in sample precision. [000128] A block partitioner 810 firstly divides the frame data 119 into CTUs, generally square in shape and configured such that a particular size for the CTUs is used. The maximum enabled size of the CTUs may be 32*32, 64*64, or 128*128 luma samples for example, configured by a ‘spslog2ctusize minus5’ syntax element presents in the ‘sequence parameter set’. The CTU size also provides a maximum CU size, as a CTU with no further splitting will contain one CU. The block partitioner 810 further divides each CTU into one or more CBs according to a luma coding tree and a chroma coding tree. The luma channel may also be referred to as a primary colour channel. Each chroma channel may also be referred to as a secondary colour channel. The CBs have a variety of sizes and may include both square and non-square aspect ratios. However, in the VVC standard, CBs, CUs, PUs, and TUs always have side lengths that are powers of two. Thus, a current CB, represented as 812, is output from the block partitioner 810, progressing in accordance with an iteration over the one or more blocks of the CTU, in accordance with the luma coding tree and the chroma coding tree of the CTU. [000129] The CTUs resulting from the first division of the frame data 119 may be scanned in raster scan order and may be grouped into one or more ‘slices’. A slice may be an ‘intra’ (or ‘I’) slice. An intra slice (I slice) indicates that every CU in the slice is intra predicted. Generally, the first picture in a coded layer video sequence (CLVS) contains only I slices and is referred to 2024202416   12 Apr 2024 as an ‘intra picture’. The CLVS may contain periodic intra pictures, forming ‘random access points’ (i.e., intermediate frames in a video sequence upon which decoding can commence). Alternatively, a slice may be uni- or bi-predicted (‘P’ or ‘B’ slice, respectively), indicating additional availability of uni- and bi-prediction in the slice, respectively. [000130] The video encoder 120 encodes sequences of pictures according to a picture structure. One picture structure is ‘low delay’, in which case pictures using inter-prediction may only reference pictures occurring previously in the sequence. Low delay enables each picture to be output as soon as the picture is decoded, in addition to being stored for possible reference by a subsequent picture. Another picture structure is ‘random access’, whereby the coding order of pictures differs from the display order. Random access allows inter-predicted pictures to reference other pictures that, although decoded, have not yet been output. A degree of picture buffering is needed so the reference pictures in the future in terms of display order are present in the decoded picture buffer, resulting in a latency of multiple frames. [000131] When a chroma format other than 4:0:0 is in use, in an I slice, the coding tree of each CTU may diverge below the 64x64 level into two separate coding trees, one for luma and another for chroma. Use of separate trees allows different block structure to exist between luma and chroma within a luma 64x64 area of a CTU. For example, a large chroma CB may be collocated with numerous smaller luma CBs and vice versa. In a P or B slice, a single coding tree of a CTU defines a block structure common to luma and chroma. The resulting blocks of the single tree may be intra predicted or inter predicted. [000132] In addition to a division of pictures into slices, pictures may also be divided into ‘tiles’. A tile is a sequence of CTUs covering a rectangular region of a picture. CTU scanning occurs in a raster-scan manner within each tile and progresses from one tile to the next. A slice can be either an integer number of tiles, or an integer number of consecutive rows of CTUs within a given tile. [000133] For each CTU, the video encoder 120 operates in two stages. In the first stage (referred to as a ‘search’ stage), the block partitioner 810 tests various potential configurations of a coding tree. Each potential configuration of a coding tree has associated ‘candidate’ CBs. The first stage involves testing various candidate CBs to select CBs providing relatively high compression efficiency with relatively low distortion. The testing generally involves a Lagrangian optimisation whereby a candidate CB is evaluated based on a weighted combination 2024202416   12 Apr 2024 of rate (i.e., coding cost) and distortion (i.e., error with respect to the input frame data 119). ‘Best’ candidate CBs (i.e., the CBs with the lowest evaluated rate / distortion) are selected for subsequent encoding into the bitstream 121. Included in evaluation of candidate CBs is an option to use a CB for a given area or to further split the area according to various splitting options and code each of the smaller resulting areas with further CBs, or split the areas even further. As a consequence, both the coding tree and the CBs themselves are selected in the search stage. [000134] The video encoder 120 produces a prediction block (PB), indicated by an arrow 820, for each CB, for example, CB 812. The PB 820 is a prediction of the contents of the associated CB 812. A subtracter module 822 produces a difference, indicated as 824 (or ‘residual’, referring to the difference being in the spatial domain), between the PB 820 and the CB 812. The difference 824 is a block-size difference between corresponding samples in the PB 820 and the CB 812. The difference 824 is transformed, quantised and represented as a transform block (TB), indicated by an arrow 836. The PB 820 and associated TB 836 are typically chosen from one of many possible candidate CBs, for example, based on evaluated cost or distortion. [000135] A candidate coding block (CB) is a CB resulting from one of the prediction modes available to the video encoder 120 for the associated PB and the resulting residual. When combined with the predicted PB in the video encoder 120, the TB 836 reduces the difference between a decoded CB and the original CB 812 at the expense of additional signalling in a bitstream. [000136] Each candidate coding block (CB), that is prediction block (PB) in combination with a transform block (TB), thus has an associated coding cost (or ‘rate’) and an associated difference (or ‘distortion’). The distortion of the CB is typically estimated as a difference in sample values, such as a sum of absolute differences (SAD), a sum of squared differences (SSD) or a Hadamard transform applied to the differences. The estimate resulting from each candidate PB may be determined by a mode selector 886 using the difference 824 to determine a prediction mode 887. The prediction mode 887 indicates the decision to use a particular prediction mode for the current CB, for example, intra-frame prediction or inter-frame prediction. Estimation of the coding costs associated with each candidate prediction mode and corresponding residual coding may be performed at significantly lower cost than entropy coding of the residual. Accordingly, a number of candidate modes may be evaluated to determine an optimum mode in a rate-distortion sense even in a real-time video encoder. 2024202416   12 Apr 2024 [000137] Determining an optimum mode in terms of rate-distortion is typically achieved using a variation of Lagrangian optimisation. [000138] Lagrangian or similar optimisation processing can be employed to both select an optimal partitioning of a CTU into CBs (by the block partitioner 810) as well as the selection of a best prediction mode from a plurality of possibilities. Through application of a Lagrangian optimisation process of the candidate modes in the mode selector module 886, the intra prediction mode with the lowest cost measurement is selected as the ‘best’ mode. The lowest cost mode includes a selected secondary transform index 888, which is also encoded in the bitstream 121 by an entropy encoder 838. [000139] In the second stage of operation of the video encoder 120 (referred to as a ‘coding’ stage), an iteration over the determined coding tree(s) of each CTU is performed in the video encoder 120. For a CTU using separate trees, for each 64x64 luma region of the CTU, a luma coding tree is firstly encoded followed by a chroma coding tree. Within the luma coding tree, only luma CBs are encoded and within the chroma coding tree only chroma CBs are encoded. For a CTU using a shared tree, a single tree describes the Cus (i.e., the luma CBs and the chroma CBs) according to the common block structure of the shared tree. [000140] The entropy encoder 838 supports bitwise coding of syntax elements using variablelength and fixed-length codewords, and an arithmetic coding mode for syntax elements. Portions of the bitstream such as ‘parameter sets’, for example, sequence parameter set (SPS) and picture parameter set (PPS) use a combination of fixed-length codewords and variablelength codewords. Slices, also referred to as contiguous portions, have a slice header that uses variable length coding followed by slice data, which uses arithmetic coding. The slice header defines parameters specific to the current slice, such as slice-level quantisation parameter offsets. The slice data includes the syntax elements of each CTU in the slice. Use of variable length coding and arithmetic coding requires sequential parsing within each portion of the bitstream. The portions may be delineated with a start code to form ‘network abstraction layer units’ or ‘NAL units’. Arithmetic coding is supported using a context-adaptive binary arithmetic coding process. [000141] Arithmetically coded syntax elements consist of sequences of one or more ‘bins’. Bins, like bits, have a value of ‘0’ or ‘1’. However, bins are not encoded in the bitstream 121 as discrete bits. Bins have an associated predicted (or ‘likely’ or ‘most probable’) value and an 2024202416   12 Apr 2024 associated probability, known as a ‘context’. When the actual bin to be coded matches the predicted value, a ‘most probable symbol’ (MPS) is coded. Coding a most probable symbol is relatively inexpensive in terms of consumed bits in the bitstream 121, including costs that amount to less than one discrete bit. When the actual bin to be coded mismatches the likely value, a ‘least probable symbol’ (LPS) is coded. Coding a least probable symbol has a relatively high cost in terms of consumed bits. The bin coding techniques enable efficient coding of bins where the probability of a ‘0’ versus a ‘1’ is skewed. For a syntax element with two possible values (i.e., a ‘flag’), a single bin is adequate. For syntax elements with many possible values, a sequence of bins is needed. [000142] The presence of later bins in the sequence may be determined based on the value of earlier bins in the sequence. Additionally, each bin may be associated with more than one context. The selection of a particular context may be dependent on earlier bins in the syntax element, the bin values of neighbouring syntax elements (i.e., those from neighbouring blocks) and the like. Each time a context-coded bin is encoded, the context that was selected for that bin (if any) is updated in a manner reflective of the new bin value. As such, the binary arithmetic coding scheme is said to be adaptive. [000143] Also supported by the entropy encoder 838 are bins that lack a context, referred to as “bypass bins”. Bypass bins are coded assuming an equiprobable distribution between a ‘0’ and a ‘1’. Thus, each bin has a coding cost of one bit in the bitstream 121. The absence of a context saves memory and reduces complexity, and thus bypass bins are used where the distribution of values for the particular bin is not skewed. One example of an entropy coder employing context and adaption is known in the art as CABAC (context adaptive binary arithmetic coder) and many variants of this coder have been employed in video coding. [000144] The entropy encoder 838 encodes a quantisation parameter 892 and, if in use for the current CB, the LFNST index 888, using a combination of context-coded and bypass-coded bins. The quantisation parameter 892 is encoded using a ‘delta QP’ generated by a QP controller module 890. The delta QP is signalled at most once in each area known as a ‘quantisation group’. The quantisation parameter 892 is applied to residual coefficients of the luma CB. An adjusted quantisation parameter is applied to the residual coefficients of collocated chroma CBs. The adjusted quantisation parameter may include mapping from the luma quantisation parameter 892 according to a mapping table and a CU-level offset, selected from a list of offsets. The secondary transform index 888 is signalled when the residual 2024202416   12 Apr 2024 associated with the transform block includes significant residual coefficients only in those coefficient positions subject to transforming into primary coefficients by application of a secondary transform. [000145] Residual coefficients of each TB associated with a CB are coded using a residual syntax. The residual syntax is designed to efficiently encode coefficients with low magnitudes, using mainly arithmetically coded bins to indicate significance of coefficients, along with lower-valued magnitudes and reserving bypass bins for higher magnitude residual coefficients. Accordingly, residual blocks comprising very low magnitude values and sparse placement of significant coefficients are efficiently compressed. Moreover, two residual coding schemes are present. A regular residual coding scheme is optimised for TBs with significant coefficients predominantly located in the upper-left corner of the TB, as is seen when a transform is applied. A transform-skip residual coding scheme is available for TBs where a transform is not performed and is able to efficiently encode residual coefficients regardless of their distribution throughout the TB. [000146] A multiplexer module 884 outputs the PB 820 from an intra-frame prediction module 864 according to the determined best intra prediction mode, selected from the tested prediction mode of each candidate CB. The candidate prediction modes need not include every conceivable prediction mode supported by the video encoder 120. Intra prediction falls into three types, first, “DC intra prediction”, which involves populating a PB with a single value representing the average of nearby reconstructed samples; second, “planar intra prediction”, which involves populating a PB with samples according to a plane, with a DC offset and a vertical and horizontal gradient being derived from nearby reconstructed neighbouring samples. The nearby reconstructed samples typically include a row of reconstructed samples above the current PB, extending to the right of the PB to an extent and a column of reconstructed samples to the left of the current PB, extending downwards beyond the PB to an extent; and, third, “angular intra prediction”, which involves populating a PB with reconstructed neighbouring samples filtered and propagated across the PB in a particular direction (or ‘angle’). In VVC, sixty-five (65) angles are supported, with rectangular blocks able to utilise additional angles, not available to square blocks, to produce a total of eighty-seven (87) angles. [000147] A fourth type of intra prediction is available to chroma PBs, whereby the PB is generated from collocated luma reconstructed samples according to a ‘cross-component linear model’ (CCLM) mode. Three different CCLM modes are available, each mode using a 2024202416   12 Apr 2024 different model derived from the neighbouring luma and chroma samples. The derived model is used to generate a block of samples for the chroma PB from the collocated luma samples. Luma blocks may be intra predicted using a matrix multiplication of the reference samples using one matrix selected from a predefined set of matrices. This matrix intra prediction (MIP) achieves gain by using matrices trained on a large set of video data, with the matrices representing relationships between reference samples and a predicted block that are not easily captured in angular, planar, or DC intra prediction modes. [000148] The module 864 may also produce a prediction unit by copying a block from nearby the current frame using an ‘intra block copy’ (IBC) method. The location of the reference block is constrained to an area equivalent to one CTU, divided into 64x64 regions known as VPDUs, with the area covering the processed VPDUs of the current CTU and VPDUs of the previous CTU(s) within each row or CTUs and within each slice or tile up to the area limit corresponding to one 128*128 luma samples, regardless of the configured CTU size for the bitstream. This area is known as an ‘IBC virtual buffer’ and limits the IBC reference area, thus limiting the required storage. The IBC buffer is populated with reconstructed samples 854 (i.e., prior to loop filtering), and so a separate buffer to a frame buffer 872 is needed. When the CTU size is 128*128 the virtual buffer includes samples only from the CTU adjacent and to the left of the current CTU. When the CTU size is 32*32 or 64*64 the virtual buffer includes CTUs from up to the four or sixteen CTUs to the left of the current CTU. Regardless of the CTU size, access to neighbouring CTUs for obtaining samples for IBC reference blocks is constrained by boundaries such as edges of pictures, slices, or tiles. Especially for feature maps of FPN layers having smaller dimensions, use of a CTU size such as 32*32 or 64*64 results in a reference area more aligned to cover a set of previous feature maps. Where feature map placement is ordered based on SAD, SSE or other difference metric, access to similar feature maps for IBC prediction offers coding efficient advantage. [000149] The residual for a predicted block when encoding feature map data is different to the residual seen for natural video. Such natural video is typically captured by an imaging sensor, or screen content, as generally seen in operating system user interfaces and the like. Feature map residuals tend to contain much detail, which is amenable to transform skip coding more than predominantly low-frequency coefficients of various transforms. Experiments show that the feature map residual has enough local similarity to benefit from transform coding. However, the distribution of feature map residual coefficients is not clustered towards the DC (top-left) coefficient of a transform block. In other words, sufficient correlation exists for a 2024202416   12 Apr 2024 transform to show gain when encoding feature map data and this is true also for when intra block copy is used to produce prediction blocks for the feature map data. Accordingly, a Hadamard cost estimate may be used when evaluating residuals resulting from candidate block vectors for intra block copy when encoding feature map data, instead of relying solely on a SAD or SSD cost estimate. SAD or SSD cost estimates tend to select block vectors with residuals more amenable to transform skip coding and may miss block vectors with residuals that would be compactly encoded using transforms. The multiple transform selection (MTS) tool of the VVC standard may be used when encoding feature map data so that, in addition to the DCT-2 transform, combinations of DST-7 and DCT-8 transforms are available horizontally and vertically for residual encoding. [000150] An intra-predicted luma coding block may be partitioned into a set of equal-sized prediction blocks, either vertically or horizontally, which each block having a minimum area of sixteen (16) luma samples. This intra sub-partition (ISP) approach enables separate transform blocks to contribute to prediction block generation from one sub-partition to the next subpartition in the luma coding block, improving compression efficiency. [000151] Where previously reconstructed neighbouring samples are unavailable, for example at the edge of the frame, a default half-tone value of one half the range of the samples is used. For example, for 10-bit video a value of five-hundred and twelve (512) is used. As no previous samples are available for a CB located at the top-left position of a frame, angular and planar intra-prediction modes produce the same output as the DC prediction mode (i.e. a flat plane of samples having the half-tone value as magnitude). [000152] For inter-frame prediction a prediction block 882 is produced using samples from one or two frames preceding the current frame in the coding order frames in the bitstream by a motion compensation module 880 and output as the PB 820 by the multiplexer module 884. Moreover, for inter-frame prediction, a single coding tree is typically used for both the luma channel and the chroma channels. The order of coding frames in the bitstream may differ from the order of the frames when captured or displayed. When one frame is used for prediction, the block is said to be ‘uni-predicted’ and has one associated motion vector. When two frames are used for prediction, the block is said to be ‘bi-predicted’ and has two associated motion vectors. For a P slice, each CU may be intra predicted or uni-predicted. For a B slice, each CU may be intra predicted, uni-predicted, or bi-predicted. 2024202416   12 Apr 2024 [000153] Frames are typically coded using a ‘group of pictures’ (GOP) structure, enabling a temporal hierarchy of frames. Frames may be divided into multiple slices, each of which encodes a portion of the frame. A temporal hierarchy of frames allows a frame to reference a preceding and a subsequent picture in the order of displaying the frames. The images are coded in the order necessary to ensure the dependencies for decoding each frame are met. An affine inter prediction mode is available where instead of using one or two motion vectors to select and filter reference sample blocks for a prediction unit, the prediction unit is divided into multiple smaller blocks and a motion field is produced so each smaller block has a distinct motion vector. The motion field uses the motion vectors of nearby points to the prediction unit as ‘control points’. Affine prediction allows coding of motion different to translation with less need to use deeply split coding trees. A bi-prediction mode available to VVC performs a geometric blend of the two reference blocks along a selected axis, with angle and offset from the centre of the block signalled. This geometric partitioning mode (“GPM”) allows larger coding units to be used along the boundary between two objects, with the geometry of the boundary coded for the coding unit as an angle and centre offset. Motion vector differences, instead of using cartesian (x, y) offset, may be coded as a direction (up / down / left / right) and a distance, with a set of power-of-two distances supported. The motion vector predictor is obtained from a neighbouring block (‘merge mode’) as if no offset is applied. The current block will share the same motion vector as the selected neighbouring block. [000154] The samples are selected according to a motion vector 878 and reference picture index. The motion vector 878 and reference picture index applies to all colour channels and thus inter prediction is described primarily in terms of operation upon PUs rather than PBs. The decomposition of each CTU into one or more inter-predicted blocks is described with a single coding tree. Inter prediction methods may vary in the number of motion parameters and their precision. Motion parameters typically comprise a reference frame index, indicating which reference frame(s) from lists of reference frames are to be used plus a spatial translation for each of the reference frames, but may include more frames, special frames, or complex affine parameters such as scaling and rotation. In addition, a pre-determined motion refinement process may be applied to generate dense motion estimates based on referenced sample blocks. [000155] Having determined and selected the PB 820 and subtracted the PB 820 from the original sample block at the subtractor 822, the residual with lowest coding cost, represented as 824, is obtained and subjected to lossy compression. The lossy compression process comprises the steps of transformation, quantisation and entropy coding. A forward primary 2024202416   12 Apr 2024 transform module 826 applies a forward transform to the difference 824, converting the difference 824 from the spatial domain to the frequency domain, and producing primary transform coefficients represented by an arrow 828. The largest primary transform size in one dimension is either a 32-point DCT-2 or a 64-point DCT-2 transform, configured by a ‘spsmax luma transformsize 64flag’ in the sequence parameter set. If the CB being encoded is larger than the largest supported primary transform size expressed as a block size (e.g. 64^64 or 32x32), the primary transform 826 is applied in a tiled manner to transform all samples of the difference 824. Where a non-square CB is used, tiling is also performed using the largest available transform size in each dimension of the CB. For example, when a maximum transform size of thirty-two (32) is used, a 64x16 CB uses two 32x16 primary transforms arranged in a tiled manner. When a CB is larger in size than the maximum supported transform size, the CB is filled with TBs in a tiled manner. For example, a 128x128 CB with 64-pt transform maximum size is filled with four 64x64 TBs in a 2x2 arrangement. A 64x128 CB with a 32-pt transform maximum size is filled with eight 32x32 TBs in a 2x4 arrangement. [000156] Application of the transform 826 results in multiple TBs for the CB. Where each application of the transform operates on a TB of the difference 824 larger than 32x32, e.g. 64x64, all resulting primary transform coefficients 828 outside of the upper-left 32x32 area of the TB are set to zero (i.e., discarded). The remaining primary transform coefficients 828 are passed to a quantiser module 834. The primary transform coefficients 828 are quantised according to the quantisation parameter 892 associated with the CB to produce primary transform coefficients 832. In addition to the quantisation parameter 892, the quantiser module 834 may also apply a ‘scaling list’ to allow non-uniform quantisation within the TB by further scaling residual coefficients according to their spatial position within the TB. The quantisation parameter 892 may differ for a luma CB versus each chroma CB. The primary transform coefficients 832 are passed to a forward secondary transform module 830 to produce the transform coefficients represented by the arrow 836 by performing either a non-separable secondary transform (NSST) operation or bypassing the secondary transform. The forward primary transform is typically separable, transforming a set of rows and then a set of columns of each TB. The forward primary transform module 826 uses either a type-II discrete cosine transform (DCT-2) in the horizontal and vertical directions, or bypass of the transform horizontally and vertically, or combinations of a type-VII discrete sine transform (DST-7) and a type-VIII discrete cosine transform (DCT-8) in either horizontal or vertical directions for luma 2024202416   12 Apr 2024 TBs not exceeding 16 samples in width and height. Use of combinations of a DST-7 and DCT-8 is referred to as ‘multi transform selection set’ (MTS) in the VVC standard. [000157] The forward secondary transform of the module 830 is generally a non-separable transform, which is only applied for the residual of intra-predicted CUs and may nonetheless also be bypassed. The forward secondary transform operates either on sixteen (16) samples (arranged as the upper-left 4*4 sub-block of the primary transform coefficients 828) or fortyeight (48) samples (arranged as three 4*4 sub-blocks in the upper-left 8*8 coefficients of the primary transform coefficients 828) to produce a set of secondary transform coefficients. The set of secondary transform coefficients may be fewer in number than the set of primary transform coefficients from which they are derived. Due to application of the secondary transform to only a set of coefficients adjacent to each other and including the DC coefficient, the secondary transform is referred to as a ‘low frequency non-separable secondary transform’ (LFNST). Such secondary transforms may be obtained through a training process and due to their non-separable nature and trained origin, exploit additional redundancy in the residual signal not able to be captured by separable transforms such as variants of DCT and DST. Moreover, when the LFNST is applied, all remaining coefficients in the TB are zero, both in the primary transform domain and the secondary transform domain. [000158] The quantisation parameter 892 is constant for a given TB and thus results in a uniform scaling to produce residual coefficients in the primary transform domain for a TB. The quantisation parameter 892 may vary periodically with a signalled ‘delta quantisation parameter’. The delta quantisation parameter (delta QP) is signalled once for CUs contained within a given area, referred to as a ‘quantisation group’. If a CU is larger than the quantisation group size, delta QP is signalled once with one of the TBs of the CU. That is, the delta QP is signalled by the entropy encoder 838 once for the first quantisation group of the CU and not signalled for any subsequent quantisation groups of the CU. A non-uniform scaling is also possible by application of a ‘quantisation matrix’, whereby the scaling factor applied for each residual coefficient is derived from a combination of the quantisation parameter 892 and the corresponding entry in a scaling matrix. The scaling matrix may have a size that is smaller than the size of the TB, and when applied to the TB a nearest neighbour approach is used to provide scaling values for each residual coefficient from a scaling matrix smaller in size than the TB size. The residual coefficients 836 are supplied to the entropy encoder 838 for encoding in the bitstream 121. Typically, the residual coefficients of each TB with at least one significant residual coefficient of the TU are scanned to produce an ordered list of values, according to a 2024202416   12 Apr 2024 scan pattern. The scan pattern generally scans the TB as a sequence of 4*4 ‘sub-blocks’, providing a regular scanning operation at the granularity of 4*4 sets of residual coefficients, with the arrangement of sub-blocks dependent on the size of the TB. The scan within each subblock and the progression from one sub-block to the next typically follow a backward diagonal scan pattern. Additionally, the quantisation parameter 892 is encoded into the bitstream 121 using a delta QP syntax element, and a slice QP for the initial value in a given slice or subpicture and the secondary transform index 888 is encoded in the bitstream 121. [000159] As described above, the video encoder 120 needs access to a frame representation corresponding to the decoded frame representation seen in the video decoder. Thus, the residual coefficients 836 are passed through an inverse secondary transform module 844, operating in accordance with the secondary transform index 888 to produce intermediate inverse transform coefficients, represented by an arrow 842. The intermediate inverse transform coefficients 842 are inverse quantised by a dequantiser module 840 according to the quantisation parameter 892 to produce inverse transform coefficients, represented by an arrow 846. The dequantiser module 840 may also perform an inverse non-uniform scaling of residual coefficients using a scaling list, corresponding to the forward scaling performed in the quantiser module 834. The inverse transform coefficients 846 are passed to an inverse primary transform module 848 to produce residual samples, represented by an arrow 850, of the TU. The inverse primary transform module 848 applies DCT-2 transforms horizontally and vertically, constrained by the maximum available transform size as described with reference to the forward primary transform module 826. The types of inverse transform performed by the inverse secondary transform module 844 correspond with the types of forward transform performed by the forward secondary transform module 830. The types of inverse transform performed by the inverse primary transform module 848 correspond with the types of primary transform performed by the primary transform module 826. A summation module 852 adds the residual samples 850 and the PU 820 to produce reconstructed samples (indicated by the arrow 854) of the CU. [000160] The reconstructed samples 854 are passed to a reference sample cache 856 and an inloop filters module 868. The reference sample cache 856, typically implemented using static RAM on an ASIC to avoid costly off-chip memory access, provides minimal sample storage needed to satisfy the dependencies for generating intra-frame PBs for subsequent CUs in the frame. The minimal dependencies typically include a ‘line buffer’ of samples along the bottom of a row of CTUs, for use by the next row of CTUs and column buffering the extent of which is set by the height of the CTU. The reference sample cache 856 supplies reference samples 2024202416   12 Apr 2024 (represented by an arrow 858) to a reference sample filter 860. The sample filter 860 applies a smoothing operation to produce filtered reference samples (indicated by an arrow 862). The filtered reference samples 862 are used by an intra-frame prediction module 864 to produce an intra-predicted block of samples, represented by an arrow 866. For each candidate intra prediction mode the intra-frame prediction module 864 produces a block of samples, that is 866. The block of samples 866 is generated by the module 864 using techniques such as DC, planar or angular intra prediction. The block of samples 866 may also be produced using a matrix-multiplication approach with neighbouring reference sample as input and a matrix selected from a set of matrices by the video encoder 120, with the selected matrix signalled in the bitstream 121 using an index to identify which matrix of the set of matrices is to be used by the video decoder 144. [000161] The in-loop filters module 868 applies several filtering stages to the reconstructed samples 854. The filtering stages include a ‘deblocking filter’ (DBF) which applies smoothing aligned to the CU boundaries to reduce artefacts resulting from discontinuities. Another filtering stage present in the in-loop filters module 868 is an ‘adaptive loop filter’ (ALF), which applies a Wiener-based adaptive filter to further reduce distortion. A further available filtering stage in the in-loop filters module 868 is a ‘sample adaptive offset’ (SAO) filter. The SAO filter operates by firstly classifying reconstructed samples into one or multiple categories and, according to the allocated category, applying an offset at the sample level. [000162] Filtered samples, represented by an arrow 870, are output from the in-loop filters module 868. The filtered samples 870 are stored in the frame buffer 872. The frame buffer 872 typically has the capacity to store several (e.g., up to sixteen (16)) pictures and thus is stored in the memory 206. The frame buffer 872 is not typically stored using on-chip memory due to the large memory consumption required. As such, access to the frame buffer 872 is costly in terms of memory bandwidth. The frame buffer 872 provides reference frames (represented by an arrow 874) to a motion estimation module 876 and the motion compensation module 880. [000163] The motion estimation module 876 estimates a number of ‘motion vectors’ (indicated as 878), each being a Cartesian spatial offset from the location of the present CB, referencing a block in one of the reference frames in the frame buffer 872. A filtered block of reference samples (represented as 882) is produced for each motion vector. The filtered reference samples 882 form further candidate modes available for potential selection by the mode selector 886. Moreover, for a given CU, the PU 820 may be formed using one reference block 2024202416   12 Apr 2024 (‘uni-predicted’) or may be formed using two reference blocks (‘bi-predicted’). For the selected motion vector, the motion compensation module 880 produces the PB 820 in accordance with a filtering process supportive of sub-pixel accuracy in the motion vectors. As such, the motion estimation module 876 (which operates on many candidate motion vectors) may perform a simplified filtering process compared to that of the motion compensation module 880 (which operates on the selected candidate only) to achieve reduced computational complexity. When the video encoder 120 selects inter prediction for a CU the motion vector 878 is encoded into the bitstream 121. [000164] Although the video encoder 120 of Fig. 8A is described with reference to versatile video coding (VVC), other video coding standards or implementations may also employ the processing stages of modules 810-890. The frame data 119 (and bitstream 121) may also be read from (or written to) memory 206, the hard disk drive 210, a CD-ROM, a Blu-ray diskTM or other computer readable storage medium. Additionally, the frame data 119 (and bitstream 121) may be received from (or transmitted to) an external source, such as a server connected to the communications network 220 or a radio-frequency receiver. The communications network 220 may provide limited bandwidth, necessitating the use of rate control in the video encoder 120 to avoid saturating the network at times when the frame data 119 is difficult to compress. Moreover, the bitstream 121 may be constructed from one or more slices, representing spatial sections (collections of CTUs) of the frame data 119, produced by one or more instances of the video encoder 120, operating in a co-ordinated manner under control of the processor 205. The bitstream 121 may also contain one slice that corresponds to one subpicture to be output as a collection of subpictures forming one picture, each being independently encodable and independently decodable with respect to any of the other slices or subpictures in the picture. [000165] Fig. 8B is a schematic block diagram 8150 showing the bitstream 121 or 143 holding encoded packed feature maps and associated metadata. The bitstream 121 or 143 contains groups of syntax elements each prefaced by a ‘network abstraction layer’ (NAL) unit header. For example, a NAL unit header 8008 precedes a sequence parameter set (SPS) 8010. The SPS 8010 specifies the layout of the picture, e.g., the feature map or picture 720 of Fig. 7, depending on the bottleneck or encoder compressor used at the method 600 and a decompressor to be used in decoding, for example at step 1120 (to be described). The layout of the picture includes the positions and sizes of any subpictures (e.g., 710), using subpicture information 8011. The SPS 8010 also indicates the chroma format, the bit depth, the resolution of the frame data represented by the bitstream 121. A coded subpicture 8024 encoding 2024202416   12 Apr 2024 subpicture 720, includes a slice header 8030 followed by slice data 8040. The slice data 8040 includes a sequence of CTUs, providing the coded representation of the frame data. The coded subpicture 8024 forms a portion of a picture 8014. Coded subpicture 8024 encodes subpicture 720, corresponding to the coefficients. The bitstream 121 includes an SEI message 8013, written by the metadata encoder. The SEI message 8013 encodes metadata needed to convert a decoded frame into a set of tensors suitable for use by the CNN head 150. The SEI message 8013 contains a complexity indication 8091, a decoder topology 8092, weights information 8093, packing information 8094, tensor information 8095, quantisation ranges 8096, and additional data 8097. The additional data 8097 may contain a plurality of pieces of machine task information 8110 and a plurality of pieces of split point information 8120. The plurality of pieces of machine task information 8110 may comprise task information indicating a type of task to be performed. Example tasks include detection, which detects objects in a video or an images, segmentation, which segments objects in a video or images, task tracking, which tracks objects, e.g., pedestrian, in a video and other similar machine vision tasks. The split point information indicates a plurality of split points of the CNN backbone 114 used to generate the tensors 117. [000166] The complexity indication 8091 provides an indication representative of a worst-case complexity for any decoder network topology that could be signalled for the bitstream 121. The decoder topology 8092 comprises information of a tensor decompressor technology which can be used as 148. The packing information 8094 provides packing information, for example if horizontal packing or the like is used. The tensor information 8095 specifies dimensionality of the tensors output by the bottleneck encoder 116. The quantisation ranges 8096 indicate quantisation ranges used in encoding at the source device 110 to enable inverse quantisation to the correct range by the destination device 140. [000167] Table 2 shows an example of the additional data 8097 including a plurality of pieces of task information and a plurality of pieces of split point information. Each piece or set of split point information indicates spatial size and channel count of each of one or more reconstructed tensors to be input to CNN head 150 for a machine task, a previous layer to produce an intermediate tensor in the backbone portion 114, and a subsequent layer to which the one or more reconstructed tensors are inputted in the head portion 150. For example, in relation to Fig. 4, the split point information 8120 for “P-Layers” comprises the tensor sizes (i) (h,w,256) indicating the tensor size of P5 471, (ii) (2h, 2w, 256) indicating the size of P4 473, (iii) (4h, 4w, 256) indicating the size of P3 475, and (iv) (8h, 8w, 256) indicating size of P2 477. The 2024202416   12 Apr 2024 split point information 8120 for “P-Layers” may further comprise information specifying previous layers (470, 472, 474, 476) to produce each of the tensors 115 (P5 471, P4 473, P3 475, and P2 477) respectively. The split point information 7120 for “P-Layers” may further comprise information specifying a subsequent layer (for example a layer 1328 shown in Fig. 13) of a particular implementation of the CNN head 150 to which one or more tensors produced in a decoder are to be input. [000168] Fig. 3E shows an example “YOLOv3” CNN architecture 370. The CNN 370 comprises convolutional, downsampling, summation, concatenation, sampling and detection layers, as indicated by a key table 390. Two different sets of split points are show in in Fig. 3E, each for a different backbone / head configuration. One configuration referred to as “ALT1” comprises split points 371, 372 and 373. The split point 371 is after layer 75 and before layer 376 of the CNN 370. The split point 372 is after layer 90 and before layer 91 of the CNN 370. The split point 373 is after layer 105 and before layer 106 of the CNN 370. A second configuration known as “DN53” comprises splits points 381, 382 and 383. The split point 381 is after layer 36 and before layer 37 of the CNN 370. The split point 382 is after layer 60 and before layer 61 of the CNN 370. The split point 383 is after layer 74 and before layer 75 of the CNN 370. [000169] In one example corresponding to the first set of split points (381, 382, 383) of Fig. 3E, the split point information 8120 for the “DN53” implementation of the CNN 370 comprises the tensor sizes (4h, 4w, 256), (2h, 2w, 512), and (h, w, 1024) The example split point information 8120 for “DN53” further comprises information specifying previous layers (470, 472, 474, 476) to produce the tensors 115 (P5 471, P4 473, P3 475, and P2 477). In this example, the encoder is assumed to encode tensors (P5, P4, P3 and P2) as described later in Fig15. Hence, in this example, the previous layers are the same for each one of the pieces of plurality layers. The example split point information 8120 for “DN53” further comprises a subsequent layer (37, 61, 75) to which one or more tensors produced in a decoder are to be input. [000170] The split point 371, 372 and 373 relate to the “ALT1” implementation of the CNN 370. The split point information 8120 for the “ALT1”example comprises the tensor sizes (h, w, 512), (2h, 2w, 256) and (h, w, 128). In an end-to-end network for single task, the split point information 8120 for “ALT1” may further comprise information specifying previous layers (470, 472, 474, 476) to produce the tensors 115 (. The split point information for “ALT1” may 2024202416   12 Apr 2024 further comprise a subsequent layer (76, 91, 106) to which one or more tensors produced in a decoder are to be input. [000171] In some end-to-end architectures where the backbone 114 and MSFC encoder at 116 are not shared, the previous layers of the split point DN53 may be 320, 324, and 328 (Fig. 3A), the previous layers of the split point ALT1 may be 75, 90, 105 as shown in Fig. 3E. Additional data Index Task information (8110) Split point Split point information (8120) 1 Segmentation P-layers Tensors size: (h,w,256); (2h, 2w, 256);( 4h, 4w, 256); (8h, 8w, 256) Previous layers: 470, 472, 474, 476 Subsequent layers: 1328 2 Detection P-layers Tensors size: (h,w,256); (2h, 2w, 256); (4h, 4w, 256); (8h, 8w, 256) Previous layers: 470, 472, 474, 476 Subsequent layers: 1328 3 Tracking DN53 Tensors size: (4h, 4w, 256); (2h, 2w, 512); (h, w, 1024) Previous layers: 470, 472, 474, 476 Subsequent layers: 37, 61, 75 (Fig. 3E) 4 Tracking ALT1 Tensors size: (h, w, 512); (2h, 2w, 256); (h, w,128) Previous layers: 470, 472, 474, 476 Subsequent layers: 76, 91, 106 (Fig. 3E) Table 2: Additional data in SEI message 2024202416   12 Apr 2024 [000172] An example implementation of the additional data 8097 is described in Appendix A. In Appendix A, multitask flag, number oftask info, taskinfo[i] , number ofsplitpoints, and splitpoint[i] in fcm decoder info provide the additional data 8097. [000173] For example, at step 680, the feature map encoder 120 encodes a multitask flag from the fcm deocder info syntax structure into the bitstream 121. If the value of the multitask flag is equal to 1, the feature map encoder 120 encodes the task information 8110 as the number oftask info, and then encodes one or more taskinfo syntax structures, each specifying task information. Further, if the value of the multitask flag is set to 1, the feature map encoder 120 encodes the split point information 8120 as the number ofsplitpoints, and encodes one or more splitpoint syntax structures each specifying split point information into the bitstream. [000174] Correspondingly, for example, the feature map decoder 144 decodes multitask_flag from the fcm deocder info syntax structure. If the value of the multitask flag is equal to 1, the feature map decoder 144 decodes the number oftask info, and decodes one or more taskinfo syntax structures, each specifying task information. Further, if the value of the multitask flag is equal to 1, the feature map decoder 144 decodes the number ofsplitpoints, and decodes one or more splitpoint syntax structures each specifying split point information. [000175] As shown in Table 2, in some cases each of the pieces of split point information can be associated with at least one of the plurality of pieces of task information. For example, the split point information for index 1 and index 2 is the same and may be relevant to either task. In other cases, each of the pieces of split point information is associated with a corresponding one of the plurality of pieces of task information. For example, the split point for each of index 3 and index 4 is different. [000176] Table 2, and similarly use of multitask flag as described in Appendix A, allows plurality of pieces of task information including multiple tasks that are different from each other to be signalled or encoded. Further, a plurality of pieces of split point information can be signalled or encoded, each pieces of split point information including at least one of (a) a number of one or more tensors to be decoded, (b) a spatial size of each of the tensors to be decoded, (c) a channel count of each of the tensors to be decoded, (d) a previous layer of the 2024202416   12 Apr 2024 portion of the neural network which produced the one or more tensors encoded to the bitstream (for example, one of the previous layers for P-layers), (e) a subsequent layer of the neural network to which at least one of the decode tensors is input (for example, one of the subsequent layers for P-layers). Depending on the association between the task information and the split point information, the subsequent layer is typically a first layer of a second portion of the neural network (the head 150), and the relevant second portion of the neural network is determined using a corresponding one of the plurality of pieces of task information. [000177] An implementation 900 of the video decoder 144, also referred to as a feature map decoder, is shown in Fig. 9. Although the video decoder 144 of Fig. 9 is an example of a versatile video coding (VVC) video decoding pipeline, other video codecs may also be used to perform the processing stages described herein. As shown in Fig. 9, the bitstream 143 is input to the video decoder 144. The bitstream 143 may be read from memory 206, the hard disk drive 210, a CD-ROM, a Blu-ray diskTM or other non-transitory computer readable storage medium. Alternatively, the bitstream 143 may be received from an external source such as a server connected to the communications network 220 or a radio-frequency receiver. The bitstream 143 contains encoded syntax elements representing the captured frame data to be decoded. [000178] The bitstream 143 is input to an entropy decoder module 920. The entropy decoder module 920 extracts syntax elements from the bitstream 143 by decoding sequences of ‘bins’ and passes the values of the syntax elements to other modules in the video decoder 144. The entropy decoder module 920 uses variable-length and fixed length decoding to decode SPS, PPS or slice header an arithmetic decoding engine to decode syntax elements of the slice data as a sequence of one or more bins. Each bin may use one or more ‘contexts’, with a context describing probability levels to be used for coding a ‘one’ and a ‘zero’ value for the bin. Where multiple contexts are available for a given bin, a ‘context modelling’ or ‘context selection’ step is performed to choose one of the available contexts for decoding the bin. The process of decoding bins forms a sequential feedback loop, thus each slice may be decoded in the slice’s entirety by a given entropy decoder 920 instance. A single (or few) high-performing entropy decoder 920 instances may decode all slices or subpictures for a frame or picture from the bitstream 143 multiple lower-performing entropy decoder 920 instances may concurrently decode the slices for a frame from the bitstream 143. 2024202416   12 Apr 2024 [000179] The entropy decoder module 920 applies an arithmetic coding algorithm, for example ‘context adaptive binary arithmetic coding’ (CABAC), to decode syntax elements from the bitstream 143. The decoded syntax elements are used to reconstruct parameters within the video decoder 144. Parameters include residual coefficients (represented by an arrow 924), a quantisation parameter 974, a secondary transform index 970, and mode selection information such as an intra prediction mode (represented by an arrow 958). The mode selection information also includes information such as motion vectors, and the partitioning of each CTU into one or more CBs. Parameters are used to generate PBs, typically in combination with sample data from previously decoded CBs. [000180] The residual coefficients 924 are passed to an inverse secondary transform module 936 where either a secondary transform is applied or no operation is performed (bypass) according to the secondary transform index 970. The inverse secondary transform module 936 produces reconstructed transform coefficients 932, that is primary transform domain coefficients, from secondary transform domain coefficients. The reconstructed transform coefficients 932 are input to a dequantiser module 928. The dequantiser module 928 performs inverse quantisation (or ‘scaling’) on the residual coefficients 932, that is, in the primary transform coefficient domain, to create reconstructed intermediate transform coefficients, represented by an arrow 940, according to the quantisation parameter 974. The dequantiser module 928 may also apply a scaling matrix to provide non-uniform dequantization within the TB, corresponding to operation of the dequantiser module 840. Should use of a nonuniform inverse quantisation matrix be indicated in the bitstream 143, the video decoder 144 reads a quantisation matrix from the bitstream 143 as a sequence of scaling factors and arranges the scaling factors into a matrix. The inverse scaling uses the quantisation matrix in combination with the quantisation parameter to create the reconstructed intermediate transform coefficients 940. [000181] The reconstructed transform coefficients 940 are passed to an inverse primary transform module 944. The module 944 transforms the coefficients 940 from the frequency domain back to the spatial domain. The inverse primary transform module 944 applies inverse DCT-2 transforms horizontally and vertically, constrained by the maximum available transform size as described with reference to the forward primary transform module 826. The result of operation of the module 944 is a block of residual samples, represented by an arrow 948. The block of residual samples 948 is equal in size to the corresponding CB. The residual samples 948 are supplied to a summation module 950. 2024202416   12 Apr 2024 [000182] At the summation module 950 the residual samples 948 are added to a decoded PB (represented as 952) to produce a block of reconstructed samples, represented by an arrow 956. The reconstructed samples 956 are supplied to a reconstructed sample cache 960 and an in-loop filtering module 988. The in-loop filtering module 988 produces reconstructed blocks of frame samples, represented as 992. The frame samples 992 are written to a frame buffer 996. [000183] The reconstructed sample cache 960 operates similarly to the reference sample cache 856 of the video encoder 120. The reconstructed sample cache 960 provides storage for reconstructed samples needed to intra predict subsequent CBs without the memory 206 (e.g., by using the data 232 instead, which is typically on-chip memory). Reference samples, represented by an arrow 964, are obtained from the reconstructed sample cache 960 and supplied to a reference sample filter 968 to produce filtered reference samples indicated by arrow 972. The filtered reference samples 972 are supplied to an intra-frame prediction module 976. The module 976 produces a block of intra-predicted samples, represented by an arrow 980, in accordance with the intra prediction mode parameter 958 signalled in the bitstream 143 and decoded by the entropy decoder 920. The intra prediction module 976 supports the modes of the module 864, including IBC and MIP. The block of samples 980 is generated using modes such as DC, planar or angular intra prediction. [000184] When the prediction mode of a CB is indicated to use intra prediction in the bitstream 143, the intra-predicted samples 980 form the decoded PB 952 via a multiplexor module 984. Intra prediction produces a prediction block (PB) of samples, which is a block in one colour component, derived using ‘neighbouring samples’ in the same colour component. The neighbouring samples are samples adjacent to the current block and by virtue of being preceding in the block decoding order have already been reconstructed. Where luma and chroma blocks are collocated, the luma and chroma blocks may use different intra prediction modes. However, the two chroma CBs share the same intra prediction mode. [000185] When the prediction mode of the CB is indicated to be inter prediction in the bitstream 143, a motion compensation module 934 produces a block of inter-predicted samples, represented as 938. The block of inter-predicted samples 938 are produced using a motion vector, decoded from the bitstream 143 by the entropy decoder 920, and reference frame index to select and filter a block of samples 998 from the frame buffer 996. The block of samples 998 is obtained from a previously decoded frame stored in the frame buffer 996. For bi-prediction, two blocks of samples are produced and blended together to produce samples for the decoded 2024202416   12 Apr 2024 PB 952. The frame buffer 996 is populated with filtered block data 992 from the in-loop filtering module 988. As with the in-loop filtering module 868 of the video encoder 120, the inloop filtering module 988 applies any of the DBF, the ALF and SAO filtering operations. Generally, the motion vector is applied to both the luma and chroma channels, although the filtering processes for sub-sample interpolation in the luma and chroma channel are different. Frames from the frame buffer 996 are output as decoded frames 145. [000186] Not shown in Figs. 8A and 9 is a module for pre-processing video prior to encoding and postprocessing video after decoding to shift sample values such that a more uniform usage of the range of sample values within each chroma channel is achieved. A multi-segment linear model is derived in the video encoder 120 and signalled in the bitstream for use by the video decoder 144 to undo the sample shifting. This linear-model chroma scaling (LMCS) tool provides compression benefit for particular colour spaces and content that have some nonuniformity, especially utilisation of a limited range, in their utilisation of the sample space that may result in higher quality loss from application of quantisation. [000187] Fig. 10A is a schematic block diagram showing a cross-layer bottleneck decoder 1000, an example implementation of the decoder 148, for restoring tensor dimensionality after compression. Fig. 10B shows a structure of a DBP (Deconvolutional, Batch-normalization, PReLU) 1090 and Fig. 10C shows a structure of a reconstruction block 1080 used in the decoder 1000. Fig. 11 shows a method 1100 for decoding a bitstream, reconstructing decorrelated feature maps, and performing a second portion of the CNN. [000188] The bottleneck decoder shown in Fig. 10A can be used when decoded task information 8110 is segmentation or detection and decoded split point 8120 is split point information for “P-layers”. [000189] The method 1100 may be implemented using apparatus such as a configured FPGA, an ASIC, or an ASSP. Alternatively, as described below, the method 1100 may be implemented by the destination device 140, as one or more software code modules of the application programs 233, under execution of the processor 205. The software code modules of the application programs 233 implementing the method 1100 may be resident, for example, in the hard disk drive 210 and / or the memory 206. The method 1100 is repeated for each frame of compressed data in the bitstream 143. The method 1100 may be stored on computer-readable 2024202416   12 Apr 2024 storage medium and / or in the memory 206. The method 1100 begins at a decode bitstream step 1110. [000190] At the step 1110 the video decoder 144, under execution of the processor 205, decodes one frame 145 from the bitstream 143 as described with reference to Fig. 9. The frame 145 contains feature maps packed as described with reference to Fig. 7. [000191] The step 1110 can be implemented as a step of a method 1600, to be described. Control in the processor 205 progresses from the step 1110 to an extract combined tensor step 1120. [000192] At the step 1120 the unpack and inverse quantise module 146 extracts feature maps for the combined tensor, e.g., 557, from the decoded frame 145. The extracted feature maps are inverse quantised from the integer domain to the floating-point domain at step 1120, generating the tensor 147. The inverse quantised feature maps generated by operation of step 1120 are shown as tensor 1019 in Fig. 10A. [000193] The inverse quantised tensor 1019 corresponds to the decoded bottleneck (compressed) tensors 147. Control in the processor 205 progresses from the step 1120 to an SSFC decode combined tensor step 1140, as shown in Fig. 11. [000194] Steps 1110 and 1120 operate to decode a frame of the bitstream to obtain a unit of information for the frame. The tensor 1019 provides a unit of information. The unit of information corresponds to feature maps of the frame encoded in the bitstream. [000195] At the step 1140, an SSFC decoder 1030 is implemented under execution of the processor 205. The SSFC decoder 1030 performs neural network layers to decompress the decoded compressed tensor 1019 to produce a decoded combined tensor 1057. In the arrangements described, a set of three functional blocks are used to generate the tensor 1057. The functional blocks are shown as blocks 1052, 1054, and 1056. The blocks 1052, 1054, and 1056 are each instances of a combination of a Deconvolutional layer, a Batch normalization, and an Activation. In performing SSFC decoding, the step 1140 can be considered to implement three sub-steps. In executing the step 1140, the method 1100 firstly executes a perform a first deconvolutional step 1141, implemented by the block 1052. The block 1052 performs a deconvolutional function, being an inverse function of a convolutional function. Ordinarily, the convolutional function extracts strong features in images, the deconvolutional 2024202416   12 Apr 2024 function restores images with full details. The convolutions performed in the SSFC encoder 530 result in a reduction in spatial resolution, i.e., a loss of information. Accordingly, a deconvolution performed in the SSFC decoder 1030 will not be able to fully (i.e., losslessly) recover the tensor provided to the SSFC encoder 530. Use of convolutions in CBP in the SSFC encoder 530 and deconvolutions in functional blocks in the SSFC decoder 1030 can compensate for the loss of information by virtue of use of trainable implementations for the convolutions and deconvolutions. [000196] Fig. 10B shows structure of a DBP block 1090. The DBP (Deconvolutional, Batch normalization, Activation) block 1090 can be used to implements the block 1052, the block 1054, or the block 1056. In the DBP block 1090, an input tensor 1031 is input to a deconvolutional layer 1032. A deconvolutional layer in DBP block can have a stride more than 1, typically 2, to increase spatial size of the input tensor to twice in height and width. If the stride of the deconvolutional layer 1032 is 2, the output tensor 1033 is four times larger (two times larger in each dimension of height and width) than the input tensor 1031. A deconvolutional layer can be set so that the number of output channels of the deconvolutional layer is higher as the number of input channels of the input tensor. Or a deconvolutional layer can be set so that the number of output channels of the deconvolutional layer is the same as the number of input channels of the input tensor. A batch-normalization layer 1034 is performed to produce tensor 1035 from the tensor 1033. The tensor 1035 has same size as tensor 1033. An activation layer, typically PReLU, 1036 is the last layer in the block 1090. The PReLU layer 1036 operates to produce the output tensor 1038 of the DBP block 1090 from the output 1035 of the batch-normalization layer 1034. The tensor 1038 has the same size as the tensors 1035 and 1033. As a result of a DBP block with a deconvolutional layer, the output tensor 1038 has a larger spatial size, e.g., 2 times, compared to the input tensor 1031 of the DBP block 1090. [000197] In the described example arrangement, the DBP block 1052 receives the tensor 1019 having C = 64 channels and outputs a tensor 1033 having C2=128 channels. At the step 1141, the block DBP 1052 results in an upsampling of the input tensor 1019 from a spatial size of h, w to a spatial size 2h, 2w, producing the tensor 1053. The deconvolution layer in the DBP 1052, e.g., 1032, may be implemented using a ‘transpose convolution’ (transposed matrix multiplication) that operates to produce a result approximating the inverse of the convolution layer 536. Using a transposed matrix multiplication results in an approximation of a deconvolution operation. Regardless of the interpretation of convolutions in the SSFC encoder 530 and deconvolutions in the SSFC decoder 1030 as being related as ‘inverse’ or 2024202416   12 Apr 2024 using transposed matrices, kernel weights in the respective convolutions and deconvolutions are derived from an end-to-end training operation and are not required to form a true inverse or an exact transpose. Accordingly, notwithstanding the use of ‘deconvolution’ operations the SSFC decoder 1030 may operate to decode bitstreams containing tensors produced from another implementation of the source device 110 that does not use a corresponding convolution (by training or by mathematical relationship). The output tensor 1053 of the first DBP block has height and width of double the height, width of the input tensor 1019 as the blocks 1052, 1054 and 1056 each output a tensor having a larger channel count and / or larger spatial size than the tensor input to the respective block. The channel count of the output tensor 1053 is the same as the channel count of the input tensor 1019. On generating the tensor 1053, the method 1100 continues to perform a second DBP block at step 1142. [000198] As shown in Fig. 10A, the blocks 1052, 1054 and 1056 are implemented in series, similarly to the blocks 532, 534 and 536 being implemented in series, such that the tensor 1057 is produced based on at least three deconvolutional layers being progressively applied to the tensor 1019. As described above, two of the blocks 1052, 1054 and 1056 increase both channel count and spatial size and one block increases at least one of channel count and spatial size. The order of the blocks increasing both or least one of channel count and spatial size can be varied. [000199] At the step 1142, the SSFC decoder 1030 performs a second DBP block, the block 1054 to further upsample the input tensor after the first upsampling from the first DBP block 1052. The DBP 1054 results in a deconvolution of the tensor 1053, and produces an overall output tensor 1055 having a larger spatial size than the tensor 1053. The block 1054 may be implemented as the block 1090 using a ‘transpose convolution’ function and results in an inverse (or approximate inverse due to constraint on kernel size) of the convolution layer 533. A function can be used to upsample the spatial size of the tensor 1053, such as a transpose convolutional layer with the parameter stride equal to 2. As a result, the output tensor 1055 of the DBP block 1054 has a spatial size of 4h, 4w and the channel count is C2, typically 128. The output tensor 1055 has height, width, and channel count dimensions equal to four times the height, width, and channel count of the input tensor 1019. On generating the tensor 1055, the method 1100 continues to perform a third DBP block at step 1144. [000200] At the step 1144, the SSFC decoder 1030 performs a third DBP block, the block 1056, to further upsample the input tensor after the second upsampling from the second DBP block 1054. The DBP 1056 results in a block of deconvolutional, batch-normalization and activation 2024202416   12 Apr 2024 of the tensor 1055 to produce an output tensor 1057 having a larger spatial size than the tensor 1055. The block 1056 may be implemented as an instance of the block 1090 using a ‘transpose convolution’ function and results in an inverse (or approximate inverse due to constraint on kernel size) of the convolution layer 532. The block 1056 can function to upsample the spatial size of the tensor 1055, such as using a transpose convolutional layer with the parameter stride equal to 2. As a result, the output tensor 1057 of the DBP block 1056 has a spatial size of 8h, 8w and the channel count is F, typically 256. The output tensor 1057 has height and width dimensions equal to 8 (eight) times the height and width of the input tensor 1019, and channel count equal to four times channel count of the input tensor 1019. [000201] The step 1140 operates to perform at least one deconvolutional layer on the tensor 1019. In the example described, the steps 1141, 1142 and 1144 operate to perform a first block DBP 1052 on the tensor 1019 to generate the intermediate tensor 1053, and perform the second DBP block 1054 using the intermediate tensor 1053 to generate the tensor 1055, then perform the third DBP block 1056 using the second intermediate tensor 1055 to generate the tensor 1057. In the described arrangement, in the SSFC decoder 1030, the first DBP block increases its input tensor in spatial size, the two last DBP blocks increase their input tensors in both spatial size and channel count. [000202] Table 3 is a summary of size of tensors that input and output each DBP block in the SSFC decoder 1030. The spatial size and channel count of the input tensor 1019 are increased gradually after each DBP block. The tensor size is in a format of height, width, channel count, where h and w are the height and width of the smallest tensor, e.g., P5. DBP1 DBP2 DBP3 Input tensor size h, w, 64 2h, 2w, 64 4h, 4w, 128 Output tensor size 2h, 2w, 64 4h, 4w, 128 8h, 8w, 256 Table 3. Summary of tensor size in SSFC decoder [000203] In other arrangements, a different number of DBP blocks may be applied. Examples include implementing a single deconvolutional layer, implementing additional deconvolutional layers with a stride of one or another single deconvolutional layer with a stride greater than or equal to one. As discussed in relation to Fig. 5, use of three blocks of convolutions can be 2024202416   12 Apr 2024 particularly beneficial in terms of end-to-end performance. On completion of the step 1144, control in the processor 205 continues from step 1140 to a reconstruct tensors step 1160. [000204] Returning to Fig. 10A, at the step 1160 a multi-scale feature reconstruction (MSFR) module 1060, under execution of the processor 205, produces a set of tensors, as described in relation to Fig. 10A. At the step 1160 the MSFR module 1060 receives the tensor 1057 generated by the SSFC decoder 1030 at the step 1140. The tensor 1057 is passed to the MSFR module 1060, which generates decoded tensors 1059, 1062, 1064, and 1066. [000205] At the early state of the MSFR 1060, a convolutional layer with kernel size 1x1 pixels 1058 can be used to rearrange all features from the input tensor 1057 of the MSFR 1060 before to be reconstructed by expensive convolutional layers with kernel 3x3 in the later steps of the MSFR. The convolutional layer 1038 does not change spatial size or channel count of the input tensor 1057 and produces the output tensor 1059 having the same size as the tensor 1057. The spatial size and channel count of the tensor 1059 is the same as the channel count for the tensor 1057. The convolutional kernel 1x1 518 produces input to the SSFC encoder 530, the output of the SSFC encoder 530 output the SSFC decoder 1030, and the output of the SSFC decoder 1030 is input to convolutional layer having a kernel size of 1x1 pixels 1058. The two convolutional layers, which have a kernel size of 1x1 pixels, 518 and 1058 work together with the SSFC encoder and decoder to provide a Bottleneck SSFC solution. [000206] The output tensor 1059 having a size of 8h, 8w, 256 is passed along as the output P’2. The tensor P’2 is the largest reconstructed tensor from output tensors of the MSFR, and the tensor P’2 is the reconstructed tensor of the input tensor P2 502 of the MSFF 510. [000207] A cascaded series of reconstruction (“Reconst”) blocks 1080 are performed in the MSFR module 1060 to produce tensors P’3 1062, P’4 1064, and P’5 1066. Each of the tensors P’3 1062, P’4 1064, and P’5 1066 are produced from the tensor 1059 using convolutional layers due to the series of reconstruction blocks. The reconstruction block 1080 is shown in Fig. 10C. The Reconst block 1080 receives a tensor 1081 and performs a convolutional layer 1082 has a parameter stride is equal 2. The convolutional layer 1082 outputs a tensor 1083 having a spatial size double in height and width and the same channel count compared with the input tensor 1081. A downsampling module 1084 operates to perform a downsampling operation on the input tensor 1081, such as an interpolation function, resulting in outputting a tensor 1085 having reduced spatial size, such as half the width and half the height, compared to the input tensor 1081. The two tensors 1083 and 1085 have the same spatial size and channel count. An 2024202416   12 Apr 2024 addition function 1086 in the module 1080 sums the two tensors 1083 and 1085 to produce an output tensor 1087 of the Reconst block 1080. A Reconst block 1061, implemented as the block 1080, receives the tensor 1059 and produces the tensor P’3 1062 having a spatial size smaller than P’2 and the same channel count as the tensor P’2. The tensor P’3 is the reconstructed tensor of the input tensor P3 503. [000208] A reconst block 1063 receives the tensor P’3 1062 and performs a spatial reduction as described with reference to the reconst block 1080, to produce the tensor 1064 P’4 having a smaller size than the tensor P’3. The tensor P’4 is the reconstructed tensor of the tensor P4 504. A reconst block 1065 receives the tensor P’4 and performs a spatial reduction to produce the tensor 1066 P’5 having a smaller size than the tensor P’4. The tensor P’5 is the reconstructed tensor of the tensor P5 505. [000209] The step 1160 can be considered to include two steps. A first step 1162 of the step 1160 performs a convolutional layer kernel of size 1x1 by implementation of the block 1058 to reconstruct a largest tensor. A second step 1164 of the step 1160 reconstructs the tensors P’3 to P’5 by operation of the blocks 1061, 1063 and 1065, effectively reconstructing tensors except the largest tensor. [000210] The tensors 1059, 1062, 1064 and 1066, i.e., the decoded tensors P’2-P’5, corresponding to the tensors P2-P5, form the tensors 149. As shown in Fig. 10A, each respective tensor of the set plurality of tensors (1059, 1062, 1064, 1066) has a resolution (or size) forming an exponential sequence with a doubling in width and height between successive tensors. For example, width and height of 1059 is double width and height of 1062. [000211] The MSFR 1060 operates to deriving successive tensors based on the tensor 1059 that is the output of performing convolutional kernel of 1x1 pixels on the tensor 1057. Each successive tensor (1062, 1064, 1066) is derived from a current tensor (such as 1059, 1062, 1064) by summing a result of performing a convolutional layer (1082) on the current tensor and a result of down-sampling (1084) of the current tensor, such that each successive tensor has smaller spatial size than the current tensor. The derived tensors (1059, 1062, 1064, 1066) form the hierarchical representation of an FPN. The tensor 1059 is used as the current tensor in deriving a first of the successive tensors and each derived successive tensor is used as the current tensor in deriving the next successive tensor. 2024202416   12 Apr 2024 [000212] Control in the processor 205 continues from step 1160 to a perform neural network second portion step 1180. At the step 1180 the CNN head 150, under execution of the processor 205, takes the tensors 149 as input on which to perform the remainder of the neural network implemented by the system 100. The method 1100 terminates on implementing the step 1180, having processed tensors associated with one frame of video data. The method 1100 is re-invoked for each frame of video data encoded in the bitstream 143. [000213] Fig. 12A is a schematic block diagram showing a head portion 1200 of a CNN for object detection, as implemented for the CNN head 150. Depending on the task to be performed in the destination device 140, different networks may be substituted for the CNN head 150. Incoming tensors 149 are separated into the tensor of each layer (i.e., tensors 1210, 1220, and 1234). The tensor 1210 is passed to a CBL module 1212 to produce tensor 1214. The tensor 1214 is passed to a detection module 1216 and an upscaler module 1222. The detection module 1216 operates to detect bounding boxes 1218. The bounding boxes 1218 are in the form of a detection tensor. The bounding boxes 1218 are passed to a non-maximum suppression (NMS) module 1248. The NMS module 1248 selects one of multiple inputs generated by detection modules to produce the detection result 151. To produce bounding boxes addressing co-ordinates in the original video data 113, prior to resizing for the backbone portion of the network 114, scaling by the original video width and height is performed. The upscaler module 1222 produces an upscaled tensor 1224 scaled by original video width and height. The upscaled tensor 1224 is passed to a CBL module 1226. The CBL module 1226 produces tensor 1228 as output. The tensor 1228 is passed to a detection module 1230 and an upscaler module 1236. The detection module 1230 operates in a similar manner to the detection module 1216 and produces a detection tensor 1232. The detection tensor 1232 is supplied to the NMS module 1248. [000214] The upscaler module 1236 operates in the same manner as the module 1222 and outputs an upscaled tensor 1238. The upscaled tensor 1238 is passed to a CBL module 1240. The CBL module 1240 operates in the same manner as the modules 1212 and 1226 to output a tensor 1242 to a detection module 1244. The detection module 1244 operates in a similar manner to the detection modules 1216 and 1230 and produces a detection tensor 1246. The detection tensor 1246 is supplied to the NMS module 1248. 2024202416   12 Apr 2024 [000215] The CBL modules 1212, 1226, and 1240 each contain a concatenation of five CBL module, for example as shown in Fig. 3D. The upscaler modules 1222 and 1236 are each instance of an upscaler module 1260 as shown in Fig. 12B. [000216] The upscaler module 1260 accepts a tensor 1262 and a tensor 1264 as inputs. The tensor 1262 is passed to a CBL module 1266 to produce a tensor 1268. The tensor 1268 is passed to an upsampler 1270 to produce an upsampled tensor 1272, using upsampling methods such as nearest-neighbour interpolation or bilinear interpolation. A concatenation module 1274 produces a tensor 1276 by concatenating the upsampled tensor 1272 with the input tensor 1264. [000217] The detection modules 1216, 1230, and 1244 are instances of a detection module 1280 as shown in Fig. 12C. The detection module 1280 receives a tensor 1282, which is passed to a CBL module 1284 to produce a tensor 1286. The tensor 1286 is passed to a convolution module 1288, which implements a detection kernel. A detection kernel a 1 x 1 kernel applied to produce the output on feature maps at the three layers. The detection kernel is 1 x 1 x (B x (5 + C)), where B is the number of bounding boxes a particular cell can predict, typically three (3), and C is the number of classes, which may be eighty (80), resulting in a kernel size of two-hundred and fifty five (255) detection attributes. The module 1288 outputs tensor 1290. The constant “5” represents four boundary box attributes (box centre x, y and size scale x, y) and one object confidence level (“objectness”). The result of a detection kernel has the same spatial dimensions as the input feature map, but the depth of the output corresponds to the detection attributes. The detection kernel is applied at each layer, typically three layers, resulting in a large number of candidate bounding boxes. A process of non-maximum suppression is applied by the NMS module 1048 to the resulting bounding boxes to discard redundant boxes, such as overlapping predictions at similar scale, resulting in a final set of bounding boxes as output for object detection. [000218] Fig. 13 is a schematic block diagram showing an alternative head portion 1300 of a CNN, which can be implemented as the CNN head 150. The head portion 1300 forms part of an overall network known as ‘Faster RCNN’ and includes a feature network (i.e., backbone portion 400), a region proposal network, and a detection network. Input to the head portion 1300 are the tensors 149. The tensors 149 include P2-P6 layer tensors 1310, 1312, 1314, 1316, and 1318. The P2-P6 tensors 1310, 1312, 1314, 1316, and 1318 are input to a region proposal network (RPN) head module 1320. The RPN head module 1320 performs a convolution on the input tensors, producing an intermediate tensor. The intermediate tensor is 2024202416   12 Apr 2024 fed into two subsequent sibling layers in the module 1320, one for classifications and one for bounding box, or ‘region of interest’ (ROI), regression. The module 1320 generates classification and bounding boxes 1322. The classification and bounding boxes 1322 are passed to an NMS module 1324. The NMS module prunes out redundant bounding boxes by removing overlapping boxes with a lower score to produce pruned bounding boxes 1326. The bounding boxes 1326 are passed to a region of interest (ROI) pooler 1328. The ROI pooler 1328 also receives the tensors P2 1310 (corresponding to 477), P3 1312 (corresponding to 475), P4 1314 (corresponding to 473), P5 1316 (corresponding to 471), and P6 1318 (corresponding to 429) and produces fixed-size feature maps from various input size maps using max pooling operations. In the operations performed by 1328, a subsampling takes the maximum value in each group of input values to produce one output value in the output tensor. [000219] In an arrangement of the CNN backbone 400 and the CNN head 1300, the ‘P6’ layer tensor 429 is omitted from the output tensors 115 and in the CNN head 1300, the P6 input tensor 1318 is produced by performing a ‘Maxpool’ operation with stride equal to two on the P5 tensor 1316. Since the P6 layer can be reconstructed from the P5 layer, there is no need to separately encode and decode the P6 layer as an explicit FPN layer among the first set of tensors or the second set of tensors. [000220] Input to the ROI pooler 1328 are the P2-P5 feature maps 1310, 1312, 1314, and 1316 (corresponding to 1059, 1062, 1064 and 1066 of Fig. 10, respectively), and region of interest proposals 1326. Each proposal (ROI) from 1326 is associated with a portion of the feature maps (1310-1316) to produce a fixed-size map. The fixed-size map is of a size independent of the underlying portion of the feature map 1310-1316. One of the feature maps 1310-1316 is selected such that the resulting cropped map has sufficient detail, for example, according to the following rule: floor(4 + log2(sqrt(boxarea) / 224)), where 224 is the canonical box size. The ROI pooler 1328 thus crops incoming feature maps according to the proposals 1326 producing a tensor 1330. The tensor 1330 is fed into a fully connected (FC) neural network head 1332. The FC head 1332 performs two fully connected layers to produce class score and bounding box predictor delta tensor 1334. The class score is generally an 80-element tensor, each element corresponding to a prediction score for the corresponding object category. The bounding box prediction deltas tensor is an 80x4 = 320 element tensor, containing bounding boxes for the corresponding object categories. Final processing is performed by an output layers module 1336, receiving the tensor 1334 and performing a filtering operation to produce a filtered tensor 1338. Low-scoring (low classification) objects are removed from further 2024202416   12 Apr 2024 consideration. A non-maximum suppression module 1340 receives the filtered tenor 1334 and removes overlapping bounding boxes by removing the overlapped box with a lower classification score, resulting in inference output tensor 151. [000221] In the example implementations described in relation to Figs. 5 and 10, the input tensors (for example P2 502, P3 503, P4 504, and P5 505) have different spatial dimensions but the same number of channels. In an arrangement of the bottleneck encoder 116 and the bottleneck decoder 148 having a three-layer FPN, such as a split YOLOv3 network as described with reference to Figs. 3A-3D and Figs. 12A-12C, input tensors, such as 329, 326, and 322, are different in the number of channels and spatial dimension. [000222] Fig. 14A is a schematic block diagram showing a cross-layer bottleneck decoder 1400, providing another example implementation of the decoder 1000 or the decoder 148, for decoding tensor dimensionality after compression. The block diagram shown in Fig. 14A implements an MSFC network for a “DN53 decoder” (as described in relation to Fig. 3E and Fig. 8B) requested for a task, for example by a user. The associated neural network portion can be used when task information 8110 is used for tracking and split point information selected based on parameters such as a user request or operation. For example, decoded split point information 8120 may indicate “DN53” split points (such as 36, 60, 74). The MSFC DN53 decoder 1400 can be used to produce (generate)t the three tensors 322, 326, and 329. An input tensor 1419 is a decompressed tensor from tensor 557 encoded by an inner codec such as a VVC encoder described in relation to Fig. 8A and decoded by a VVC decoder as described in relation to Fig. 9. A DBP block 1452 is perform on the input tensor 1419 to produce tensor 1453. The block DBP can be implemented with a deconvolutional layer, a batch-normalization, and a PReLU layer such as DBP block 1090 shown in Fig. 10B. The output tensor 1453 has double size in height, width, and channel count compared to the input tensor 1419. A second DBP block 1454 is performed on the intermediate tensor 1453 to produce tensor 1459 being double size in height, width, and channel count compares to the height, width, and channel count of the input tensor 1453. The tensor 1459 has the same size as the largest spatial size of the three tensors 329, 326, and 322. The blocks 1452 and 1454 form an SSFC decoder 1430. [000223] The output tensor 1459 having a size of 4h, 4h, 256 is passed to a reconstruction block MSFR 1460 and used as an output L’2. The tensor L’2 is the largest spatial reconstructed tensor from output tensors of the MSFR 1460. The tensor 1459 is input to a convolutional layer 1461 with stride equal 2 and a higher number of output channel to produce a second output 2024202416   12 Apr 2024 tensor L’1 1462. The tensor 1462 has a half size in height and width, and double size in channel count compared to height, width, and channel count of the tensor 1459. The tensor L'1 1462 has the same size as the tensor 326. The tensor 1459 is input to a convolutional layer 1463 with stride equal 4 and a higher number of output channel to produce a third output tensor L’0 1464. The tensor 1464 has one-quarter size in height and width, and four times in channel count compared to height, width, and channel count of the tensor 1459. The tensor L'0 1464 has the same size as the tensor 329. [000224] The three reconstructed tensors L’2 1459, L’1 1462, and L’0 1464 can be inputted to layers 37, 62, 75 respectively in a YOLOv3 network 370 as shown in Fig. 3E. The A CNN head portion 150 performs the rest layers of the network 370 from layer 37 on the reconstructed tensors and produces a tracking result. [000225] Fig. 14B is a schematic block diagram showing a cross-layer bottleneck decoder 1401, an example alternative implementation of the decoder 148 or 1000, for restoring tensor dimensionality after compression. The block diagram shown in Fig. 14B implements a MSFC network for an “ALT1” decoder (as described in relation to Fig. 3E and Fig. 8B) requested for a task, for example by a user. The associated neural network portion can be used when task information 8110 is decoded and selected based on a user operation for tracking. and split point information 8120 is decoded and a selection based on user operation is split point information for “ALT1” split point. The MSFC ALT1 decoder 1401 can be used to reconstruct the three tensors 371, 372, and 373 output from layers 75, 90, and 105 respectively in a YOLOv3 shown in Fig. 3E. An input tensor 1471 is a decompressed tensor from tensor 557 by an inner codec such as a VVC described in relation to Fig. 8A and a corresponding decoder described in relation to Fig. 9. A DBP block 1472 is performed on the input tensor 1471 to produce tensor 1473. The block DBP can be implemented with a deconvolutional layer, a batch-normalization, and a PReLU layer such as DBP block 1090 shown in Fig. 10B. The output tensor 1473 has double size in height, width, and a higher channel count, e.g., 96, compared to the input tensor 1471. A second DBP block 1474 is performed on the intermediate tensor 1473 to produce tensor 1489 having double size in height, width, and a higher channel count, e.g., 128 compared to the height, width, and channel count of the input tensor 1471. The tensor 1489 has the same size as the largest spatial size of the three tensors 371, 372, and 373. The blocks 1472 and 1474 form an SSFC decoder 1470. 2024202416   12 Apr 2024 [000226] The output tensor 1489 having a size of 4h, 4h, 128 is passed along a reconstruction block MSFR 1480 as an output L’’2 of the split points for ALT1. The tensor L’’2 is the largest spatial reconstructed tensor from output tensors of the MSFR 1480 and has a same size as the tensor 371. The tensor 1489 is input to a convolutional layer 1481 with stride equal 2 and a double number of output channel to produce a second output tensor L’’1 1482. The tensor 1482 has a half size in height and width, and double size in channel count compared to height, width, and channel count of the tensor 1499. The tensor L’'1 1482 has the same size as the tensor 372. The tensor 1489 is input to a convolutional layer 1483 with stride equal 4 and a double number of output channel to produce a third output tensor L’’0 1484. The tensor 1484 having one-quarter size in height and width, and four times in channel count compares to height, width, and channel count of the tensor 1489. The tensor L’’0 1464 has the same size as the tensor 373. [000227] The three reconstructed tensors L’’2 1489, L’’1 1482, and L’’0 1484 can be inputted to layers 76, 91, and 106 respectively in a YOLOv3 network as shown in Fig. 3E. A CNN head portion 150 performs the rest layers from layer 76 on the reconstructed tensors and produces a tracking result from the split point of ALT1. [000228] Fig. 15 is a schematic block diagram showing a cross-layer tensor encoder to decoder system 1500 from bottleneck encoder to bottleneck decoder with multiple choices of bottleneck decoders available. An input 1501 representing the input 115 may contain four tensors P2 1502 (502), P3 1503(503), P4 1504 (504), and P5 1505 (505). The four input tensors are input to a MSFC encoder 1510 to produce a encoded tensor 1533. The MSFC encoder 1510 contains a block MSFF 1520 configured (for example similarly to 510) to fuse all input tensors with different size and produce one output tensor 1531 having the same size as the largest tensor, e.g., P2 at 8h, 8w, 256. A SSFC encoder block 1530 compresses the tensor 1531 to tensor 1533 in both spatial size and channel count. For example, the block 1530 may operate in a similar manner to the block 530. The output tensor 1533 of MSFC encoder 1510 has a size of h, w, 64. [000229] Features from the four tensors from P2-P5 having size of (8h, 8w, 256), (4h, 4w, 256), (2h, 2w, 256), and (h, w, 256) are compressed by MSFC encoder 1510 to produce one tensor 1533 having size of h, w, 64. An inner codec 1534, e.g., VVC encoder 800, is used to continue compress and decompress the tensor 1533 to produce a tensor 1535. The tensor 1535 has a same size as the tensor 1533. The inner code 1534 also produces an SEI message, such as the SEI message 8013. The SEI message contains additional data 15344 (corresponding to the additional information 8097). The additional data 15344 may contain information 2024202416   12 Apr 2024 corresponding to task information 8110 and split point information 8120. For example, the additional data 15344 may contain information described in Table 2 which comprising index 1, index 2, index 3, and index 4. Another arrangement of the additional information is described in Appendix A. The SEI messages can be included in bitstream 121 and received in bitstream 143. The additional data 15344 can have structure and contain information such as the additional data message 8097. In the example described arrangement, as described in Table 2, the index 1 1540 indicates a machine task (task information) is segmentation and split point information is P-Layers, index 2 indicates a machine task is detection and split point information is P-Layers, index 3 shows a machine task is tracking and split point information is DN53, and index 4 indicates a machine task is tracking and split point information is ALT1. Each index in additional data 15344 contains a different combination of task information and set of split point. A corresponding one of MSFC decoders 1540, 1550, 1560 and 1570 is implemented for indices 1, 2, 3 and 4 respectively. Each of the MFSC implementations 1540, 1550, 1560 and 1570 has an associated set of output tensors corresponding to the tensors 149, being tensors 1547, 1557, 1567 and 1577 respectively. [000230] Fig. 16 shows the method 1600 for decoding additional data and a bitstream, reconstructing decorrelated feature maps, and performing a second portion of the CNN. The method 1600 may be implemented by the destination device 140. The method 1600 is another arrangement of the method 1100 when a multiple task model is used. The method 1600 may be implemented using apparatus such as a configured FPGA, an ASIC, or an ASSP. Alternatively, as described below, the method 1600 may be implemented by the destination device 140, as one or more software code modules of the application programs 233, under execution of the processor 205. The software code modules of the application programs 233 implementing the method 1600 may be resident, for example, in the hard disk drive 210 and / or the memory 206. The method 1600 is repeated for each frame of compressed data in the bitstream 143. The method 1600 may be stored on computer-readable storage medium and / or in the memory 206. The method 1600 begins at a decode additional data step 1610. [000231] At step 1610, the feature map decoder 144 decodes the additional data 8097 from the bitstream. The additional data 8097 may contain a plurality of pieces of task information and a plurality of pieces of split point information such as the example shown in Table 2 for 8110 and 8120. In some implementations, the additional data 8097 may comprise multitask flag, number oftask info, taskinfo[i], number ofsplitpoints, splitpoint[i] in fcm decoder info as described in Appendix A for 8110 and 8120. 2024202416   12 Apr 2024 [000232] In implementations using the structure of Appendix A, the feature map decoder 144 decodes multitask flag from the fcm deocder info syntax structure. If the value of the multitask flag is equal to 1, the feature map decoder 144 decodes the number oftask info, and decodes one or more taskinfo syntax structure each specifying task information. Further, if the value of the multitask flag is equal to 1, the feature map decoder 144 decodes the number ofsplitpoints, and decodes one or more splitpoint syntax structures each specifying split point information. [000233] The step 1610 operates to decode the additional information from the bitstream, the additional information (8097) comprising (a) a plurality of pieces of task information, each indicating a machine task capable of being performed on later decoded tensors and (b) a plurality of pieces of split point information. The method 1600 continues under execution of the processor 205 from step 1610 to a selecting step 1620. [000234] In some implementations, the step 1610 further decodes weight information from the bitstream 143, being the weight information 8093. The weight information identifies information of weight parameters used for the corresponding machine task. For example, 8093 may indicate weight parameters used for one or more of index values 1-4 of Table 2. The weight information may be decoded directly from the bitstream 143. In other implementations, information may be decoded into the bitstream by which the weight parameters may be determined, e.g. a pointer to a memory location for stored weights or a flag indicating a corresponding set of weights. [000235] At step 1620, the feature map decoder 144 selects one of the plurality of pieces of task information decoded at step 1610, based on operator input. Further, at step 1620, the feature map decoder 144 selects one of the plurality of pieces of split point information decoded at step 1610, for example based on operator input. In an example using information in Table 2 as the additional data, one of additional data index is selected by or based on a user operation or based on a preset task selection, and task information and split point information corresponding to the selected additional data index are selected at step 1620. If weight information is decoded at step 1610, weight parameters may also be selected for the selected task at step 1624. The step 1620 operates to select one of the pieces of task information and one of the pieces of split point information decoded at step 1610. [000236] The method 1600 continues under execution of the processor 205 from step 1620 to a selecting step 1624. At step 1624, the feature map decoder selects an MSFC decoder based on 2024202416   12 Apr 2024 the selected task information and the selected split point information at step 1620. For example, one of MSFC structures 1000, 1400 and 1401 may be selected, or one of the options 1540, 1550, 1560 and 1570 shown in Fig. 15 may be selected. The method 1600 continues under execution of the processor 205 from step 1624 to a decoding step 1630. [000237] At 1630, the feature map decoder 144 decodes a tensor (1535) from the bitstream. The step 1630 operates in the same manner as the step 1110 of Fig. 11. For a given frame, the steps 1110 and 1630 can be considered to decode initial tensors from the bitstream. [000238] The method 1600 continues under execution of the processor 205 from step 1630 to an MSFC decoder step 1640. At 1640, the selected feature map decoder performs the MSFC decoder selected at the step 1614. At step 1640, the selected MSFC decoder produces one or more decoded tensors from the tensor decoded at step 1630. The step 1640 operates in a similar manner to the implementing each of the extract a combined tensor step 1120, the perform SSFC decoder step 1140, and the reconstructed tensors step 1160. The MSFC decoder network may be implemented by a decoder network such as method 1000 shown in Fig. 10A, method MSFC DN53 decoder 1400 shown in Fig. 14A, or method MSFC ALT1 decoder 1401 shown in Fig. 14B. [000239] If the feature map decoder 144 select task information specifying a segmentation and split point information for P-Layers among the additional data 15344 (8097) based on user operation at the step 1620, the feature map decoder 144 selects the MSFC decoder 1540 for segmentation at the step 1614. The selected MSFC decoder 1540 produces the output 1547 (P’5-P’2 tensors) from the input tensor 1535. The MSFC decoder 1540 can be implemented by the method 1000 shown in Fig. 10A. The output 1547 is input to the corresponding head portion network for the segmentation. [000240] If the feature map decoder 144 selects task information specifying detection and split point information for P-Layers among the additional data 15344 (8097) based on user operation at the step 1620, the feature map decoder 144 selects the MSFC decoder 1550 for detection. The selected MSFC decoder 1550 produces output 1557 (P’5-P’2 tensors) from the input tensor 1535. The MSFC decoder 1550 can be implemented by the method 1000 shown in Fig. 10A. The output 1557 is input to the corresponding head portion network for the detection. [000241] If the feature map decoder 144 selects task information specifying tracking and split point information for DN53 among the additional data 15344 (8097) based on user operation at 2024202416   12 Apr 2024 the step 1620, the feature map decoder 144 selects the MSFC decoder 1560 for tracking at the step 1624. The MSFC decoder 1560 produces output 1567 (three tensors L'0 1464, L'1 1462, L'2 1459) from the input tensor 1535. The MSFC decoder 1560 can be implemented by the method 1400 shown in Fig. 14A. The output 1567 is input to a corresponding head portion network for tracking. [000242] If the feature map decoder 144 selects task information specifying tracking and split point information for ALT1 among the additional data 15344 based on user operation at step 1620, the feature map decoder 144 selects the MSFC decoder 1570 for tracking. The MSFC decoder 1570 produces output 1577 (three tensors L''0 1484, L''1 1482, L''2 1489). The MSFC decoder 1570 can be implemented by the method 1401 shown in Fig. 14B. The output 1577 is input to a corresponding head portion network for tracking. [000243] The step 1640 operates to produce (or generate) decoded tensors from the tensors decoded at step 1630 tensors based on the task information and the split point information selected at 1624. [000244] Returning Fig. 16, the method 1600 continues under execution of the processor 205 from step 1640 to perform neural network second portion step 1660. The reconstructed features output from the step 1640 are input to the neural network head portion 150 step 1660. The step 1660 performs the second portion of the neural network, in the same manner as implemented at step 1180) and outputs a result for each task such as detection, segmentation or tracking. The step 1660 may use weight parameters for the head network 150 selected at step 1624 in some implementations. The method 1600 ends for a given frame after execution of the step 1660. [000245] Regarding training, at a first stage of training, an end-to-end network with the MSFC encoder for sharing 1510 and a MSFC decoder can be used to produce weights for the MSFC encoder for sharing 1510. A dataset for segmentation with high accurate ground truth and high number of samples such as OpenImage dataset can be trained using that end-to-end network. If the first stage uses a network for segmentation, weights for MSFC encoder for sharing 1510 and weights for MSFC decoder for segmentation 1540 can be obtained from this stage. In a second stage of training, an end-to-end network with MSFC encoder 1510 and a MSFC decoder for detection, e.g., 1550 can be used to produce MSFC decoder weights for detection. In a third stage of training, an end-to-end network with MSFC encoder 1510 and a MSFC decoder for tracking, e.g., 1560 or 1570 can be used to produce MSFC decoder weights for tracking at split points DN53 or ALT1. As a result, multiple tasks can be implemented sharing the same MSFC 2024202416   12 Apr 2024 encoder weights during training. Sharing the same MSFC encoder weights can save time and resources for training, and in some implementations take advantage of a good training dataset such as Openimage dataset. In the testing or inference phase, multiple tasks are shared a same input tensor, e.g., 1535. [000246] In the described arrangements, four different tasks as shown in Table 2 can be implemented using one source device enabled for multiple tasks where the encoder part is shared among the four tasks. The information of weights for encoder or decoder and for task indexes, e.g, index1, 2, 3, and 4 shown in Table 2, are encoded in weights information 8093 in SEI message 8013. For example, a first weight parameters used for first machine task (e.g. the segmentation) and a second weight parameters used for second machine task (e.g. tracking) may be encoded in weights info 8093 and may be different to each other. The weights information 8093 can be encoded at the step 680 and decoded at the step 1610 as well as the additional data 8097 for example. The first weight parameters are decoded to be used for the first machine task, the second weight parameters are decoded and to be used for the second machine task and so on. Depending on a task chosen by a user’s operation, a network portion for encoder or decoder and the corresponding weights are activated to perform the required task. INDUSTRIAL APPLICABILITY [000247] The arrangements described are applicable to the computer and data processing industries and particularly for the digital signal processing for the encoding and decoding of signals such as video and image signals, achieving high compression efficiency. [000248] The arrangements describe use MSFC in a manner that includes encoding by additively combining tensors from a hierarchical structure (such as feature pyramid network) of a frame of data into a tensor at the largest spatial resolution of the incoming tensors and then spatially reducing the combined tensor using at least one downsampling operation (implemented typically using a convolution operation of stride greater than one). The resulting downsampled tensor can be compressed when converted to a packed frame using video compression techniques to achieve high compression efficiency while retaining fidelity after decoding for performance of network layers following the FPN. Such arrangements facilitate performance of neural networks segmented into stages, to be executed in different devices with 2024202416   12 Apr 2024 some bandwidth or otherwise cost-constrained communication link in use between the stages. When decoding the combined tensor, prior to separation into FPN layers an upsampling (or ‘deconvolution) stage is implemented to restore the spatial resolution to that of the largest of the tensors forming the FPN. For tasks requiring preservation of greater spatial detail, such as instance segmentation, the resulting mAP from using a combined tensor having the spatial resolution of the largest of the tensors among the FPN layers is higher than that achieved using conventional techniques, typically operating on a combined tensor having the smallest spatial resolution of the tensors of the FPN. [000249] The SSFC encode uses a number of convolutional layers to decrease channel count and spatial size of different degrees. While different numbers of convolutional layers can be used at the SSFC stage, use of at least three convolutional layers can provide a particular benefit of improved end-to-end network performance by allowing sufficiently high performance to be achieved through without an overly large encoded feature map being required though dimensionality reduction. The dimensionality reduction can in turn reduce training complexity and time. [000250] In contrast to previous solutions, the arrangements described implement MSFF such that the resulting fused or combined tensor (i.e., 519) is based on height and width of a largest tensor (for example using P2 502 and upsampled tensors from other smaller tensors such as P3 503, P4 504, and P5 505 in Fig. 5, i.e., 512, 511, 510). The combined tensor (i.e., 519) is transformed using at least one subsequent convolutional stage (for example at 652, 654, and 656) to reduce height and width of the final tensor (such as 557) output from the bottleneck decoder 500. [000251] The decoder stage in the arrangements described similarly uses multiple convolutional layers at SSFC (for example at steps 1141, 1142, and 1144) to generate a tensor (such as 1057) to be rearranged by a convolutional kernel if size 1x1 pixels and then reconstructed using convolutional layers with a stride of 2 (for example at MSFR 1060). Use of the height and width of the largest tensor in MSFF, followed by multiple convolutions in SSFC, (and corresponding arrangements in terms of decoding) were found to decrease computational complexity (i.e., improved compression efficiency) while maintaining mAP. [000252] Further, as described above, the arrangements described allow a single encoder stage to be compatible for use with multiple CNN heads with decreased training resources, for 2024202416   12 Apr 2024 example by encoding information indicating tasks and / or split point information into the bitstream. [000253] The foregoing describes only some embodiments of the present invention, and modifications and / or changes can be made thereto without departing from the scope and spirit of the invention, the embodiments being illustrative and not restrictive. [000254] In the context of this specification, the word “comprising” means “including principally but not necessarily solely” or “having” or “including”, and not “consisting only of”. Variations of the word “comprising”, such as “comprise” and “comprises” have correspondingly varied meanings. 2024202416   12 Apr 2024 APPENDIX A An example SEI message format and associated semantics for representing metadata associated with tensor decompressor structure, tensor packing, and complexity indication in a bitstream are as follows: fcm_decoder_info( payloadSize ) { Descriptor set_level_flag u(1) if( set_level_flag ) fcm_level u(8) update_decoder_flag u(1) if( update_decoder_flag = = 1 ) { no_weights_flag u(1) explicit_signal_decoder_flag u(1) if( explicit_signal_decoder_flag ) { explicit_decoder_compression_idc u(2) explicit_decoder_format_idc u(4) explicit_decoder_format_version_idc ue(v) explicitdecoderpayloadlen ue(v) for( i = 0; i < explicit_decoder_payload_len; i ++ ) decoderpayload[ i ] u(8) register_decoder_idc_flag u(1) if( register_decoder_idc_flag ) decoderidc ue(v) } else { registered_decoder_idc ue(v) } } if( !no_weights_flag ) { update_weights_flag u(1) if( update_weights_flag ) { explicit_signal_weights_flag u(1) if( explicit_signal_weights_flag ) { 2024202416   12 Apr 2024 explicit_weights_idc ue(v) explicit_weights_payload_len ue(v) for(       i       =       0;       i       < explicit_weights_payload_len; i++ ) weightspayload[ i ] u(8) } } } set_region_cnt_flag u(1) if( set_region_cnt_flag ) region_cnt ue(v) set_region_packing_flag u(1) if( set_region_packing_flag ) { for( i = 0; i < region_cnt; i++ ) { top_left_rsctuaddr[ i ] u(v) top_right_rsctuaddr[ i ] u(v) bottom_left_rsctuaddr[ i ] u(v) bottom_right_rsctuaddr[ i ] u(v) horizontal_packing_flag[ i ] u(1) } } set_reduced_tensor_info_flag u(1) if( set_reduced_tensor_info_flag ) for( i = 0; i < region_cnt; i++ ) { region_tensor_cnt[ i ] ue(v) for( j = 0; j < region_tensor_cnt[ i ]; j++ ) { reduced_tensor_batch_size[ i ][ j ] ue(v) reduced_tensor_max_channels[ i ][ j ] ue(v) reducedtensorwidth[ i ] [ j ] ue(v) reduced_tensor_height[ i ] [ j ] ue(v) } } } 2024202416   12 Apr 2024 cropping_enabled_flag u(1) set_restored_tensor_info_flag u(1) if( set_restored_tensor_info_flag ) { restored_tensor_cnt ue(v) for( i = 0; i < restored_tensor_cnt; i++ ) { restored_tensor_batch_size[ i ] ue(v) restored_tensor_channels[ i ] ue(v) restored_tensor_width[ i ] ue(v) restored_tensor_height[ i ] ue(v) explicitcroppingoffset[ i ] u(1) if(explicit_cropping_offset ) { crop_left_offset[ i ] u(1) croptopoffset[ i ] u(1) } else { crop_horizontal_placement[ i ] u(2) or ue(v) crop_vertical_placement[ i ] u(2) or ue(v) } } } update_tensor_channels_flag u(1) if( update_tensor_channels_flag ) for( i = 0; i < region_cnt; i++ ) for( j = 0; j < region_tensor_cnt[ i ] ) { update_tensor_channel_flag[ i ][ j ] u(1) if( update_tensor_channel_flag ) tensor_channel_cnt ue(v) } quantization_range_update_flag u(1) if( quantization_range_update_flag ) { qr_mantissa_len ue(v) for( i = 0; i < region_cnt; i++ ) for( j = 0; j < region_tensor_cnt[ i ] ) { qr_min_exp[ i ][ j ] ue(v) 2024202416   12 Apr 2024 qrminexpsign[ i ][ j ] u(1) qr_min_mantissa[ i ][ j ] u(b) qr_min_mantissa_sign[ i ][ j ] u(1) qrmaxexp[ i ][ j ] ue(v) qr_max_exp_sign[ i ][ j ] u(1) qr_max_mantissa[ i ][ j ] u(b) qr_max_mantissa_sign[ i ][ j ] u(1) } output_datatype_update_flag u(1) if(output_datatype_update_flag) { output_datatype_idc ue(v) if( output_datatype_idc == 0 ) { output_datatype_exponent_len ue(v) output_datatype_mantissa_len ue(v) output_datatype_implicit_mantissa_flag u(1) if( output_implicit_mantissa_flag ) { output_data_implicit_mantissa_value u(b) } } output_scaling_enable_flag u(1) if( output_scaling_enable_flag ) { for( i = 0; i < restored_tensor_cnt; i++ ) { qrsecondminexp[ i ] ue(v) qr_second_min_exp_sign[ i ] u(1) qr_second_min_mantissa[ i ] u(b) qr_second_min_mantissa_sign[ i ] u(1) qr_second_max_exp[ i ] ue(v) qr_second_max_exp_sign[ i ] u(1) qr_second_max_mantissa[ i ] u(b) qr_second_max_mantissa_sign[ i ] u(1) } } multitask_flag ue(1) 2024202416   12 Apr 2024 if (multitask_flag) { number_of_task_info ue(v) for (i=0, i < number_of_task_info, i++){ task_information[i] ue(v) } number_of_split_points ue(v) for (i=0, I < number_of_split_points, i++) { split_point[i] ue(v) } } } Where u(n) refers to a fixed-length codeword n bits in length and ue(v) refers to an unsigned exponential Golomb variable-length codeword. FCM decoder info semantics: setlevelflag equal to one indicates that the tensor decompression complexity indication is to be signalled in this instance of the FCM decoder info SEI message. fcm level signals the complexity indication for any tensor decompressors to be performed in the decoder. The complexity indication provides a worst-case limit on the complexity of any instantiated tensor decompressor. It is a requirement of bitstream conformance that the tensor decompression complexity indication is signalled prior to use of the FCM decoder, e.g., signalled with the first frame of packed tensor data in the bitstream. The following table shows permitted maximum values for complexity aspects for given fcvcmlevel values: fcm_level MAC count Weight count 0 <5M <1M 1 <15M <5M 2 <50M <10M 3-254 (reserved for future use) (reserved for future use) 255 2024202416   12 Apr 2024 update decoder flag equal to one indicates that the FCM decoder is to be updated, effective from this instance of the FCMM decoder info SEI message onwards. noweightsflag equal to one indicates that the FCM decoder does not include any trained elements (e.g., convolutions) and therefore does not require any weights. explicitsignal decoder flag equal to one indicates that the FCM decoder architecture is signalled explicitly in this instance of the FCM decoder info SEI message. When equal to zero, this instance of the FCM decoder info SEI message instead references a previously signalled FCM decoder architecture or references an FCM decoder architecture obtained by external means, e.g., a predetermined architecture or an architecture available from a publicly accessible registry. explicitdecoder compression idc specifies the compression technique (if any) applied to the payload containing the representation of the FCM decoder architecture, in accordance with the following table: explicit_decoder_compression_idc Compression method 0 None 1 DEFLATE 2 LZMA 3 Reserved for future use explicitdecoder format idc specifies the format in which the FCM decoder architecture is encoded, with the following formats supported: explicit_decoder_format_idc Decoder representation format 0 ONNX 1 NNEX 2 Pytorch 2024202416   12 Apr 2024 3 Variable-length scheme 4-15 Reserved for future use explicitdecoder formatversion idc specifies the version of the format in which the FCVCM decoder architecture is encoded. For each supported format, a separate enumeration of explicitdecoder formatversion idc values to versions Of the format is specified. explicitdecoder payload len specifies the length of the payload containing the FCM decoder representation in bytes, after application of Compression (if applicable). decoder payload[ i ] specifies the ith byte of the FCM decoder representation. register decoder idc flag equal to one indicates that the FCM decoder representation signalled in this instance of the FCM decoder info SEI message is to be registered (retained) in the decoder for potential future reference. decoder idc specifies an index value for addressing the FCM decoder representation in a registry of retained FCM decoder architectures. explicitsignalweights flag equal to one indicates that weights associated with the signalled FCM decoder representation are included in this instance of the FCM decoder info SEI message. explicitweightsidc specifies an index for the weights signalled in this instance of the FCM decoder info SEI message. explicitweightspayload len specifies the length of the weights payload in the FCM decoder info SEI message. weightspayload[ i ] specifies the ith byte of the weights payload in the FcM decoder info SEI Message. register weightsidc flag equal to one specifies that the weights signalled in this instance of the FCVCM decoder info SEI message are stored in the FCM decoder for potential future reference. registered decoder idc specifies an index to address an FCM decoder representation that is either known to the decoder by external means or was registered with the FCM decoder in an earlier instance of the FCM decoder info SEI Message. set region cntflag equal to one indicates that this instance of the FCM decoder info SEI message signals a count of regions into which the current and subsequent pictures are to be divided. 2024202416   12 Apr 2024 region cnt indicates a count of regions into which the current and subsequent pictures are to be divided. Each region is rectangular in shape and aligned to CTU boundaries. Each region is populated with feature Maps from one or more Tensors. set region packingflag equal to one indicates that this instance of the FCM decoder info SEI message specifies a division of the current picture into one or more rectangular regions. This division remains in effect until the next instance of an FCM decoder info SEI message with setregionpackingflag equal to one. topleft rsctuaddr[ i ] specifies the address in raster-scan order of the CTU in the top-left position in the ith region. toprightrsctuaddr[ i ] specifies the address in raster-scan order of the CTU in the top-right position in the ith region. bottom left rsctuaddr[ i ] specifies the address in raster-scan order of the CTU in the bottomleft position in the ith region. bottom rightrsctuaddr[ i ] specifies the address in raster-scan order of the CTU in the bottomright position in the ith region. horizontalpackingflag[ i ] equal to one specifies when the packing or unpacking progresses from one feature maps of one tensor to feature maps of the next tensor within the ith region, packing will continue long in a left-to-right manner. When equal to zero, upon progressing from feature maps of one tensor to feature maps of the next tensor, packing of feature maps advances to the leftmost position in the current region and below the previously packed feature maps within the current region. The value one may be used where multiple tensors, each containing few (e.g., one) feature maps are to be packed, requiring a region generally larger in width than in height and generally smaller frame area for the regions. croppingenabledflag equal to one specifies that the FCM decoder may crop the decoded tensors from the feature restoration module 1250 to match the required dimensions of the restored-domain tensors according to the crop* syntax elements (the ‘cropping parameters’). When croppingenabled flag is equal to zero, it is a requirement of bitstream conformance that the tensors resulting from the feature restoration module 1250 match the required dimensions of the restored-domain tensors. set reduced tensor info flag equal to one specifies that the number of tensors in the defined regions and dimensions of the reduced-domain tensors is signalled in this instance of the FCM decoder info SEI message. 2024202416   12 Apr 2024 region tensor cnt[ i ] specifies the number of reduced-domain tensors to be packed in the ith region. reduced tensor batch size[ i ][ j ] specifies the batch size of the jth tensor in the reduced domain being packed into the ith region. reduced tensor maxchannels[ i ] [ j ] specifies the maximum number of feature maps (i.e., channels) of the jth tensor in the reduced domain being packed in the ith region. reduced tensor width[ i ] [ j ] specifies the width of feature maps of the jth tensor in the reduced domain being packed in the ith region. reduced tensor height[ i ][ j ] specifies the height of feature maps of the jth tensor in the reduced domain being packed in the ith region. setrestored tensor infoflag equal to one specifies that the number of tensors output from the FCM decoder, i.e., tensors in the restored domain, is specified in this instance of the FCM decoder info SEI message. restored tensor cnt specifies the number of restored-domain tensors output from the FCM decoder. restored tensor batch size[ i ] specifies the batch size in the ith restored-domain tensor output from the FCM decoder. restored tensor channels[ i ] specifies the number of channels in the ith restored-domain tensor output from the FCM decoder. restored tensor width[ i ] specifies the width of the ith restored-domain tensor output from the FCM decoder. restored tensor height[ i ] specifies the height of the ith restored-domain tensor output from the FCM decoder. explicitcroppingoffset equal to one specifies an explicit cropping offset applied to restored tensors 1252 when cropping to produce output tensors 149. An explicit offset indicates an element offset corresponding to the top-left element in each feature map of the restored tensors 1252. A feature map of width and height restoredtensor width and restoredtensor height is extracted forming output tensors 149. crop leftoffset[ i ] indicates the column of the top-left element of each feature map to be extract from the ith output of 1252 of the feature restoration module 1250. crop topoffset[ i ] indicates the row of the top-left element of each feature map to be extract 2024202416   12 Apr 2024 from the ith output of 1252 of the feature restoration module 1250. crop horizontalplacement[ i ] indicates that the output feature map, of width indicated by restoredtensorwidth[ i ] is obtained from (0) the leftmost region, (1) the centre region, or (2) the rightmost region of the ith tensor of tensors 1252, whose size is determined by operation of the feature restoration module 1250. crop verticalplacement[ i ] indicates that the output feature map, of height indicated by restoredtensorheight[ i ] is obtained from (0) the topmost region, (1) the centre region, or (2) the lowermost region of the ith tensor of tensors 1252, whose size is determined by operation of the feature restoration module 1250. update tensor channelsflag equal to one indicates that the flags to update packed number of feature maps for each tensor in each region are to be signalled in this instance of the FCM decoder info SEI message. update tensor channelflag[ i ][ j ] equal to one indicates that the packed number of feature maps for the jth tensor in the ith region is to be signalled in this instance of the FCM decoder info SEI message. tensor channel cnt[ i ][ j ] specifies the packed number of feature maps (i.e., channels) for the jth tensor of the ith region. When tensor channelcnt[ i ][ j] is not signalled and tensor maxchannels[ i ][ j] is signalled, the value is inferred to be equal to the corresponding tensormaxchannels[ i ][ j ]. When tensorchannelcnt[ i ][ j ] is not signalled or inferred in the current instance of the FCM decoder info SEI message, the value remains in effect from the previous instance of the FCM decoder SEI message (if available), otherwise the value is inferred as 0. qr mantissalen specifies the number of bits to be used to encode the mantissa portion of the reduced-domain quantisation range. qr min exp[ i ][ j ] specifies the exponent portion of the lower bound of the reduced-domain quantisation range for the jth tensor in the reduced domain being packed in the ith region. qr min exp sign[ i ][ j ] specifies the sign of the exponent portion of the lower bound of the reduced-domain quantisation range for jth tensor in the reduced domain being packed in the ith region. qr min mantissa[ i ][ j ] specifies the fraction portion of the lower bound of the reduced-domain quantisation range for the jth tensor in the reduced domain being packed in the ith region, with a bit width as specified by qrmantissalen. 2024202416   12 Apr 2024 qr min mantissasign[ i ][ j ] specifies the sign of the fraction portion of the lower bound of the reduced-domain quantisation range for the jth tensor in the reduced domain being packed in the ith region. qr max exp[ i ][ j ] specifies the exponent portion of the upper bound of the reduced-domain quantisation range for the jth tensor in the reduced domain being packed in the ith region. qr max exp sign[ i ][ j ] specifies the sign of the exponent portion of the upper bound of the reduced-domain quantisation range for jth tensor in the reduced domain being packed in the ith region. qr max mantissa[ i ][ j ] specifies the fraction portion of the upper bound of the reduced-domain quantisation range for the jth tensor in the reduced domain being packed in the ith region, with a bit width as specified by qrmantissa_len. qr max mantissasign[ i ][ j ] specifies the sign of the fraction portion of the upper bound of the reduced-domain quantisation range for the jth tensor in the reduced domain being packed in the ith region.outputdatatype update flag equal to one specifies that this instance of the FCM decoder SEI message updates the datatype of the FCM decoder output tensors and / or their range. outputdatatype idc equal to zero specifies a custom data format for the FCM decoder output and other values indicating floating-point or integer data formats. outputdatatype exponentlen specifies the length of the exponent for a custom output data format, with a value of zero indicating an integer rather than floating-point output format. outputdatatype mantissalen specifies the length of the mantissa for a custom output data format when the exponent length is nonzero or the number of bits for a custom output data format when the exponent length is equal to zero. output datatype implicitmantissaflag equal to one specifies that tensors output from the FCM decoder all use a mantissa value rather than using a mantissa signalled on a per-element basis for each output tensor. outputdataimplicit mantissavalue when present signals the implicit mantissa used for all elements of all output tensors from the FCM decoder. outputscalingenableflag equal to one indicates that this instance of the FCM decoder SEI message updates the quantisation min and max (or lower and upper bound) for the quantisation stage performed after the feature restoration stage. qr second mantissalen specifies the number of bits to be used to encode the mantissa portion 2024202416   12 Apr 2024 of the quantisation range. qr second min exp[ i ] specifies the exponent portion of the lower bound of the output quantisation range for the ith tensor in the restored domain. qr second min exp sign[ i ] specifies the sign of the exponent portion of the lower bound of the output quantisation range for ith tensor in the restored domain. qr second min mantissasign[ i ] specifies the sign of the fraction portion of the lower bound of the output quantisation range for the ith tensor in the restored domain. qr second max exp[ i ] specifies the exponent portion of the upper bound of the output quantisation range for the ith tensor in the restored domain. qr second max exp sign[ i ] specifies the sign of the exponent portion of the upper bound of the output quantisation range for ith tensor in the restored domain. qr second max mantissa[ i ] specifies the fraction portion of the upper bound of the output quantisation range for the ith tensor in the restored domain, with a bit width as specified by qrsecondmantissa len. qr second max mantissasign[ i ] specifies the sign of the fraction portion of the upper bound of the output quantisation range for the ith tensor in the restored domain. multitask flag equal to one indicates that this instance of the FCM decoder SEI message updates with multitask network. number oftask info specifies the total number of taskinfo[ i ] syntax element are present in the fcmdecoderinfo syntax structure. The value of number oftaskinfo may be equal 3 if machine tasks used are detection, segmentation, and tracking. taskinfo[i] specifies specific i-th task information. The value of taskinfo [i] is from 0 to 2 in this description. The following table shows task information for each value of the task_info syntax element. task_info Task information 0 Segmentation 1 Detection 2024202416   12 Apr 2024 2 Tracking number_of_split_points specifies the total number of split_point[ i ] syntax elements present in the fcm_decoder_info syntax sturcture. The value of the number_of_split_points may be equal 3 if split point information for P-Layers, split point information for DN53, and split point information for ALT1 are used. split_point[i] specifies specific i-th split point information. The following table shows split pint information for each value of the split_point syntax element. In the following table, split_point equal to 0 specifies split pint information for P-layers, split_point equal to 1 specifies split point information for DN53, and split_point equal to 2 specifies split point information for ALT1. split_point Split point information 0 Tensors size: (h,w,256); (2h, 2w, 256);( 4h, 4w, 256); (8h, 8w, 256) Previous layers: 470, 472, 474, 476 Subsequent layers: 1328 1 Tensors size: (4h, 4w, 256); (2h, 2w, 512); (h, w, 1024) Previous layers: 470, 472, 474, 476 Subsequent layers: 37, 61, 75 (Fig. 3E) 2 Tensors size: (h, w, 512); (2h, 2w, 256); (h, w,128) Previous layers: 470, 472, 474, 476 Subsequent layers: 76, 91, 106 (Fig. 3E)

Claims

1. A method of encoding tensors, the method comprising:encoding one or more tensors produced by a portion of a neural network into a bitstream; andencoding additional information into the bitstream, the additional information comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network.

2. The method according to claim 1, wherein each of the plurality of pieces of split point information is associated with at least one of the plurality of pieces of task information.

3. The method according to claim 1, wherein each of the plurality of pieces of split pointinformation is associated with a corresponding one of the plurality of pieces of task information.

4. The method according to claim 1, whereinthe plurality of pieces of task information comprises (a) first task information indicating a first machine task and (b) the second task information indicating a second machine task different from the first machine task, andthe plurality of pieces of split point information comprises:first split point information comprising at least one of (a) a number of first decoded tensors for the first machine task, (b) a spatial size of each of the one or more first tensors, (c) a channel count of each of the one or more first tensors, and (d) a previous layer of the portion of the neural network which produced the one or more tensors encoded to the bitstream, (e) a subsequent layer of the neural network to which at least one of the one or more first tensors is input, andsecond split point information comprising at least one of (a) tensor counts of one or more second decoded tensors for the second machine task, (b) a spatial size of each of the one or more second tensors, (c) a channel count of each of the one or more second tensors, and (d) a previous layer of the portion of the neural network which produced the one or more tensors encoded to the bitstream, (e) a subsequent layer of the neural network to which at least one of the one or more second produced tensors is input.

5. The method according to claim 4, wherein the subsequent layer is a first layer of a second2024202416   12 Apr 2024portion of the neural network, and the second portion of the neural network is determined using a corresponding one of the plurality of pieces of task information.

6. The method according to claim 4, further comprising encoding weight information intothe bitstream, the weight information identifying first weight parameters used for the first machine task and second weight parameters used for second machine task.

7. A method of decoding one or more tensors from a bitstream, the method comprising:decoding one or more initial tensors from the bitstream;decoding additional information from the bitstream, the additional data comprising (a) a plurality of pieces of task information each of which indicates machine task which is capable of being performed for the one or more decoded tensors and (b) a plurality of pieces of split point information;selecting one of the plurality of pieces of task information and one of the plurality of pieces of split point information; andproducing the decoded one or more tensors from the initial one or more tensors based on the selected task information and the selected split point information.

8. The method according to claim 7, whereinthe plurality of pieces of task information comprises (a) first task information indicating a first machine task and (b) the second task information indicating a second machine task different from the first machine task, andwherein the plurality of pieces of split point information comprises:first split point information comprising at least one of (a) tensor counts of one or more first decoded tensors for the first machine task, (b) a spatial size of each of the one or more first tensors, (c) a channel count of each of the one or more first tensors, and (d) a previous layer of a portion of a neural network which generated the initial tensors encoded in the bitstream, (e) a subsequent layer of the neural network to which at least one of the one or more first tensors is input,second split point information comprising at least one of (a) tensor counts of one or more second decoded tensors for the second machine task, (b) a spatial size of each of the one or more second tensors, (c) a channel count of each of the one or more second tensors, and (d) a previous layer of a portion of a neural network which generated the initial tensors encoded in the bitstream, (e) a subsequent layer of the neural network to which at least one of the one or more second2024202416   12 Apr 2024tensors is input.

9. The method according to claim 8, further comprising decoding weight information from the bitstream, the weight information indicating first weight parameters used for the first machine task and second weight parameters used for second machine task.

10. An encoder for encoding tensors, the encoder configured to:encode a one or more tensors produced by a portion of a neural network into a bitstream; andencode additional information into the bitstream, the additional information comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network. 11.A computer-implemented medium non-transitory computer-readable storage medium which stores a program for executing a method of encoding tensors, the method comprising:encoding one or more tensors produced by a portion of a neural network into a bitstream; andencoding additional information into the bitstream, the additional information comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network.

12. A system comprising:a memory; anda processor, wherein the processor is configured to execute code stored on the memory for implementing a method of encoding tensors, the method comprising:encoding one or more tensors produced by a portion of a neural network into a bitstream; andencoding additional information into the bitstream, the additional information comprising (a) a plurality of pieces of task information each of which indicates machine task process which is capable of being performed for the one or more tensors upon being decoded from the bitstream and (b) a plurality of pieces of split point information for the portion of the neural network.2024202416   12 Apr 202413. A decoder for decoding one or more tensors from a bitstream, the decoder configured to:decode one or more initial tensors from the bitstream;decode additional information from the bitstream, the additional data comprising (a) a plurality of pieces of task information each of which indicates machine task which is capable of being performed for the one or more decoded tensors and (b) a plurality of pieces of split point information;select one of the plurality of pieces of task information and one of the plurality of pieces of split point information; andproduce the decoded one or more tensors from the initial one or more tensors based on the selected task information and the selected split point information.

14. A computer-implemented medium non-transitory computer-readable storage medium which stores a program for executing a method of decoding one or more tensors from a bitstream, the method comprising:decoding one or more initial tensors from the bitstream;decoding additional information from the bitstream, the additional data comprising (a) a plurality of pieces of task information each of which indicates machine task which is capable of being performed for the one or more decoded tensors and (b) a plurality of pieces of split point information;selecting one of the plurality of pieces of task information and one of the plurality of pieces of split point information; andproducing the decoded one or more tensors from the initial one or more tensors based on the selected task information and the selected split point information.

15. A system comprising:a memory; anda processor, wherein the processor is configured to execute code stored on the memory for implementing a method of decoding one or more tensors from a bitstream, the method comprising:decoding one or more initial tensors from the bitstream;decoding additional information from the bitstream, the additional data comprising (a) a2024202416   12 Apr 2024plurality of pieces of task information each of which indicates machine task which is capable of being performed for the one or more decoded tensors and (b) a plurality of pieces of split point information;selecting one of the plurality of pieces of task information and one of the plurality of pieces of split point information; andproducing the decoded one or more tensors from the initial one or more tensors based on the selected task information and the selected split point information.CANON KABUSHIKI KAISHAPatent Attorneys for the ApplicantSpruson & Ferguson

Citation Information

Patent Citations

  • High-level syntax for signaling neural networks within a media bitstream

    US20220256227A1

  • System and method for encoding and decoding data

    WO2023039627A1