Method, apparatus and system for encoding and decoding tensor

By applying video compression standards and feature encoding technology, the intermediate tensor data of convolutional neural networks is compressed, which solves the problems of high computational complexity and low processing efficiency in the existing technology, and realizes efficient tensor data processing.

CN120019661AInactive Publication Date: 2025-05-16CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071919.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-07-28
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress intermediate tensor data in convolutional neural networks, resulting in high computational complexity and low processing efficiency.

Method used

By using video compression standards such as VVC, combined with feature encoding technology, the intermediate tensor data of convolutional neural networks is compressed, and the dimension of the tensor is reduced using bottleneck encoder and multi-scale feature compression technology.

Benefits of technology

It realizes efficient compression of tensor data in convolutional neural networks, reduces computational complexity and processing requirements, and improves processing efficiency and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019661A_ABST
    Figure CN120019661A_ABST
Patent Text Reader

Abstract

Systems and methods for encoding tensors related to image data in a bitstream. The method comprises: acquiring a first tensor for image data, the first tensor derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, each layer of the first type having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor that has performed a convolution module of one of the plurality of layers of the first type but is not performed a batch normalization module of the one of the plurality of layers of the first type. The method further includes performing a predetermined process on the first tensor to derive a second tensor; and encoding a second tensor in the bitstream, wherein the number of dimensions of the data structure of the first tensor is greater than the number of dimensions of the data structure of the second tensor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit under 35 U.S.C. §119 of the filing date of Australian patent application 2022252785, filed on October 13, 2022, which is incorporated herein by reference in its entirety as if fully set forth herein. Technical Field

[0003] The present invention generally relates to digital video signal processing, and in particular, to methods, devices and systems for encoding and decoding tensors from convolutional neural networks. The present invention also relates to a computer program product comprising a computer-readable medium having recorded thereon a computer program for encoding and decoding tensors from convolutional neural networks using video compression techniques. Background Art

[0004] Convolutional Neural Networks (CNNs) are an emerging technology for addressing use cases involving machine vision, such as object detection, instance segmentation, object tracking, human pose estimation, and action recognition. Applications of CNNs may involve the use of an "edge device" with sensors and some processing capabilities coupled to an application server as part of the "cloud". CNNs may require relatively high computational complexity, which exceeds what can typically be provided in terms of utilizing the computational capacity or power consumption of edge devices. Executing CNNs in a distributed manner has become a solution for running cutting-edge networks using limited-capacity edge devices. In other words, distributed processing enables traditional edge devices to still provide the capabilities of cutting-edge CNNs by allocating processing between edge devices and external processing components such as cloud servers.

[0005] CNNs typically include many layers such as convolutional layers and fully connected layers, where data is passed from one layer to the next in the form of "tensors". Splitting the network across different devices introduces the need to compress the intermediate tensor data passed from one layer to the next within the CNN, such compression can be called "feature compression" because the intermediate tensor data is often called "features" or "feature maps" and represents a partially processed form of an input such as an image frame or video frame. The International Organization for Standardization / International Electrotechnical Commission Joint Technical Committee 1 / Subcommittee 29 / Working Group 2-8 (ISO / IEC JTC1 / SC29 / WG2-8) (also known as the "Moving Picture Experts Group" (MPEG)) has been assigned the task of studying video-related compression techniques. WG2 "MPEG Technical Requirements" has established the "Machine Video Compression" (VCM) ad-hoc group, which has been commissioned to study video compression for machine consumption, as well as feature compression. The Feature Compression Commission is in the exploratory phase with the expectation of issuing a Call for Evidence (CfE) to solicit techniques that can significantly outperform feature compression results achieved using state-of-the-art standardized techniques.

[0006] CNN requires weights for each layer to be determined in a training phase, in which a very large amount of training data is passed through the CNN and the determined results are compared with the ground truth associated with the training data. A process for updating the network weights (such as stochastic gradient descent, etc.) is applied to iteratively refine the network weights until the network performs with the desired accuracy. In the case of a "stride" greater than 1 in the convolution phase, the output tensor from the convolution has a lower spatial resolution than the corresponding input tensor. The pooling operation obtains an output tensor with a smaller dimension than the input tensor. An example of a pooling operation is "max pooling" (or "Maxpool"), which reduces the spatial size of the output tensor compared to the input tensor. Maximum pooling generates an output tensor by dividing the input tensor into groups of data samples (e.g., 2×2 groups of data samples) and selecting the maximum value from each group as the output for the corresponding value in the output tensor. The process of executing a CNN with input and gradually transforming the input into output is often referred to as "inference".

[0007] Typically, a tensor has four dimensions, namely: batch, channel, height, and width. The first dimension "batch", which is of size "1" when performing inference on video data, indicates that one frame is passed through the CNN at a time. When training the network, the value of the batch dimension can increase so that multiple frames are passed through the network before updating the network weights according to a predetermined "batch size". Multiple frames of video can be passed through as a single tensor with a batch dimension whose size increases according to the number of frames of a given video. However, for practical considerations related to memory consumption and access, inference on video data is typically performed frame by frame. The "channel" dimension indicates the number of concurrent "feature maps" for a given tensor, and the height and width dimensions indicate the size of the feature map at a particular stage of the CNN. The number of channels through the CNN varies depending on the network architecture. The size of the feature map also varies depending on the subsampling that occurs in a particular network layer.

[0008] The input to the first layer of a CNN is a batch of one or more images (e.g., a single image or video frame), which is typically sized to be compatible with the dimensions of the tensor input to the first layer. Images or video frames can also be fed in batches of size greater than 1. The dimensions of the tensor depend on the CNN architecture, which generally has some dimensions related to the input width and height and a further "channel" dimension.

[0009] By slicing or reducing the tensor into a collection of two-dimensional arrays, a collection of two-dimensional "feature maps" is obtained based on the channel dimension of the tensor. They are called two-dimensional "feature maps" because each slice of the tensor has some relationship with the corresponding input image, thereby capturing properties such as various edge types. At layers farther away from the input to the network, the properties can be more abstract. The "task performance" of a CNN is measured by comparing its results when performing a task with a specific input to a provided ground truth value that is usually prepared by a human and is considered to indicate the "correct" result.

[0010] Once the network topology is determined, the network weights can be updated over time as more training data becomes available. The overall complexity of CNNs tends to be relatively high when performing a relatively large number of multiplication-accumulation (MAC) operations and writing and reading a large number of intermediate tensors relative to memory. In some applications, CNNs are implemented entirely in the "cloud", which results in the need for high and expensive processing power. In other applications, CNNs are implemented in edge devices such as cameras or mobile phones, which results in less flexibility but more distributed processing loads. Emerging architectures involve splitting the network into multiple parts, one of which runs in the edge device and the other runs in the cloud. This distributed network architecture can be referred to as "collaborative intelligence" and provides benefits such as reusing partial results from the first part of the network for several different second parts (possibly each part is optimized for different tasks). Collaborative intelligence architectures introduce the need for efficient compression of tensor data for transmission over networks such as WANs. Intermediate CNN features can be compressed instead of raw video data, which can be referred to as "feature encoding". The feasibility of feature coding (especially its competitiveness with video coding) often depends on two main factors: the size of the features relative to the size of the original video data; and the ability of the feature encoder to discover and exploit redundancy in the features and the network.

[0011] As described below, video compression standards can be used for feature compression. Various methods can be used to shrink or reduce the data being presented for compression. However, some methods for shrinking or reducing the data being presented for compression may result in a decrease in accuracy that is not suitable for some tasks implemented by CNN.

[0012] Feature compression can benefit from existing video compression standards, such as the universal video coding (VVC) developed by the Joint Video Experts Group (JVET). It is expected that VVC will address the continuous demand for even higher compression performance, and address the growing market demand for service delivery over WAN (wherein the bandwidth cost is relatively high), especially with the increase in the capabilities of video formats (e.g., with higher resolution and higher frame rate). VVC can be implemented in contemporary silicon processes and provides an acceptable compromise between the achieved performance and the implementation cost. The implementation cost can be considered to be one or more than one aspect in terms of silicon area, CPU processor load, memory utilization, and bandwidth. Part of the versatility of the VVC standard is the wide selection of tools that can be used to compress video data, and the wide range of applications that VVC is suitable for. Other video compression standards (such as high-efficiency video coding (HEVC) and AV-1, etc.) can also be used for feature compression applications.

[0013] The video data comprises a sequence of frames of image data, each frame comprising one or more color channels. Typically, one primary color channel and two secondary color channels are required. The primary color channel is often referred to as the "luminance" channel, and the (one or more) secondary color channels are often referred to as the "chrominance" channel. Although video data is typically displayed in an RGB (red-green-blue) color space, this color space has a high correlation between the three corresponding components. The video data representation seen by an encoder or decoder is often using a color space such as YCbCr. YCbCr concentrates the brightness mapped to "luma" according to a transfer function in the Y (primary) channel, and concentrates the chrominance in the Cb and Cr (secondary) channels. Due to the use of decorrelated YCbCr signals, the statistics of the luminance channel are significantly different from those of the chrominance channels. The main difference is that after quantization, the chrominance channel contains relatively fewer significant coefficients for a given block compared to the coefficients of the corresponding luminance channel block. Furthermore, the Cb and Cr channels may be spatially sampled (subsampled) at a lower rate than the luma channel, e.g., half horizontally and half vertically (known as a "4:2:0 chroma format"). The 4:2:0 chroma format is commonly used in "consumer" applications such as Internet video streaming, broadcast television, and Blu-Ray. TM ) storage on disk, etc.). When only luma samples are present, the resulting monochrome frame is considered to use the "4:0:0 chroma format".

[0014] The VVC standard specifies a "block-based" architecture in which a frame is first partitioned into a square array of regions called "coding tree units" (CTUs). A CTU generally occupies a relatively large area such as 128×128 luminance samples. Other possible CTU sizes when using the VVC standard are 32×32 and 64×64. However, the CTUs at the right and lower edges of each frame may be smaller in area, where implicit splitting occurs to ensure that the CB remains in the frame. Associated with each CTU is a "coding tree" ("shared tree") for both the luminance channel and the chrominance channel, or a separate tree for each luminance channel and chrominance channel. The coding tree defines the decomposition of the area of ​​the CTU into a set of blocks (also called "coding blocks" (CBs)). When a shared tree is in use, a single coding tree specifies blocks for both the luminance channel and the chrominance channel, in which case a set of collocated coding blocks is called a "coding unit" (CU) (i.e., each CU has a coding block for each color channel). CBs are processed to be encoded or decoded in a specific order. As a result of using the 4:2:0 chroma format, a CTU with a luma coding tree for a 128×128 luma sample region has a corresponding chroma coding tree for a 64×64 chroma sample region collocated with the 128×128 luma sample region. When a single coding tree is used for the luma channel and the chroma channels, the set of collocated blocks for a given region is often referred to as a "unit", e.g., the CU described above, as well as the "prediction unit" (PU) and "transform unit" (TU). A single tree with a CU spanning the color channels of 4:2:0 chroma format video data results in chroma blocks that are half the width and height of the corresponding luma blocks. When separate coding trees are used for a given region, the CB described above, as well as the "prediction block" (PB) and "transform block" (TB) are used.

[0015] Despite the above distinction between "unit" and "block", the term "block" may be used as a general term for an area or region of a frame for which operations are applied to all color channels.

[0016] For each CU, a prediction unit (PU) ("prediction unit") is generated for the content (sample values) of the corresponding region of the frame data. In addition, a representation of the difference (or "spatial domain" residual) between the prediction and the content of the region seen at the input of the encoder is formed. The differences in each color channel can be transformed and encoded into a sequence of residual coefficients, thereby forming one or more TUs for a given CU. The applied transform can be a discrete cosine transform (DCT) or other transform applied to each block of residual values. The transform is applied separately (i.e., a two-dimensional transform is performed twice). The block is first transformed by applying a one-dimensional transform to each row of samples in the block. The partial result is then transformed by applying a one-dimensional transform to each column of the partial result to produce a final block of transform coefficients that substantially decorrelate the residual samples. The VVC standard supports transforms of various sizes, including transforms of rectangular blocks whose dimensions on each side are powers of 2. The transform coefficients are quantized for entropy encoding into the bitstream.

[0017] VVC has features of intra-frame prediction and inter-frame prediction. Intra-frame prediction involves using previously processed samples in the frame being used to generate a prediction of the current block of data samples in the frame. Inter-frame prediction involves using a block of samples obtained from a previously decoded frame to generate a prediction of the current block of samples in the frame. The block of samples obtained from the previously decoded frame is offset relative to the spatial position of the current block according to a motion vector, which is usually filtered. The intra-frame prediction block can be: (i) uniform sample values ​​("DC intra-frame prediction"), (ii) a plane with an offset and horizontal and vertical gradients ("plane intra-frame prediction"), (iii) a population of blocks with neighboring samples applied in a specific direction ("angular intra-frame prediction"), or (iv) the result of a matrix multiplication using neighboring samples and selected matrix coefficients. Further differences between the predicted block and the corresponding input sample can be corrected to some extent by encoding a "residual" into the bitstream. The residual is generally transformed from the spatial domain to the frequency domain to form residual coefficients in the "main transform" domain. The residual coefficients may be further transformed by applying a "secondary transform" to produce residual coefficients in a "secondary transform domain". The residual coefficients are quantized according to a quantization parameter, which results in a loss of accuracy in the reconstruction of the samples produced at the decoder, but a reduction in the bit rate in the bitstream. A sequence of pictures may be encoded according to a specified structure of pictures using intra-frame prediction and pictures using intra-frame or inter-frame prediction, and a specified dependency on the previous picture in a coding order (which may be different from the display or delivery order). The "random access" configuration results in periodic intra-frame pictures, forming the entry point for the decoder to start decoding the bitstream. According to the hierarchical structure of the specified depth, other pictures in the random access configuration generally use inter-frame prediction to predict content based on pictures before and after the current picture in the display or delivery order. Using pictures after the current picture in the display order to predict the current picture requires a certain degree of picture buffering, as well as a delay between the decoding of a given picture and the display (and removal from the buffer) of the given picture. Summary of the invention

[0018] It is an object of the present invention to substantially overcome or at least ameliorate one or more disadvantages of existing arrangements.

[0019] One aspect of the present disclosure provides a method for encoding a tensor related to image data into a bitstream, the method comprising: obtaining a first tensor for the image data, the first tensor being derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, each layer of the first type having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor for which a convolution module of one of the plurality of layers of the first type has been performed but a batch normalization module of the one of the plurality of layers of the first type has not been performed; performing predetermined processing on the first tensor to derive a second tensor, wherein the number of dimensions of a data structure of the first tensor is greater than the number of dimensions of a data structure of the second tensor; and encoding the second tensor into the bitstream.

[0020] Another aspect of the present disclosure provides a method for deriving a tensor based on a bitstream associated with image data, the derived tensor being used for processing using a portion of a neural network, the method comprising: decoding the tensor from the bitstream; and performing predetermined processing on the decoded tensor to generate the derived tensor, the number of dimensions of the data structure of the derived tensor being greater than the number of dimensions of the data structure of the decoded tensor, wherein the neural network includes at least a plurality of layers of a first type, each layer of the first type having at least a convolution module and a batch normalization module, and the derived tensor corresponds to a tensor in which the convolution module of one of the plurality of layers of the first type has been processed but the batch normalization module of the one of the plurality of layers of the first type has not been processed.

[0021] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing a program for executing a method for encoding a tensor related to image data into a bitstream, the method comprising: obtaining a first tensor for the image data, the first tensor being derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, the first type of layers having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor for which a convolution module of one of the plurality of layers of the first type has been performed but a batch normalization module of the one of the plurality of layers of the first type has not been performed; performing predetermined processing on the first tensor to derive a second tensor, wherein the number of dimensions of a data structure of the first tensor is greater than the number of dimensions of a data structure of the second tensor; and encoding the second tensor into the bitstream.

[0022] Another aspect of the present disclosure provides an encoder configured to: obtain a first tensor related to image data, the first tensor being derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, the layers of the first type having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor for which the convolution module of one of the plurality of layers of the first type has been performed but the batch normalization module of the one of the plurality of layers of the first type has not been performed; perform predetermined processing on the first tensor to derive a second tensor, wherein the number of dimensions of the data structure of the first tensor is greater than the number of dimensions of the data structure of the second tensor; and encode the second tensor into a bit stream.

[0023] Another aspect of the present disclosure provides a system, comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory for implementing a method for encoding a tensor related to image data into a bitstream, the method comprising: obtaining a first tensor for the image data, the first tensor being derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, the first type of layers having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor for which a convolution module of one of the plurality of layers of the first type has been performed but a batch normalization module of the one of the plurality of layers of the first type has not been performed; performing predetermined processing on the first tensor to derive a second tensor, wherein the number of dimensions of a data structure of the first tensor is greater than the number of dimensions of a data structure of the second tensor; and encoding the second tensor into the bitstream.

[0024] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing a program for executing a method for deriving a tensor based on a bitstream associated with image data, the derived tensor being used for processing using a portion of a neural network, the method comprising: decoding a tensor from the bitstream; and performing predetermined processing on the decoded tensor to generate the derived tensor, the number of dimensions of the data structure of the derived tensor being greater than the number of dimensions of the data structure of the decoded tensor, wherein the neural network includes at least a plurality of layers of a first type, the first type of layers having at least a convolution module and a batch normalization module, and the derived tensor corresponds to a tensor in which the convolution module of one of the plurality of layers of the first type has been processed but the batch normalization module of the one of the plurality of layers of the first type has not been processed.

[0025] Another aspect of the present disclosure provides a decoder configured to: decode a tensor from a bitstream associated with image data; and perform predetermined processing on the decoded tensor to generate an derived tensor for processing using a portion of a neural network, wherein the number of dimensions of the data structure of the derived tensor is greater than the number of dimensions of the data structure of the decoded tensor, wherein the neural network includes at least a plurality of layers of a first type, the layers of the first type have at least a convolution module and a batch normalization module, and the derived tensor corresponds to a tensor in which the convolution module of one of the plurality of layers of the first type has been processed but the batch normalization module of the one of the plurality of layers of the first type has not been processed.

[0026] Another aspect of the present disclosure provides a system, comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory for implementing a method for deriving a tensor based on a bitstream associated with image data, the derived tensor being used for processing using a portion of a neural network, the method comprising: decoding a tensor from the bitstream; and performing predetermined processing on the decoded tensor to generate the derived tensor, the number of dimensions of the data structure of the derived tensor being greater than the number of dimensions of the data structure of the decoded tensor, wherein the neural network comprises at least a plurality of layers of a first type, the layers of the first type having at least a convolution module and a batch normalization module, and the derived tensor corresponds to a tensor in which the convolution module of one of the plurality of layers of the first type has been processed but the batch normalization module of the one of the plurality of layers of the first type has not been processed.

[0027] Other aspects are also disclosed. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] At least one embodiment of the present invention will now be described with reference to the following drawings and appendices, in which:

[0029] Figure 1 is a schematic block diagram illustrating a distributed machine task system;

[0030] Figure 2A and Figure 2B Form a practical Figure 1 A schematic block diagram of a general computer system for a distributed machine task system;

[0031] Figure 3A is a schematic block diagram showing the functional modules of the main body of CNN;

[0032] Figure 3B It is shown Figure 3A A schematic block diagram of a residual block;

[0033] Figure 3C It is shown Figure 3A A schematic block diagram of a residual unit of ;

[0034] Figure 3D It is shown Figure 3A Schematic block diagram of the CBL module;

[0035] Figure 4 is a schematic block diagram showing the backbone of a CNN for object tracking;

[0036] Figure 5 is a schematic block diagram showing a portion of a CNN including a multi-scale feature encoder portion;

[0037] Fig. 6A Show a method for performing the first part of a CNN, shrinking using a bottleneck encoder, and encoding the resulting shrunken feature map;

[0038] Figure 6B Show a method for performing the first part of a CNN, shrinking using a bottleneck encoder, and encoding the resulting shrunken feature map;

[0039] Figure 7 is a schematic block diagram illustrating a packing arrangement for a plurality of compressed tensors;

[0040] Figure 8 is a schematic block diagram showing the functional modules of a video encoder;

[0041] Fig. 9 is a schematic block diagram showing the functional modules of a video decoder;

[0042] Fig.10 is a schematic block diagram showing a portion of a CNN including a multi-scale feature decoder portion;

[0043] Fig.11 A method for decoding a bitstream, reconstructing a decorrelated feature map, and performing the second part of a CNN is shown;

[0044] Fig.12 A method for decoding a bitstream, reconstructing a decorrelated feature map, and performing the second part of a CNN is shown;

[0045] Fig.13 is a schematic block diagram showing a head of a CNN for object tracking;

[0046] Fig.14A is a schematic block diagram showing an alternative multi-scale feature fusion module;

[0047] Fig. 14B is a schematic block diagram showing an alternative multi-scale feature reconstruction module;

[0048] Fig.15A is a schematic block diagram illustrating an alternative bottleneck encoder; and

[0049] Fig. 15B is a schematic block diagram illustrating an alternative bottleneck decoder. DETAILED DESCRIPTION

[0050] Where steps and / or features having the same reference numerals are referenced in any one or more of the accompanying drawings, these steps and / or features have the same function(s) or operation(s) for the purposes of this specification, unless otherwise intended.

[0051] A distributed machine task system may include edge devices, such as a web camera or smartphone that generates intermediate compressed data. A distributed machine task system may also include end devices, such as server farm ("cloud") based applications that operate on the intermediate compressed data to generate some task result. In addition, edge device functionality may be embodied in the cloud, and the intermediate compressed data may be stored for later processing (potentially for multiple different tasks as needed). An example of a machine task is object tracking, with a mean object tracking accuracy (MOTA) score as a typical task result. Other examples of machine tasks include object detection and instance segmentation, both of which generate task results measured as "mean average precision" (mAP) for detections that exceed a threshold of intersection over union (IoU), such as 0.5.

[0052] A convenient form of intermediate compressed data is a compressed video bitstream, due to the availability of high performance compression standards and their implementations. Video compression standards typically operate on integer samples of some given bit depth (such as 10 bits, etc.) arranged in a planar array. Color video has, for example, three planar arrays corresponding to color components Y, Cb, Cr or R, G, B, depending on the application. CNNs typically operate on floating point data in the form of tensors. Tensors generally have much smaller spatial dimensions than the input video data on which CNNs operate, but have many more channels than the typical three channels of color video data.

[0053] Tensors typically have the following dimensions: frame, channel, height, and width. For example, a tensor of dimension [1,256,76,136] would be considered to contain two hundred and fifty-six (256) feature maps, each of size 136 × 76. For video data, inference is typically performed on one frame at a time, rather than using a tensor containing multiple frames.

[0054] VVC supports partitioning a picture into multiple sub-pictures, each of which can be independently encoded and independently decoded. In one approach, each sub-picture is encoded as a "slice" or a continuous sequence of encoded CTUs. A "tile" mechanism can also be used to partition a picture into multiple independently decodable areas. Sub-pictures can be specified in a slightly flexible manner, where various rectangular sets of CTUs are encoded as corresponding sub-pictures. The flexible definition of sub-picture dimensions allows data types that require different areas to be efficiently maintained in one picture, thereby avoiding large "unused" areas (i.e., areas in the frame that are not used for reconstruction of tensor data).

[0055] Figure 1 is a schematic block diagram illustrating the functional modules of a distributed machine task system 100. The concept of distributing machine tasks across multiple systems is sometimes referred to as "collaborative intelligence" (CI). The system 100 can be used to implement methods for decorrelating, packing, and quantizing feature maps into planar frames to encode feature maps and decoding feature maps from encoded data. These methods can be implemented so that the associated overhead data is not too heavy, and the task performance on the decoded feature maps is resilient to varying bit rates of the bitstream, and the quantized representation of tensors does not unnecessarily consume bits where the bits do not provide a commensurate benefit in terms of task performance.

[0056] The system 100 includes a source device 110 for generating encoded tensor data 115 from a CNN backbone 114 in the form of an encoded video bitstream 121. The system 100 also includes a destination device 140 for decoding the tensor data in the form of an encoded video bitstream 143. A communication channel 130 is used to communicate the encoded video bitstream 121 from the source device 110 to the destination device 140. In some arrangements, one or both of the source device 110 and the destination device 140 may include a respective mobile phone handset (e.g., a "smart phone") or a web camera and a cloud application. The communication channel 130 may be a wired connection such as Ethernet or a wireless connection such as WiFi or 5G, including a connection across a wide area network (WAN) or across an ad hoc connection. In addition, the source device 110 and the destination device 140 may include an application that captures the encoded video data on some computer-readable storage media (such as a hard drive or memory in a file server, etc.).

[0057] like Figure 1As shown, source device 110 includes video source 112, CNN backbone 114, bottleneck encoder 116, quantization and packing module 118, feature map encoder 120 and transmitter 122. Video source 112 typically includes a source of captured video frame data (denoted as 113), such as a camera sensor, a previously captured video sequence stored on a non-transitory recording medium, or a video feed from a remote camera sensor. Video source 112 can also be the output of a computer graphics card, for example, displaying the video output of an operating system and various applications executed on a computing device (e.g., a tablet computer). Examples of source devices 110 that can include a camera sensor as a video source 112 include smart phones, video camcorders, professional video cameras, and network video cameras.

[0058] The system 100 uses a "bottleneck" (i.e., an additional network layer that limits the tensor dimensionality on the encoding side and restores the tensor dimensionality on the decoder side) to reduce the dimensionality of the tensor at the interface between the first network portion and the second network portion. An approach called "multi-scale feature compression" (MSFC) is used to "fuse" the multi-scale representations produced by the feature pyramid network (FPN) together into a single tensor. MSFC is typically used to merge all FPN layers into a single tensor. Merging all FPN layers into a single tensor is achieved at the expense of spatial detail of the less decomposed (larger) layers of the FPN. The loss of spatial detail can result in an unacceptable decrease in accuracy for some tasks or operations implemented by the system 100.

[0059] The described arrangement separates the tensors of FPN into groups and applies the MSFC technique separately, rather than merging all FPN layers into a single tensor. Applying the MFSC technique separately allows a certain degree of cross-layer fusion without severe degradation of spatial details. For tasks such as instance segmentation that require preservation of greater spatial details, using the separate MSFC technique results in a higher mAP than when all FPN layers are merged into a single tensor with low spatial resolution.

[0060] The CNN backbone 114 receives the video frame data 113 and processes specific layers of the overall CNN (such as layers corresponding to the "backbone" of the CNN, etc.), thereby outputting a tensor 115. The backbone layers of the CNN can produce as output multiple tensors (e.g., corresponding to different spatial scales of the input image represented by the video frame data 113 (sometimes referred to as a "feature pyramid network" (FPN) architecture)). The tensors produced by the FPN backbone form a hierarchical representation of the frame data 113 including data of feature maps. Each successive layer of the hierarchical representation has half the width and height of the previous layer. Later layers generated further into the backbone network tend to contain features with more abstract representations of the frame data 113. Less decomposed layers generated earlier in the backbone network tend to contain features representing less abstract features of the frame data 113, such as various geometric properties (such as edges at various angles, etc.), etc. When system 100 is performing a "YOLOv3" network, FPN can obtain three tensors corresponding to three layers output from backbone 114 as tensor 115, where tensor 115 has different spatial resolutions and numbers of channels. When system 100 is performing a network such as "Faster RCNN X101-FPN" or "Mask RCNN X101-FPN", tensor 115 includes tensors for four layers P2 to P5. Bottleneck encoder 116 receives tensor 115. Bottleneck encoder 116 is used to compress one or more internal layers of the overall CNN. The internal layers of the overall CNN provide the output of the CNN backbone 114, which is compressed or contracted by bottleneck encoder 116 using a trained set of neural network layers to convert into a lower number of channels and a smaller spatial resolution than required by tensor 115. Bottleneck encoder 116 outputs bottleneck tensor 117.

[0061] The bottleneck tensor 117 is passed to a quantization and packing module 118. Each feature map of the bottleneck tensor 117 is quantized from floating point to integer precision by the module 118 and packed into a monochrome frame to produce a packed frame 119. The packed frame 119 is input to a feature map encoder 120. The feature map encoder 120 encodes the packed frame 118 to generate a bitstream 121. The bitstream 121 is supplied to a transmitter 122 for transmission over a communication channel 130, or the bitstream 121 is written to a storage unit 132 for later use.

[0062] The source device 110 supports a specific network for the CNN backbone 114. However, the destination device 140 may use one of several networks for the CNN head 150. In this way, partially processed data in the form of packed feature maps may be stored for later use in performing various tasks without having to repeat the operation of the CNN backbone 114.

[0063] The bitstream 121 is transmitted as encoded video data (or "encoded video information") by a transmitter 122 over a communication channel 130. In some implementations, the bitstream 121 may be stored in a storage 132 (where the storage 132 is a non-transitory storage device such as a "Flash" memory or a hard drive) until (or in lieu of) later transmission over the communication channel 130. For example, the encoded video data may be provided to consumers on demand over a wide area network (WAN) for use in video analytics applications.

[0064] The destination device 140 includes a receiver 142, a feature map decoder 144, an unpacking and inverse quantization module 146, a bottleneck decoder 148, a CNN head 150, and a CNN task result buffer 152. The receiver 142 receives the encoded video data from the communication channel 130 and passes the video bitstream 143 to the feature map decoder 144. The feature map decoder decodes the bitstream to generate decoded packed frames 145. The decoded frames are input to the unpacking and inverse quantization module 146. The module 146 unpacks and inverse quantizes the tensors of the frame 145 to generate dequantized tensors, which are output as decoded bottleneck tensors 147. The decoded bottleneck tensors 147 are fed to the bottleneck decoder 148. The bottleneck decoder 148 performs the inverse operation of the bottleneck encoder 116 to produce the extraction tensor 149. The extraction tensor 149 is passed to the CNN head 150. CNN head 150 performs the later layers of the task started from CNN trunk 114 to produce task results 151, which are stored in task result buffer 152. The contents of task result buffer 152 can be presented to the user, for example, via a graphical user interface, or provided to an analysis application that decides an action based on the task results, which may include a summary-level presentation of the aggregated task results to the user. It is also possible that the functionality of each of source device 110 and destination device 140 is embodied in a single device, examples of which include mobile phones and tablet computers and cloud applications.

[0065] Although example devices are described above, source device 110 and destination device 140 may each typically be configured within a general purpose computer system via a combination of hardware and software components. Figure 2ASuch a computer system 200 is illustrated, and includes: a computer module 201; input devices such as a keyboard 202, a mouse pointer device 203, a scanner 226, a camera 227 that can be configured as a video source 112, and a microphone 280; and output devices including a printer 215, a display device 214 (which can be configured as a display device for presenting task results 151), and a speaker 217. The computer module 201 can use an external modulator-demodulator (modem) transceiver device 216 to communicate with a communication network 220 via a connection 221. The communication network 220, which can represent a communication channel 130, can be a (WAN), such as the Internet, a cellular telecommunications network, or a private WAN. In the case where the connection 221 is a telephone line, the modem 216 can be a traditional "dial-up" modem. Alternatively, in the case where the connection 221 is a high-capacity (e.g., cable or optical) connection, the modem 216 can be a broadband modem. A wireless modem may also be used to make a wireless connection to the communication network 220. The transceiver device 216 may provide the functionality of the transmitter 122 and the receiver 142, and the communication channel 130 may be embodied in the connection 221.

[0066] The computer module 201 typically includes at least one processor unit 205 and a memory unit 206. For example, the memory unit 206 may have a semiconductor random access memory (RAM) and a semiconductor read-only memory (ROM). The computer module 201 also includes a plurality of input / output (I / O) interfaces, wherein the plurality of input / output (I / O) interfaces include: an audio-video interface 207 coupled to a video display 214, a speaker 217, and a microphone 280; an I / O interface 213 coupled to a keyboard 202, a mouse 203, a scanner 226, a camera 227, and an optional joystick or other human interface device (not illustrated); and an interface 208 for an external modem 216 and a printer 215. The signal from the audio-video interface 207 to the computer monitor 214 is typically the output of a computer graphics card. In some implementations, the modem 216 may be built into the computer module 201, such as built into the interface 208. The computer module 201 also has a local network interface 211, which allows the computer system 200 to be coupled to a local area communication network 222, known as a local area network (LAN), via a connection 223. Figure 2A As shown, the local area communication network 222 may also be connected to the wide area network 220 via a connection 224, wherein the local area communication network 222 will typically include a so-called "firewall" device or a device with similar functionality. The local network interface 211 may include an Ethernet TM)Circuit card, Bluetooth TM ) wireless arrangement or IEEE 802.11 wireless arrangement; however, for the interface 211, a variety of other types of interfaces can be implemented. The local network interface 211 can also provide the functionality of the transmitter 122 and the receiver 142, and the communication channel 130 can also be embodied in the local communication network 222.

[0067] I / O interfaces 208 and 213 may provide either or both serial connectivity and parallel connectivity, wherein the former is typically implemented according to the Universal Serial Bus (USB) standard and has a corresponding USB connector (not illustrated). A storage device 209 is provided, and the storage device 209 typically includes a hard disk drive (HDD) 210. Other storage devices (not illustrated) such as floppy disk drives and tape drives may also be used. An optical disk drive 212 is typically provided to serve as a non-volatile source of data. For example, optical disks (e.g., CD-ROM, DVD, Blu-ray Disc, etc.) may be used. TM )), portable memory devices such as USB-RAM, portable external hard drives and floppy disks as suitable sources of data for computer system 200. Typically, any of HDD 210, optical drive 212, networks 220 and 222 may also be configured to operate as video source 112, or as a destination for decoded video data to be stored for reproduction via display 214. Source device 110 and destination device 140 of system 100 may be embodied in computer system 200.

[0068] The components 205 to 213 of the computer module 201 typically communicate via an interconnect bus 204 and in a manner that results in conventional modes of operation of the computer system 200 known to those skilled in the relevant art. For example, the processor 205 is coupled to the system bus 204 using a connection 218. Likewise, the memory 206 and the optical drive 212 are coupled to the system bus 204 via a connection 219. Examples of computers on which the described arrangement may be practiced include IBM-PC and compatibles, Sun SPARCstations, Apple Mac™ or similar computer systems.

[0069] Where appropriate or desired, the source device 110 and the destination device 140 and the methods described below may be implemented using the computer system 200. In particular, the source device 110, the destination device 140 and the methods to be described may be implemented as one or more software applications 233 executable within the computer system 200. The instructions 231 (see FIG. 233 ) executed within the computer system 200 may be used to execute the source device 110, the destination device 140 and the methods to be described. Figure 2B) to implement the source device 110, the destination device 140 and the steps of the method. The software instructions 231 may be formed into one or more code modules, each for performing one or more specific tasks. The software may also be split into two separate parts, with a first part and corresponding code modules performing the method, and a second part and corresponding code modules managing a user interface between the first part and a user.

[0070] For example, the software may be stored in a computer-readable medium including a storage device described below. The software is loaded from the computer-readable medium into the computer system 200 and then executed by the computer system 200. A computer-readable medium having such software or a computer program recorded on the computer-readable medium is a computer program product. The use of the computer program product in the computer system 200 preferably implements an advantageous device for implementing the source device 110 and the destination device 140 and the method.

[0071] The software 233 is typically stored in the HDD 210 or the memory 206. The software is loaded into the computer system 200 from a computer readable medium and executed by the computer system 200. Thus, for example, the software 233 may be stored on an optically readable disk storage medium (e.g., a CD-ROM) 225 read by the optical drive 212.

[0072] In some examples, the application 233 is provided to the user in a manner encoded on one or more CD-ROMs 225 and read via corresponding drives 212, or alternatively, the application 233 can be read by the user from the network 220 or 222. Furthermore, the software can also be loaded into the computer system 200 from other computer-readable media. Computer-readable storage media refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to the computer system 200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray discs, hard disk drives, ROMs or integrated circuits, USB memories, magneto-optical disks, or computer-readable cards such as PCMCIA cards, etc., regardless of whether these devices are internal or external to the computer module 201. Examples of transient or non-tangible computer-readable transmission media that may also participate in the provision of software, applications, instructions and / or video data or encoded video data to the computer module 201 include: radio or infrared transmission channels and network connections to another computer or networked device, and the Internet or Intranet including e-mail transmissions and information recorded on websites.

[0073] The second part of the above-mentioned application program 233 and the corresponding code module can be executed to implement one or more graphical user interfaces (GUIs) to be rendered or otherwise represented on the display 214. By typically manipulating the keyboard 202 and the mouse 203, users and applications of the computer system 200 can manipulate the interface in a functionally applicable manner to provide control commands and / or input to the applications associated with these (one or more) GUIs. Other forms of user interfaces that are functionally applicable can also be implemented, such as audio interfaces that utilize voice prompts output via the speaker 217 and user voice commands input via the microphone 280, etc.

[0074] Figure 2B is a detailed schematic block diagram of the processor 205 and the "memory" 234. The memory 234 represents Figure 2A A logical aggregation of all memory modules (including storage device 209 and semiconductor memory 206) that can be accessed by computer module 201 in.

[0075] When the computer module 201 is initially powered on, a power-on self-test (POST) program 250 is executed. The POST program 250 is typically stored in Figure 2A 249 of the semiconductor memory 206. Hardware devices such as ROM 249 storing software are sometimes referred to as firmware. POST program 250 checks the hardware within computer module 201 to ensure proper operation, and typically checks processor 205, memory 234 (209, 206) and basic input-output system software (BIOS) module 251, which is also typically stored in ROM 249, for correct operation. Once POST program 250 runs successfully, BIOS 251 activates Figure 2A Activating the hard disk drive 210 causes the bootstrap loader program 252 residing on the hard disk drive 210 to be executed via the processor 205. This loads the operating system 253 into the RAM memory 206, where the operating system 253 begins to operate on the RAM memory 206. The operating system 253 is a system-level application executable by the processor 205 to implement various high-level functions including processor management, memory management, device management, storage management, software application interface, and general user interface.

[0076] The operating system 253 manages the memory 234 (209, 206) to ensure that each process or application running on the computer module 201 has sufficient memory to execute without conflicting with the memory allocated to another process. Figure 2AThe different types of memory available in the computer system 200 are described so that each process can run efficiently. Therefore, the aggregate memory 234 is not intended to illustrate how specific segments of memory are allocated (unless otherwise specified), but rather to provide an overview of the memory accessible to the computer system 200 and how such memory is used.

[0077] like Figure 2B As shown, the processor 205 includes a plurality of functional modules, wherein the plurality of functional modules include a control unit 239, an arithmetic logic unit (ALU) 240, and a local or internal memory 248 sometimes referred to as a cache memory. The cache memory 248 typically includes a plurality of storage registers 244 to 246 in a register section. One or more internal buses 241 functionally connect these functional modules to each other. The processor 205 typically also has one or more interfaces 242 for communicating with external devices via the system bus 204 using a connection 218. The memory 234 is coupled to the bus 204 using a connection 219.

[0078] The application program 233 includes an instruction sequence 231 that may include conditional branch instructions and loop instructions. The program 233 may also include data 232 used when executing the program 233. The instructions 231 and data 232 are stored in memory locations 228, 229, 230 and 235, 236, 237, respectively. Depending on the relative size of the instructions 231 and the memory locations 228 to 230, a particular instruction may be stored in a single memory location, as described by the instruction shown in memory location 230. Alternatively, the instruction may be segmented into multiple portions, each stored in a separate memory location, as described by the instruction segments shown in memory locations 228 and 229.

[0079] Typically, a processor 205 is given a set of instructions, which are executed within the processor 205. The processor 205 awaits a subsequent input, which the processor 205 reacts to by executing another set of instructions. Each input may be provided from one or more of a plurality of sources, including data generated by one or more of the input devices 202, 203, data received from an external source across one of the networks 220, 202, data retrieved from one of the storage devices 206, 209, or data retrieved from a storage medium 225 inserted into a corresponding reader 212 (all of which are described in detail in the accompanying drawings). Figure 2A Execution of the instruction set may result in outputting data in some cases. Execution may also involve storing data or variables to memory 234.

[0080] The bottleneck encoder 116, the bottleneck decoder 148, and the method may use input variables 254 stored in corresponding memory locations 255, 256, 257 within the memory 234. The bottleneck encoder 116, the bottleneck decoder 148, and the method produce output variables 261 stored in corresponding memory locations 262, 263, 264 within the memory 234. The intermediate variables 258 may be stored in memory locations 259, 260, 266, and 267.

[0081] refer to Figure 2B The processor 205, registers 244, 245, 246, arithmetic logic unit (ALU) 240 and control unit 239 work together to perform the micro-operation sequence required to perform the "fetch, decode and execute" cycle for each instruction in the instruction set that constitutes the program 233. Each fetch, decode and execute cycle includes:

[0082] A fetch operation for fetching or reading an instruction 231 from a memory location 228 , 229 , 230 ;

[0083] a decode operation in which the control unit 239 determines which instruction was fetched; and

[0084] An execution operation, wherein in the execution operation, the control unit 239 and / or the ALU 240 executes the instruction.

[0085] Thereafter, a further fetch, decode and execute cycle for the next instruction may be performed. Similarly, a store cycle may be performed, whereby the control unit 239 stores or writes a value to the memory location 232 .

[0086] To be explained FIG. 3A to FIG. 6B as well as Figures 8 to 13 Each step or sub-process in the method is associated with one or more segments of program 233, and is typically performed by register segments 244, 245, 247, ALU240 and control unit 239 in processor 205 working together to perform fetch, decode and execute cycles for each instruction in the instruction set of the segment of program 233.

[0087] Figure 3A is a schematic block diagram showing the architecture 300 of the functional modules of the feature extraction part 310 of the CNN. The feature extraction part 310 forms part of the implementation of the CNN backbone 114 and may be followed by "skip connections" in networks such as YOLOv3. The CNN part 310 may be referred to as "Dark-Net 53". Different feature extractions are also possible, resulting in different numbers and dimensions of layers of the tensor 330 output for each frame.

[0088] like Figure 3A As shown in FIG. 1 , the video data 113 is passed to a resizer module 304. The resizer module 304 resizes the frame to a resolution suitable for processing by the CNN portion 310, thereby generating resized frame data 312. If the resolution of the frame data 113 is already suitable for the CNN backbone 310, then the operation of the resizer module 304 is not required. The resized frame data 312 is passed to a convolutional batch normalization leaky rectified linear (CBL) module 314 (also referred to as a CBL layer) to generate a tensor 316. The CBL 314 contains reference Figure 3D The modules described by the CBL module 360 ​​are shown.

[0089] refer to Figure 3D , CBL module 360 ​​takes tensor 361 as input. The tensor 361 is passed to convolution layer 362 to produce tensor 363. When convolution layer 362 has a stride of 1 and padding is set to k samples (where the size of the convolution kernel is 2k+1), tensor 363 has the same spatial dimensions as tensor 361. When convolution layer 362 has a larger stride (such as 2, etc.), tensor 363 has a smaller spatial dimension than tensor 361, for example, for a stride of 2, the size of tensor 363 is halved. Regardless of the stride, for a particular CBL block, the size of the channel dimension of tensor 363 may vary compared to the channel dimension of tensor 361. Tensor 363 is passed to batch normalization module 364, which outputs tensor 365. Batch normalization module 364 normalizes the input tensor 363, applies a scaling factor and an offset value to produce an output tensor 365. The scaling factors and offset values ​​are derived from the training process. Tensor 365 is passed to a leaky rectified linear activation ("Leaky ReLU") module 366 to produce tensor 367. Module 366 provides a "leaky" activation function whereby positive values ​​in the tensor are passed through and negative values ​​are severely reduced in magnitude, e.g., to 0.1 times their previous value.

[0090] Return to Figure 3A , tensor 316 is passed from CBL block 314 to residual block 11 (Res11) module 320. Module 320 comprises a sequential cascade of three residual blocks internally containing 1, 2 and 8 residual units respectively.

[0091] References Figure 3BResBlock 340 is shown describing a residual block such as that present in module 320. ResBlock 340 receives tensor 341. Tensor 341 is zero-filled by zero-filling module 342 to produce tensor 343. Tensor 343 is passed to CBL module 344 to produce tensor 345. Tensor 345 is passed to residual unit 346, where residual block 340 contains a series of cascaded residual units. The last residual unit in residual unit 346 outputs tensor 347.

[0092] References Figure 3C 350 is used to describe a residual unit such as unit 346. ResUnit 350 takes tensor 351 as input. Tensor 351 is passed to CBL module 352 to produce tensor 353. Tensor 353 is passed to a second CBL unit 354 to produce tensor 355. Addition module 356 sums tensor 355 with tensor 351 to produce tensor 357. Addition module 356 may also be referred to as a "shortcut" because input tensor 351 substantially affects output tensor 357. For an untrained network, ResUnit 350 acts to pass through tensors. When training, CBL modules 352 and 354 act to deviate tensor 357 from tensor 351 based on training data and ground truth data.

[0093] Return to Figure 3A , the Res11 module 320 outputs a tensor 322. The tensor 322 is output from the backbone module 310 as one of the layers and is also provided to the Res8 module 324. The Res8 module 324 is a residual block (i.e., 340) including eight residual units (i.e., 350). The Res8 module 324 generates a tensor 326. The tensor 326 is passed to the Res4 module 328 and is output from the backbone module 310 as one of the layers. The Res4 module 328 is a residual block (i.e., 340) including four residual units (i.e., 350). The Res4 module 328 generates a tensor 329. The tensor 329 is output from the backbone module 310 as one of the layers. In general, the layer tensors 322, 326, and 329 are output as tensors 115.

[0094] Figure 44 is a schematic block diagram of an architecture 400 of functional modules of a CNN backbone network 4001 that can be used as an implementation of a CNN backbone 114. The backbone 4001 implements a portion of a skip connection 4004 and a connection to a feature extraction network Darknet-53 310. The frame data 113 is input to the Darknet-53 network 310 to generate three tensors 329, 326, and 322. The three tensors 329, 326, and 322 can have different numbers of channels and / or different spatial resolutions or scales. The tensors 329, 326, and 322 are passed to the skip connection 4004, which generates tensors 451, 461, 491, 429, and 469. As described below, the tensors 429 and 469 from the skip connection are used in the skip connection 4004 in addition to being outputs from the skip connection 4004 to provide support for other layers in the skip connection 4004. Tensors 451 , 461 , and 491 may be used as tensor 115 .

[0095] Skip connections 400 include three sequences (or "tracks") of CBL layers. The CBL layers within each sequence of CBL layers in architecture 400 block can be implemented using the structure of CBL 360 and include sequentially linked convolutional layer modules, batch normalization layer modules, and leaky ReLU layer modules. The first track obtains a first tensor 329 from Darknet-53 network 310. Tensor 329 is passed through five CBL layers 420, 422, 424, 426, and 428. The first CBL 420 outputs a tensor 421. Tensor 421 is passed through a sequence of CBLs 422, 424, 426, and 428 in succession. CBL 428 produces a first tensor 451 in tensor 115 and a tensor 429 of skip connection section 4004. The second track takes the output tensor 429 of the first track, performs a CBL layer 430, and then performs an upsampling block 423 to produce a tensor 433. Tensor 433 is concatenated with the second tensor 326 via a concatenation block 434 to produce a tensor 435. Tensor 435 passes through five blocks of CBL 440, 442, 444, 446, and 468. The first CBL 440 outputs a tensor 441. Tensor 441 is continuously passed through a sequence of CBL 442, 444, 446, and 468. CBL 468 produces a second tensor 461 in tensor 115 and a tensor 469 of the skip connection part 4004. The third track takes the output tensor 469 of the second track and performs a CBL block 470, followed by an upsampling block 472 to produce a tensor 473. Tensor 473 is concatenated with the third feature map 322 via a concatenation block 474 to produce a tensor 475. Tensor 475 passes through four blocks of CBL 480, 482, 484, and 486. The first CBL 480 outputs a tensor 481. The tensor 481 is passed successively through a sequence of CBLs 482, 484, and 486 to produce a tensor 487. The tensor 487 is input to a convolutional layer 490 of a CBL 488 to produce a third tensor 491 in the tensor 115.

[0096] The split point between the CNN backbone 114 and the CNN head 150 in the previous arrangement can be selected (i) at the output of Darknet-53CNN310 (tensors 329, 326, and 322), (ii) after the first of the five CBL blocks in each track (tensors 421, 441, and 481), or (iii) after the last of the five CBL blocks (tensors 429 and 469, and the tensor output of CBL 488). In the described arrangement, the split point of the CNN backbone 114 can also be taken at the output of each convolution module in each of the CBL layers 428, 448, and 488 (output tensors 451, 461, and 491).

[0097] In the example CNN backbone 4001, the split point is selected at the output of each convolution module in the last CBL layer 428, 448 and 488 of each track of the CBL layer within the skip connection 4004. The tensors output from the split point are 451, 461 and 491. The three tensors 451, 461 and 491 provide tensor 115. In the described example, the split point is after the convolution module inside the CBL layer. The CBL layers 428, 468 and 488 can be considered as the last layers of the three tracks of the portion of the neural network that includes the convolution layers. The split point is effectively in the last three layers of each track of the skip connection section 4004.

[0098] The CBL layer 428 includes three modules, namely, a convolution module 450, a batch normalization module 452, and a leaky ReLU module 454. The tensor 427 generated by the CBL 426 is input to the convolution layer 450 to generate a tensor 451. The tensor 451 provides the first tensor in the tensor 115. The CBL layer 468 includes three modules, namely, a convolution module 460, a batch normalization module 462, and a leaky ReLU module 464. The tensor 447 output by the CBL 446 is input to the convolution layer 460 to generate a tensor 461. The tensor 461 provides the second tensor in the tensor 115. The layer 488 includes three modules, namely, a convolution module 490, a batch normalization module (not shown), and a leaky ReLU module (not shown). Only the module 490 of the CBL 488 is performed in the backbone CNN 114.

[0099] The first tensor 451 from the split point needs to be passed through batch normalization 452 and leaky ReLU 454 to produce tensor 429. Tensor 429 is returned to the skip connection for input to CBL 430, so the backbone 114 needs to include modules 452 and 454. The second tensor 461 from the split point needs to be passed through batch normalization 462 and leaky ReLU 464 to produce tensor 469. Tensor 469 is returned to the skip connection for input to CBL 470. Therefore, the backbone CNN 114 needs to include modules 462 and 464.

[0100] The CNN backbone 4001 can take a video frame of resolution 1088×608 as input and generate three tensors corresponding to three layers, the three tensors having the following dimensions: [1,256,76,136], [1,512,38,68], [1,1024,19,34]. Another example of three tensors corresponding to three layers can be [1,512,34,19], [1,256,68,38], [1,128,136,76] separated at the 75th network layer, the 90th network layer, and the 105th network layer in the CNN backbone 4001, respectively (corresponding to the outputs of CBL layers 420, 440, and 480, respectively (that is, tensors 421, 441, and 481, respectively)). Each tensor can have a different resolution from the next tensor. The resolution of the tensor can form an exponential sequence in which the height and width double between each consecutive tensor in the tensor. In forming the output tensor 115 , modules 322 , 326 , and 329 provide a hierarchical representation of the frame data including data for feature maps encoded into a bitstream. The separation point depends on the CNN 310 .

[0101] The output tensor 429 of the first track and the output tensor 469 of the second track of the skip connection stage are used successively in the second and third tracks of the skip connection 4004 in the source device 110. Therefore, there are some "overlapping" layers (or modules) running in both the trunk (114) in the source device 110 and the CNN head (150) in the destination device 140. In the third track (starting from the CBL 470), there are no overlapping layers because the output tensor 491 of the third track is only used in the head 150 and is not fed back to another track in the skip connection 4004.

[0102] Overlapping layers are layers or modules that are (i) after the first split point (after the output of module 450) until the end of the first track of the skip connection (modules 452 and 454), and (ii) after the second split point (after the output of module 460) until the end of the second track of the skip connection (modules 462 and 464). If the split point is close to the end of the skip connection section 4004, the number of overlapping layers or modules is reduced. If the split point is taken after each CBL layer in CBL layers 420, 440 and 480, the overlapping layers are CBL 422, 424, 426 and 428 in the first track and CBL 442, 444, 446 and 468 in the second track. As a result, selecting the split point to output tensors 451, 461 and 491 reduces execution redundancy and reduces the execution time of the end-to-end network of CNNs 114 and 150.

[0103] Figure 5460 and 490. The multi-scale feature encoder 550 receives the tensor [L2, L1, L0] as input and generates (one or more) output tensors 560. Block 550 corresponds to bottleneck encoder 116, and output 560 corresponds to output bottleneck tensor 117. Output 560 is consistent with Fig. 6A and Fig.14A A tensor in the arrangement of the bottleneck encoder 550 and is consistent with Fig.15A The outputs of modules 454 and 464 are not encoded by encoder 550.

[0104] The multi-scale feature encoder 550 can use a convolutional neural network that can learn, train, create and select features from the input tensors of the FPN (such as L0, L1, and L2, etc.) to produce a representation with reduced dimensions in terms of the number of tensors, the number of channels, and the width and height within (one or more than one) compressed tensors. For example, MFSC can be used. The functional module 550 can also use non-trainable methods such as PCA (principal component analysis). The PCA implementation uses an orthogonal transform to transform the number of feature channels (such as 256, etc.) to a smaller number (such as 25, etc.) by obtaining the 25 strongest features (feature vectors or basis vectors) from the transformed features and using coefficients to represent the contents of the 256 feature maps as the weighted sum of the basis vectors. In order to utilize the inter-feature map transformation in PCA, a smaller spatial resolution tensor (such as L2, etc.) can be upsampled before being cascaded with a larger spatial resolution tensor (such as L1, etc.). The feature vector is obtained from the cascaded tensor transformation. If the PCA method is used, the coefficients also need to be encoded and transmitted to the header 150 for decoding purposes.

[0105] Using the split points described, CNN head 114 outputs tensors 451, 461, and 491, each produced by a convolution module. Typically, tensors produced from convolution modules are more stable when training with multi-scale feature compression. Using other split points described can result in tensors 115 (such as (329, 326, 322) or (421, 441, 481), etc.) produced from activation functions such as Leaky ReLU (which are less stable in trainable transformations such as MSFC).

[0106] Fig. 6A is a schematic block diagram showing an example of functional modules of a multi-scale feature encoder 600 corresponding to the multi-scale feature encoder 550. The feature maps or tensors output at the split point may include three feature maps L2, L1, and L0 of sizes (B, C2, h / 32, 2 / 32), (B, C1, h / 16, 2 / 16), and (B, C0, h / 8, 2 / 8), respectively. When describing the size, B is the batch size, C2, C1, and C0 are the number of channels, and h and w are the height and width of the input image or video frame. The CNN backbone 4001 may take a video frame of resolution 1088×608 as input and produce three tensors corresponding to three layers, which have the following dimensions for the YOLOv3 network: [1, 1024, 19, 34], [1, 512, 38, 68], [1, 256, 76, 136]. Another example of three tensors corresponding to three layers may be [1,512,34,19], [1,256,68,38], [1,128,136,76] separated at the 79th network layer, the 94th network layer, and the 109th network layer in CNN 4001, respectively. Tensor 115 is generated using split points from each of the layers 79, 94, and 109 of the neural network including the backbone 114 and the head 150. In another example for YOLOv3, tensor 115 is generated using split points after the convolution layer at the CBL layer for each of the layers in the layers 79, 94 and 109 of the neural network. In a software implementation of the JDE network (such as the software implementation used in the VCM Ad Hoc Group Object Tracking Feature Anchor (ISO / IEC JTC 1 / SC 29 / WG 2m59940)), all layers of the network are enumerated, and the indices of the CBL layers 428, 468, 488 are 79, 94, and 109 in the enumeration, respectively. Each tensor can have a different resolution than the other tensors in tensors 115. The resolution of each tensor can be doubled in height and width between corresponding tensors. In forming the output tensor 115, tensors 514, 524, and 534 provide a hierarchical representation of frame data including data for feature maps encoded into a bitstream. The separation point depends on CNN310.

[0107] The multi-scale feature encoder 600 includes two blocks. The first block is a multi-scale feature fusion MSFF block 608 and the second block is a single-scale feature compression (SSFC) encoder 650. The first tensor L2 451 is passed to an upsampler 610. The upsampler 610 upsamples the tensor 451 to produce a tensor 612 with a larger spatial scale, i.e., L2 451 at (h / 32, w / 32, 512) is upsampled to match the spatial scale of the larger tensor L1 461. The third feature map L0 491 is passed through a downsampler 612. The downsampler 612 downsamples the tensor 491 and produces a tensor 614 with a smaller spatial scale, i.e., L0 491 at (h / 8, w / 8, 128) is downsampled to match the spatial scale of the smaller tensor 461.

[0108] The three tensors 612, 461, 614 have the same spatial size and are passed through a cascade module 620, which merges all tensors into a single tensor 622 along the channel dimension. The merged tensor 622 is input to a squeeze and excitation (SE) block 624. The SE block 624 is trained to adaptively change the weighting of different channels in the tensor 622 based on the first fully connected layer output. The first fully connected layer output reduces each feature map of each channel to a single value. The single value is passed through a non-linear activation unit (ReLU) to create a conditional representation of the unit suitable for the weighting of other channels. The second fully connected layer of block 624 performs the restoration of the conditional channel to the full channel number. Thus, the SE block 624 is able to extract nonlinear inter-channel correlations when generating tensor 627 from tensor 5622 to a greater extent than is possible using a pure convolutional (linear) layer. Tensors 612, 461, and 614 contain 512, 256, and 128 channels, and tensor 622 contains 896 channels. The decorrelation implemented by SE block 624 spans tensor 627, which contains 896 channels. Tensor 627 is passed to convolution layer 628. Convolution layer 628 implements one or more convolution layers to produce a combined tensor 629 in which the number of channels is reduced to F channels (typically 256 channels). After MSFF block 608, SSFC encoder 650 is implemented under execution of processor 205. The operation of SSFC encoder 550 reduces the dimensionality of combined tensor 629 to produce compressed tensor 657. Combined tensor 629 is passed to convolution layer 652 to produce tensor 653. Tensor 653 has a number of channels reduced from 256 to a smaller value C' (such as 64, etc.). The value 96 can also be used for C', which results in reference to Figure 7653 is passed to a batch normalization module 654 to produce a tensor 655. The batch normalized tensor 655 has the same dimensions as the tensor 653. The tensor 655 is passed to a TanH layer 656. The TanH layer 656 implements a hyperbolic tangent (TanH) layer to produce a compressed tensor 557. The use of the hyperbolic tangent (TanH) layer compresses the dynamic range of the values ​​within the tensor 657 to [-1, 1], thereby removing outliers. The compressed tensor 657 has the same dimensions as the tensors 653 and 655. The tensor 657 corresponds to the bottleneck tensor 117.

[0109] Figure 6B The method 670 is shown as being implemented using a device such as a configured FPGA, ASIC, or ASSP. Alternatively, the method 670 may be implemented by the source device 110 as one or more software code modules of the application 233 under execution by the processor 205, as described below. The software code modules of the application 233 implementing the method 600 may reside, for example, in the hard drive 210 and / or the memory 206. The method 670 is repeated for each frame of video data generated by the video source 112. The method 670 may be stored on a computer readable storage medium and / or in the memory 206. The method 670 begins by performing a neural network first portion step 672.

[0110] At step 672, the CNN backbone 114, under execution of the processor 205, performs the first part of the neural network ( Figure 4 Step 672 effectively implements the CNN backbone 114. For example, the following may be performed as in reference Figure 5 The CNN layers are described to generate L2, L1, L0 tensors that form tensor 115. CNN backbone 114 performs other modules (BN and LR modules) of CBL modules 428, 468 to generate tensors (429 and 469) for the next backbone stage. Method 600 continues from step 672 to select tensor step 676 under the control of processor 205.

[0111] At step 676, feature maps (tensors) L2, L1, L0 are extracted or selected. Tensors L2, L1, and L0 are selected from the outputs of the first modules (convolution modules) of CBL modules 428, 468, and 488. In other words, tensors 451, 461, and 491 are selected.

[0112] Steps 672 and 676 operate to obtain a first tensor (115) associated with the image data. Each tensor in the first tensor is obtained using the output of the backbone 114 of the neural network. As described above, the neural network includes at least one layer of the first type that is a CBL layer. Each CBL layer has at least a convolution module and a batch normalization module, and each derived tensor is a tensor that has undergone a convolution module of a CBL layer but has not undergone a batch normalization module.

[0113] Method 670 continues from step 676 to a combine tensors step 680 under the control of processor 205 .

[0114] In step 680, Fig. 6A The MSFF module 608 of the processor 205 combines the tensors in the set of tensors (i.e., 451, 461, and 481) to produce a combined tensor 629. The downsampling module 612 operates on the tensor with the larger spatial scale (i.e., L0 491 at B, 128, h / 8, w / 8) to downsample to match the spatial scale of the smaller tensor (i.e., L1 524 at B, 256, h / 8, w / 8) to produce a downscaled L0 tensor 614. The upsampling module 610 operates on the tensor with the smaller spatial scale (i.e., L2 451 at B, 512, h / 62, w / 32) to downsample to match the spatial scale of the larger tensor (i.e., L1 461 at B, 256, h / 8, w / 8) to produce an upscaled L2 tensor 612. The concatenation module 620 performs channel-by-channel concatenation of tensors 612, 461, and 614 to produce a concatenated tensor 622 of dimensions B, 896, h / 16, w / 16. The concatenated tensor 522 is passed to a squeeze and excitation (SE) module 624 to produce a tensor 627. The SE module 624 sequentially performs global pooling, a fully connected layer with reduced channel number, a rectified linear unit activation, a second fully connected layer with restored channel number, and a sigmoid activation function to produce a scaled tensor. The method 670 continues from step 680 to the SSFC encoding combined tensor step 680 under the control of the processor 205.

[0115] In step 684, as described in relation to Fig. 6A As described above, the SSFC encoder 650 is implemented under execution of the processor 205. The operation of the SSFC encoder 650 reduces the dimension of the combined tensor 629 to produce a compressed tensor 657.

[0116] Steps 680 and 684 operate to perform predetermined processing (such as MFSC, etc.) on the first tensor 115 to derive the tensor 117, the number of dimensions of the data structure of the tensor 115 being greater than the number of dimensions of the data structure of the tensor 117. The method 670 continues from step 684 to a pack compressed tensor step 688 under the control of the processor 205.

[0117] At step 688, the quantization and packing module 118, under execution of the processor 205, quantizes the compressed tensor 657 (117) from the floating point domain to the integer (sample) domain and packs the quantized tensor into a single monochrome video frame. Figure 7 An example single monochrome video frame 700 is shown in FIG. 700 corresponds to frame data 119. Due to the use of the TanH activation function at 656, the range of the compressed tensor 657 is [-1, 1]. The property of TanH in removing outliers results in a distribution suitable for linear quantization of the bit depth of frame 700. The channels of the compressed tensor 657 are packed into feature maps of a certain size, such as feature map 710 in frame 700. The channels of the compressed tensor 657 are packed into feature maps in frame 700. In a reference such as FIG. Fig.15A and Fig. 15B In the arrangement described above with multiple compressed tensors in 657, the packed feature maps of different tensors may differ in width and height. The method 670 continues from step 688 to a compress frame step 692 under the control of the processor 205.

[0118] At step 692, the feature map encoder 120, under execution of the processor 205, encodes the packed frame 119 (eg, frame 700) to generate a bitstream 121. Figure 8 688 and 692 operate to encode the tensor 117 into the bitstream 121. The method 670 terminates when step 692 is executed, wherein the FPN layer of the image frame 312 is reduced in dimension and compressed into the video bitstream 121.

[0119] The described arrangement effectively divides the tensor generated by the CNN backbone 114 into a first set and a second set of several tensors (also referred to as multiple tensors), the tensors in each set having different spatial resolution feature maps from each other. At step 676, the bottleneck encoder 116 selects several adjacent tensors in the tensor 602 as multiple tensors under the execution of the processor 205.

[0120] Figure 8 1 is a schematic block diagram showing the functional modules of the video encoder 120 (also called feature map encoder). The video encoder 120 performs the packetization of the frame 119 (in Figure 7 700) to produce a bitstream 121. Typically, data is passed between functional modules within the video encoder 120 in groups of samples or coefficients (such as a partition of a block into fixed-size sub-blocks, etc.) or as an array. Figure 2A and Figure 2BAs shown, the video encoder 120 can be implemented using a general-purpose computer system 200, wherein various functional modules can be implemented using dedicated hardware within the computer system 200, using software executable within the computer system 200 (such as one or more software code modules of a software application 233 residing on the hard disk drive 205 and controlled by the processor 205 for execution, etc.). Alternatively, the video encoder 120 can be implemented using a combination of dedicated hardware and software executable within the computer system 200. The video encoder 120 and the method can alternatively be implemented in dedicated hardware such as one or more integrated circuits that perform the functions or sub-functions of the method. Such dedicated hardware may include a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific standard product (ASSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or one or more microprocessors and associated memory. In particular, the video encoder 120 includes modules 810 to 890, wherein each of these modules can be implemented as one or more software code modules of the software application 233.

[0121] although Figure 8 The video encoder 120 is an example of a Versatile Video Coding (VVC) video encoding pipeline, but other video codecs may also be used to perform the processing stages described herein. The frame data 119 may be in any chroma format and bit depth supported by the profile in use, for example, 4:0:0, 4:2:0, with a sample precision of eight (8) to ten (10) bits for the "Main 10" profile of the VVC standard.

[0122] The block partitioner 810 first partitions the frame data 119 into CTUs, which are typically square in shape and configured so that a specific size of the CTU is used. The maximum effective size of the CTU can be, for example, 32×32, 64×64, or 128×128 luminance samples, configured by the "sps_log2_ctu_size_minus5" syntax element present in the "sequence parameter set". The CTU size also provides the maximum CU size, because a CTU that is not further split will contain one CU. The block partitioner 810 also partitions each CTU into one or more CBs based on the luminance coding tree and the chrominance coding tree. The luminance channel may also be referred to as the primary color channel. Each chrominance channel may also be referred to as a secondary color channel. CBs have various sizes and may include both square and non-square aspect ratios. However, in the VVC standard, CBs, CUs, PUs, and TUs always have side lengths that are powers of two. Thus, based on the luma coding tree and chroma coding tree of the CTU, a current CB denoted as 812 is output from the block partitioner 810 (advancing based on iteration over one or more blocks of the CTU).

[0123] The CTUs resulting from the first partitioning of the frame data 119 may be scanned in a raster scan order, and may be grouped into one or more "slices". A slice may be an "intra" (or "I") slice. An intra slice (I slice) indicates that each CU in the slice is intra predicted. Typically, the first picture in a coding layer video sequence (CLVS) contains only I slices and is referred to as an "intra picture". A CLVS may contain periodic intra pictures that form "random access points" (i.e., intermediate frames in a video sequence from which decoding may begin). Alternatively, a slice may be uni-predicted or bi-predicted ("P" or "B" slices, respectively), indicating the additional availability of uni-prediction and bi-prediction in a slice, respectively.

[0124] The video encoder 120 encodes a sequence of pictures according to a picture structure. One picture structure is "low delay", in which a picture using inter-frame prediction may only reference pictures that appeared previously in the sequence. Low delay allows each picture to be output as soon as it is decoded, and is also stored for possible reference by subsequent pictures. Another picture structure is "random access", in which the encoding order of the pictures is different from the display order. Random access allows inter-frame predicted pictures to reference other pictures that have been decoded but have not yet been output. A certain degree of picture buffering is required so that future reference pictures in display order are present in the decoded picture buffer, resulting in multiple frame delays.

[0125] When using a chroma format other than 4:0:0, in an I slice, the coding tree for each CTU can diverge below the 64×64 level into two separate coding trees, one for luma and the other for chroma. The use of separate trees allows different block structures to exist between luma and chroma within the luma 64×64 region of the CTU. For example, a large chroma CB can be co-located with many smaller luma CBs, and vice versa. In a P or B slice, a single coding tree for a CTU defines a block structure common to luma and chroma. The resulting blocks of a single tree can be intra-predicted or inter-predicted.

[0126] In addition to partitioning a picture into slices, a picture can also be partitioned into "tiles". A tile is a sequence of CTUs covering a rectangular area of ​​the picture. CTU scanning is performed within each tile in a raster scan, advancing from one tile to the next. A slice can be an integer number of tiles, or an integer number of consecutive CTU rows within a given tile.

[0127] For each CTU, the video encoder 120 operates in two stages. In the first stage (referred to as the "search" stage), the block partitioner 810 tests various potential configurations of the coding tree. Each potential configuration of the coding tree has an associated "candidate" CB. The first stage involves testing various candidate CBs to select a CB that provides relatively high compression efficiency and relatively low distortion. The test typically involves Lagrangian optimization, whereby the candidate CB is evaluated based on a weighted combination of rate (i.e., coding cost) and distortion (i.e., error relative to the input frame data 119). The "best" candidate CB (i.e., the CB with the lowest evaluated rate / distortion) is selected for subsequent encoding into the bitstream 121. Included in the evaluation of the candidate CB is the option of using a CB for a given region, or further splitting the region according to various splitting options and encoding each smaller resulting region with a further CB, or even further splitting the region. Therefore, both the coding tree and the CB itself are selected in the search stage.

[0128] The video encoder 120 generates a prediction block (PB) indicated by arrow 820 for each CB (e.g., CB 812). PB 820 is a prediction of the content of the associated CB 812. Subtractor module 822 generates a difference (or "residual", meaning the difference is in the spatial domain) between PB 820 and CB 812, represented as 824. Difference 824 is the block size difference between corresponding samples in PB 820 and CB 812. Difference 824 is transformed, quantized, and represented as a transform block (TB) indicated by arrow 836. PB 820 and associated TB 836 are typically selected from one of a plurality of possible candidate CBs, for example, based on an assessed cost or distortion.

[0129] A candidate coding block (CB) is a CB obtained from one of the prediction modes available to the video encoder 120 for the associated PB and the resulting residual. When combined with a predicted PB in the video encoder 120, the TB 836 reduces the difference between the decoded CB and the original CB 812 at the expense of additional signaling in the bitstream.

[0130] Thus, each candidate coding block (CB) (i.e., a prediction block (PB) in combination with a transform block (TB)) has an associated coding cost (or "rate") and an associated difference (or "distortion"). The distortion of the CB is typically estimated as a difference in sample values, such as a sum of absolute differences (SAD), a sum of squared differences (SSD), or a Hadamard transform applied to the difference, etc. The mode selector 886 uses the difference 824 to determine the estimates obtained from each candidate PB to determine a prediction mode 887. The prediction mode 887 indicates a decision to use a particular prediction mode (e.g., intra-frame prediction or inter-frame prediction) for the current CB. The estimation of the coding cost associated with each candidate prediction mode and the corresponding residual encoding can be performed at a significantly lower cost than entropy encoding of the residual. Therefore, even in a real-time video encoder, multiple candidate modes can be evaluated to determine the best mode in terms of rate-distortion.

[0131] Determining the best mode based on rate-distortion is typically implemented using a variation of Lagrangian optimization.

[0132] A Lagrangian or similar optimization process may be employed to select both the best partitioning of a CTU into CBs (using the block partitioner 810) and the selection of the best prediction mode from among multiple possibilities. The intra prediction mode with the lowest cost measure is selected as the "best" mode by applying a Lagrangian optimization process of the candidate modes in the mode selector module 886. The lowest cost mode includes the selected secondary transform index 888, which is also encoded into the bitstream 121 by the entropy encoder 838.

[0133] In the second stage of the operation of the video encoder 120 (referred to as the "encoding" stage), the determined coding tree(s) for each CTU are iterated in the video encoder 120. For CTUs using separate trees, the luma coding tree is encoded first, followed by the chroma coding tree, for each 64x64 luma region of the CTU. Within the luma coding tree, only the luma CBs are encoded, and within the chroma coding tree, only the chroma CBs are encoded. For CTUs using shared trees, a single tree describes the CU (i.e., luma CBs and chroma CBs) according to the common block structure of the shared tree.

[0134] The entropy encoder 838 supports bitwise encoding of syntax elements using variable length and fixed length codewords, as well as arithmetic coding modes of syntax elements. Portions of the bitstream such as "parameter sets" (e.g., sequence parameter sets (SPS) and picture parameter sets (PPS)) use a combination of fixed length codewords and variable length codewords. Slices (also called continuous portions) have a slice header using variable length encoding, followed by slice data using arithmetic encoding. The slice header defines parameters specific to the current slice, such as slice-level quantization parameter offsets, etc. The slice data includes syntax elements for each CTU in the slice. The use of variable length encoding and arithmetic coding requires sequential parsing within each portion of the bitstream. These portions can be described with start codes to form "network abstraction layer units" or "NAL units". Arithmetic coding is supported using context-adaptive binary arithmetic coding processing.

[0135] The arithmetically coded syntactic elements consist of a sequence of one or more "bins" (binary files). Like bits, bin values ​​are "0" or "1". However, bins are not encoded in the bitstream 121 as discrete bits. Bins have associated predicted (or "likely" or "maximum probability") values ​​and associated probabilities (known as "context"). When the actual bin to be encoded matches the predicted value, the "maximum probability symbol" (MPS) is encoded. Encoding the maximum probability symbol is relatively cheap in terms of consumed bits in the bitstream 121, including a total cost of less than one discrete bit. When the actual bin to be encoded does not match the possible value, the "minimum probability symbol" (LPS) is encoded. Encoding the minimum probability symbol has a relatively high cost in terms of consumed bits. The bin encoding technique enables efficient encoding of bins that skew the probability of "0" vs "1". For syntactic elements with two possible values ​​(i.e., "flag"), a single bin is sufficient. For syntactic elements with many possible values, a sequence of bins is required.

[0136] The presence of a later bin in a sequence can be determined based on the value of an earlier bin in the sequence. In addition, each bin can be associated with more than one context. The selection of a particular context can depend on the bin values ​​of an earlier bin in a syntax element, and adjacent syntax elements (i.e., adjacent syntax elements from adjacent blocks), etc. Each time a context coding bin is encoded, the context selected for that bin (if present) is updated to reflect the new bin value. Therefore, the binary arithmetic coding scheme is called adaptive.

[0137] The entropy encoder 838 also supports bins lacking context (referred to as "bypass bins"). Bypass bins are encoded assuming an equal probability distribution between "0" and "1". Thus, each bin has a coding cost of one bit in the bitstream 121. The absence of context saves memory and reduces complexity, thus using bypass bins where the distribution of values ​​for a particular bin is not skewed. An example of an entropy encoder that employs context and adaptation is known in the art as CABAC (Context Adaptive Binary Arithmetic Coder), and many variations of this encoder have been employed in video encoding.

[0138] The entropy encoder 838 encodes a quantization parameter 892 using a combination of context-encoded and bypass-encoded bins, and, if used for the current CB, an LFNST index 888. The quantization parameter 892 is encoded using a "delta QP" generated by the QP controller module 890. The delta QP is signaled at most once in each region known as a "quantization group". The quantization parameter 892 is applied to the residual coefficients of the luma CB. The adjusted quantization parameter is applied to the residual coefficients of the collocated chroma CB. The adjusted quantization parameter may include mapping from the luma quantization parameter 892 according to a mapping table and a CU-level offset selected from an offset list. The secondary transform index 888 is signaled when the residual associated with the transform block includes valid residual coefficients only in those coefficient positions that are transformed into primary coefficients by applying a secondary transform.

[0139] The residual coefficients of each TB associated with a CB are encoded using a residual syntax. The residual syntax is designed to efficiently encode coefficients with low amplitudes, primarily using arithmetic coded bins to indicate the significance of the coefficients and the amplitude of lower values, and reserving bypass bins for residual coefficients of higher amplitudes. Therefore, residual blocks consisting of very low amplitude values ​​and sparse placement of significant coefficients are efficiently compressed. In addition, there are two residual coding schemes. As seen when the transform is applied, the conventional residual coding scheme is optimized for TBs where significant coefficients are primarily located in the upper left corner of the TB. The transform skip residual coding scheme can be used for TBs that are not transformed, and is able to efficiently encode residual coefficients regardless of their distribution throughout the TB.

[0140] The multiplexer module 884 outputs the PB 820 from the intra prediction module 864 according to the determined best intra prediction mode selected from the test prediction modes of each candidate CB. The candidate prediction modes need not include every conceivable prediction mode supported by the video encoder 120. Intra prediction is divided into three types: first, "DC intra prediction", which involves filling the PB with a single value representing the average of nearby reconstructed samples; second, "plane intra prediction", which involves filling the PB with samples according to a plane, using a DC offset and vertical and horizontal gradients derived from nearby reconstructed neighboring samples. The nearby reconstructed samples typically include a row of reconstructed samples above the current PB (extending to a certain extent to the right of the PB) and a column of reconstructed samples to the left of the current PB (extending to a certain extent downward beyond the PB); and third, "angle intra prediction", which involves filling the PB with reconstructed neighboring samples filtered and propagated across the PB in a particular direction (or "angle"). In VVC, sixty-five (65) angles are supported, with rectangular blocks being able to take advantage of additional angles not available to square blocks, yielding a total of eighty-seven (87) angles.

[0141] A fourth type of intra prediction can be used for chroma PBs, whereby the PBs are generated from collocated luma reconstructed samples according to a "cross component linear model" (CCLM) mode. Three different CCLM modes are available, each using a different model derived from neighboring luma and chroma samples. The derived model is used to generate a block of samples for the chroma PBs from collocated luma samples. Matrix multiplication of reference samples can be used to intra predict luma blocks using a matrix selected from a predefined set of matrices. This matrix intra prediction (MIP) achieves gains by using matrices trained on a large set of video data, where the matrices represent the relationship between reference samples and prediction blocks that is not easily captured in angular, planar, or DC intra prediction modes.

[0142] Module 864 can also generate prediction units by copying blocks from the vicinity of the current frame using the "intra-block copy" (IBC) method. The location of the reference block is constrained to be an area equivalent to one CTU, which is divided into 64×64 regions called VPDUs, where the area covers the processed VPDU of the current CTU, and the VPDUs of (one or more) previous CTUs within each row or CTU and within each slice or tile, regardless of the configured CTU size for the bitstream, until the area corresponding to one 128×128 luma samples is limited. This area is called the "IBC virtual buffer" and limits the IBC reference area, thereby limiting the required storage. The IBC buffer is filled with reconstructed samples 854 (i.e., before loop filtering), so a buffer separate from the frame buffer 872 is required. When the CTU size is 128×128, the virtual buffer includes samples only from the CTU adjacent to the left of the current CTU. When the CTU size is 32×32 or 64×64, the virtual buffer includes CTUs from up to four or sixteen CTUs to the left of the current CTU. Regardless of the CTU size, access to neighboring CTUs to obtain samples of the IBC reference block is constrained by boundaries (such as the edges of pictures, slices, or tiles). In particular, for feature maps of FPN layers with smaller sizes, using CTU sizes such as 32×32 or 64×64 results in more aligned reference areas to cover the set of previous feature maps. Accessing similar feature maps for IBC prediction provides the advantage of efficient coding when feature map placement is sorted based on SAD, SSE, or other difference metrics.

[0143] The residual of the prediction block when encoding feature map data is different from the residual seen for natural video. Such natural video is typically captured by a camera sensor or screen content, as commonly seen in operating system user interfaces, etc. Feature map residuals tend to contain many details, which are more suitable for transform skip coding than the significant low-frequency coefficients of various transforms. Experiments show that feature map residuals have sufficient local similarity to benefit from transform coding. However, the distribution of feature map residual coefficients does not cluster toward the DC (upper left) coefficient of the transform block. In other words, when encoding feature map data, there is enough correlation to make the transform show gain, and this also applies when using intra-block replication to generate prediction blocks of feature map data. Therefore, when encoding feature map data, when evaluating the residuals generated by candidate block vectors for intra-block replication, Hadamard cost estimates can be used instead of relying solely on SAD or SSD cost estimates. SAD or SSD cost estimates tend to select block vectors with residuals that are more suitable for transform skip coding, and may miss block vectors with residuals that will be compactly encoded using transforms. When encoding feature map data, the Multiple Transform Selection (MTS) tool of the VVC standard may be used so that, in addition to the DCT-2 transform, a combination of DCT-7 and DCT-8 transforms may also be used for residual coding horizontally and vertically.

[0144] An intra-predicted luma coding block can be partitioned vertically or horizontally into a set of equally sized prediction blocks, each with a minimum area of ​​sixteen (16) luma samples. This intra sub-partitioning (ISP) approach enables a single transform block to contribute to the prediction block generation from one sub-partition to the next in the luma coding block, thereby improving compression efficiency.

[0145] In cases where previously reconstructed neighboring samples are not available, such as at the edge of a frame, a default halftone value of half the sample range is used. For example, for 10-bit video, a value of five hundred twelve (512) is used. Since no previous samples are available for the CB located at the top left position of the frame, the angular and planar intra prediction modes produce the same output as the output of the DC prediction mode (i.e., a flat plane of samples with halftone values ​​as amplitudes).

[0146] For inter-frame prediction, a prediction block 882 is generated by a motion compensation module 880 using samples from one or two frames before the current frame in the coding order frame in the bitstream, and the prediction block 882 is output by a multiplexer module 884 as a PB820. In addition, for inter-frame prediction, a single coding tree is typically used for both the luminance channel and the chrominance channel. The order of the coded frames in the bitstream may be different from the order of the frames when captured or displayed. When one frame is used for prediction, the block is called "single prediction" and has one associated motion vector. When two frames are used for prediction, the block is called "double prediction" and has two associated motion vectors. For P slices, each CU can be intra-predicted or uni-predicted. For B slices, each CU can be intra-predicted, uni-predicted, or bi-predicted.

[0147] Frames are typically encoded using a "group of pictures" (GOP) structure to achieve a temporal hierarchy of frames. Frames can be divided into multiple slices, each of which encodes a portion of a frame. The temporal hierarchy of frames allows frames to reference previous and next pictures in the order in which the frames are displayed. Images are encoded in the necessary order to ensure that the correlation for decoding each frame is met. Affine inter-frame prediction mode is available, where instead of using one or two motion vectors to select and filter the reference sample blocks of a prediction unit, the prediction unit is divided into multiple smaller blocks and a motion field is generated so that each smaller block has a different motion vector. The motion field uses the motion vectors of nearby points of the prediction unit as "control points". Affine prediction allows encoding of motions other than translation, with less need for coding trees that use depth splitting. The dual prediction mode available with VVC performs geometric blending of two reference blocks along the selected axis, and signals the angle and offset relative to the center of the block. This geometric partitioning mode ("GPM") allows the use of larger coding units along the boundary between two objects, and the geometry of the boundary for the coding of the coding unit is used as the angle and center offset. Motion vector differences can be encoded as direction (up / down / left / right) and distance (a set of power-of-2 distances are supported) instead of using Cartesian (x,y) offsets. Motion vector predictors are obtained from neighboring blocks ("merge mode") as if no offset was applied. The current block will share the same motion vector as the selected neighboring block.

[0148] Samples are selected based on the motion vector 878 and the reference picture index. The motion vector 878 and the reference picture index are applied to all color channels, so inter-frame prediction is described mainly in terms of operations on PUs rather than PBs. The decomposition of each CTU into one or more inter-frame prediction blocks is described with a single coding tree. The inter-frame prediction method can vary in the number of motion parameters and their precision. The motion parameters typically include a reference frame index that indicates which reference frame(s) in the reference frame list will be used plus the spatial translation of each reference frame, but may include more frames, specific frames, or complex affine parameters (such as scaling and rotation, etc.). In addition, a predetermined motion refinement process can be applied to generate a dense motion estimate based on the reference sample block.

[0149] The PB 820 has been determined and selected, and the PB 820 is subtracted from the original sample block at the subtractor 822 to obtain a residual with the lowest coding cost (denoted as 824), and the residual is lossily compressed. The lossy compression process includes the steps of transform, quantization and entropy coding. The forward main transform module 826 applies a forward transform to the difference 824, converts the difference 824 from the spatial domain to the frequency domain, and produces the main transform coefficients represented by the arrow 828. The maximum main transform size in one dimension is a 32-point DCT-2 or 64-point DCT-2 transform configured by the "sps_max_luma_transform_size_64_flag" in the sequence parameter set. If the CB being encoded is larger than the maximum supported main transform size (e.g., 64×64 or 32×32) represented as a block size, the main transform 826 is applied in a tiled manner to transform all samples of the difference 824. In the case of using a non-square CB, the tiling is also performed using the maximum available transform size in each dimension of the CB. For example, when using a maximum transform size of thirty-two (32), a 64×16 CB uses two 32×16 primary transforms arranged in a tiled manner. When the size of the CB is larger than the maximum supported transform size, the CB is filled with TBs in a tiled manner. For example, a 128×128 CB with a 64-pt transform maximum size is filled with four 64×64 TBs in a 2×2 arrangement. A 64×128 CB with a 32-pt maximum size is filled with eight 32×32 TB transforms in a 2×4 arrangement.

[0150] The application of the transform 826 results in multiple TBs for the CBs. When each application of the transform operates on a TB of difference 824 greater than 32×32 (e.g., 64×64), all resulting main transform coefficients 828 outside the upper left 32×32 region of the TB are set to zero (i.e., discarded). The remaining main transform coefficients 828 are passed to a quantizer module 834. The main transform coefficients 828 are quantized according to a quantization parameter 892 associated with the CB to produce main transform coefficients 832. In addition to the quantization parameter 892, the quantizer module 834 may also apply a "scaling list" to allow non-uniform quantization within a TB by further scaling the residual coefficients according to their spatial position within the TB. The quantization parameter 892 may be different for the luma CB and each chroma CB. The main transform coefficients 832 are passed to a forward secondary transform module 830 to produce transform coefficients represented by arrow 836 by performing a non-separable secondary transform (NSST) operation or bypassing the secondary transform. The forward main transform is typically separable, transforming a set of rows and then a set of columns of each TB. For luma TBs with a width and height not exceeding 16 samples, the forward main transform module 826 uses a type II discrete cosine transform (DCT-2) in the horizontal and vertical directions, or a bypass of the transform in the horizontal and vertical directions, or a combination of a type VII discrete sine transform (DST-7) and a type VIII discrete cosine transform (DCT-8) in the horizontal or vertical directions. In the VVC standard, the combination of using DST-7 and DCT-8 is called a "multi-transform selection set" (MTS).

[0151] The forward secondary transform of module 830 is typically a non-separable transform that is applied only to the residual of the intra-predicted CU and can nevertheless be bypassed. The forward secondary transform operates on sixteen (16) samples (arranged as the upper left 4×4 sub-block of the main transform coefficients 828) or forty-eight (48) samples (arranged as three 4×4 sub-blocks of the upper left 8×8 coefficients of the main transform coefficients 828) to produce a set of secondary transform coefficients. The set of secondary transform coefficients can be fewer in number than the set of main transform coefficients from which they are derived. Since the secondary transform is only applied to a set of coefficients that are adjacent to each other and include the DC coefficient, the secondary transform is called a "low-frequency non-separable secondary transform" (LFNST). This secondary transform can be obtained through a training process and, due to its non-separable nature and trained origin, can exploit additional redundancy in the residual signal that cannot be captured by a separable transform (such as a variant of DCT and DST, etc.). In addition, when LFNST is applied, all remaining coefficients in the TB are zero in both the main transform domain and the secondary transform domain.

[0152] The quantization parameter 892 is constant for a given TB, and thus results in uniform scaling of the generation of residual coefficients in the main transform domain for the TB. The quantization parameter 892 may vary periodically with a signaled "delta quantization parameter". The delta quantization parameter (delta QP) is signaled once for the CU contained in a given region (referred to as a "quantization group"). If the CU is larger than the quantization group size, the delta QP is signaled once using one of the TBs of the CU. That is, the delta QP is signaled once for the first quantization group of the CU by the entropy encoder 838, and not for any subsequent quantization groups of the CU. Non-uniform scaling may also be achieved by applying a "quantization matrix", whereby the scaling factor applied for each residual coefficient is derived from a combination of the quantization parameter 892 and the corresponding entry in the scaling matrix. The scaling matrix may have a size smaller than the size of the TB, and when applied to a TB, a nearest neighbor approach is used to provide a scaled value for each residual coefficient from a scaling matrix of a size smaller than the size of the TB. The residual coefficients 836 are supplied to an entropy encoder 838 for encoding in the bitstream 121. Typically, the residual coefficients of each TB of a TU having at least one valid residual coefficient are scanned according to a scan pattern to produce an ordered list of values. The scan pattern typically scans the TB as a sequence of 4×4 "sub-blocks", thereby providing a conventional scanning operation with a granularity of 4×4 sets of residual coefficients, where the arrangement of the sub-blocks depends on the size of the TB. The scanning within each sub-block and the progression from one sub-block to the next typically follows a reverse diagonal scan pattern. In addition, a quantization parameter 892 is encoded into the bitstream 121 using a differential QP syntax element, and a slice QP and a secondary transform index 888 of an initial value in a given slice or sub-picture are encoded into the bitstream 121.

[0153] As described above, the video encoder 120 needs to access a frame representation that corresponds to the decoded frame representation seen in the video decoder. Therefore, the residual coefficients 836 are passed through an inverse secondary transform module 844, which operates according to the secondary transform index 888 to produce intermediate inverse transform coefficients represented by arrows 842. The intermediate inverse transform coefficients 842 are inversely quantized by a dequantizer module 840 according to a quantization parameter 892 to produce inverse transform coefficients represented by arrows 846. The dequantizer module 840 can also use a scaling list to perform inverse non-uniform scaling of the residual coefficients, which corresponds to the forward scaling performed in the quantizer module 834. The inverse transform coefficients 846 are passed to an inverse main transform module 848 to produce residual samples for the TU (represented by arrows 850). The inverse main transform module 848 applies a DCT-2 transform horizontally and vertically, which is constrained by the maximum available transform size described with reference to the forward main transform module 826. The type of inverse transform performed by inverse secondary transform module 844 corresponds to the type of forward transform performed by forward secondary transform module 830. The type of inverse transform performed by inverse main transform module 848 corresponds to the type of main transform performed by main transform module 826. Summation module 852 adds residual samples 850 and PU 820 to produce reconstructed samples of the CU (indicated by arrow 854).

[0154] The reconstructed samples 854 are passed to the reference sample cache 856 and the in-loop filter module 868. The reference sample cache 856, which is typically implemented using static RAM on an ASIC to avoid expensive off-chip memory accesses, provides the minimum sample storage required to meet the dependencies for generating intra PBs for subsequent CUs in a frame. The minimum dependencies typically include a "line buffer" of samples along the bottom of a row of CTUs for use by the next row of CTUs and a column buffer (whose range is set by the height of the CTU). The reference sample cache 856 supplies reference samples (represented by arrow 858) to the reference sample filter 860. The sample filter 860 applies a smoothing operation to produce filtered reference samples (indicated by arrow 862). The filtered reference samples 862 are used by the intra prediction module 864 to produce an intra prediction block of samples represented by arrow 866. For each candidate intra prediction mode, the intra prediction module 864 generates a sample block (i.e., 866). The sample block 866 is generated by the module 864 using techniques such as DC, planar or angular intra prediction. A matrix multiplication approach may also be used to generate sample block 866, with neighboring reference samples as input and a matrix selected by video encoder 120 from a set of matrices, with the selected matrix signaled in bitstream 121 using an index to identify which matrix in the set of matrices is to be used by video decoder 144.

[0155] The in-loop filter module 868 applies several filtering stages to the reconstructed samples 854. The filtering stages include a "deblocking filter" (DBF), which applies smoothing aligned with CU boundaries to reduce artifacts caused by discontinuities. Another filtering stage present in the in-loop filter module 868 is an "adaptive loop filter" (ALF), which applies a Wiener-based adaptive filter to further reduce distortion. Another available filtering stage in the in-loop filter module 868 is a "sample adaptive offset" (SAO) filter. The SAO filter operates by first classifying the reconstructed samples into one or more categories and applying an offset at the sample level according to the assigned category.

[0156] Filtered samples, represented by arrow 870, are output from the in-loop filter module 868. The filtered samples 870 are stored in a frame buffer 872. The frame buffer 872 typically has the capacity to store several (e.g., up to sixteen (16)) pictures and is therefore stored in the memory 206. Due to the large memory consumption required, on-chip memory is typically not used to store the frame buffer 872. As such, access to the frame buffer 872 is expensive in terms of memory bandwidth. The frame buffer 872 provides a reference frame (represented by arrow 874) to the motion estimation module 876 and the motion compensation module 880.

[0157] The motion estimation module 876 estimates a plurality of "motion vectors" (denoted as 878), each of which is a Cartesian spatial offset relative to the position of the current CB, thereby referencing a block in one of the reference frames in the frame buffer 872. A filtered block of reference samples (denoted as 882) is generated for each motion vector. The filtered reference samples 882 form a further candidate mode for potential selection by the mode selector 886. In addition, for a given CU, the PB 820 may be formed using one reference block ("uni-prediction"), or may be formed using two reference blocks ("bi-prediction"). For the selected motion vector, the motion compensation module 880 generates the PB 820 according to a filtering process that supports sub-pixel precision in the motion vector. In this way, the motion estimation module 876 (which operates on many candidate motion vectors) can perform a simplified filtering process compared to the motion compensation module 880 (which operates only on the selected candidate) to achieve reduced computational complexity. When the video encoder 120 selects inter-frame prediction for the CU, the motion vector 878 is encoded into the bitstream 121.

[0158] Although the reference to Versatile Video Coding (VVC) describes Figure 8 The video encoder 120 of FIG. 8 is shown, but other video coding standards or implementations may also use the processing stages of modules 810 to 890. The video encoder 120 of FIG. 8 may also be used from the memory 206, the hard disk drive 210, the CD-ROM, the Blu-ray Disc TM) or other computer-readable storage medium (or write to the memory 206, hard drive 210, CD-ROM, Blu-ray disc or other computer-readable storage medium). In addition, the frame data 119 (and the bitstream 121) can be received from (or sent to) an external source (such as a server or radio frequency receiver connected to the communication network 220). The communication network 220 may provide limited bandwidth, so that rate control needs to be used in the video encoder 120 to avoid saturating the network when the frame data 119 is difficult to compress. In addition, the bitstream 121 can be constructed from one or more slices, which represent spatial portions (a set of CTUs) of the frame data 119, which are generated by one or more instances of the video encoder 120 operating in a coordinated manner under the control of the processor 205. The bitstream 121 may also include a slice corresponding to a picture to be output as a collection of sub-pictures forming a picture (each sub-picture is independently coded and independently decodable relative to other slices or any slice or sub-picture in the picture).

[0159] exist Fig. 9 The video decoder 144 (also called a feature map decoder) is shown in FIG. Fig. 9 The video decoder 144 of is an example of a Versatile Video Coding (VVC) video decoding pipeline, but other video codecs may also be used to perform the processing stages described herein. Fig. 9 As shown, the bitstream 143 is input to the video decoder 144. The bitstream 143 can be read from the memory 206, the hard drive 210, the CD-ROM, the Blu-ray disc or other non-transitory computer-readable storage medium. Alternatively, the bitstream 143 can be received from an external source (such as a server or a radio frequency receiver connected to the communication network 220). The bitstream 143 contains encoded syntax elements representing the captured frame data to be decoded.

[0160] The entropy decoder module 920 applies an arithmetic coding algorithm, such as "context adaptive binary arithmetic coding" (CABAC), to decode syntax elements from the bitstream 143. The decoded syntax elements are used to reconstruct parameters within the video decoder 144. The parameters include residual coefficients (represented by arrow 924), quantization parameters 974, secondary transform indexes 970, and mode selection information (represented by arrow 958) such as intra-frame prediction modes. The mode selection information also includes information such as motion vectors and partitioning of each CTU into one or more than one CB. The parameters are used to generate PBs, typically in combination with sample data from previously decoded CBs.

[0161] The residual coefficients 924 are passed to the inverse secondary transform module 936, where the secondary transform is applied or not performed (bypassed) according to the secondary transform index. The inverse secondary transform module 936 generates reconstructed transform coefficients 932, i.e., main transform domain coefficients, from the secondary transform domain coefficients. The reconstructed transform coefficients 932 are input to the dequantizer module 928. The dequantizer module 928 inverse quantizes (or "scales") the residual coefficients 932 (i.e., in the main transform coefficient domain) to create a reconstructed intermediate transform coefficient represented by arrow 940 according to the quantization parameter 974. The dequantizer module 928 can also apply a scaling matrix to provide non-uniform dequantization within the TB, which corresponds to the operation of the dequantizer module 840. If the use of a non-uniform inverse quantization matrix is ​​indicated in the bitstream 143, the video decoder 144 reads the quantization matrix from the bitstream 143 as a sequence of scaling factors and arranges the scaling factors into a matrix. Inverse scaling uses the quantization matrix in combination with the quantization parameter to create the reconstructed intermediate transform coefficients 940.

[0162] The reconstructed transform coefficients 940 are passed to an inverse main transform module 944. Module 944 transforms the coefficients 940 from the frequency domain back to the spatial domain. The inverse main transform module 944 applies an inverse DCT-2 transform horizontally and vertically, subject to the constraints of the maximum available transform size as described with reference to the forward main transform module 726. The result of the operation of module 944 is a block of residual samples represented by arrow 948. The residual sample block 948 is equal in size to the corresponding CB. The residual samples 948 are supplied to a summation module 950.

[0163] At summation module 950, residual samples 948 are added to the decoded PB (denoted as 952) to produce a block of reconstructed samples represented by arrow 956. Reconstructed samples 956 are supplied to a reconstructed sample cache 960 and an in-loop filtering module 988. The in-loop filtering module 988 produces a reconstructed block of frame samples denoted as 992. Frame samples 992 are written to a frame buffer 996.

[0164] The reconstructed sample cache 960 operates in a manner similar to the reconstructed sample cache 856 of the video encoder 120. The reconstructed sample cache 960 provides storage for the reconstructed samples required for intra prediction of subsequent CBs without the memory 206 (e.g., by using data 232, which is typically on-chip memory, instead). Reference samples, represented by arrows 964, are obtained from the reconstructed sample cache 960 and are supplied to a reference sample filter 968 to produce filtered reference samples, represented by arrows 972. The filtered reference samples 972 are supplied to an intra prediction module 976. The module 976 produces a block of intra prediction samples, represented by arrows 980, based on the intra prediction mode parameters 958 signaled in the bitstream 143 and decoded by the entropy decoder 920. The intra prediction module 976 supports the modes of module 764, including IBC and MIP. The sample block 980 is generated using a mode such as DC, planar or angular intra prediction.

[0165] When the prediction mode of the CB is indicated in the bitstream 143 to use intra prediction, the intra prediction samples 980 form the decoded PB 952 via the multiplexer module 984. Intra prediction produces a prediction block (PB) of samples, which is a block in one color component that is derived using "neighboring samples" in the same color component. Neighboring samples are samples that are adjacent to the current block and have been reconstructed because they are at the front in the block decoding order. In the case where the luma block and the chroma block are collocated, the luma block and the chroma block can use different intra prediction modes. However, the two chroma CBs share the same intra prediction mode.

[0166] When the prediction mode of the CB is indicated as inter-prediction in the bitstream 143, the motion compensation module 934 generates a block of inter-prediction samples, denoted as 938. The inter-prediction sample block 938 is generated by selecting and filtering the sample block 998 from the frame buffer 996 using the motion vector and reference frame index decoded from the bitstream 143 by the entropy decoder 920. The sample block 998 is obtained from a previously decoded frame stored in the frame buffer 996. For bi-prediction, two sample blocks are generated and mixed together to generate samples for the decoded PB 952. The frame buffer 996 is filled with filter block data 992 from the in-loop filtering module 988. Like the in-loop filtering module 868 of the video encoder 120, the in-loop filtering module 988 applies any of the DBF, ALF, and SAO filtering operations. Generally, the motion vector is applied to both the luma and chroma channels, although the filtering process for sub-sample interpolation in the luma and chroma channels is different. Frames from frame buffer 996 are output as decoded frames 145 .

[0167] exist Figure 8 and Fig. 9Not shown is a module for pre-processing the video prior to encoding and post-processing the video after decoding to shift sample values ​​so that more uniform use of the sample value range within each chroma channel is achieved. A multi-segment linear model is derived in the video encoder 120 and signaled in the bitstream for use by the video decoder 804 to undo the sample shifting. The linear model chroma scaling (LMCS) tool provides compression benefits for specific color spaces and content that have some non-uniformity in their utilization of sample space (especially limited range utilization) that may result in higher quality loss from the application of quantization.

[0168] Fig.10 1 is a schematic block diagram of an example architecture 1000 showing functional modules for performing decoding as implemented in bottleneck decoder 148 and CNN head 150. Architecture 1000 includes a multi-scale feature decoder 1004, a portion of a CBL decoder block 1050, a portion of a CBL decoder block 1052, and a portion of a CBL decoder block 1054. CBL 1050 can be combined with portion 450 to perform CBL layer 428 in an end-to-end network. Decoder block 1052 can be combined with block 460 to perform CBL layer 448 in an end-to-end network. Decoder block 1054 can be combined with block 490 to perform CBL layer 488 in an end-to-end network.

[0169] The input tensor 1002 from the unpacking and inverse quantization 146 corresponds to the tensor 147. The input 1002 is passed through the multi-scale feature decoder 1004 to produce three tensors L'2 1012, L'1 1022, and L'0 1032. The three tensors L'2, L'1, and L'0 have the same dimensions as the tensors L2 451, L1 461, and L0 491, respectively.

[0170] 1014 of block 1050 to produce tensor 1015. Tensor 1015 is input to Leaky ReLU module 1016 of block 1050 to produce tensor 1018. Tensor L'1 is passed through batch normalization 1024 of block 1052 to produce tensor 1025. Tensor 1025 is input to Leaky ReLU block 1026 of block 1052 to produce tensor 1028. Tensor L'0 is passed through batch normalization 1034 of block 1054 to produce tensor 1035. Tensor 1035 is input to Leaky ReLU module 1036 of block 1054 to produce tensor 1038. The two tensors 1018, 1028 are related to tensors 429, 469, respectively.

[0171] If the PCA decoder method is used in the multi-scale feature decoder 1004, the input tensor 1002 includes feature vectors and coefficients. Function module 1004 is performed to restore three tensors through the PCA algorithm. If the tensor L2 is upsampled before PCA encoding, in block 1004, the decoded PCA tensor is downsampled to reconstruct a tensor L'2 having the same size as L2.

[0172] Fig.11 1 is a schematic block diagram showing the functional modules of the multi-scale feature decoder 1100, thereby providing an implementation example of the multi-scale feature decoder 1004. Block 1100 includes two blocks, SSFC decoder 1110 and MSFR 1130. Input tensor 1111 is related to output tensor 657 from multi-scale feature encoder 600 and corresponds to unpacking and inverse tensor 147. Tensor 1111 is input to single-scale feature compression SSFC decoder 1110. Tensor 1111 having a size of (B, C', h / 16, w / 16) is passed through convolution module 1112 to recover the number of channels F and produce tensor 1113. Tensor 1113 has a size of (B, F, h / 16, w / 16). Tensor 1113 is passed through batch normalization 1114 to produce tensor 1115. The size of tensor 1115 is the same as that of tensor 1113. Tensor 1115 is passed through activation block PReLU 1116. Block 1116 implements an activation function such as PReLU or TanH, or Leaky ReLU, etc.

[0173] The output of block 1110 is tensor 1117. Tensor 1117 is input to multi-scale feature reconstruction MSFR 1130 to produce a reconstruction tensor of tensors L2, L1 and L0. Tensor 1117 may have the same spatial dimensions as those of tensor L2, and tensor 1117 corresponds to the reconstruction tensor L'1 (1022) of tensor L1. In order to produce a tensor with a higher number of channels and a smaller value of spatial dimension, a convolution module with a stride value greater than 1 may be used. Tensor L2 has a spatial size that is doubled in height and width compared to the spatial size of tensor L1. A convolution module 1136 with a stride of 2 is used. The output of module 1136 is tensor 1012L'2. Tensor L'2 has the same size as that of L2.

[0174] In order to produce a tensor with a smaller number of channels and a higher value of spatial dimension, a transpose (deconvolution) module with a stride value greater than 1 can be used. Tensor L0 has a spatial size that is halved in height and width compared to the spatial size of tensor L1. A transpose module 1146 with a stride of 2 is used. The output of module 1146 is tensor 1032L'0. Tensor L'0 has the same size as that of L0. Tensor 1160 is three tensors L'2, L'1 and L'0 that provide tensor 149.

[0175] Fig.12 A method for decoding a bitstream, reconstructing a decorrelated feature map, and performing the second part of a CNN is shown. The method 1200 may be implemented using a device such as a configured FPGA, ASIC, or ASSP. Alternatively, as described below, the method 1200 may be implemented by the destination device 140 as one or more software code modules of the application 233 under execution of the processor 205. The software code modules of the application 233 implementing the method 1200 may reside, for example, in the hard drive 210 and / or the memory 206. The method 1200 is repeated for each compressed data frame in the bitstream 143. The method 1200 may be stored on a computer-readable storage medium and / or in the memory 206. The method 1200 begins with a decode bitstream step 1210.

[0176] At step 1210, feature decoder 144 receives bitstream 143 associated with image data encoded at source device 110. Decoder 144 is described in detail with respect to Fig. 9 The operation is to decode the frame 145 of packed data from the bitstream. The method 1200 continues from step 1210 to extract combined tensor step 1220 under the control of the processor 205 .

[0177] At step 1220, the unpacking and inverse quantization module 146 extracts feature maps for the combined tensor (e.g., 657) from the decoded frame 145 and combines the feature maps to produce the output tensor 149. Each feature map is assigned to a different channel within the tensor 149. At step 1220, the extracted feature maps are inverse quantized from the integer domain to the floating point domain. The inverse quantized feature maps generated by the operation of step 1220 are Fig.11 1111. The method 1200 continues from step 1220 to a SSFC decoding step 1240 under the control of the processor 205.

[0178] At step 1240, the SSFC decoder 1130 is implemented under execution of the processor 205. The SSFC decoder performs a neural network layer to decompress the decoded compressed tensor 1111 to produce a decoded combined tensor 1117 (tensor 149). The convolution layer 1112 receives the tensor 1111 with C'=64 channels and outputs a tensor 1113 with F=256 channels. The tensor 1113 is passed to the batch normalization layer 1114. The batch normalization layer outputs a tensor 1115. The tensor 1115 is passed to the parameterized leaky rectified linear (PReLU) layer 1116. The PReLU layer 1116 outputs a tensor 1117.

[0179] Steps 1240 and 1260 perform predetermined processing (implementing bottleneck decoder 148) on each of the tensors decoded in step 1210 to derive each of tensors 1012, 1022, and 1032. Tensors 1012, 1022, and 1032 correspond to tensors 429, 469, and 491, respectively, as generated by the convolution modules (450, 460, 490) of the CBL modules without performing the batch processing modules of the CBL modules. Fig.11 As described above, the tensor 1160 derived by the operation of the MSFC decoder 150 has a smaller number of dimensions than the decoded tensor 147 .

[0180] Method 1200 continues from step 1240 to a reconstruct tensor step 1260 under the control of processor 205. At step 1260, multi-scale feature reconstruction (MSFR) module 1130 receives tensor 1117 generated by the operation of step 1240 under the execution of processor 205. Fig.11 As described, step 1260 is executed to generate tensors L'2, L'1, and L'0 as outputs of tensor 149. Method 1200 continues from step 1260 under the control of processor 205 to perform a second neural network portion step 1260.

[0181] At step 1280, CNN head 1300 (150), under execution by processor 205, receives tensors 149 as input and performs the remainder of the neural network implemented by system 100 on these tensors 149. Method 1200 terminates upon implementation of step 1280, thereby having processed tensors associated with a frame of video data. Method 1200 is re-invoked for each frame of video data encoded in bitstream 143.

[0182] Step 1280 includes two main steps 1282 and 1284. In the first step 1282, the reconstructed tensor includes three tensors that may have different spatial resolutions. Thus, the first step 1282 performs a set of modules for batch normalization and leaky ReLU (activation) for each of the tensors L'0, L'1, and L'0. The batch normalization and leaky ReLU modules belong to the last three CBL layers of the three tracks in the skip connection stage. Batch normalization and leaky ReLU modules 1014, 1016, 1024, and 1026 corresponding to modules 452, 452, 462, and 464 (where 452 and 454 are part of CBL layer 428 and 462 and 464 are part of CBL layer 468) need to be performed in step 1282 to produce tensors 1018 and tensor 1028. Tensors 1018, 1028 are supplied to the next stage of the CNN head 150. Only the batch normalization and Leaky ReLU modules (i.e., modules 1034 and 1036) corresponding to the CBL layer 488 need to be performed in the head at step 1282. Since the split point is implemented at the output of the convolution 490 (that is, at the tensor 491), the batch normalization and Leaky ReLU stages of the CBL layer 488 do not need to be performed in the source device 110, and since the layer 488 corresponds to the last FPN layer, the output of the CBL layer 488 is not used in the skip connection 4004. The method 1200 continues from step 1282 to perform the head step 1284 under the control of the processor 205. At step 1284, the output of each Leaky ReLU module is subjected to a series of head layers to generate detection results and tracking results at different scales of the input image or video frame.

[0183] Fig.13 1 is a schematic block diagram showing the functional modules of a head 1300 of a CNN (which can be used as the CNN head 150). The head 1300 is implemented when executing step 1280. The input tensor of the head is a tensor 1160 corresponding to the tensor 149. The tensor 149 contains three tensors 1012 (L'2), 1022 (L'1) and 1032 (L'0).

[0184] Three tensors 1012, 1022, and 1032 are input to the CNN head 1300. Each tensor is input to the batch normalization module which is the first module in the CNN head 150. At step 1282, the first tensor 1012 is input to the batch normalization 1014 and then passed through the Leaky ReLU module 1016 to produce the tensor 1018. All of the following stages are implemented when step 1284 is executed. The tensor 1018 is passed through the CBL 1350 and then passed through the convolution module 1352 to produce the tensor 1353. The embedding block 1354 contains two modules, the first module is the convolution embedding 1357 and the second module is the cascader 1359. The "embedding" process involves associating the classes of detected objects (such as "person" etc.) in different frames as specific instances of each person. The embedding operates on the slices of the tensor where the detected objects are located and attempts to create associations based on similarity. Typically, the "size" of an embedding refers to the number (and selection) of feature maps used to generate the tensor for that association. Using fewer feature maps to create the association reduces computational complexity, but also reduces the ability to distinguish between different instances of an object. The embedding block 1354 takes tensor 1018 as input and outputs a tensor 1357 containing the number of channels equal to the size of the embedding. In tracking tasks, the size of the embedding should be equal to or higher than the number of objects being tracked. The concatenation block 1359 concatenates the two tensors 1353 and 1357 to produce tensor 1355. Tensor 1355 is input to head Y1 1356.

[0185] At step 1282, the second tensor 1022 is input to batch normalization 1024, and the output is passed through the LeakyReLU module 1026 to produce tensor 1028. All of the following stages are implemented when step 1284 is executed. Tensor 1028 is passed through CBL 1360 and then input to convolution module 1362 to produce tensor 1363. Embedding block 1364 contains two modules, the first module is convolution embedding 1367 and the second module is cascader 1369. Convolution embedding 1367 takes tensor 1028 as input and outputs tensor 1367 containing the number of channels equal to the size of the embedding. Cascade block 1369 cascades the two tensors 1363 and 1367 to produce tensor 1365. Tensor 1365 is input to head Y2 1366.

[0186] At step 1282, the third tensor 1032 is input to batch normalization 1034, and the output is passed through the LeakyReLU module 1036 to produce tensor 1038. Fig.13 As shown, CNN head 1300 performs the batch normalization module of CBL modules 428 , 468 , and 488 , but does not perform the convolution module of CBL modules 428 , 468 , and 488 .

[0187] All of the following stages are implemented when executing step 1284. Tensor 1038 is passed through CBL 1390 and then input to convolution module 1392 to produce tensor 1393. Embedding block 1394 contains two modules, the first module is convolution embedding 1397 and the second module is cascader 1399. Convolution embedding 1397 takes tensor 1038 as input and outputs tensor 1397 containing the number of channels equal to the size of the embedding. Concatenation block 1399 concatenates the two tensors 1393 and 1397 to produce tensor 1395. Tensor 1395 is input to head Y3 1396.

[0188] Each Yolo head Y1, Y2, and Y3 may contain values ​​for a mask, anchor, and number of classes. Each of the heads Y1 1356, Y2 1366, and Y3 1396 includes a suitable YOLO network (e.g., YOLOv3 or YOLOv4), and operates to determine predictions (bounding boxes) Y'1, Y'2, and Y'3. The outputs Y'1, Y'2, and Y'3 are input to a MOTA generator 1310. The MOTA generator uses thresholds of some values ​​such as IoU, NMS (non-maximum suppression), etc. to make a decision on whether the detected box belongs to an object, and generates some tracking metrics such as MOTA, MOTP, etc. The output of the generator 1310 provides an inference result 151.

[0189] like Fig.13 As shown, the neural network head 150 includes multiple detection (CBL) blocks (1350, 1360, and 1390), followed by output convolution blocks (1352, 1362, and 1392, respectively) and the head network, with the first tensor derived at a split point in the network before the detection block. The three detection blocks operate at resolutions and detect objects with different receptive fields (or "anchors") within each resolution to detect objects at different "scales". The output convolution block flattens the output from the detection block and prepares for the final detection decision.

[0190] The described arrangement uses a skip connection network with split points at the output of the convolutional layers in the CBL modules 428, 468, and 488. As a result, both the backbone 4001 and the head 1300 undergo batch normalization modules (452, 462) and leaky ReLU modules (454, 464) from the CBL 428 layer and the CBL 468 layer. Figure 4 It is shown that backbone 4001 performs modules 452 and 454 to produce tensor 429, and performs modules 462 and 464 to produce tensor 469. Tensors 429 and 469 are used to skip the next step in the connection stage. Fig.10 and Fig.13It is shown that the batch normalization module and leaky ReLU are performed again, but in the head, to produce tensors 1018, 1028 for the next step in the head.

[0191] Each convolution module in the embedding block uses a linear function as the activation function. The output of these layers is the probability that each predicted box belongs to one of the objects that need to be tracked or detected.

[0192] In another arrangement of the multi-scale feature encoder 600, as shown in FIG. Fig.14A and Fig. 14B As described, the tensor is downscaled to the spatial size of an L2 tensor 451.

[0193] Fig.14A 14 is a schematic block diagram 1400 showing an alternative multi-scale feature fusion module 1410. The MSFF module 1410 is used in place of the MSFF module 608 at step 680 of the method 670. Module 1410 operates to reduce the spatial area of ​​the tensor 629 to the spatial area of ​​the L2 layer (i.e., tensor 451). Therefore, the spatial area of ​​the tensor 117 (and thus each feature map 710) is also reduced to the spatial area of ​​the tensor 451, thereby reducing the required area of ​​the frame 700 and improving the compression efficiency achieved by the feature map encoder 120. The L0 tensor 491 is passed to a downsampling module 1422, which performs a 4 to 1 downsampling operation horizontally and vertically. The downsampling operation produces a tensor 1423 having 1 / 16 of the area of ​​the tensor 491. The downsampling module 1422 can use methods such as interval elimination or filtering to perform the downsampling operation.

[0194] Downsampling module 1420 also downsamples L1 tensor 461 horizontally and vertically by 2 to 1 using methods such as culling or filtering to produce tensor 1421 having 1 / 4 the area of ​​tensor 461. Concatenation module 1426 concatenates tensors 451, 1421, and 1423 along the channel dimension to produce tensor 1428 having 128+256+512=896 channels (all at the width and height of L2 tensor 451). Tensor 1428 is passed to squeeze and excite module 1430, which is described in reference to Fig. 6A The SE module 624 of FIG. 1434 performs the operations described above to produce a tensor 1432. The convolution module 1434 applies a convolution operation to the tensor 1432 to produce a tensor 1429 corresponding to the tensor 629. The tensor 1429 has F channels, where F is typically 256. The tensor 629 is passed to the SSFC encoder 650, where the remaining operations used for bottleneck encoding are as described in reference. Fig. 6A described.

[0195] As a result of the application of MSFF 1410, smaller feature maps are produced and the bitrate of bitstream 121 is smaller, while still retaining adequate performance for the task of object tracking of people. When tracking people, the object detection of the JDE neural network only needs to recognize one object type (that is, "person"), so higher degrees of compression are possible in the bottleneck encoder and bottleneck decoder without unduly degrading task performance.

[0196] Fig. 14B 14 is a schematic block diagram 1450 illustrating an alternative multi-scale feature reconstruction module 1460. The MSFR module 1460 is used as an alternative to the MSFR module 1130 and is used in conjunction with the use of an alternative MSFF module 1410 in the bottleneck encoder 116 that may be operated at the reconstruct tensor step 1260 of the method 1200. Due to the use of the alternative MSFF module 1410, the tensor 1117 has a width and height corresponding to the L2 layer tensor 451. The tensor 1117 is passed to the convolution module 1464, which produces an L'2 tensor 1012 having 512 channels and having the same width and height as the L2 layer tensor 451. The tensor 1117 is passed to the transposed convolution module 1462 to produce an L'1 tensor 1022 having 256 channels and twice the width and height of the tensor 1117. The transposed convolution 1462 operates as a convolution with half the stride of module 1462, which results in oversampling tensor 1117. Thus, module 1462 produces tensor 1022 having a higher width and height than tensor 1117. Transposed convolutions may also be referred to as "fractional stride convolutions" and may be trained to act as the inverse of the earlier performed correlation to which a stride greater than 1 is applied. The L'1 tensor 1022 is also passed to the transposed convolution module 1469 with half the stride to produce an L'0 tensor 1032 having twice the width and height of the L'1 tensor 1032 and 128 channels. The operation of the MSFR module 1460 results in a reconstruction of the L0 to L2 layers suitable for use by the CNN head 150 (identified as the reconstructed versions of L'0 to L'2).

[0197] Fig.15A 1 is a schematic block diagram illustrating an alternative bottleneck encoder 1500. In an arrangement of source device 110, bottleneck encoder 1500 is used in place of bottleneck encoder 600. Bottleneck encoder 1500 provides a "double-scale" operation, thereby producing two compressed tensors (that is, tensors 657a and 657b). Tensor 657b has twice the width and height of tensor 657a, and provides higher fidelity in tasks that require spatial detail, such as detection of people occupying a small portion of frame 113. In applications where video source 112 has a wide field of view (e.g., due to a high mounting point), the additional spatial detail provided by bottleneck encoder 1500 is beneficial.

[0198] The bottleneck encoder 1500 includes a multi-scale feature fusion module 1510. The L0 tensor 491 is passed to a downsampling module 1530, which produces a tensor 1531 by halving the width and height of the tensor 491 (such as by applying filtering or culling). The L1 tensor 461 is passed to a downsampling module 1512, which produces a tensor 1513 by halving the width and height of the tensor 461, also by using a technique such as culling or filtering. The cascade module 1514 cascades the L2 tensor 451 and the tensor 1513 along the channel dimension to produce a tensor 1515 with 768 channels. The tensor 1515 is passed to a reference Fig. 6A The SE module 624 operates the squeeze and excitation module 1516 to produce a tensor 1517. The tensor 1517 is passed to the convolution module 1518 to produce a tensor 1519 as the first output from the MSFF module 1510, which has 256 channels and the same width and height as the tensor 1517. The cascade module 1532 cascades the L1 tensor 461 and the tensor 1531 along the channel dimension to produce a tensor 1533 with 386 channels. The tensor 1533 is passed to the convolution module 1518 according to the reference Fig. 6A The SE module 624 operates the squeeze and pump module 1534 to produce a tensor 1535. The tensor 1535 is passed to the convolution module 1536 to produce a tensor 1537 having 256 channels and forming the second output of the MSFF module 1510.

[0199] Tensor 1519 is passed to SSFC encoder module 1520 to generate tensor 657a with 64 channels. Tensor 1537 is passed to SSFC encoder module 1540 to generate tensor 657b with 64 channels. SSFC encoder modules 1520 and 1540 can each be based on Fig. 6A The SSFC encoder 650 operates.

[0200] In an arrangement using the bottleneck encoder 1500, step 680 of method 670 is operable to perform the MSFF module 1510 (that is, modules 1512, 1530, 1514, 1532, 1516, 1518, 1534, and 1536). Step 684 of method 670 is operable to perform the SSFC encoders 1520 and 1540, thereby generating tensors 657a and 657b. Step 688 of method 670 is operable to quantize and pack both tensors 657a and 657b into separate non-overlapping regions of frame 700.

[0201] Fig. 15Bis a schematic block diagram showing an alternative bottleneck decoder 1550. The bottleneck decoder 1550 is used to reconstruct the L0 to L2 layers in an arrangement of the destination device 140 corresponding to the arrangement of the source device 110, wherein the compressed tensors are generated by the bottleneck encoder 1500. If the source device 110 uses the encoder 1500, the decoder 1550 can be used instead of the decoder 1100. At step 1220 of the method 1200, extraction and inverse quantization of two tensors 147a and 147b from the frame 700 corresponding to the tensors 657a and 657b generated by the bottleneck encoder 1500 are performed. At step 1240, the SSFC decoder 1560 decodes the tensor 1117a and generates a tensor 1562 having 256 channels. The SSFC decoder 1570 decodes the tensor 1117b and generates an L'1 tensor 1022 having 256 channels. The SSFC decoders 1560 and 1570 can be described as referring to Fig.11 The SSFC decoder 1110 operates as described.

[0202] The multi-scale feature reconstruction module 1580 generates L'0, L'1, and L'2 tensors (that is, 1032, 1022, and 1012), noting that the output of the SSFC decoder 1570 is directly fed as the output L'1 tensor 1022. Tensor 1562 is passed to a convolution module 1564, which generates an L'2 tensor 1012 with 512 channels. L'1 tensor 1022 is passed to a transposed convolution module 1572, which generates a tensor L'0 1032 with twice the width and height of tensor 1022 by using half the fractional stride. The application of module 1580 (that is, modules 1564 and 1572) is performed at step 1260 of method 1200. As a result of using bottleneck encoder 1500 and bottleneck decoder 1550, a greater degree of spatial detail is preserved in tensors 1012, 1022, and 1032 relative to tensors 451, 461, and 491, thereby improving the maximum achievable task performance for tasks such as human object tracking.

[0203] Industrial Applicability

[0204] The described arrangement is suitable for use in the computer and data processing industry, and in particular in digital signal processing for encoding and decoding signals such as video and image signals, thereby achieving high compression efficiency.

[0205] The ability to select different split points of CNN and encode these split points into the bitstream allows flexibility in compression efficiency, because suitable split points can be identified for the trade-off between the desired complexity, bit rate and task performance. In addition, if the backbone neural network is embedded in the edge device, the efficiency of encoding can also be achieved. Since different machine vision tasks can be implemented by different head CNNs, the complexity of the decoder side can be reduced and flexibility can be increased. The efficiency of the decoder side can also be improved by reducing the tensor dimension output by the backbone CNN by selecting the split point. The described arrangement also allows different CNNs (e.g., YOLOv3, YOLOv4 or JDE) to be selected and the selection is encoded in the bitstream, again allowing increased options and flexibility. For example, selecting a lower complexity CNN architecture such as YOLOv3, YOLOv4 and JDE and selecting an appropriate split point can make feature encoding more competitive than traditional encoding solutions. Selecting the split point within the CBL module as the output of the convolutional layer provides the additional benefit of reducing computational complexity (increasing competitiveness) because the values ​​(such as 429 and 469, etc.) that need to be fed back to the head can be considered without having to fully include the CBL blocks (such as 428, 468, and 488, etc.) in both the backbone and head networks. The arrangement described herein uses the example of a YOLOv3 neural network. However, other types of neural networks such as YOLOv4 and JDE networks may also be used.

[0206] The foregoing describes only some embodiments of the present invention, and modifications and / or changes may be made thereto without departing from the scope and spirit of the present invention, the embodiments being illustrative rather than restrictive.

Claims

1. A method for encoding a tensor associated with image data into a bitstream, the method comprising: Obtaining a first tensor for the image data, the first tensor derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, each layer of the first type having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor for which the convolution module of one of the plurality of layers of the first type has been performed but the batch normalization module of the one of the plurality of layers of the first type has not been performed; Performing predetermined processing on the first tensor to derive a second tensor, wherein the number of dimensions of a data structure of the first tensor is greater than the number of dimensions of a data structure of the second tensor; and The second tensor is encoded into the bitstream.

2. The method according to claim 1, wherein: The last layer of the portion of the neural network includes the convolution module.

3. The method according to claim 1, wherein: The neural network also includes multiple detection blocks, each of which is followed by an output convolution block and a head network, and the first tensor is derived at a split point in the neural network before the detection block.

4. The method according to claim 1, wherein: The first portion of the neural network includes a skip connection section including three tracks, and the split point is in the last three layers of the three tracks of the skip connection section.

5. The method according to claim 1, wherein: The neural network is a JDE network, and the first tensor is generated using split points from each of the 79th layer of the neural network, the 94th layer of the neural network, and the 109th layer of the neural network.

6. The method according to claim 1, wherein: The neural network is a JDE network, the first type of layer is a convolutional batch normalization leaky rectified linear layer, i.e., a CBL layer, and the first tensor is generated using a split point after a convolutional layer at a CBL layer for each of the 79th, 94th, and 109th layers of the neural network.

7. A method for deriving a tensor based on a bitstream associated with image data, the derived tensor for processing using a portion of a neural network, the method comprising: decoding a tensor from the bitstream; as well as performing predetermined processing on the decoded tensor to generate a derived tensor, the number of dimensions of the data structure of the derived tensor being greater than the number of dimensions of the data structure of the decoded tensor, The neural network comprises at least a plurality of layers of a first type, each layer of the first type having at least a convolution module and a batch normalization module, and The derived tensor corresponds to a tensor in which the convolution module of one layer of the first type of layers has been processed but the batch normalization module of the one layer of the first type of layers has not been processed.

8. The method according to claim 7, further comprising: The derived tensor is subjected to a batch normalization module of the one of the plurality of layers of the first type without being subjected to a convolution module of the one of the plurality of layers using the portion of the neural network.

9. The method according to claim 7, wherein: The two first layers of the portion of the neural network include a batch normalization and an activation function corresponding to the one of the plurality of layers of the first type.

10. The method according to claim 7, wherein: The portion of the neural network includes a plurality of detection blocks followed by an output convolution block, an embedding block, and a head network, and the first tensor is generated at a split point in the neural network before the detection block.

11. The method according to claim 10, wherein: The neural network further includes a skip connection section, and the split point is in the last three layers of three tracks of the skip connection section.

12. The method according to claim 7, further comprising: The derived tensor is input to the portion of the neural network that generates a result for the machine task.

13. A non-transitory computer-readable storage medium storing a program for executing a method for encoding a tensor associated with image data into a bitstream, the method comprising: Obtaining a first tensor for the image data, the first tensor being derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, the layers of the first type having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor for which the convolution module of one of the plurality of layers of the first type has been performed but the batch normalization module of the one of the plurality of layers of the first type has not been performed; Performing predetermined processing on the first tensor to derive a second tensor, wherein the number of dimensions of a data structure of the first tensor is greater than the number of dimensions of a data structure of the second tensor; and The second tensor is encoded into the bitstream.

14. An encoder configured to: Obtaining a first tensor associated with the image data, the first tensor being derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, the layers of the first type having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor for which the convolution module of one of the plurality of layers of the first type has been performed but the batch normalization module of the one of the plurality of layers of the first type has not been performed; Performing predetermined processing on the first tensor to derive a second tensor, wherein the number of dimensions of a data structure of the first tensor is greater than the number of dimensions of a data structure of the second tensor; and Encode the second tensor into a bitstream.

15. A system comprising: Memory; as well as A processor, wherein the processor is configured to execute code stored on the memory for implementing a method for encoding a tensor associated with image data into a bitstream, the method comprising: Obtaining a first tensor for the image data, the first tensor being derived using a portion of a neural network, the neural network comprising at least a plurality of layers of a first type, the layers of the first type having at least a convolution module and a batch normalization module, wherein the first tensor corresponds to a tensor for which the convolution module of one of the plurality of layers of the first type has been performed but the batch normalization module of the one of the plurality of layers of the first type has not been performed; Performing predetermined processing on the first tensor to derive a second tensor, wherein the number of dimensions of a data structure of the first tensor is greater than the number of dimensions of a data structure of the second tensor; and The second tensor is encoded into the bitstream.

16. A non-transitory computer readable storage medium storing a program for executing a method for deriving a tensor based on a bitstream associated with image data, the derived tensor for processing using a portion of a neural network, the method comprising: decoding a tensor from the bitstream; as well as performing predetermined processing on the decoded tensor to generate a derived tensor, the number of dimensions of the data structure of the derived tensor being greater than the number of dimensions of the data structure of the decoded tensor, The neural network comprises at least a plurality of layers of a first type, wherein the layers of the first type have at least a convolution module and a batch normalization module, and The derived tensor corresponds to a tensor in which the convolution module of one layer of the first type of layers has been processed but the batch normalization module of the one layer of the first type of layers has not been processed.

17. A decoder configured to: decoding a tensor from a bitstream associated with the image data; and performing predetermined processing on the decoded tensor to generate a derived tensor for processing using a portion of a neural network, the derived tensor having a data structure with a greater number of dimensions than the data structure of the decoded tensor, in, The neural network comprises at least a plurality of layers of a first type, the layers of the first type having at least a convolution module and a batch normalization module, and The derived tensor corresponds to a tensor in which the convolution module of one layer of the first type of layers has been processed but the batch normalization module of the one layer of the first type of layers has not been processed.

18. A system comprising: Memory; as well as a processor, wherein the processor is configured to execute code stored on the memory for implementing a method for deriving a tensor based on a bitstream associated with image data, the derived tensor for processing using a portion of a neural network, the method comprising: decoding a tensor from the bitstream; and performing predetermined processing on the decoded tensor to generate a derived tensor, the number of dimensions of the data structure of the derived tensor being greater than the number of dimensions of the data structure of the decoded tensor, The neural network comprises at least a plurality of layers of a first type, wherein the layers of the first type have at least a convolution module and a batch normalization module, and The derived tensor corresponds to a tensor in which the convolution module of one layer of the first type of layers has been processed but the batch normalization module of the one layer of the first type of layers has not been processed.