Neural network decoding method, device, computer equipment and computer readable medium
Through block division technology and 3D pyramid encoding method, the problem of low compression and decompression efficiency of neural network models in the prior art is solved, and efficient resource utilization and model deployment are achieved.
Patent Information
- Application Number
- CN202180005471.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-15
- Filing Date
- 2021-04-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-04-19
AI Technical Summary
The prior art is difficult to effectively compress and decompress neural network models, resulting in waste of storage and computing resources.
The block division technology is adopted to divide the tensors in the neural network model through coding tree unit (CTU) blocks, and compress and decompress using 3D pyramid encoding method and a unified encoding method.
It realizes efficient compression and decompression of neural network models, reduces the consumption of storage and computing resources, and improves the deployment and operation efficiency of the model.
Smart Images

Figure CN114450692B_ABST
Abstract
Description
[0001] Incorporation by reference
[0002] The present disclosure claims priority to U.S. Patent Application No. 17 / 232,069, filed on April 15, 2021, entitled “Neural Network Model Compression Using Block Partitioning,” which claims priority to U.S. Provisional Application No. 63 / 015,213, filed on April 24, 2020, entitled “Block Definitions and Uses for Neural Network Model Compression,” and to U.S. Provisional Application No. 63 / 042,303, filed on June 22, 2020, entitled “3D Pyramid Coding Method for Neural Network Model Compression,” and to U.S. Provisional Application No. 63 / 079,310, filed on September 16, 2020, entitled “Unification-Based Coding Method for Neural Network Model Compression,” and to U.S. Provisional Application No. 63 / 079,706, filed on September 17, 2020, entitled “Unification-Based Coding Method for Neural Network Model Compression.” The disclosures of these prior applications are incorporated herein by reference in their entirety. Technical Field
[0003] This disclosure describes embodiments generally related to neural network model compression / decompression. Background Art
[0004] The background description provided herein is intended to generally present the context of the present disclosure. To the extent described in this background section, the works of the currently named inventors and aspects of the description that do not qualify as prior art at the time of filing are neither explicitly nor implicitly admitted to be prior art to the present disclosure.
[0005] Various applications in the fields of computer vision, image recognition, and speech recognition rely on neural networks to achieve performance improvements. Neural networks are based on a collection of connected nodes (also called neurons) that loosely mimic neurons in a biological brain. Neurons can be organized into multiple layers. Neurons in one layer can be connected to neurons in the immediately previous layer and the immediately following layer.
[0006] A connection between two neurons, such as a synapse in a biological brain, can pass a signal from one neuron to another. The neuron that receives the signal then processes the signal and can send signals to other connected neurons. In some examples, to find the output of a neuron, the inputs to the neuron are weighted by the weights of the connections from the inputs to the neuron, and the weighted inputs are summed to generate a weighted sum. A bias can be added to the weighted sum. In addition, the weighted sum is then passed through an activation function to produce an output. Summary of the invention
[0007] Various aspects of the present invention provide methods and devices for neural network model compression / decompression. In some examples, a neural network model decompression device includes a processing circuit. The processing circuit can be configured to receive a first syntax element in an NNR aggregation unit header of a compressed NNR aggregation unit from a bitstream of a compressed neural network representation (NNR) of a neural network. The first syntax element can indicate a coding tree unit (CTU) scanning order for processing tensors in the NNR aggregation unit. The tensors in the NNR aggregation unit can be reconstructed based on the CTU scanning order indicated by the first syntax element.
[0008] In one embodiment, a first value of the first syntax element may indicate that the CTU scan order is a first raster scan order along a horizontal direction, and a second value of the first syntax element indicates that the CTU scan order is a second raster scan order along a vertical direction. In one embodiment, a second syntax element in an NNR aggregation unit header of the NNR aggregation unit may be received from a bitstream. The second syntax element may indicate a maximum bit depth of quantized coefficients of tensors in the NNR aggregation unit.
[0009] In one embodiment, a third syntax element may be received, the third syntax element indicating whether CTU block partitioning is enabled for tensors in the NNR aggregation unit. The third syntax element may be a model-related syntax element or a tensor-related syntax element, the model-related syntax element is used to specify whether CTU block partitioning is enabled for a layer of a neural network, and the tensor-related syntax element is used to specify whether CTU block partitioning is enabled for a tensor in the NNR aggregation unit.
[0010] In one embodiment, a fourth syntax element related to the model or to the tensor may be received, the fourth syntax element indicating the CTU dimension of the tensor in the NNR aggregation unit. In one embodiment, the NNR unit may be received before receiving any NNR aggregation unit. The NNR unit may include a fifth syntax element indicating whether CTU partitioning is enabled.
[0011] In some examples, another apparatus for decompressing a neural network model includes a processing circuit. The processing circuit may be configured to receive one or more first syntax elements from a bitstream represented by a compressed neural network, the first syntax elements being associated with a three-dimensional coding unit (CU3D) divided from a first three-dimensional coding tree unit (CTU3D). The first CTU3D may be obtained from a tensor partition in a neural network. The one or more first syntax elements may indicate that the CU3D is partitioned based on a 3D pyramid tree structure including multiple depths. Each depth corresponds to one or more nodes. Each node has a node value. A second syntax element corresponding to a node value of a node in a 3D pyramid tree structure may be received from the bitstream in a breadth-first scanning order for scanning nodes in the 3D pyramid tree structure. The model parameters of the tensor may be reconstructed based on the received second syntax element corresponding to the node value of each node in the 3D pyramid tree structure. In various embodiments, the 3D pyramid tree structure is one of an octree structure, a single-fork tree structure, a label tree structure, and a single-fork label tree structure.
[0012] In one embodiment, a second syntax element is received starting from a starting depth in the depth of the 3D pyramid tree structure. The starting depth may be indicated in the bitstream or inferred at the decoder. In one embodiment, a third syntax element may be received, the third syntax element indicating a starting depth for receiving a second syntax element, the second syntax element indicating a node value of a node in the 3D pyramid tree structure. When the starting depth is the last depth of the 3D pyramid tree structure, a model parameter of a tensor may be decoded from the bitstream using a decoding method based on a non-3D pyramid tree.
[0013] In one embodiment, a third syntax element may be received, the third syntax element indicating a starting depth for receiving a second syntax element, the second syntax element indicating a node value of each node in the 3D pyramid tree structure. When the starting depth is the last depth of the 3D pyramid tree structure, the second syntax element is received starting from the second to last depth in the depth of the 3D pyramid tree structure.
[0014] In another embodiment, a third syntax element may be received, the third syntax element indicating a starting depth for receiving a second syntax element, the second syntax element indicating a node value of a node in a 3D pyramid tree structure. When the starting depth is the last depth of the 3D pyramid tree structure and the 3D pyramid tree structure is a single-branch tag tree structure associated with single-branch tree partial encoding and tag tree partial encoding, for the single-branch tree partial encoding, the second syntax element is received starting from the second-to-last depth in the depth of the 3D pyramid tree structure; and for the tag tree partial encoding, the second syntax element is received starting from the last depth in the depth of the 3D pyramid tree structure.
[0015] In one embodiment, when the one or more first syntax elements indicate that the CU3D is partitioned based on a 3D pyramid tree structure, dependent quantization is disabled. In one embodiment, when the one or more first syntax elements indicate that the CU3D is partitioned based on a 3D pyramid tree structure, a dependent quantization construction process may be performed. Model parameters of tensors skipped during the encoding process based on the 3D pyramid tree structure are excluded from the dependent quantization construction process.
[0016] In one embodiment, a fourth syntax element associated with the CU3D may be received, the fourth syntax element indicating whether all model parameters of the CU3D are unified. In one embodiment, a zero value may be used as a value of a forward neighbor of a first coefficient in a kernel of a tensor to determine a context model for entropy decoding the first coefficient in the kernel. In one embodiment, one or more fifth syntax elements may be received in a bitstream, the fifth syntax element indicating a width or height of a second CTU3D of the tensor. When the width, the height, or both the width and the height are a model parameter, it may be determined that the model parameters of the second CTU3D are decoded based on the baseline encoding method.
[0017] In some examples, another apparatus for decompressing a neural network model includes a processing circuit. The processing circuit may be configured to receive a first syntax element in a bitstream of a compressed neural network representation of the neural network, the first syntax element being associated with a CTU3D obtained by partitioning a tensor in a layer of the neural network. The first syntax element may indicate whether all child nodes at a bottom depth of a pyramid tree structure associated with the CTU3D are unified. When the first syntax element indicates that all child nodes at a bottom depth of a pyramid tree structure associated with the CTU3D are unified, the CTU3D may be decoded based on a three-dimensional singular tree (3D singular tree) encoding method.
[0018] In one embodiment, a second syntax element associated with a layer of a neural network may be received in a bitstream. The second syntax element may indicate whether the layer is encoded using a pyramid tree structure-based encoding method. In one embodiment, child nodes that do not share the same parent node at a bottom depth have different uniform values.
[0019] In one embodiment, the starting depth of the 3D single tree coding method can be inferred to be the bottom depth of the pyramid tree structure. In one embodiment, the unified flag of the node at the bottom depth of the pyramid tree structure is not encoded in the bitstream. In one embodiment, a unified value encoded in the bitstream for all child nodes sharing the same parent node at the bottom depth can be received. A sign bit of all child nodes sharing the same parent node at the bottom depth can be received. The sign bit follows the unified value in the bitstream.
[0020] In one embodiment, a uniform value for each group of child nodes sharing the same parent node at the bottom depth may be received from the bitstream.Sign bits for child nodes in each group of child nodes sharing the same parent node at the bottom depth may be received.
[0021] In one embodiment, in response to the first syntax element indicating that all child nodes at the bottom depth of the pyramid tree structure associated with the CTU3D are not all uniform, the CTU3D may be decoded based on a three-dimensional label tree (3D label tree) encoding method. In one embodiment, the starting depth of the 3D label tree encoding method may be inferred to be the bottom depth of the pyramid tree structure.
[0022] In one embodiment, the value of the node at the bottom depth of the pyramid tree structure may be decoded according to one of the following: receiving the value of the node at the bottom depth of the pyramid tree structure, the value of each node being encoded in a bitstream based on a predetermined scanning order; receiving the absolute value of each node at the bottom depth of the pyramid tree structure in the bitstream based on a predetermined scanning order, and then receiving the sign of each node if the absolute value is not zero, or receiving the absolute value of each node at the bottom depth of the pyramid tree structure in the bitstream based on a predetermined scanning order, and then receiving the sign of each node at the bottom depth of the pyramid tree structure in the bitstream based on a predetermined scanning order if the node has a non-zero value.
[0023] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for neural network model decompression, cause the computer to perform a method for neural network model decompression. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Other features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0025] Figure 1 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0026] Figure 2An exemplary compressed neural network representation (NNR) unit is shown.
[0027] Figure 3 Exemplary polymeric NNR units are shown.
[0028] Figure 4 An exemplary NNR bitstream is shown.
[0029] Figure 5 An exemplary NNR unit syntax is shown.
[0030] Figure 6 An exemplary NNR unit header syntax is shown.
[0031] Figure 7 An exemplary NNR aggregation unit header syntax is shown.
[0032] Figure 8 An exemplary NNR unit payload syntax is shown.
[0033] Fig. 9 An exemplary NNR model parameter set payload syntax is shown.
[0034] Fig.10 An exemplary syntax for a model parameter set is shown.
[0035] Fig.11 An exemplary syntax of an aggregation unit header is shown, which includes signaling of a coding tree unit (CTU) scanning order for one or more weight tensors.
[0036] Fig.12 shows an example of the syntax for scanning weight coefficients in a weight tensor.
[0037] Fig.13 An example for decoding the absolute value of a quantized weight coefficient according to some embodiments of the present disclosure is shown.
[0038] Fig.14 Two examples of adaptive three-dimensional CTU (CTU3D) / three-dimensional coding unit (CU3D) partitioning using raster scanning along the vertical direction are shown.
[0039] Fig.15 An exemplary partitioning process based on a 3D pyramid tree structure is shown.
[0040] Fig.16 Two scalar quantizers according to embodiments of the present disclosure are shown.
[0041] Fig.17 The decoding process of CTU partitioning according to one embodiment of the present disclosure is shown.
[0042] Fig.18 The decoding process of 3D pyramid coding according to one embodiment of the present disclosure is shown.
[0043] Fig.19 The decoding process based on unified coding according to one embodiment of the present disclosure is shown.
[0044] Fig. 20 is a schematic diagram of a computer system according to one embodiment of the present disclosure. DETAILED DESCRIPTION
[0045] Various aspects of the present disclosure provide various techniques for neural network model compression / decompression. These techniques involve coding tree unit (CTU) block partitioning, 3D pyramid tree structure-based encoding, and unified-based encoding.
[0046] Artificial neural networks can be used for a wide range of tasks in multimedia analysis and processing, media coding, data analysis, and many other fields. The success of using artificial neural networks is based on the feasibility of processing much larger and more complex neural networks (deep neural networks, DNNs) than in the past and the availability of large-scale training data sets. Therefore, the trained neural network can contain a large number of model parameters, resulting in a considerable capacity (e.g., hundreds of MB). The model parameters may include coefficients of the trained neural network, such as weights, biases, scaling factors, batch normalization (batch norm, batchnorm) parameters, etc. These model parameters can be organized into model parameter tensors. Model parameter tensors are used to refer to multidimensional structures (e.g., arrays or matrices) that combine the relevant model parameters of the neural network. For example, when available, the coefficients of the layers in the neural network can be divided into weight tensors, bias tensors, scaling factor tensors, batch norm tensors, etc.
[0047] Many applications require the potential deployment of a particular trained network instance to a larger number of devices that may have limitations in processing power and storage (e.g., mobile devices or smart cameras) and limitations in communication bandwidth. These applications can benefit from the neural network compression / decompression techniques disclosed herein.
[0048] I. Neural Network-Based Devices and Applications
[0049] Figure 1A block diagram of an electronic device (130) according to one embodiment of the present disclosure is shown. The electronic device (130) may be configured to run a neural network-based application. In some embodiments, the electronic device (130) receives and stores a compressed (encoded) neural network model (e.g., a compressed representation of a neural network in the form of a bitstream). The electronic device (130) may decompress (or decode) the compressed neural network model to restore the neural network model, and may run an application based on the neural network model. In some embodiments, the compressed neural network model is provided from a server such as an application server (110).
[0050] exist Figure 1 In an example, the application server (110) includes a processing circuit (120), a memory (115), and an interface circuit (111) coupled together. In some examples, a neural network is appropriately generated, trained, or updated. The neural network can be stored in the memory (115) as a source neural network model. The processing circuit (120) includes a neural network model codec (121). The neural network model codec (121) includes an encoder that can compress the source neural network model and generate a compressed neural network model (a compressed representation of the neural network). In some examples, the compressed neural network model is in the form of a bit stream. The compressed neural network model can be stored in the memory (115). The application server (110) can provide the compressed neural network model to other devices, such as an electronic device (130), in the form of a bit stream through the interface circuit (111).
[0051] It should be noted that the electronic device (130) may be any suitable device, such as a smartphone, a camera, a tablet computer, a laptop computer, a desktop computer, a gaming headset, etc.
[0052] exist Figure 1 In some examples, the electronic device (130) includes a processing circuit (140), a cache (150), a main memory (160), and an interface circuit (131) coupled together. In some examples, the compressed neural network model is received by the electronic device (130) through the interface circuit (131), for example, in the form of a bit stream. The compressed neural network model is stored in the main memory (160).
[0053] The processing circuit (140) includes any suitable processing hardware, such as a central processing unit (CPU), a graphics processing unit (GPU), etc. The processing circuit (140) includes components suitable for executing a neural network-based application, and includes components suitable for being configured as a neural network model codec (141). The neural network model codec (141) includes a decoder that can decode a compressed neural network model received, for example, from an application server (110). In one example, the processing circuit (140) includes a single chip (e.g., an integrated circuit) on which one or more processors are disposed. In another example, the processing circuit (140) includes multiple chips, each of which can include one or more processors.
[0054] In some embodiments, the main memory (160) has a relatively large storage space and can store various information, such as software code, media data (e.g., video, audio, image, etc.), compressed neural network models, etc. The cache (150) has a relatively small storage space, but has a much faster access speed than the main memory (160). In some examples, the main memory (160) may include a hard disk drive, a solid-state drive, etc., and the cache (150) may include a static random access memory (SRAM), etc. In one example, the cache (150) may be an on-chip memory disposed on, for example, a processor chip. In another example, the cache (150) may be an off-chip memory disposed on one or more memory chips separated from the processor chip. Generally, the on-chip memory has a faster access speed than the off-chip memory.
[0055] In some embodiments, when the processing circuit (140) executes an application using a neural network model, the neural network model codec (141) may decompress the compressed neural network model to restore the neural network model. In some examples, the cache (150) is large enough so that the restored neural network model can be buffered in the cache (150). The processing circuit (140) can then access the cache (150) to use the restored neural network model in the application. In another example, the cache (150) has a limited memory space (e.g., on-chip memory), the compressed neural network model can be decompressed layer by layer or block by block, and the cache (150) can buffer the restored neural network model layer by layer or block by block.
[0056] It should be noted that the neural network model codec (121) and the neural network model codec (141) can be implemented by any suitable technology. In some embodiments, the encoder and / or decoder can be implemented by an integrated circuit. In some embodiments, the encoder and decoder can be implemented as one or more processors that execute a program stored in a non-transitory computer-readable medium. The neural network model codec (121) and the neural network model codec (141) can be implemented according to the encoding and decoding features described below.
[0057] The present disclosure provides a technique for compressing neural network representation (NNR), which can be used to encode and decode neural network models such as deep neural network (DNN) models to save storage and computation. Deep neural networks (DNN) can be used in a wide range of video applications, such as semantic classification, object detection / recognition, object tracking, video quality enhancement, etc.
[0058] A neural network (or artificial neural network) typically includes multiple layers between an input layer and an output layer. In some examples, a layer in a neural network corresponds to a mathematical transformation that converts the input of the layer into the output of the layer. The mathematical transformation can be a linear relationship or a nonlinear relationship. The neural network can move through each layer to calculate the probability of each output. In this way, each mathematical transformation is considered to be a layer, and a complex DNN can have multiple layers. In some examples, the mathematical transformation of a layer can be represented by one or more tensors (e.g., a weight tensor, a bias tensor, a scaling factor tensor, a batch norm tensor, etc.).
[0059] II. Block Definition and Use of Neural Network Model Compression
[0060] 1. Advanced Syntax
[0061] In some embodiments, a high-level syntax for a bitstream carrying a neural network (model) in a compressed or encoded representation may be defined based on the concept of an NNR unit. An NNR unit is a data structure for carrying neural network data and related metadata. An NNR unit carries compressed or uncompressed information related to neural network metadata, topology information, all or part of layer data, filters, kernels, biases, quantized weights, tensors, etc.
[0062] Figure 2 An exemplary NNR unit (200) is shown. As shown, the NNR unit (200) may include the following data elements:
[0063] - NNR Unit Size: This data element signals the total byte size of the NNR Unit including the NNR Unit Size itself.
[0064] - NNR unit header: This data element contains information related to the NNR unit type and related metadata.
[0065] -NNR Unit Payload: This data element contains compressed or uncompressed data related to the neural network.
[0066] Figure 3 An exemplary aggregated NNR unit (300) is shown. An aggregated NNR unit (300) is an NNR unit that can carry multiple NNR units in its payload. The aggregated NNR unit provides a grouping mechanism for several NNR units that are related to each other and benefit from aggregation under a single NNR unit. Figure 4 An exemplary NNR bitstream (400) is shown. The NNR bitstream (400) may include a sequence of NNR units. The first NNR unit in the NNR bitstream may be an NNR start unit (ie, an NNR unit of type NNR_STR).
[0067] Figure 5 An exemplary NNR unit syntax is shown. Figure 6 An exemplary NNR unit header syntax is shown. Figure 7 An exemplary NNR aggregation unit header syntax is shown. Figure 8 An exemplary NNR unit payload syntax is shown. Fig. 9 An exemplary NNR model parameter set payload syntax is shown.
[0068] 2. Reshaping and scanning sequence
[0069] In some examples, the dimension of the weight tensor is greater than 2 (e.g., in a convolutional layer, the dimension is 4), and the weight tensor can be reshaped into a two-dimensional (2D) tensor. In one example, if the dimension of the weight tensor does not exceed 2 (e.g., a fully connected layer or a bias layer), no reshaping is performed. In order to encode the weight tensor, the weight coefficients in the weight tensor are scanned in a certain order. In some examples, the weight coefficients in the weight tensor can be scanned from left to right for each row and from the top row to the bottom row, for example, in a row-first manner.
[0070] 3. Block division mark
[0071] In some embodiments, the weight tensor may be reshaped into a 2D tensor, which is then divided into blocks called coding tree units (CTUs). The coefficients of the 2D tensor may then be encoded based on the resulting CTU blocks. For example, a scan order may be defined based on these CTU blocks, and encoding may be performed according to the scan order. In some other embodiments, the weight tensor may be reshaped into a 3D tensor, for example, with the number of input channels and output channels as the first and second dimensions, respectively, and the number of elements in the kernel (filter) as the third dimension. Block partitioning may then be performed along the planes of the first and second dimensions, resulting in a 3D block called a CTU3D. Thus, encoding of the tensor may be performed based on the scan order of these CTU3Ds. In some examples, CTU blocks or CTU3D blocks may be blocks of equal size.
[0072] In some embodiments, a model-related syntax element ctu_partition_flag is used to specify whether block partitioning (CTU partitioning) is enabled for the weight tensor of each layer of the neural network. For example, a first value of the syntax element (e.g., a value of 0) indicates that block partitioning is disabled, and a second value of the syntax element (e.g., a value of 1) indicates that block partitioning is enabled. In one embodiment, the syntax element is a 1-bit flag. In another embodiment, the syntax element may be represented by multiple bits. One value of the syntax element indicates whether block partitioning is performed. For example, a zero value may indicate that partitioning is not performed. Other values of the syntax element may be used to indicate the size of a CTU or CTU3D block.
[0073] Fig.10 An exemplary syntax of a model parameter set is shown. The model parameter set includes ctu_partition_flag. The ctu_parition_flag can be used to enable or disable block partitioning of the neural network model controlled by the model parameter set.
[0074] In some embodiments, a syntax element ctu_partition_flag associated with a tensor may be used to specify whether block partitioning (CTU partitioning) is enabled for one or more individual weight tensors of a neural network. For example, a first value of the syntax element (e.g., a value of 0) indicates that block partitioning is disabled for the corresponding one or more tensors, and a second value of the syntax element (e.g., a value of 1) indicates that block partitioning is enabled for the corresponding one or more tensors. In one embodiment, the syntax element is a 1-bit flag. In another embodiment, the syntax element may be represented by multiple bits. One value of the syntax element indicates whether block partitioning is performed on the corresponding one or more tensors. For example, a zero value may indicate that partitioning is not performed on the corresponding one or more tensors. Other values of the syntax element may be used to indicate the size of the CTU or CTU3D block partitioned from the corresponding one or more tensors.
[0075] In one example, the syntax element ctu_partition_flag related to the tensor is included in the compressed data unit header. In one example, the syntax element ctu_partition_flag related to the tensor is included in the aggregation unit header.
[0076] 4. Signal representation in CTU dimension
[0077] In one embodiment, when the model-related flag ctu_partition_flag has a value indicating that block partitioning is enabled, the model-related 2-bit max_ctu_dim_flag can be used to specify the model-related maximum CTU dimension (denoted as gctu_dim) of the weight tensor of the neural network. For example, gctu can be determined according to the following formula:
[0078] gctu_dim=(64>>max_ctu_dim_flag).
[0079] For example, corresponding to a max_ctu_dim_flag value of 0, 1, 2, or 3, gctu_dim may have values of 64, 32, 16, and 8.
[0080] In one embodiment, for a 2D reshaped tensor, the maximum CTU width associated with the tensor may be scaled proportionally to the kernel size of each convolution tensor as follows:
[0081] max_ctu_height=gctu_dim,
[0082] max_ctu_width=gctu_dim*kernel_size.
[0083] The height / width of the right / bottom CTU may be less than max_ctu_height / max_ctu_width. It should be noted that the number of bits of max_ctu_dim_flag may be changed to other values (e.g., greater than 2 bits). Other mapping functions involving max_ctu_dim_flag may be used to calculate gctu_dim. max_ctu_width may not scale proportionally with kernel_size (e.g., for dividing CTU 3D blocks). Alternatively, max_ctu_width may scale proportionally with any arbitrary value to form suitable 2D or 3D blocks of various sizes.
[0084] In another embodiment, when the model-related or tensor-related flag ctu_partition_flag has a value indicating that block partitioning is enabled, the 2-bit max_ctu_dim_flag associated with the tensor may be used to specify the maximum CTU dimension associated with the tensor of the corresponding weight tensor of the neural network according to the following formula:
[0085] gctu_dim=(64>>max_ctu_dim_flag).
[0086] In one embodiment, the maximum CTU width associated with a tensor may scale proportionally with the kernel size of each convolution tensor as follows:
[0087] max_ctu_height=gctu_dim,
[0088] max_ctu_width=gctu_dim*kernel_size.
[0089] Similarly, the height / width of the right / bottom CTU may be less than max_ctu_height / max_ctu_width. The number of bits of max_ctu_dim_flag associated with the tensor may be changed to other values (e.g., greater than 2 bits). Other mapping functions involving max_ctu_dim_flag associated with the tensor may be used to calculate gctu_dim. max_ctu_width may not scale proportionally with kernel_size (e.g., for dividing CTU 3D blocks). max_ctu_width may scale proportionally with an arbitrary value to form suitable 2D or 3D blocks of various sizes.
[0090] 5.CTU scanning order
[0091] In some embodiments, the syntax element ctu_scan_order associated with the tensor is used to specify the CTU-associated (or CTU3D) scanning order of the corresponding one or more tensors. For example, the first value (e.g., value 0) of the syntax element ctu_scan_order associated with the tensor indicates that the scanning order associated with the CTU is a raster scanning order along the horizontal direction. The second value (e.g., value 1) of the syntax element ctu_scan_order associated with the tensor indicates that the scanning order associated with the CTU is a raster scanning order along the vertical direction. In one example, the syntax element ctu_scan_order associated with the tensor is included in the compressed data unit header. In one example, the syntax element ctu_scan_order associated with the tensor is included in the aggregation unit header.
[0092] In one embodiment, when the model-related syntax element flag ctu_partition_flag has a value indicating that block (CTU) partitioning is enabled, the tensor-related syntax element ctu_scan_order is included in the syntax table.
[0093] Fig.11 An exemplary syntax of an aggregation unit header is shown, which includes signaling of a CTU scan order for one or more weight tensors. At line (1101), a FOR loop is performed on multiple NNR units. Each of the multiple NNR units may include, for example, a weight tensor. For each of the multiple NNR units, when ctu_partition_flag has a value indicating that CTU partitioning is enabled (line (1102)), ctu_scan_order[i] may be received, where i may be an index of each of the multiple NNR units. ctu_partition_flag may be a model-related syntax element.
[0094] Additionally, at line (1102), a syntax element quant_bitdepth[i] may be received for each of the one or more NNR units. quant_bitdepth[i] may specify a maximum bit depth for quantized coefficients of each tensor in the NNR aggregation unit.
[0095] In another embodiment, when the syntax element flag ctu_partition_flag related to the tensor has a value indicating that block (CTU) partitioning is enabled for the corresponding one or more tensors, the syntax element ctu_scan_order related to the tensor is included in the syntax table.
[0096] 6. Flag Dependency
[0097] In one example, ctu_partition_flag is defined as a model-dependent flag. The ctu_scan_order flag is placed in the nnr_aggregate_unit_header section (e.g., in Fig.11 ctu_partition_flag is placed in the nnr_model_parameter_set_payload section inside nnr_aggregate_unit. nnr_aggregate_unit_header may be serialized before nnr_model_parameter_set_payload, which makes it impossible to decode ctu_scan_order.
[0098] To solve this problem, in one example, the nnr_unit related to the model may be arranged serially before any nnr_aggregate_unit. ctu_partition_flag and max_ctu_dim_flag may be included in the nnr_unit. Therefore, the NNR aggregation unit following the nnr_unit may use any information defined in the nnr_unit.
[0099] III. 3D Pyramid Coding for Neural Network Model Compression
[0100] 1. Scanning order
[0101] Fig.12 An example of a syntax for scanning the weight coefficients in a weight tensor is shown. For example, the dimension of the weight tensor is greater than 2 (e.g., in a convolutional layer, the dimension is 4), and the weight tensor can be reshaped into a two-dimensional tensor. In one example, if the dimension of the weight tensor does not exceed 2 (e.g., a fully connected layer or a bias layer), no reshaping is performed. In order to encode the weight tensor, the weight coefficients in the weight tensor are scanned in a certain order. In some examples, the weight coefficients in the weight tensor can be scanned from left to right for each row and from the top row to the bottom row, for example, in a row-first manner.
[0102] exist Fig.12 In the example, the 2D integer array StateTransTab[][] specifies the state transition table for dependent scalar quantization and can be configured as follows:
[0103] StateTransTab[][]={{0, 2}, {7, 5}, {1, 3}, {6, 4}, {2, 0}, {5, 7}, {3, 1}, {4, 6}}.
[0104] 2. Quantification
[0105] In various embodiments, three types of quantization methods may be used: a baseline quantization method, a codebook-based quantization method, and a dependent scalar quantization method.
[0106] In the baseline quantization method, uniform quantization can be applied to the model parameter tensor (or parameter tensor) using a fixed step size. In one example, the fixed step size can be represented by the parameters qpDensity and qp. A flag denoted dq_flag can be used to enable uniform quantization (e.g., dq_flag is equal to 0). The reconstructed values in the decoded tensor can be integer multiples of the step size.
[0107] In the codebook-based approach, the model parameter tensor can be represented as a tensor of indices and a codebook, the indexed tensor having the same shape as the original tensor. The size of the codebook can be selected at the encoder and sent as a metadata parameter. The index has an integer value and can be further entropy encoded. In one example, the codebook consists of floating point 32-bit values. The reconstructed values in the decoded tensors are the values of the codebook elements referred to by the index values of these decoded tensors.
[0108] In the dependent scalar quantization method, dependent scalar quantization can be applied to the parameter tensor using a fixed step size, for example represented by the parameters qpDensity and qp, and a state transition table of size 8. A flag represented by dq_flag equal to 1 can be used to enable dependent scalar quantization. The reconstructed values in the decoded tensor are integer multiples of the step size.
[0109] 3. Entropy Coding
[0110] To encode the quantized weight coefficients, entropy coding techniques may be used.In some embodiments, the absolute values of the quantized weight coefficients are encoded in a sequence comprising a unary sequence, which may be followed by a fixed length sequence.
[0111] In some examples, the distribution of weight coefficients in the layer generally follows a Gaussian distribution, and the proportion of weight coefficients with large values is very small, but the maximum value of the weight coefficient can be very large. In some embodiments, very small values can be encoded using unary coding, and larger values can be encoded based on Golomb coding. For example, when Golomb coding is not used, an integer parameter called maxNumNoRem is used to indicate the maximum number. When the quantized weight coefficient is not greater than (e.g., equal to or less than) maxNumNoRem, the quantized weight coefficient can be encoded by unary coding. When the quantized weight coefficient is greater than maxNumNoRem, the part of the quantized weight coefficient equal to maxNumNoRem is encoded by unary coding, and the rest of the quantized weight coefficient is encoded by Golomb coding. Therefore, the unary sequence includes a first part of the unary coding and a second part for encoding the exponential Golomb residual bits.
[0112] In some embodiments, the quantized weight coefficients may be encoded by the following two steps.
[0113] In the first step, for the quantized weight coefficient, the binary syntax element sig_flag is encoded. The binary syntax element sig_flag specifies whether the quantized weight coefficient is equal to 0. If sig_flag is equal to 1 (indicating that the quantized weight coefficient is not equal to 0), the binary syntax element sign_flag is further encoded. The binary syntax element sign_flag indicates whether the quantized weight coefficient is positive or negative.
[0114] In a second step, the absolute value of the quantized weight coefficient may be encoded into a sequence comprising a unary sequence, which may be followed by a fixed length sequence. When the absolute value of the quantized weight coefficient is equal to or less than maxNumNoRem, the sequence comprises a unary encoding of the absolute value of the quantized weight coefficient. When the absolute value of the quantized weight coefficient is greater than maxNumNoRem, the unary sequence may comprise a first portion for encoding maxNumNoRem using a unary encoding and a second portion for encoding the exponential Golomb remainder, and the fixed length sequence is used to encode the fixed length remainder.
[0115] In some examples, unary coding is applied first. For example, a variable such as j is initialized to 0, and another variable X is set to j+1. The syntax element abs_level_greater_X is encoded. In one example, when the absolute value of the quantized weight level is greater than the variable X, abs_level_greater_X is set to 1 and the unary coding continues; otherwise, abs_level_greater_X is set to 0 and the unary coding is completed. When abs_level_greater_X is equal to 1 and the variable j is less than maxNumNoRem, the variable j is incremented by 1, and the variable X is also incremented by 1. Then, another syntax element abs_level_greater_X is encoded. This process continues until abs_level_greater_X is equal to 0 or the variable j is equal to maxNumNoRem. When the variable j is equal to maxNumNoRem, the coded bits are the first part of the unary sequence.
[0116] When abs_level_greater_X equals 1 and j equals maxNumNoRem, continue to encode using Golomb coding. Specifically, variable j is reset to 0, and X is set to 1 << j. The unary coding remainder can be calculated by subtracting maxNumNoRem from the absolute value of the quantized weight coefficient. Encode the syntax element abs_level_greater_thanX. In one example, when the unary coding remainder is greater than variable X, abs_level_greater_X is set to 1; otherwise, abs_level_greater_X is set to 0. If abs_level_greater_X equals 1, variable j is incremented, 1 << j is added to X, and another abs_level_greater_X is encoded. This process continues until abs_level_greater_X equals 0, thus encoding the second part of the unary sequence. When abs_level_greater_X equals 0, the unary coding remainder can be one of these values (X, X - 1, … X - (1 << j)+1). A code of length j can be used to encode an index pointing to one of the values in (X, X - 1, … X - (1 << j)+1), and this code can be referred to as the fixed - length remainder.
[0117] Fig.13 Examples for decoding the absolute value of the quantized weight coefficient according to some embodiments of the present disclosure are shown. In Fig.13 the example, QuantWeight[i] represents the quantized weight coefficient at the i - th position in the array; sig_flag specifies whether the quantized weight coefficient QuantWeight[i] is non - zero (e.g., sig_flag being 0 indicates QuantWeight[i] is 0); sign_flag specifies whether the quantized weight coefficient QuantWeight[i] is positive or negative (e.g., sign_flag being 1 indicates QuantWeight[i] is negative); abs_level_greater_x[j] indicates whether the absolute level of QuantWeight[i] is greater than j + 1 (e.g., the first part of the unary sequence); abs_level_greater_x2[j] includes the unary part of the exponential Golomb remainder (e.g., the second part of the unary sequence); and abs_remainder indicates the fixed - length remainder.
[0118] According to one aspect of the present disclosure, a context modeling approach may be used to encode the three flags sig_flag, sign_flag, and abs_level_greater_X. Thus, flags with similar statistical behavior may be associated with the same context model, so that the probability estimator (inside the context model) may be adapted to the underlying statistics.
[0119] In one example, the context modeling method uses three context models for sig_flag according to whether the weight coefficient of the left adjacent quantization is 0, less than 0, or greater than 0.
[0120] In another example, the context model method uses three other context models for sign_flag according to whether the left neighboring quantized weight coefficient is 0, less than 0, or greater than 0.
[0121] In another example, for each of the abs_level_greater_X flags, the context modeling method uses one or two separate context models. In one example, when X≤maxNumNoRem, two context models are used depending on sign_flag. In one example, when X>maxNumNoRem, only one context model is used.
[0122] 4. CTU3D and recursive CU3D block partitioning
[0123] In some embodiments, the model parameter tensor may be divided into CTU3D blocks, each of which is further divided into 3D coding unit (CU3D) blocks. CU3D may be further divided and encoded based on a pyramid tree structure. For example, the pyramid tree structure may be a 3D octree, a 3D single-branch tree, a 3D label tree, or a 3D single-branch label tree structure. After a specific training / retraining operation, the weight coefficient may have a local structure. Coding methods using 3D octrees, 3D single-branch trees, 3D label trees, and / or 3D single-branch label tree structures may generate more efficient representations by using local distributions of CTU3D / CU3D blocks. These pyramid tree structure-based methods may be coordinated with baseline methods (i.e., coding methods based on non-pyramid tree structures).
[0124] Typically, for convolutional layers with a layout of [R][S][C][K], the dimension of the weight tensor (or model parameter tensor) can be 4; for fully connected layers with a layout of [C][K], the dimension can be 2; and for bias and batch norm layers, the dimension can be 1. R and S represent the convolution kernel size (width and height), C represents the input feature size, and K represents the output feature size.
[0125] In one embodiment, for a convolutional layer, the 2D [R] [S] dimension may be reshaped into a 1D [RS] dimension, such that a 4D tensor [R] [S] [C] [K] is reshaped into a 3D tensor [RS] [C] [K]. A fully connected layer is considered as a special case of a 3D tensor where R = S = 1.
[0126] In one embodiment, the 3D tensor [RS][C][K] may be divided along the [C][K] plane by non-overlapping smaller blocks (CTU3D). Each CTU3D has a shape of [RS][ctu3d_height][ctu3d_width], where, in one example, ctu3d_height=max_ctu3d_height and ctu3d_width=max_ctu3d_width. For a CTU3D located on the right and / or bottom of the tensor, its ctu3d_height is the remainder of C / max_ctu3d_height, and its ctu3d_width is the remainder of K / max_ctu3d_width.
[0127] In one embodiment, the values of max_ctu3d_height and max_ctu3d_width may be explicitly signaled in the bitstream, or may be implicitly inferred. In one example, when max_ctu3d_height=C and max_ctu3d_width=K, block partitioning is disabled.
[0128] In one embodiment, a quadtree structure may be used to perform a simplified block structure, where a CTU3D / CU3D is recursively divided into smaller CU3Ds until a maximum recursion depth is reached. Starting from the CTU3D node, this quadtree of CU3D blocks may be scanned and processed using a depth-first quadtree scan order. Child nodes under the same parent node may be scanned and processed using a raster scan order in either the horizontal or vertical direction.
[0129] Fig.14 Two examples of adaptive CTU3D / CU3D partitioning using raster scanning along the vertical direction are shown.
[0130] In one embodiment, for CU3Ds of a given quadtree depth, the max_cu3d_height / max_cu3d_width of these CU3Ds are calculated using the following formula.
[0131] max_cu3d_height=max_ctu3d_height>>depth
[0132] max_cu3d_width=max_ctu3d_width>>depth
[0133] The maximum recursion depth is reached when both max_cu3d_height and max_cu3d_width are less than or equal to a predetermined threshold. The threshold may be explicitly included in the bitstream, or may be a predetermined number (e.g., 8) that may be implicitly inferred by the decoder. When the predetermined threshold is the size of a CTU3D, the recursive partitioning is disabled.
[0134] In one embodiment, a rate-distortion (RD)-based coding algorithm decides whether to split a parent CU3D into multiple smaller sub-CU3Ds. If the combined RD of these smaller sub-CU3Ds is less than the RD from the parent CU3D, the parent CU3D is split into multiple smaller sub-CU3Ds. Otherwise, the parent CU3D will not be split further. A split flag is defined to record the corresponding split decision. This flag can be skipped at the last depth of CU partitioning.
[0135] In one embodiment, a recursive CU3D block partitioning operation is performed based on a quadtree structure to partition a CTU3D into CU3D blocks, and a split flag is defined to record each split decision at a node in the quadtree structure.
[0136] In another embodiment, no recursive CU3D block partitioning operation is performed in a CTU3D block, and no split flag for recording the split decision is defined. In this case, the CU3D block is the same as the CTU3D block.
[0137] 5.3D Pyramid Tree Structure
[0138] In various embodiments, the pyramid tree structure (or 3D pyramid tree structure) may be a tree data structure in which each internal node may have 8 child nodes. The 3D pyramid tree structure may be used to partition a 3D tensor (or a sub-block such as a CTU3D or CU3D) into 8 octants recursively along the z, y, and x axes.
[0139] Fig.15 An exemplary partitioning process based on a 3D pyramid tree structure is shown. Fig.15 As shown, the 3D octree (1505) is a tree data structure in which each internal node (1510) has exactly 8 child nodes (1515). The 3D octree (1505) is used to partition the three-dimensional tensor (1520) by recursively subdividing the three-dimensional tensor (1520) into 8 octants (1525) along the z, y, and x axes. The nodes at the last depth of the 3D octree (1505) can be blocks of size 2×2×2 coefficients.
[0140] In various embodiments, different methods may be adopted to construct a 3D pyramid tree structure to represent coefficients in a CU3D at the encoder side or the decoder side.
[0141] In one embodiment, the 3D octree for CU3D may be constructed as follows. A node value of 1 at a 3D octree position at the last depth indicates that the codebook index (in the case of using a codebook encoding method) or the coefficient (in the case of using a direct quantization encoding method) in the corresponding node is not 0. A node value of 0 at a 3D octree position at the bottom depth indicates that the codebook index or coefficient in the corresponding node is 0. The node value of a 3D octree position at other depths is defined as the maximum value of its 8 child nodes.
[0142] In one embodiment, a 3D singular tree for CU3D may be constructed as follows: a node value of 1 at a 3D singular tree position at a depth other than the last depth indicates that its child nodes (and child nodes of child nodes, including the node at the last depth) have non-uniform (different) values; a node value of 0 at a 3D singular tree position at a depth other than the last depth indicates that all its child nodes (and child nodes of child nodes, including the node at the last depth) have uniform (same) values.
[0143] In one embodiment, the 3D label tree for CU3D may be constructed as follows. The node value of the 3D label tree position at the last depth indicates that the absolute value of the codebook index (in the case of using the codebook encoding method) or the absolute coefficient (in the case of using the direct quantization encoding method) in the corresponding CU3D is not 0. The node value of the 3D label tree position at other depths is defined as the minimum value of its 8 child nodes. In another embodiment, the node value of the 3D label tree position at other depths may be defined as the maximum value of its 8 child nodes.
[0144] In one embodiment, the 3D singular tag tree of CU3D is constructed by combining the 3D tag tree and the 3D singular tree.
[0145] It should be noted that for some CU3D blocks with different depths / heights / widths, there may be insufficient coefficients to construct a complete 3D pyramid in which all 8 child nodes of all parent nodes are available. If a parent node does not have all 8 child nodes, scanning and encoding of these non-existent child nodes may be skipped.
[0146] 6.3D Pyramid Scanning Order
[0147] After the 3D pyramid is constructed, all nodes may be traversed using a predetermined scan order to encode node values at the encoder side or decode node values at the decoder side.
[0148] In one embodiment, starting from the top node, a depth-first search scan order may be used to pass through all nodes. The scan order of child nodes sharing the same parent node may be arbitrarily defined, for example, (0, 0, 0) -> (0, 0, 1) -> (0, 1, 0) -> (0, 1, 1) -> (1, 0, 0) -> (1, 0, 1) -> (1, 1, 0) -> (1, 1, 1).
[0149] In another embodiment, starting from the top node, a breadth-first search can be used to walk through all nodes. Because each pyramid depth is a 3D shape, the scan order in each depth can be arbitrarily defined. In one embodiment, the scan order is defined using the following pseudo code to align with the pyramid encoding method:
[0150]
[0151]
[0152] In another embodiment, the encoding_start_depth syntax element may be used to indicate the first depth participating in the encoding or decoding process. When all nodes are traversed using a predetermined scan order, if the depth of the node is greater than encoding_start_depth, the encoding of the current node value is skipped. Multiple CU3Ds, CTU3Ds, layers, or models may share one encoding_start_depth. This syntax element may be explicitly signaled in the bitstream, or predefined and implicitly inferred.
[0153] In one embodiment, encoding_start_depth is explicitly signaled in the bitstream. In another embodiment, encoding_start_depth is predefined and implicitly inferred. In another embodiment, encoding_start_depth is set to the last depth of the 3D pyramid tree structure and implicitly inferred.
[0154] 7.3D Pyramid Coding Method
[0155] At the decoder side, after constructing the 3D pyramid tree structure, the corresponding encoding method can be executed to pass through all nodes and encode the coefficients represented by different 3D trees. At the decoder side, corresponding to different encoding methods, the encoded coefficients can be decoded accordingly.
[0156] For a 3D octree, if the value of the parent node is 0, the scanning and encoding of its child nodes (and the child nodes of the child nodes) are skipped, because the value of the child node should always be 0. If the value of the parent node is 1, and the values of all child nodes except the last child node are 0, the last child node can be scanned, but the encoding of the value of the last child node can be skipped, because the value of the last child node should always be 1. If the current depth is the last depth of the pyramid and if the current node value is 1, the sign of the map value is encoded when the codebook method is not used, and then the map value (quantized value) itself is encoded.
[0157] For a 3D singular tree, in one embodiment, the value of a given node may be encoded. If the node value is 0, the corresponding uniform value may be encoded, and the encoding of its child nodes (and children of child nodes) may be skipped, since the absolute value of the child nodes should always be equal to the uniform value. The child nodes may be scanned until the bottom depth is reached, where the sign bit of each child node may be encoded if the node value is not 0.
[0158] For a 3D singular tree, in another embodiment, the value of a given node may be encoded. If the node value is 0, the uniform value corresponding thereto may be encoded, and the encoding of its child nodes (and child nodes of child nodes) may be skipped, because the absolute value of the child nodes should always be equal to the uniform value. And after processing all nodes in the CU3D, the pyramid tree structure may be scanned again, and if the node value is not 0, the sign bit of each child node at the bottom depth may be encoded.
[0159] For a 3D label tree, if a node is a top node without a parent node, the value of the node can be encoded. For any child node, the difference between the parent node and the child node can be encoded. If the value of the parent node is X and the values of all child nodes except the last child node are greater than X, the last child node can be scanned, but the encoding of the value of the last child node can be skipped because the value of the last child node should always be X.
[0160] For a 3D single-branch label tree, the value of a given node from a single-branch tree can first be encoded. Then, if the node is a top node without a parent node, the label tree value can be encoded using a label tree encoding method, or the difference between the label tree value of the parent node and the child node can be encoded. The node skipping method introduced in the label tree encoding section is also used. If the single-branch tree node value is 0, the scanning and encoding of its child nodes (and child nodes of child nodes) can be skipped because the value of the child node should always be equal to a uniform value.
[0161] In one embodiment, when encoding_start_depth is the last depth, the coefficient skipping methods described herein may be disabled to encode all coefficients. In one example, a syntax element may be received in a bitstream at a decoder side, indicating a starting depth in a 3D pyramid tree structure. When the starting depth is the last depth of the 3D pyramid tree structure, a non-3D pyramid tree based decoding method may be used to decode the model parameters of the model parameter tensor from the bitstream.
[0162] In another embodiment, when encoding_start_depth is the last depth, in order to utilize the coefficient skipping methods described herein, the 3D pyramid tree may be encoded by adjusting the starting depth so that the starting depth is the second-to-last depth. In one example, a syntax element indicating the starting depth in the 3D pyramid tree structure may be received in a bitstream at a decoder. When the starting depth is the last depth of the 3D pyramid tree structure, the decoding process may start at the decoder from the second-to-last depth in the depth of the 3D pyramid tree structure.
[0163] In another embodiment, when encoding_start_depth is the last depth, for a 3D single-branch tag tree, the single-branch tree portion of the 3D pyramid tree may be encoded by adjusting encoding_start_depth so that encoding_start_depth is the second-to-last depth. The tag tree portion of the 3D pyramid tree may be encoded without adjusting encoding_start_depth. In one example, a syntax element indicating a starting depth in a 3D pyramid tree structure may be received in a bitstream of a decoder. When the starting depth is the last depth of the 3D pyramid tree structure and the 3D pyramid tree structure is a single-branch tag tree structure associated with single-branch tree portion encoding and tag tree portion encoding, a decoding process may be performed at the decoder as follows. For single-branch tree portion encoding, the decoding process may start from the second-to-last depth in the depth of the 3D pyramid tree structure. For tag tree portion encoding, the decoding process may start from the last depth in the depth of the 3D pyramid tree structure.
[0164] 8. Dependence on Quantification
[0165] In some embodiments, a dependent scalar quantization method is used for neural network parameter approximation. A related entropy coding method can be used to cooperate with the quantization method. The method introduces dependencies between quantized parameter values, which reduces distortion during parameter approximation. In addition, the dependencies can be exploited in the entropy coding stage.
[0166] In dependent quantization, the allowed reconstruction values of neural network parameters (e.g., weight parameters) depend on the selected quantization index of the neural network parameters preceding in the reconstruction order. The main effect of this approach is that the allowed reconstruction vectors (given by all reconstructed neural network parameters of a layer) are more densely packed in the N-dimensional vector space (N represents the number of parameters in the layer) than in conventional scalar quantization. This means that for a given average number of allowed reconstruction vectors per N-dimensional unit volume, the average distance (e.g., mean squared error (MSE) or mean absolute error (MAE) distortion) between the input vector and the closest reconstruction vector is reduced (for a typical distribution of input vectors).
[0167] In the dependent quantization process, the parameters can be reconstructed in a scanning order (the scanning order is the same as the order in which the parameters are entropy decoded) due to the dependency between the reconstructed values. Then, the method of dependent scalar quantization can be implemented by defining two scalar quantizers with different reconstruction levels and defining a process for switching between the two scalar quantizers. Therefore, for each parameter, there may be two available scalar quantizers, such as Fig.16 shown.
[0168] Fig.16 Two scalar quantizers used according to an embodiment of the present disclosure are shown. The first quantizer Q0 maps the neural network parameter levels (numbers from -4 to 4 below the dot) to even multiples of the quantization step size Δ. The second quantizer Q1 maps the neural network parameter levels (numbers from -5 to 5) to odd multiples of the quantization step size Δ or to 0.
[0169] For quantizers Q0 and Q1, the positions of the available reconstruction levels are uniquely specified by the quantization step size Δ. The two scalar quantizers Q0 and Q1 are characterized as follows:
[0170] Q0: The reconstruction level of the first quantizer Q0 is given by an even multiple of the quantization step size Δ. When this quantizer is used, the reconstructed neural network parameter t′ is calculated according to the following formula:
[0171] t′=2·k·Δ,
[0172] where k represents the associated parameter level (the quantization index sent).
[0173] Q1: The reconstruction level of the second quantizer Q1 is given by an odd multiple of the quantization step size Δ and a reconstruction level equal to zero. The mapping of the neural network parameter level k to the reconstruction parameter t′ is specified by:
[0174] t′=(2·k-sgn(k))·Δ,
[0175] where sgn(·) represents the sign function
[0176]
[0177] Instead of explicitly signaling the quantizer (Q0 or Q1) used by the current weight parameter in the bitstream, the quantizer is determined by the parity of the weight parameter level that precedes the current weight parameter in the encoding / reconstruction order. Switching between quantizers is achieved by a state machine represented by Table 1. The state has 8 possible values (0, 1, 2, 3, 4, 5, 6, 7) and is uniquely determined by the parity of the weight parameter level that precedes the current weight parameter in the encoding / reconstruction order. For each layer, the state variable is initially set to 0. When the weight parameters are reconstructed, the state is then updated according to Table 1, where k represents the value of the transform coefficient level. The next state depends on the current state and the parity of the current weight parameter level k (k&1). Therefore, the state update can be obtained by the following formula:
[0178] state = sttab[state][k&1]
[0179] Where sttab represents Table 1.
[0180] Table 1 shows a state transition table for determining a scalar quantizer for a neural network parameter, where k represents the value of the neural network parameter:
[0181] Table 1
[0182]
[0183] The state uniquely specifies the scalar quantizer to be used. If the state value of the current weight parameter is an even number (0, 2, 4, 6), the scalar quantizer Q0 is used. Otherwise, if the state value is an odd number (1, 3, 5, 7), the scalar quantizer Q1 is used.
[0184] In some embodiments, a baseline encoding method may be used (in which the encoding / decoding method based on the 3D pyramid tree structure is not used). In the baseline encoding method, all coefficients of the model parameter tensor may be scanned according to a scan order and entropy encoded. For a dependent quantization process used in combination with the baseline encoding method, the coefficients may be reconstructed in a scan order (the scan order is the same as the order in which the coefficients are entropy decoded).
[0185] Due to the nature of the 3D pyramid coding method described herein, certain coefficients in the model parameter tensor may be skipped from the entropy coding process. Therefore, in one embodiment, when using the 3D pyramid coding method, the dependent quantization process (which operates on all coefficients in the model parameter tensor) may be disabled.
[0186] In another embodiment, when the 3D pyramid coding method is used, a dependent quantization process may be enabled. For example, the dependent quantization construction process may be modified so that if these coefficients are skipped from the entropy coding process, these coefficients may be excluded from the construction process of the dependent quantization coefficients. In one example, when one or more syntax elements indicate that the CU3D is divided based on the 3D pyramid tree structure, the dependent quantization construction process may be performed on the CU3D. The model parameters of the CU3D skipped during the coding process based on the 3D pyramid tree structure are excluded from the dependent quantization construction process.
[0187] In another embodiment, the absolute values of the coefficients are used in dependent quantization.
[0188] 9. Entropy Coding Context
[0189] In some embodiments, when dependency quantization is not used, context modeling may be performed as follows.
[0190] For an octree node value represented as Oct_flag and a symbol represented as sign in a 3D octree-based encoding method, a context model index represented as ctx may be determined according to the following formula:
[0191] Oct_flag:
[0192] int p0=(z>=1)? oct[d][z-1][y][x]: 0;
[0193] int p1=(z>=2)? oct[d][z-2][y][x]: 0;
[0194] int ctx=(p0==p1)? ! ! p0:2;
[0195] sign:
[0196] int p0=(z>=1)? map[z-1][y][x]: 0;
[0197] int ctx=(p0==0)? 0: (p0<0)? 1:2.
[0198] In the above calculation, oct[d][z][y][x] represents the node value at position [z][y][x] at depth d; map[z][y][x] represents the quantized value at position [z][y][x].
[0199] For a non-zero flag denoted as nz_flag and a symbol denoted as sign in a 3D singular tree-based coding method, a context model index denoted as ctx may be determined according to the following formula:
[0200] nz_flag:
[0201] int p0=(last_depth&&map_z>=1)? std::abs(map[map_z-1][map_y][map_x]): 0;
[0202] int ctx=(p0==0)? 0:1;
[0203] sign:
[0204] int p0=(z>=1)? map[z-1][y][x]: 0;
[0205] int ctx=(p0==0)? 0: (p0<0)? 1:2.
[0206] For the 3D label tree based encoding method, the context model index denoted as ctx with a non-zero flag denoted as nz_flag and a sign denoted as sign may be determined according to the following formula:
[0207] nz_flag:
[0208] int p0=(z>=1)? tgt[d][z-1][y][x]: 0;
[0209] int p1=(z>=2)? tgt[d][z-2][y][x]: 0;
[0210] int ctx=(p0==p1)? ! ! p0:2;
[0211] sign:
[0212] int p0=(z>=1)? tgt[d][z-1][y][x]: 0;
[0213] int ctx=(p0==0)? 0: (p0<0)? 1:2.
[0214] where tgt[d][z][y][x] represents the value of the node at position [z][y][x] at depth d.
[0215] For a 3D single-branch label tree, the context model index denoted as ctx with a non-zero flag denoted as nz_flag and a sign denoted as sign can be determined according to the following formula:
[0216] nz_flag:
[0217] int p0=(z>=1)? tgt[d][z-1][y][x]: 0;
[0218] int p1=(z>=2)? tgt[d][z-2][y][x]: 0;
[0219] int ctx=(p0==p1)? ! ! p0:2;
[0220] sign:
[0221] int p0=(z>=1)? map[z-1][y][x]: 0;
[0222] int ctx=(p0==0)? 0: (p0<0)? 1:2.
[0223] In one example, when dependent quantization is used, the context modeling of nz_flag may be adjusted such that ctx=ctx+3*state_id.
[0224] 10. Syntax Cleanup
[0225] All coefficients in a CU3D may be made uniform. In one embodiment, uaflag may be defined in the CU3D header to indicate whether all coefficients in the CU3D are uniform. In one example, a value of uaflag = 1 indicates that all coefficients in the CU3D are uniform.
[0226] In one embodiment, ctu3d_map_mode_flag may be defined to indicate whether all CU3D blocks in a CTU3D share the same map_mode. If ctu3d_map_mode_flag = 1, map_mode is signaled. It should be noted that this flag may also be inferred implicitly (to 0).
[0227] In one embodiment, enable_start_depth may be defined to indicate whether cu3d encoding may start from a depth other than the bottom depth. If enable_start_depth=1, start_depth (or encoding_start_depth) is signaled. Note that this flag may also be inferred implicitly (to 1).
[0228] In one embodiment, an enable_zdep_reorder flag may be defined to indicate whether zdep_array reordering is allowed. Note that this flag may also be inferred implicitly (to 0).
[0229] 11. Coordination of Baseline Coding Method and Pyramid Coding Method
[0230] In the baseline method, the weight tensor can be reshaped into a 2D matrix with the shape of [output_channel][input_channel*kernel size]. The coefficients in the same kernel are stored in consecutive memory locations. When calculating the context of sig_flag and sign_flag, the neighboring coefficient is defined as the last coefficient processed before the current coefficient. For example, the neighboring coefficient of the first coefficient in one kernel is the last coefficient in the previous kernel.
[0231] In one embodiment, to calculate the context model index, for the first coefficient in a kernel, the values of its neighboring coefficients are set to 0, rather than the value of the last coefficient in the previous kernel being set to 0.
[0232] In one embodiment, during the 3D pyramid encoding process, if the quantization mode is not a codebook mode and if start_depth is the last pyramid depth, a baseline encoding method is selected and its RD is calculated and compared with other modes (eg, encoding methods based on a 3D pyramid tree structure).
[0233] In one embodiment, if ctu3d_width and / or ctu3d_height is 1, the baseline encoding method is automatically selected.
[0234] 12.3D Pyramid Coding Syntax Table
[0235] In Appendix B of the present disclosure, syntax tables Table 2 to Table 17 are listed as examples of the encoding method based on the 3D pyramid tree structure disclosed herein. Syntax elements introduced in the listed syntax tables are defined at the end of each corresponding syntax table.
[0236] 13. Unified coding method for neural network model compression
[0237] In some embodiments, a uniform-based encoding method may be used. A layer_uniform_flag flag may be defined for convolutional and fully connected layers to indicate whether the layer is encoded using a 3D pyramid tree structure-based encoding method. In one example, if the layer_uniform_flag flag is equal to a first value (e.g., 0), the layer is encoded using a baseline method.
[0238] If layer_uniform_flag is equal to a second value (e.g., 1), a coding method based on a 3D pyramid tree structure may be used. For example, the layer may be reshaped into a CTU3D layout. For each CTU3D, a ctu3d_uniform_flag flag may be defined to indicate whether all child nodes sharing the same parent node at the bottom depth are uniform (nodes that do not share the same parent node may have different uniform values).
[0239] If the ctu3d_uniform_flag flag is equal to a first value (e.g., 1) for this CTU3D, all child nodes sharing the same parent node at the bottom depth are unified (nodes that do not share the same parent node may have different uniform values), and in one embodiment, a 3D unary tree encoding method may be used to encode this CTU3D. encoding_start_depth is set to the last depth of the 3D pyramid tree structure (e.g., associated with a CU3D or CTU3D) and is implicitly inferred. The encoding of the uniform value of a node may be skipped because the uniform value of the node should always be 0.
[0240] In one embodiment, for all child nodes sharing the same parent node at the bottom depth, a uniform value may be encoded, and then the sign bits of these child nodes may be encoded if the node value is not 0. In another embodiment, for all child nodes sharing the same parent node at the bottom depth, a uniform value may be encoded. And after processing all nodes in the CU3D, if the node value is not 0, the pyramid (pyramid tree structure) may be scanned again to encode the sign bit of each child node at the bottom depth.
[0241] In one embodiment, if the ctu3d_uniform_flag flag is equal to a second value (eg, 0), the CTU3D may be encoded using a 3D tag tree encoding method. encoding_start_depth is set to the last depth of the 3D pyramid tree structure (eg, associated with a CU3D or CTU3D) and is implicitly inferred.
[0242] In one embodiment, the value of each child node may be encoded based on a predetermined scan order. In another embodiment, the absolute value of each child node may be encoded based on a predetermined scan order, and then the sign bit of each child node may be encoded. In another embodiment, the absolute values of all child nodes may be encoded based on a predetermined scan order. And after processing all nodes in the CU3D, if the node value is not 0, the sign bits of all child nodes may be encoded.
[0243] 14. Syntax table based on unified encoding
[0244] In Appendix C of the present disclosure, syntax tables Table 18 to Table 21 are listed as examples of the unified coding method disclosed herein. Syntax elements introduced in the listed syntax tables are defined at the end of each corresponding syntax table.
[0245] IV. Example of the Coding Process
[0246] Fig.17 The decoding process (1700) of CTU partitioning according to one embodiment of the present disclosure is shown. The process (1700) may start from (S1701) and proceed to (S1710).
[0247] At (S1710), a first syntax element in an NNR aggregation unit header of an NNR aggregation unit may be received from a bitstream. The first syntax element may indicate a CTU scanning order for processing a model parameter tensor transmitted in the NNR aggregation unit. For example, a first value of the first syntax element may indicate that the CTU scanning order is a first raster scanning order along a horizontal direction. A second value of the first syntax element may indicate that the CTU scanning order is a second raster scanning order along a vertical direction.
[0248] In one example, another syntax element may be received in advance to control whether CTU block partitioning is enabled for tensors in the NNR aggregation unit. For example, the other syntax element may be a syntax element related to a model or a syntax element related to a tensor, the syntax element related to the model being used to specify whether CTU block partitioning is enabled for a layer of a neural network, and the syntax element related to a tensor being used to specify whether CTU block partitioning is enabled for a tensor in the NNR aggregation unit.
[0249] At (S1720), the tensors in the NNR aggregation unit may be reconstructed based on the CTU scanning order. When CTU block partitioning is enabled for the tensors in the NNR aggregation unit, on the encoder side, the tensors may be divided into CTUs. The CTUs may be scanned and encoded according to the scanning order indicated by the first syntax element. On the decoder side, based on the indicated scanning order, the decoder may understand the order in which the CTUs are encoded and thus organize the decoded CTUs into tensors. The process (1700) may proceed to (S1799) and terminate at (S1799).
[0250] Fig.18 A decoding process (1800) of 3D pyramid coding according to one embodiment of the present disclosure is shown. The process (1800) may start from (S1801) and proceed to (S1810).
[0251] At (S1810), one or more first syntax elements may be received from a bitstream represented by a compressed neural network. The one or more first syntax elements may be associated with a CU3D divided from a CTU3D. The CTU3D may be obtained from tensor division in a neural network. The one or more first syntax elements may indicate that the CU3D is divided based on a coding mode corresponding to a 3D pyramid tree structure. The 3D pyramid tree structure may include multiple depths. Each depth corresponds to one or more nodes. Each node has a node value. For example, the 3D pyramid tree structure may be one of an octree structure, a single-branch tree structure, a label tree structure, a single-branch label tree structure, and the like.
[0252] At (S1820), a second sequence of syntax elements corresponding to node values of each node in the 3D pyramid tree structure may be received from the bitstream in a breadth-first scanning order for scanning nodes in the 3D pyramid tree structure. Thus, at the decoder side, the node values (represented by syntax elements) may be received based on a depth-first scanning order. In other embodiments (process (1800) is not implemented), the 3D pyramid tree structure may be scanned at the encoder side according to a depth-first scanning order.
[0253] At (S1830), the model parameters of the tensor may be reconstructed based on the received second syntax element corresponding to the node value of the node in the 3D pyramid tree structure. As described above, at the decoder, corresponding to the octree structure, the single-branch tree structure, the label tree structure, or the single-branch label tree structure, a 3D pyramid coding method may be used to encode the node value of the 3D pyramid tree structure and the coefficient value of the CU3D divided using the 3D pyramid tree structure. At the decoder, corresponding to the adopted 3D pyramid coding method, the node value and the coefficient value may be reconstructed accordingly. The process may proceed to (S1899) and terminate at (S1899).
[0254] Fig.19 A decoding process (1900) based on unified encoding according to one embodiment of the present disclosure is shown. The process (1900) may start from (S1901) and proceed to (S1910).
[0255] At (S1910), a syntax element associated with a CTU3D may be received. The CTU3D may be partitioned from tensors in a layer of a neural network in a bitstream of a compressed neural network representation of the neural network. The syntax element may indicate whether all child nodes at a bottom depth of a pyramid tree structure associated with the CTU3D are unified. Child nodes that do not share the same parent node at the bottom depth may have different unified values.
[0256] At (S1920), in response to the syntax element indicating that all child nodes at the bottom depth of the pyramid tree structure associated with the CTU3D are unified, the CTU3D may be decoded based on the 3D single tree encoding method. In one example, it may be inferred that the starting depth of the 3D single tree encoding method is the bottom depth of the pyramid tree structure. In one embodiment, the unified flag of the node at the bottom depth of the pyramid tree structure is not encoded in the bitstream.
[0257] In one embodiment, a uniform value encoded in a bitstream for all child nodes sharing the same parent node at the bottom depth may be received. A sign bit for all child nodes sharing the same parent node at the bottom depth may be received. The sign bit follows the uniform value in the bitstream.
[0258] In one embodiment, a uniform value for each group of child nodes sharing the same parent node at the bottom depth may be received. Then, a sign bit for a child node in each group of child nodes sharing the same parent node at the bottom depth may be received.
[0259] At (S1930), in response to the syntax element indicating that all child nodes at the bottom depth of the pyramid tree structure associated with CTU3D are not all uniform, CTU3D may be decoded based on a 3D label tree encoding method. In one example, the starting depth of the 3D label tree encoding method may be inferred to be the bottom depth of the pyramid tree structure.
[0260] In various embodiments, the values of nodes at the bottom depth of the pyramid tree structure may be decoded according to one of the following methods: In a first method, the values of nodes at the bottom depth of the pyramid tree structure may be received, and the value of each node may be encoded in a bitstream based on a predetermined scanning order.
[0261] In the second method, the absolute value of each node at the bottom depth of the pyramid tree structure may be received in the bitstream based on a predetermined scanning order, and then (if the absolute value is not 0) the sign of each node is received. In the third method, the absolute value of each node at the bottom depth of the pyramid tree structure may be received in the bitstream based on a predetermined scanning order, and then (if the node has a non-zero value) the sign of each node at the bottom depth of the pyramid tree structure is received in the bitstream based on a predetermined scanning order. The process (1900) may proceed to (S1999) and terminate at (S1999).
[0262] V.Computer System
[0263] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig. 20A computer system (2000) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0264] Computer software may be encoded using any suitable machine code or computer language that may be subjected to assembly, compilation, linking, or similar mechanisms to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0265] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0266] Fig. 20 The components of the computer system (2000) shown are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary embodiment of the computer system (2000).
[0267] The computer system (2000) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., 2D video, 3D video including stereoscopic video).
[0268] The input human-machine interface device may include one or more of the following (only one of each is shown): keyboard (2001), mouse (2002), touch pad (2003), touch screen (2010), data gloves (not shown), joystick (2005), microphone (2006), scanner (2007), camera (2008).
[0269] The computer system (2000) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (2010), a data glove (not shown), or a joystick (2005), but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers (2009), headphones (not depicted)), visual output devices (e.g., screens (2010) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting 2D visual output or output beyond 3D through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)) and printers (not depicted).
[0270] The computer system (2000) may also include human-accessible storage devices and their associated media, for example, optical media including CD / DVD ROM / RW (2020) with CD / DVD and other media (2021), thumb drives (2022), removable hard drives or solid-state drives (2023), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), and the like.
[0271] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0272] The computer system (2000) may also include an interface to one or more communication networks. The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter connected to some universal data port or peripheral bus (2049) (e.g., a USB port of the computer system (2000)); as described below, other network interfaces are typically integrated into the kernel of the computer system (2000) by attaching to the system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system (2000) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., a CANbus connected to certain CANbus devices), or bidirectional, for example, using a LAN or WAN digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of these networks and network interfaces.
[0273] The above-mentioned human-machine interface device, human-machine accessible storage device and network interface may be attached to the kernel (2040) of the computer system (2000).
[0274] The core (2040) may include one or more central processing units (CPUs) (2041), graphics processing units (GPUs) (2042), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (2043), hardware accelerators (2044) for certain tasks, etc. These devices, as well as read-only memory (ROM) (2045), random access memory (2046), internal mass storage (2047) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (2048). In some computer systems, the system bus (2048) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (2048) or to the core's system bus (2048) via a peripheral bus (2049). The architecture of the peripheral bus includes PCI, USB, etc.
[0275] The CPU (2041), GPU (2042), FPGA (2043) and accelerator (2044) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (2045) or RAM (2046). Transition data can also be stored in RAM (2046), while permanent data can be stored, for example, in internal mass storage (2047). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with the following: one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc.
[0276] The computer readable medium may have thereon computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer code may be of a type well known and available to those skilled in the art of computer software.
[0277] As a non-limiting example, a computer system having an architecture (2000), particularly a kernel (2040), may provide functionality due to one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media may be media associated with user-accessible mass storage as described above, as well as certain non-temporary memories of the kernel (2040), such as kernel internal mass storage (2047) or ROM (2045). Software implementing the various embodiments of the present disclosure may be stored in such devices and executed by the kernel (2040). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software may enable the kernel (2040), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (2046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (2044)) that may replace software or operate in conjunction with software to perform specific processes or specific portions of specific processes described herein. Where appropriate, references to portions of software may include logic and vice versa. Where appropriate, references to portions of computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0278] Although the present disclosure has described a number of exemplary embodiments, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.
[0279] Appendix A: Acronyms
[0280] DNN: Deep Neural Network
[0281] NNR: Compressed Neural Network Representation; Encoded Representation of Neural Networks
[0282] CTU: Coding Tree Unit
[0283] CTU3D: Coding Tree Unit 3D
[0284] CU: Coding Unit
[0285] CU3D: 3D Coding Unit
[0286] RD: Rate Distortion
[0287] Appendix B: Syntax table of pyramid encoding method
[0288] Table 2
[0289]
[0290]
[0291] Table 3
[0292]
[0293]
[0294] Table 4
[0295]
[0296]
[0297]
[0298]
[0299] Table 5
[0300]
[0301] nzflag Non-zero flag for quantized coefficients
[0302] sign The sign bit of the quantized coefficient
[0303] Table 6
[0304]
[0305] Table 7
[0306]
[0307]
[0308]
[0309] Table 8
[0310]
[0311] sign predicted_size-prev_predicted_size sign bit
[0312] predicted_flag 0 indicates that position n is not a predicted entry, 1 indicates that position n is a predicted entry
[0313] Table 9
[0314]
[0315]
[0316] Table 10
[0317]
[0318]
[0319] nzflag Non-zero flag for quantized coefficients
[0320] sign The sign bit of the quantized coefficient
[0321] Table 11
[0322]
[0323]
[0324]
[0325] uni_flag unidirectional tree node value
[0326] nzflag Non-zero flag for quantized coefficients
[0327] sign The sign bit of the quantized coefficient
[0328] Table 12
[0329]
[0330]
[0331] oct_flag octree node value
[0332] sign The sign bit of the quantized coefficient
[0333] Table 13
[0334]
[0335]
[0336] nz_flag Non-zero flag of the node value
[0337] sign The sign bit of the quantized coefficient
[0338] Table 14
[0339]
[0340]
[0341] uni_flag unidirectional tree node value
[0342] nz_flag Non-zero flag of the node value
[0343] sign The sign bit of the quantized coefficient
[0344] Table 15
[0345]
[0346]
[0347] nz_flag Non-zero flag of the node value
[0348] sign The sign bit of the quantized coefficient
[0349] Table 16
[0350]
[0351]
[0352]
[0353] Table 17
[0354]
[0355]
[0356] nzflag non-zero flag
[0357] uiBit Unary part of the exponential Golomb remainder
[0358] uiBits Fixed length remainder
[0359] Appendix C: Syntax table based on unified encoding
[0360] Table 18
[0361]
[0362] Ndim(arrayName[]) returns the dimension of arrayName[].
[0363] scan_order specifies the block scan order for parameters with more than one dimension according to the following table:
[0364] 0: Do not scan blocks
[0365] 1: 8×8 blocks
[0366] 2: 16×16 blocks
[0367] 3: 32×32 blocks
[0368] 4: 64×64 blocks
[0369] layer_uniform_flag specifies whether to use a uniform method to encode the quantization weights QuantParam[]. layer_uniform_flag equal to 1 indicates that QuantParam[] is encoded using a uniform method.
[0370] Table 19
[0371]
[0372] The 2D integer array StateTransTab[][] specifies the state transition table used for dependent scalar quantization, as follows:
[0373] StateTransTab[][]={{0, 2}, {7, 5}, {1, 3}, {6, 4}, {2, 0}, {5, 7}, {3, 1}, {4, 6}}
[0374] Table 20
[0375]
[0376]
[0377] ctu3d_uniform_flag specifies whether to use a uniform method to encode the quantized CTU3D weights QuantParam[]. ctu3d_uniform_flag equal to 1 indicates that the uniform method is used to encode QuantParam[].
[0378] sign_flag specifies whether the quantization weight QuantParam[i] is positive or negative. sign_flag equal to 1 indicates that QuantParam[i] is negative.
[0379] Table 21
[0380]
[0381] sig_flag specifies whether the quantization weight QuantParam[i] is non-zero. sig_flag equal to 0 indicates that QuantParam[i] is zero.
[0382] sign_flag specifies whether the quantization weight QuantParam[i] is positive or negative. sign_flag equal to 1 indicates that QuantParam[i] is negative.
[0383] abs_level_greater_x[j] indicates whether the absolute level of QuantParam[i] is greater than j+1.
[0384] abs_level_greater_x2[j] includes the unary part of the exponential golomb remainder.
[0385] abs_remainder indicates a fixed-length remainder.
Claims
1. A method for neural network decoding, comprising: receiving, from a bitstream of a compressed neural network representation NNR, a first syntax element in an NNR aggregation unit header of a compressed NNR aggregation unit, the first syntax element indicating a coding tree unit (CTU) scanning order for processing tensors in the NNR aggregation unit; as well as Reconstructing a tensor in the NNR aggregation unit based on a CTU scanning order indicated by the first syntax element; A third syntax element is received, the third syntax element indicating whether CTU block partitioning is enabled for a tensor in the NNR aggregation unit.
2. The method according to claim 1, wherein: A first value of the first syntax element indicates that the CTU scanning order is a first raster scanning order along a horizontal direction, and a second value of the first syntax element indicates that the CTU scanning order is a second raster scanning order along a vertical direction.
3. The method according to claim 1, further comprising: A second syntax element in the NNR aggregation unit header is received from the bitstream, the second syntax element indicating a maximum bit depth of quantized coefficients of tensors in the NNR aggregation unit.
4. The method according to any one of claims 1 to 3, wherein: The third syntax element is a model-related syntax element or a tensor-related syntax element, the model-related syntax element is used to specify whether the CTU block partitioning is enabled for the layer of the neural network, and the tensor-related syntax element is used to specify whether the CTU block partitioning is enabled for the tensor in the NNR aggregation unit.
5. The method according to any one of claims 1 to 3, further comprising: A fourth syntax element related to the model or to the tensor is received, the fourth syntax element indicating a CTU dimension of the tensor in the NNR aggregation unit.
6. The method according to any one of claims 1 to 3, further comprising: An NNR unit is received before receiving any NNR aggregation unit, the NNR unit including a fifth syntax element indicating whether CTU partitioning is enabled.
7. A method for neural network decoding, comprising: receiving one or more first syntax elements from a bitstream of a compressed neural network representation NNR, the first syntax elements being associated with a three-dimensional coding unit CU3D partitioned from a first three-dimensional coding tree unit CTU3D, the first CTU3D being obtained from a tensor partition in a neural network, the one or more first syntax elements indicating that the CU3D is partitioned based on a three-dimensional 3D pyramid tree structure comprising a plurality of depths, each depth corresponding to one or more nodes, each node having a node value; receiving a second syntax element corresponding to a node value of each node in the 3D pyramid tree structure, the second syntax element being received from the bitstream in a breadth-first scanning order and used for scanning the nodes in the 3D pyramid tree structure; as well as Based on the received second syntax elements corresponding to the node values of the nodes in the 3D pyramid tree structure, the model parameters of the tensor are reconstructed.
8. The method according to claim 7, wherein: The 3D pyramid tree structure is one of an octree structure, a singular tree structure, a tag tree structure and a singular tag tree structure.
9. The method according to claim 7, wherein: The second syntax element is received starting from a starting depth in the depth of the 3D pyramid tree structure, the starting depth being indicated in the bitstream or inferred at a decoder.
10. The method according to claim 7, further comprising: receiving a third syntax element, the third syntax element indicating a starting depth for receiving the second syntax element, the second syntax element indicating a node value of each node in the 3D pyramid tree structure; and When the starting depth is the last depth of the 3D pyramid tree structure, decoding from the bit stream using a non-3D pyramid tree-based decoding method to obtain the model parameters of the tensor; or When the starting depth is a last depth of the 3D pyramid tree structure, the second syntax element is received starting from a second to last depth among the depths of the 3D pyramid tree structure.
11. The method according to claim 7, further comprising: receiving a third syntax element, the third syntax element indicating a starting depth for receiving the second syntax element, the second syntax element indicating a node value of each node in the 3D pyramid tree structure, Wherein, when the starting depth is the last depth of the 3D pyramid tree structure and the 3D pyramid tree structure is a single-branch label tree structure associated with a single-branch tree partial encoding and a label tree partial encoding, For the singular tree portion encoding, receiving the second syntax element starting from a second to last depth in the depth of the 3D pyramid tree structure, and For the tag tree portion encoding, the second syntax element is received starting from a last depth among the depths of the 3D pyramid tree structure.
12. The method according to any one of claims 7 to 11, further comprising: When the one or more first syntax elements indicate that the CU3D is partitioned based on the 3D pyramid tree structure, dependent quantization is disabled.
13. The method according to any one of claims 7 to 11, further comprising: When the one or more first syntax elements indicate that the CU3D is partitioned based on the 3D pyramid tree structure, a dependent quantization construction process is performed, wherein model parameters of tensors skipped during an encoding process based on the 3D pyramid tree structure are excluded from the dependent quantization construction process.
14. The method according to any one of claims 7 to 11, further comprising: A fourth syntax element associated with the CU3D is received, the fourth syntax element indicating whether all model parameters of the CU3D are unified.
15. The method according to any one of claims 7 to 11, further comprising: A context model for entropy decoding a first coefficient in a kernel of the tensor is determined using a value of zero as a value of a forward neighbor of the first coefficient in the kernel.
16. The method according to any one of claims 7 to 11, further comprising: receiving, in the bitstream, one or more fifth syntax elements indicating a width or a height of a second CTU3D of the tensor; as well as When the width, the height, or both the width and the height are a model parameter, it is determined to decode the model parameter of the second CTU3D based on a baseline encoding method.
17. A method for neural network decoding, comprising: Receiving, in a bitstream of a compressed neural network representation NNR, a first syntax element associated with a three-dimensional coding tree unit CTU3D resulting from a tensor partition in each layer of the neural network, the first syntax element indicating whether all child nodes at a bottom depth of a pyramid tree structure associated with the CTU3D are unified; and When the first syntax element indicates that all child nodes at the bottom depth of the pyramid tree structure associated with the CTU3D are unified, the CTU3D is decoded based on a three-dimensional 3D unary tree encoding method.
18. The method according to claim 17, further comprising: A second syntax element associated with a layer of the neural network is received in the bitstream, the second syntax element indicating whether the layer is encoded using a pyramid tree structure based encoding method.
19. The method according to claim 17, wherein: Child nodes that do not share the same parent node at the bottom depth have different uniform values.
20. The method of claim 17, further comprising: The starting depth of the 3D singular tree encoding method is inferred to be the bottom depth of the pyramid tree structure.
21. The method according to claim 17, wherein: A unified flag for a node at a bottom depth of the pyramid tree structure is not encoded in the bitstream.
22. The method of claim 17, further comprising: receiving a uniform value encoded in the bitstream for all child nodes sharing a same parent node at the bottom depth; as well as Sign bits are received for all child nodes that share a same parent node at the bottom depth, the sign bits following the uniform value in the bitstream.
23. The method of claim 17, further comprising: receiving from the bitstream a uniform value for each set of child nodes sharing a common parent node at the bottom depth; as well as Sign bits of child nodes in each group of child nodes that share a common parent node at the bottom depth are received.
24. The method according to any one of claims 17 to 23, further comprising: In response to the first syntax element indicating that all child nodes at a bottom depth of a pyramid tree structure associated with the CTU3D are not all uniform, the CTU3D is decoded based on a three-dimensional 3D label tree encoding method.
25. The method according to claim 24, further comprising: It is inferred that the starting depth of the 3D label tree encoding method is the bottom depth of the pyramid tree structure.
26. The method according to claim 25, further comprising: The value of a node at the bottom depth of the pyramid tree structure is decoded according to one of the following: receiving a value of a node at a bottom depth of the pyramid tree structure, the value of each node being encoded in the bitstream based on a predetermined scanning order, After receiving the absolute value of each node at the bottom depth of the pyramid tree structure in the bitstream based on a predetermined scanning order, if the absolute value is not zero, receiving the sign of each node, or After receiving the absolute value of each node at the bottom depth of the pyramid tree structure based on a predetermined scan order in the bitstream, if the node has a non-zero value, receiving the sign of each node at the bottom depth of the pyramid tree structure based on the predetermined scan order in the bitstream.
27. A neural network decoding device, comprising: A processing circuit, the processing circuit being configured to: receiving, from a bitstream of a compressed neural network representation NNR, a first syntax element in an NNR aggregation unit header of a compressed NNR aggregation unit, the first syntax element indicating a coding tree unit (CTU) scanning order for processing tensors in the NNR aggregation unit; as well as Reconstructing a tensor in the NNR aggregation unit based on a CTU scanning order indicated by the first syntax element; A third syntax element is received, the third syntax element indicating whether CTU block partitioning is enabled for a tensor in the NNR aggregation unit.
28. A neural network decoding device, comprising: A processing circuit, the processing circuit being configured to: receiving one or more first syntax elements from a bitstream of a compressed neural network representation NNR, the first syntax elements being associated with a three-dimensional coding unit CU3D partitioned from a first three-dimensional coding tree unit CTU3D, the first CTU3D being obtained from a tensor partition in a neural network, the one or more first syntax elements indicating that the CU3D is partitioned based on a three-dimensional 3D pyramid tree structure comprising a plurality of depths, each depth corresponding to one or more nodes, each node having a node value; receiving a second syntax element corresponding to a node value of each node in the 3D pyramid tree structure, the second syntax element being received from the bitstream in a breadth-first scanning order and used for scanning the nodes in the 3D pyramid tree structure; as well as Based on the received second syntax elements corresponding to the node values of the nodes in the 3D pyramid tree structure, the model parameters of the tensor are reconstructed.
29. A neural network decoding device, comprising: A processing circuit, the processing circuit being configured to: Receiving, in a bitstream of a compressed neural network representation NNR, a first syntax element associated with a three-dimensional coding tree unit CTU3D resulting from a tensor partition in each layer of the neural network, the first syntax element indicating whether all child nodes at a bottom depth of a pyramid tree structure associated with the CTU3D are unified; and When the first syntax element indicates that all child nodes at the bottom depth of the pyramid tree structure associated with the CTU3D are unified, the CTU3D is decoded based on a three-dimensional 3D unary tree encoding method.
30. A computer device comprising: a memory configured to store computer program code; as well as At least one processor is configured to access the computer program code, and operate according to the instructions of the computer program code to execute the method according to any one of claims 1 to 6, or the method according to any one of claims 7 to 16, or the method according to any one of claims 17 to 26.
31. A non-transitory computer-readable medium storing instructions, which, when executed by a processor for neural network decoding, cause the processor to perform the method according to any one of claims 1 to 6, or the method according to any one of claims 7 to 16, or the method according to any one of claims 17 to 26.
Citation Information
Patent Citations
Multiple zone scanning order for video coding
CN103636223A
Derived disparity vector in 3D video coding
CN105027571A