Concepts for encoding and decoding neural network parameters
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
- Filing Date
- 2023-01-05
- Publication Date
- 2026-08-06
AI Technical Summary
【0026】 本発明の主な利点を理解するために、最初にニューラルネットワークのトピック及びパラメータコーディングのための関連方法について簡単に紹介する。
Smart Images

Figure 0007901433000105 
Figure 0007901433000106 
Figure 0007901433000107
Abstract
Description
Technical Field
[0001] Some embodiments relate to methods, decoders and / or encoders for entropy coding of neural network parameters and their incremental updates, and in particular to reduced set coding and history-dependent significance coding.
Background Art
[0002] Typically, neural networks have millions of parameters and thus may require hundreds of megabytes to represent. The MPEG-7 Part 17 standard [2] for compression of neural networks for multimedia content description and analysis provides different methods for quantization and integer representation of neural network parameters, such as, for example, independent scalar quantization and codebook-based integer representation. Furthermore, it specifies an entropy quantization scheme also known as deepCABAC [4].
[0003] It is desirable to provide concepts for improved compression of neural networks, for example, to reduce the bitstream of neural networks and thus the signaling cost. In addition or alternatively, it is desirable to make neural network coding more efficient, for example, from the perspective of reducing the bitrate required for coding.
[0004] This is achieved by the subject matter of the independent claims of this application. Further embodiments according to the invention are defined by the subject matter of the dependent claims of this application.
Summary of the Invention
[0005] According to a first aspect of the present invention, the inventors of this application have recognized that one problem encountered when attempting to encode / decode neural network parameters (NN parameters), for example using deepCABAC, is that currently a certain amount of syntax elements must be encoded / decoded for the neural network parameters. This can be costly in terms of memory requirements for storing representations of the neural network. According to a first aspect of this application, this difficulty is overcome by enabling the inference of predetermined syntax elements, which depend on the mapping scheme used in the quantization / dequantization of the neural network parameters. The inventors have found that inferring predetermined syntax elements, rather than encoding / decoding them to / from a data stream, is advantageous in terms of memory requirements / signaling costs and decoding / encoding efficiency. In particular, mapping schemes for mapping quantization indices to reconstruction levels, for example in decoding, or vice versa, for example in encoding, have been found to be good indicators for evaluating whether a given syntax element carries the necessary information and therefore should be encoded / decoded for one of the NN parameters, or whether a given syntax element is obsolete and therefore can be inferred instead of encoded / decoded for each NN parameter. This is based on the idea that it is possible to derive information about the possible values of a group of NN parameters to be encoded / decoded from a mapping scheme selected for a group of NN parameters. Therefore, it is not necessary to decode / encode the values of each NN parameter by decode / encode all syntax elements. Instead, it is possible to infer one or more syntax elements to decode / encode the respective values of each NN parameter.
[0006] Therefore, according to a first aspect of this application, a device for decoding neural network (NN) parameters that define a neural network from a data stream, for example called a decoder, is configured to map quantization indices to reconstruction levels and to obtain a mapping scheme from the data stream for checking whether the mapping scheme satisfies predetermined criteria. The decoding device takes one of the NN parameters, If the mapping method satisfies predetermined criteria, the state of a predetermined syntax element is estimated from the mapping method to be, for example, a predetermined state, such as one or more possible states, and the mapping method is subjected to a quantization index derived using the predetermined syntax element to obtain a reconstruction level, such as the reconstruction level of the NN parameters. If the mapping method does not meet the predetermined criteria, the predetermined syntax elements are derived from the data stream, for example, or read, decoded, To obtain the reconstruction level, for example, the reconstruction level of the NN parameters, the quantization index derived using a given syntax element is applied to the mapping scheme, It is configured to be reconstructed by [a specific method / system].
[0007] Therefore, according to a first aspect of this application, a device for encoding neural network (NN) parameters that define a neural network into a data stream, for example called an encoder, is configured to map reconstruction levels to quantization indices and to obtain a mapping scheme for encoding the mapping scheme into a data stream. The device for encoding may be configured to check whether the mapping scheme satisfies predetermined criteria. The device for encoding applies the reconstruction levels of the NN parameters to the mapping scheme to obtain quantization indices, The mapping method skips the encoding of a predetermined syntax element into the data stream if it satisfies a predetermined criterion, for example, such that the predetermined syntax element is inferred by the decoder, and skips the encoding of a predetermined syntax element if the predetermined syntax element is part of the representation of the quantization index. If the mapping method does not meet the specified criteria, the specified syntax elements are encoded into the data stream, It is configured to encode one of the NN parameters.
[0008] One embodiment relates to a method for decoding / encoding neural network (NN) parameters from / to a data stream. This method is based on the same considerations as the encoder / decoder described above. This method can, incidentally, be completed using all the features and functions described with respect to the encoder / decoder.
[0009] According to a second aspect of the present invention, the inventors of this application have recognized that one problem encountered when attempting to update neural network parameters (NN parameters) stems from the fact that most neural network parameters do not change during updates. This can result in a high bitrate for coding and signaling updates, even if updates must be provided for all NN parameters. According to a second aspect of this application, this difficulty is overcome by selecting a probabilistic model for coding the update parameters of the NN parameters, for example, depending on the current value of each NN parameter and / or the previous update parameters of each NN parameter. The inventors have found that the current NN parameters and previous update parameters of each NN parameter can provide good indication of the update parameters. In particular, they have found that by using such a specially selected probabilistic model, the coding efficiency of each update parameter can be improved. This is based on the idea that the current NN parameters or the history of updates for each NN parameter can enable better modeling of probabilities. Thus, such a selected probabilistic model provides an optimized probability estimate for each update parameter.
[0010] Accordingly, according to a second aspect of this application, a device for decoding neural network (NN) parameters defining a neural network from a data stream, for example called a decoder, is configured to receive an update parameter among the NN parameters and to update the NN parameters using the update parameter. The device is configured to entropically decode the update parameter from the data stream, and is configured to select one of a set of probabilistic models for entropy decoding of the update parameter, for example, depending on a sequence of previous update parameters (e.g., previously received; for example, a sequence includes a sequence of previous update parameters, in which a number of update parameters specialize in the update parameter currently being decoded, and for example, the update parameter may provide an incremental update with respect to the NN parameter; for example, the update parameter is an update in the sense of the concept in Section 2.2.1) and / or depending on the NN parameter.
[0011] Accordingly, according to a second aspect of this application, a device for encoding neural network (NN) parameters defining a neural network into a data stream, for example called an encoder, is configured to acquire one update parameter among the NN parameters. The device is configured to entropically encode the update parameter into the data stream, and the device is configured to select one of a set of probabilistic models for entropy encoding the update parameter, depending on a previous sequence of update parameters of the NN parameters, for example, previously encoded and transmitted, and / or depending on the NN parameter.
[0012] One embodiment relates to a method for decoding / encoding neural network (NN) parameters from / to a data stream. This method is based on the same considerations as the encoder / decoder described above. This method can, incidentally, be completed using all the features and functions described with respect to the encoder / decoder.
[0013] One embodiment relates to a data stream having a picture or video encoded using the encoding method described herein.
[0014] One embodiment relates to a computer program having program code for performing the method described herein when executed on a computer.
[0015] The drawings are not necessarily to scale, and instead, the focus is on illustrating the general principles of the present invention. In the following description, various embodiments of the present invention will be described with reference to the following drawings. [Brief explanation of the drawing]
[0016] [Figure 1] This figure shows one embodiment of a device for decoding NN parameters according to the mapping method. [Figure 2] This figure shows one embodiment of a device for encoding NN parameters according to a mapping scheme. [Figure 3] This figure shows one embodiment of a mapping scheme that includes three positive quantization indices. [Figure 4] This figure shows one embodiment of a mapping scheme that includes two positive quantization indices. [Figure 5] This figure shows one embodiment of a mapping scheme that includes negative quantization indices. [Figure 6] This figure shows one embodiment of a mapping scheme that includes positive and negative quantization indices. [Figure 7]FIG. is a diagram showing an embodiment of a mapping method including two negative quantization indices. [Figure 8] FIG. is a diagram showing an embodiment of decoding of quantization indices. [Figure 9] FIG. is a diagram showing an embodiment of an apparatus for decoding update parameters of NN parameters. [Figure 10] FIG. is a diagram showing an embodiment of an apparatus for encoding update parameters of NN parameters. [Figure 11] FIG. is a diagram showing an embodiment of decoding of quantization indices representing update parameters of NN parameters. [Figure 12] FIG. is a diagram showing an embodiment of a method for decoding NN parameters according to a mapping method. [Figure 13] FIG. is a diagram showing an embodiment of a method for encoding NN parameters according to a mapping method. [Figure 14] FIG. is a diagram showing an embodiment of a method for decoding update parameters of NN parameters. [Figure 15] FIG. is a diagram showing an embodiment of a method for encoding update parameters of NN parameters. [Figure 16] FIG. is a diagram showing a neural network. [Figure 17] FIG. is a diagram showing a uniform reconstruction quantizer. [Figure 18] FIG. is a diagram showing encoding of an integer codebook. DETAILED DESCRIPTION OF THE INVENTION
[0017] Equivalent or equal elements, or elements having equivalent or equal functions, even if they occur in different figures, are denoted by the same or equivalent reference numerals in the following description.
[0018] The embodiments described below are detailed, but it should be understood that the embodiments provide many applicable concepts that can be implemented with a wide variety of coding concepts. The specific embodiments described are merely examples of specific ways of implementing and using the concepts and do not limit the scope of the embodiments. Several details are provided in the following description to provide a more complete description of the embodiments of this disclosure. However, it will be apparent to those skilled in the art that other embodiments can be implemented without these specific details. In other examples, well-known structures and devices are shown in block diagram form rather than in detail, so as not to obscure the examples described herein. Furthermore, the features of the different embodiments described herein can be combined with each other unless otherwise specified.
[0019] Further embodiments are described in the claims and the accompanying descriptions.
[0020] Embodiments of the present invention describe a method for parameter coding of a full set of neural network parameters or for incremental updating of a set of neural network parameters, or more specifically, for encoding / decoding integer indices related to the parameters of a neural network. These integer indices may be the output of a quantization process prior to the encoding stage.
[0021] Such integer indices can, for example, indicate a quantization level that can be multiplied by the floating-point quantization step size to generate a reconstructed version of the model, or they can specify an index that is then mapped to reconstructed weight values using the codebook.
[0022] Typically, integer indices are encoded using entropy coding methods such as DeepCABAC, which is also part of the MPEG-7 Part 17 standard for neural network compression for multimedia content description and analysis.[2]
[0023] Furthermore, the methods described may be used in combination with all existing methods for neural network compression (for example, those given in the MPEG-7 Part 17 standard for neural network compression for multimedia content description and analysis [2]), provided that the previously given requirements are met.
[0024] The described method addresses the entropy coding stage, or more specifically, the binarization process and context modeling, as defined by, for example, the DeepCABAC entropy coding method based on CABAC[3]. However, the method is applicable to all entropy coding methods that use similar binarization or context modeling processes.
[0025] The methodology of the apparatus can be divided into the following distinct main parts: 1. Integer representation (quantization) of neural network parameters 2. Binarization and Lossless Coding 3. Lossless decoding
[0026] To understand the main advantages of this invention, we will first briefly introduce the topic of neural networks and related methods for parameter coding.
[0027] 1. Applicable Area In their most basic forms, neural networks consist of a chain of affine transformations followed by element-wise nonlinear functions. They can be represented as directed acyclic graphs, as shown in Figure 16. Each node has a specific value that is propagated forward to the next node by multiplication with the respective weight values of the edges. All input values are then simply aggregated. Figure 16 is a diagram of a two-layer feedforward neural network, i.e., Figure 16 shows the graph representation of a feedforward neural network. Specifically, this two-layer neural network is a nonlinear function that maps a four-dimensional input vector to real lines.
[0028] Mathematically, the neural network shown in Figure 16 calculates its output using the following method. TIFF0007901433000001.tif8105 Here, W1 and W2 are the weight parameters (edge weights) of the neural network, and sigma σ is some nonlinear function. For example, so-called convolutional layers can also be used by casting them as matrix-matrix products as described in [1]. Incremental updates are usually intended to provide updates to the weights of W1 and W2 and may be the result of an additional training process. The updated versions of W1 and W2 usually yield a modified output. Below, the procedure for calculating an output from a given input will be referred to as inference. Intermediate results will also be referred to as hidden layers or hidden activation values, which constitute linear transformations + element-wise nonlinearities, e.g., the calculation of the first dot product + nonlinearity described above.
[0029] Typically, neural networks have millions of parameters, which can require hundreds of megabytes to represent. Consequently, their inference procedures involve calculating numerous inner product operations between large matrices, thus requiring significant computational resources to execute. Therefore, reducing the complexity of performing these inner products is crucial.
[0030] 2. The concept of coding This section describes concepts that may be optionally implemented in embodiments of the present invention. These concepts may be implemented individually or in combination.
[0031] 2.1 Quantization, Integer Representation, and Entropy Coding For example, each embodiment may be implemented in accordance with the MPEG-7 Part 17 standard for Compression of Neural Networks for Multimedia Content Description and Analysis[2]. This standard provides different methods for quantization and integer representation of neural network parameters, such as independent scalar quantization and codebook-based integer representation. Furthermore, it specifies an entropy quantization scheme also known as deepCABAC[4]. These methods are briefly summarized in this section for better understanding. Further details can be found in [2].
[0032] 2.1.1 Scalar Quantizer Neural network parameters can be quantized using a scalar quantizer. As a result of quantization, the set of acceptable values for the parameters is reduced. In other words, neural network parameters are mapped to a countable set (actually a finite set) of so-called reconstruction levels. The set of reconstruction levels represents a suitable subset of the set of possible neural network parameter values. To simplify the entropy coding below, the acceptable reconstruction levels are represented by quantization indices transmitted as part of the bitstream. On the decoder side, the quantization indices are mapped to the reconstructed neural network parameters. The possible values of the reconstructed neural network parameters correspond to the set of reconstruction levels. On the encoder side, the result of scalar quantization is a set of (integer) quantization indices. Figure 17 shows a uniform reconstruction quantizer.
[0033] In embodiments of the present invention, a Uniform Reconstruction Quantizer (URQ) can be optionally used. Its basic design is shown in Figure 17. The URQ has the characteristic that reconstruction levels are arranged at equal intervals. The distance Δ between two adjacent reconstruction levels is called the quantization step size. One of the reconstruction levels is equal to 0. Thus, the complete set of available reconstruction levels is uniquely specified by the quantization step size Δ. The decoder mapping of the quantization index q to the reconstructed weight parameter t' is, in principle, given by the following simple equation. In this context, the term "independent scalar quantization" refers to the property that, given the quantization index q of any weight parameter, the associated reconstructed weight parameter t' can be determined independently of all the quantization indices of the other weight parameters.
[0034] 2.1.2 Integer Codebook Representation Instead of directly transmitting quantization levels, a codebook can be used to map them to integer indices. Each integer index specifies a location in the codebook that contains the associated quantization level. The integer indices and codebook are transmitted as part of the bitstream. On the decoder side, the integer indices are mapped to quantization levels using the codebook (table lookup), and then the quantization levels are mapped to the reconstructed neural network parameters using the same method as for scalar quantization (multiplying by the quantization step size, which is also part of the bitstream).
[0035] Codebook representation can be useful when there are few intrinsic weight parameters but the quantization is very fine-grained. In these cases, scalar quantization can result in several intrinsic but large quantization levels being transmitted, as described in Section 2.1.1, for example. Using codebook representation, the transmitted values can be mapped to smaller indices.
[0036] In other words, according to the embodiment, the neural network parameters may be quantized using scalar quantization to obtain a quantization level, which is then mapped to a quantization index using a codebook, for example, as described below.
[0037] The integer indices are encoded using the entropy encoding stage, i.e., deepCABAC, in the same way that the quantization levels output by the quantization stage are encoded.
[0038] The integer codebook is encoded as described below. `integer_codebook()` is defined as shown in Figure 18.
[0039] For example, the function integerCodebook[j] returns, for each quantization index, the associated reconstruction level or quantization level value associated with that quantization index. The quantization index forms a monotonic sequence of integers, and the values of the quantization index are associated with their positions in the sequence of integers by the variable cbZeroOffset. For example, as seen in Figure 18, the sequence of integer values to be encoded is contained in the vector integerCodebook, and they are strictly monotonic. For encoding, the encoder encodes a given integer value codebook_0_value, placed at a given position cbZeroOffset in the sequence of integer values, into the data stream. In the example, the given position is the position in the monotonic sequence of integers associated with the value 0.
[0040] In the example in Figure 18, the first for loop first encodes the positions of the sequence of integer values prior to a given position cbZeroOffset. Specifically, for each such position j, the first difference is calculated between the integer value immediately preceding each position, which is stored in previousValue as the for loop traverses these positions toward the beginning of the sequence of integer values, and the integer value integerCodebook[j] at each position, reduced by 1. This first difference is encoded into the data stream, i.e., codebook_delta_left = previousValue - integerCodebook[j] - 1. Then, in the second for loop, the position of the sequence of integer values following the given position is first encoded. For each such position in the sequence of integer values, the integer value integerCodebook[j] at that position and the integer value immediately preceding that position, which is stored in previousValue when the for loop traverses these positions toward the end of the sequence of integer values and initializes them with the given integer value, are calculated immediately before the second for loop, the given integer value codebook_0_value is calculated, the second difference with the value reduced by 1 is calculated and the second difference is encoded into the data stream, i.e., codebook_delta_right = integerCodebook[j] - previousValue - 1. As mentioned above, the order in which the differences are encoded can be switched or even alternated. To decode the sequence of integer values, the given integer value codebook_0_value located at a given position cbZeroOffset in the sequence of integer values is decoded from the data stream.Next, for each position in the sequence of integer values preceding a given position, the first difference between the integer value immediately preceding that position and the integer value at that position reduced by 1 is decoded from the data stream, i.e., codebook_delta_left = previousValue - integerCodebook[j] - 1. For each position in the sequence of integer values following the given position, the second difference between the integer value at that position and the integer value immediately preceding that position reduced by 1 is decoded from the data stream, i.e., codebook_delta_right = integerCodebook[j] - previousValue - 1.
[0041] More specifically, the number of integer values in a sequence of integer values is also encoded into a data stream, i.e., codebook_size. This is done using a variable-length code, i.e., a quadratic Exp-Golomb code.
[0042] Information about a given position in an encoded sequence of integer values, codebook_centre_offset, can also be encoded in the data stream. This encoding is done differently for the center position of the sequence. That is, cbZeroOffset - (codebook_size>>1) = codebook_centre_offset is encoded, i.e., the difference between the rank of the given position cbZeroOffset and the integer rounded half of the number of integer values, i.e., codebook_size>>1. This is done using a variable-length code, i.e., a quadratic Exp-Golomb code.
[0043] The given integer value codebook_0_value is encoded using the 7th-order Exp-Golomb code.
[0044] The difference between the first and second is coded using the k-th order Exp-Golomb code, where k is coded into the data stream as codebook_egk. This is coded as a 4-bit unsigned integer.
[0045] An integer codebook, for example, is defined by the variables cbZeroOffset and integerCodebook - a predetermined position in the integer codebook sequence, for example, z, and a predetermined integer value located at that position, for example, C(z).
[0046] The parameters that define a sequence, such as a codebook, include an exponential Golomb code parameter, such as Exp-Golomb code parameter k (codebook_egk), and the number of integer values in the sequence, such as the number of elements in the codebook (codebook_size). These parameters are decoded from a data stream, such as a bitstream, used to create a determined sequence of integer values.
[0047] The given position (cbZeroOffset) is a variable (codebook_centre_offset) calculated using the difference (codebook_centre_offset) between the rank of the given position and the integer rounded half of the number of integer values, which is encoded within the bitstream. In one embodiment, the variable codebook_centre_offset is defined as a third difference, e.g., y. The variable codebook_centre_offset specifies an offset for accessing integer values in a sequence, e.g., elements in a codebook, relative to the center of the sequence, e.g., a codebook. The difference (codebook_centre_offset) is decoded from a data stream, e.g., a bitstream, used to create a determined sequence of integer values.
[0048] The parameter `codebook_0_value`, which defines an encoded sequence, such as a codebook, specifies a predetermined integer value (`integerCodebook`) located at a given position (`cbZeroOffset`), such as the codebook value at position `CbZeroOffset`. This parameter is involved in the creation of a decoded sequence of integer values, such as the variable `codebook` (an array representing the codebook).
[0049] When generating the decoded sequence, the first difference (codebook_delta_left) and the second difference (codebook_delta_right) are decoded from the data stream, such as a bitstream.
[0050] The first difference (codebook_delta_left) specifies the difference between the integer value at each position and the integer value immediately preceding it. This difference is subtracted by 1 for each position in the sequence of integer values preceding the given position (cbZeroOffset). For example, the difference between the codebook value and its right-hand neighbor is obtained by subtracting 1 from the value to the left toward the center position. The first difference (codebook_delta_left) is involved in creating the decoded sequence of integer values, such as the variable Codebook (an array representing the codebook), as shown in Figure 18. For each position in the sequence of positions preceding the given position (cbZeroOffset), the integer value at each position is calculated by linearly summing the first difference (codebook_delta_left), the integer value immediately preceding that position (previousValue=integerCodebook[j+1]), and 1.
[0051] integerCodebook[j]=previousValue-codebook_delta_left-1 The second difference (codebook_delta_right) specifies the difference between the integer value at each position and the integer value immediately preceding that position, and is decremented by 1 for each position in the sequence of integer values following the given position (cbZeroOffset). For example, the difference between the codebook value and its left neighbor is calculated by subtracting 1 from the value to the right of the center position. The second difference is involved in creating the decoded sequence of integer values, e.g., the variable Codebook (an array representing the codebook), as shown in Figure 18. For each position in the sequence of integer values following the given position (cbZeroOffset), the integer value at each position is calculated by linearly summing the second difference (codebook_delta_right), the integer value immediately preceding that position (previousValue=integerCodebook[j-1]), and 1.
[0052] integerCodebook[j]=previousValue+codebook_delta_right+1. The exponential Golomb code parameter (codebook_egk) is used to decode the syntax elements codebook_delta_left, which define the first difference, and codebook_delta_right, which defines the second difference.
[0053] 2.1.3 Entropy coding As a result of the quantization applied in the previous step, the weight parameters are mapped to a finite set of so-called reconstruction levels. These can be represented by (integer) quantizer indices (also called quantization indices (e.g., in scalar quantization described above), integer indices (e.g., when an integer codebook is used as described above), or parameter levels or weight levels) and quantization step sizes, which may, for example, be fixed for the entire layer. To reconstruct all the quantization weight parameters of a layer, the step size and dimensions of the layer can be known by the decoder. They may, for example, be transmitted separately.
[0054] 2.1.3.1 Binarization and Encoding of Quantized or Integer Indexes Using Context-Adaptive Binary Arithmetic Coding (CABAC) The quantized index (integer representation) or (codebook) integer index is then transmitted using entropy coding techniques. Thus, the weight layers are mapped using scans to sequences of quantized weight levels, e.g., quantized indices or integer indices. For example, a row-first scan order can be used, starting from the top row of the matrix and encoding the contained values from left to right. In this way, all rows are encoded from top to bottom. It is also possible to apply any other scan. For example, the matrix can be rotated horizontally and / or vertically and / or left or right by 90 / 180 / 270 degrees and transposed or flipped before applying a row-first scan. Optionally, the quantization levels can be mapped to integer indices using codebooks. Below, we refer to the values encoded as indices, whether they are quantized weight levels or codebook integer indices.
[0055] For example, NN parameters, such as the NN parameters of one layer of a neural network, can be represented by a 2D matrix where each row contains the NN parameters associated with exactly one output neuron of the neural network. For example, for one output neuron, the row may include one or more parameters, such as weights, for each of the input neurons of the output neuron.
[0056] For example, CABAC (Context-Adaptive Binary Arithmetic Coding) is used to encode the index. See [2] for details. Therefore, the index being sent TIFF0007901433000003.tif73 can be broken down into a series of binary symbols or syntax elements and then passed to a Binary Arithmetic Encoder (CABAC).
[0057] In the first step, a binary syntax element sig_flag of the index is derived that specifies whether the corresponding index is equal to 0. If sig_flag is equal to 1, a further binary syntax element sign_flag is derived. The bin indicates whether the current index is positive (e.g., bin = 0) or negative (e.g., bin = 1).
[0058] Next, a single-term sequence of the bin is encoded, followed by encoding of a fixed-length sequence as follows.
[0059] The variable k is initialized with a non-negative integer, and X is initialized with 1 << k.
[0060] One or more syntax elements abs_level_greater_X indicating that the absolute value of the index is greater than X are encoded. If abs_level_greater_X is equal to 1, the variable k is updated (e.g., incremented by 1), then 1 << k is added to X, and further abs_level_greater_X are encoded. This procedure continues until abs_level_greater_X becomes equal to 0. Thereafter, a fixed-length code of length k is sufficient to complete the encoding of the index. For example, the variable TIFF0007901433000004.tif727 can be encoded using k bits. Alternatively, the variable TIFF0007901433000005.tif79 is encoded using k bits and can be defined as TIFF0007901433000006.tif754. Alternatively, any other mapping of the variable TIFF0007901433000007.tif78 to a fixed-length code of k bits may be used.
[0061] If we increment k by 1 after each abs_level_greater_X, this method is equivalent to applying exponential Golomb coding (if sign_flag is not considered).
[0062] 2.1.3.2 Decoding Quantized Indexes Using Context-Adaptive Binary Arithmetic Coding (CABAC) Decoding the index (integer representation) works similarly to encoding. The decoder first decodes sig_flag. If this is equal to 1, it is followed by a single-term sequence of sign_flag and abs_level_greater_X, where the update of k (and therefore the increment of X) must follow the same rules as the encoder. Finally, it decodes a fixed-length code of k bits, resulting in an integer (for example, TIFF0007901433000008.tif78 and Depending on which of the two, TIFF0007901433000009.tif79, was encoded, TIFF0007901433000010.tif78 or Interpret as TIFF0007901433000011.tif79. Then, decode the index. The absolute value of TIFF0007901433000012.tif76 can be reconstructed from X to form a fixed-length portion. For example, If TIFF0007901433000013.tif78 is used as a fixed-length portion, It is TIFF0007901433000014.tif727. Or, If TIFF0007901433000015.tif79 is encoded, The filename is TIFF0007901433000016.tif759. As a final step, the code is determined according to the decoded sign_flag. Applicable to TIFF0007901433000017.tif76, index It is necessary to generate TIFF0007901433000018.tif73. Then, whenever the index refers to the codebook index, the weight levels quantized by the codebook table lookup are used. You can obtain TIFF0007901433000019.tif713. Finally, the quantized weight level Step size in TIFF0007901433000020.tif713 The quantized weights are obtained by multiplying by TIFF0007901433000021.tif73. TIFF0007901433000022.tif74 will be reconstructed.
[0063] Context Modeling In CABAC entropy coding, most syntax elements of the quantized weight levels are coded using binary probabilistic modeling. Each binary decision (bin) is associated with a context, which represents a probabilistic model of the class of coded bins. The probability of one of two possible bin values is estimated for each context based on the bin values already coded in the corresponding context. Different context modeling techniques can be applied depending on the application. Typically, for several bins associated with coding the quantized weights, the context used for coding is selected based on the syntax elements already submitted. Depending on the actual application, different probabilistic estimators can be selected, e.g., SBMP[4], or HEVC[5] or VTM-4.0[6]. The selection affects, for example, compression efficiency and complexity.
[0064] The following describes a context modeling method suitable for a wide range of neural networks. It involves the level of quantized weights at a specific position (x,y) within the weight matrix (layer). To decode TIFF0007901433000023.tif73, a local template is applied to the current position. This template contains several other (ordered) positions, such as (x-1,y), (x,y-1), (x-1,y-1), etc. For each position, a state identifier is derived.
[0065] In the implemented deformation form (indicated as Si1), the position (x,y) state identifier is used. TIFF0007901433000024.tif77 is derived as follows: When position (x,y) points outside the matrix, or the level of quantized weights at position (x,y) If TIFF0007901433000025.tif77 has not yet been decrypted or is equal to 0, the status identifier The filename is TIFF0007901433000026.tif715. Otherwise, the status identifier is Let's assume the filename is TIFF0007901433000027.tif740.
[0066] For a given template, a sequence of state identifiers is derived, and each possible constellation of the state identifier values is mapped to a context index to identify the context in which it is used. The template and mapping may differ for different syntax elements. For example, from a template containing (ordered) positions (x-1,y), (x,y-1), (x-1,y-1), an ordered sequence of state identifiers is derived. TIFF0007901433000028.tif711, TIFF0007901433000029.tif711, TIFF0007901433000030.tif715 is derived. For example, this sequence is a context index It may also be mapped to TIFF0007901433000031.tif768. For example, context index TIFF0007901433000032.tif73 can be used to identify the number of contexts for sig_flag.
[0067] In one modified form (shown in Method 1), the level of quantized weights at position (x,y) The local template for sig_flag or sign_flag in TIFF0007901433000033.tif77 consists of only one position (x-1, y) (i.e., the leftmost position). Related state identifiers TIFF0007901433000034.tif711 is derived according to the implemented modification form Si1.
[0068] Regarding sig_flag, Depending on the value of TIFF0007901433000035.tif711, one of three contexts is selected, or for sign_flag, Depending on the value of TIFF0007901433000036.tif711, one of the other three contexts will be selected.
[0069] In another variant implementation (shown in Method 2), the local template for the sig flag includes three ordered positions (x-1, y), (x-2, y), and (x-3, y). State identifier TIFF0007901433000037.tif711, The associated sequence for TIFF0007901433000038.tif723 is derived according to the implemented variant Si2.
[0070] Regarding sig_flag, context index TIFF0007901433000039.tif73 is derived as follows: In the case of TIFF0007901433000040.tif719, It is TIFF0007901433000041.tif711. Otherwise, In the case of TIFF0007901433000042.tif719, It is TIFF0007901433000043.tif711. Otherwise, In the case of TIFF0007901433000044.tif719, It is TIFF0007901433000045.tif711. Otherwise, The filename is TIFF0007901433000046.tif711.
[0071] This can also be expressed by the following formula. Similarly to TIFF0007901433000047.tif13116, the number of left neighbors can be increased or decreased such that the context index C is equal to the distance to the next non-zero weight (not exceeding the template size) to the left.
[0072] 2.2 Incremental Neural Network Encoding 2.2.1 Concepts of Base Model and Update Model This concept introduces the neural network model from Section 1, which can be considered a complete model in the sense that it can compute an output for a given input. It is denoted as TIFF0007901433000048.tif76. Each base model is the base layer It consists of layers, represented as TIFF0007901433000049.tif729. The base layer includes, for example, base values that can be selected so that they can be efficiently represented or compressed / transmitted in the first step. Furthermore, this concept is applied to the update model ( TIFF0007901433000050.tif735 introduces an architecture that may be similar to, or even identical to, the base model. The updated model does not have to be a complete model in the sense described above. Instead, the updated model is combined with the base model using a synthesis method, thereby making them a new complete model. TIFF0007901433000051.tif78 can be formed. This model itself can serve as a base model for further updated models. Updated model TIFF0007901433000052.tif78 is a refresh layer It consists of layers called TIFF0007901433000053.tif737. The update layer contains base values that can be selected, for example, so that they can be represented or compressed / transmitted separately and efficiently.
[0073] The updated model may be the result of an (additional) training process applied to the base model on the encoder side. Depending on the type of update provided by the updated model, several synthesis methods can be applied. It should be noted that the methods described herein are not limited to any particular type of update / synthesis method and are applicable to any architecture using a base model / updated model technique.
[0074] In a preferred embodiment, TIFF0007901433000054.tif73rd updated model TIFF0007901433000055.tif78 is a new model layer according to the following: Base model to form TIFF0007901433000056.tif710 A layer with a difference value (also called an incremental update) is added to the corresponding layer of TIFF0007901433000057.tif76. Includes TIFF0007901433000058.tif79 (index j identifies individual layers). TIFF0007901433000059.tif754 The new model layer forms a new (updated) model, which then serves as the base model for subsequent incremental updates that are sent separately.
[0075] In a more preferred embodiment, TIFF0007901433000060.tif73 The third updated model is the new model according to the following Corresponding base layer to form TIFF0007901433000061.tif79 A layer having a scaling factor value multiplied by the value of TIFF0007901433000062.tif77 Includes TIFF0007901433000063.tif79. TIFF0007901433000064.tif752 The new model layer forms a new (updated) model, which then serves as the base model for subsequent incremental updates that are sent separately.
[0076] In some cases, the update model may also update one or more existing layers (i.e., for layer k) instead of updating the layers as described above. Please note that it is possible to include a new layer that replaces TIFF0007901433000065.tif742).
[0077] 2.2.2 Neural Network Parameter Encoding for Incremental Updates The base model concept and one or more incremental updates can be utilized in the entropy coding stage to improve coding efficiency. Layer parameters are typically represented by multidimensional tensors. For the coding process, all tensors are mapped to a 2D matrix, typically as entities such as rows and columns. This 2D matrix is then traversed in a predetermined order, and the parameters are coded / transmitted. Note that the method described below is not limited to 2D matrices. This method is applicable to all representations of neural network parameters that provide parameter entities of known size, such as rows, columns, blocks, and / or combinations thereof. The 2D matrix representation is used below to better understand the method.
[0078] In a preferred embodiment, the layer parameters are represented as a 2D matrix providing value entities such as rows and columns.
[0079] 2.2.2.1 Improved Context Modeling for Base Model Update Model Structure The concepts of a base model and one or more update models can be utilized in the entropy coding stage. The methods described herein are applicable to any entropy coding scheme that uses a context model, such as those described in Section 2.1.4.
[0080] Typically, separate update models (and base models) are available correlated on the encoder and decoder sides. This can be used in the context modeling phase to improve coding efficiency by providing new context models and methods for context model selection.
[0081] In a preferred embodiment, binarization (sig_flag, sign_flag, etc.), context modeling, and encoding schemes according to Section 2.1.3.1 are applied.
[0082] In another preferred embodiment, a given number of context models (context sets) of symbols to be encoded are duplicated to form two or more sets of context models. Then, a set of context models is selected based on the values of parameters located in the same place within the corresponding layers of a particular previously encoded update or base model. This is because the parameters located in the same place are first threshold values. The first set is selected if it is lower than TIFF0007901433000066.tif75, and the value is the threshold. If TIFF0007901433000067.tif75 or higher, the second set is selected, and the value is the threshold. This means that a third set is selected if the value is TIFF0007901433000068.tif75 or higher, for example. This procedure can be applied with higher or lower thresholds.
[0083] In a preferred embodiment equivalent to the above-described embodiment, a single threshold The file TIFF0007901433000069.tif713 will be used.
[0084] In a more preferred embodiment, a given number of context models (context sets) of symbols to be encoded are duplicated to form two or more sets of context models. Then, a set of context models is selected based on the absolute value of a parameter located in the same place within the corresponding layer of a particular previously encoded update or base model. This is where the absolute value of the parameter located in the same place is a first threshold. The first set is selected if it is lower than TIFF0007901433000070.tif75, and the absolute value is another threshold. The second set is selected if TIFF0007901433000071.tif75 or greater, and the absolute value is the threshold. This means that a third set is selected if the value is TIFF0007901433000072.tif75 or higher, for example. This procedure can be applied with higher or lower thresholds.
[0085] In a preferred embodiment equivalent to the previous embodiment, a sig_flag indicating whether the current value to be encoded is equal to 0 is encoded, which employs a set of context models. The embodiment uses a single threshold Use TIFF0007901433000073.tif713.
[0086] Another preferred embodiment is equivalent to the previous embodiment, but instead of sig_flag, sign_flag, which indicates the sign of the current value to be encoded, is encoded.
[0087] A more preferred embodiment is the same as the previous embodiment, but instead of sig_flag, abs_level_greater_X is encoded, which indicates whether the current value being encoded is greater than X.
[0088] In a more preferred embodiment, a given number of context models (context sets) of symbols to be encoded is doubled to form two sets of context models. Next, a set of context models is selected depending on whether or not there is a corresponding previously encoded update (or base) model. If there is no corresponding previously encoded update (or base) model, the first set of context models is selected; otherwise, the second set is selected.
[0089] In another preferred embodiment, a context model from a set of context models for syntax elements is selected based on the value of a parameter located in the same place in a particular corresponding previously encoded updated (or base) model. This is because the parameter located in the same place is a threshold. The first model is selected if the value is lower than TIFF0007901433000074.tif75, and the value is the threshold. If TIFF0007901433000075.tif75 or higher, the second model is selected, and the value is a different threshold. This means that a third set is selected if the value is TIFF0007901433000076.tif75 or higher, for example. This procedure can be applied with higher or lower thresholds.
[0090] In a preferred embodiment equivalent to the previous embodiment, sign_flag, which indicates the sign of the current value to be encoded, is encoded. The first threshold for the context model selection process is The second threshold is TIFF0007901433000077.tif713. The filename is TIFF0007901433000078.tif713.
[0091] In another preferred embodiment, a context model from a set of context models for syntax elements is selected based on the absolute value of a parameter located in the same place in a particular corresponding previously encoded updated (or base) model. This is where the absolute value of the parameter located in the same place is a threshold. The first model is selected if the value is lower than TIFF0007901433000079.tif75, and the value is the threshold. If TIFF0007901433000080.tif75 or higher, the second model is selected, and the value is the threshold. This means that a third model is selected if the value is TIFF0007901433000081.tif75 or higher. This procedure can be applied with more or fewer thresholds.
[0092] In a preferred embodiment equivalent to the previous embodiment, a sig_flag indicating whether the current value to be encoded is equal to 0 is encoded. The first threshold set in TIFF0007901433000082.tif713 and The second threshold set in TIFF0007901433000083.tif713 will be adopted.
[0093] In another preferred embodiment, equivalent to the previous embodiment, instead of sig_flag, the abs_level_greater_X flag, which indicates whether the current value being encoded is greater than X, is encoded. Furthermore, Only one threshold, set in TIFF0007901433000084.tif713, will be used.
[0094] It should be noted that any of the above embodiments can be combined with one or more of the other embodiments.
[0095] 3. Embodiment 1 of the Invention: Encoding of a Reduced Set of Values This invention describes a modified binarization scheme for encoding quantized weight levels or integer indices, respectively, given information about the possible values to be encoded / decoded available to the encoder and decoder. For example, all the indices to be encoded / decoded may be non-positive or non-negative. In such cases, sign_flag can be considered obsolete as it does not carry information. Another example is when the maximum absolute level of all transmitted values is given. In this case, the number of abs_level_greater_X_flags may be reduced.
[0096] The next section (Section 3.1) describes the modified binarization scheme, assuming that the necessary information is available to the encoder and decoder. Section 3.2 then describes the concepts of how to derive the information, more specifically, the value of sign_flag and the maximum absolute level on the decoder side.
[0097] Figure 1 shows a device 10 for decoding NN parameters from a data stream 14, for example, referred to as a decoder 10, according to one embodiment. The decoder 10 comprises a mapping module 40, which is used to map the quantization index 32, which encodes the NN parameters in the data stream 14, to the reconstruction level 42 of the NN parameters. The reconstruction level 42 may be dequantized to obtain the NN parameters (for example, a scalar quantizer that can be optionally characterized by quantization parameters that can be signaled in the data stream, e.g., the scalar quantizer described in Section 2.1.1), but is not necessarily so, and in other examples, the mapping 40 can directly map the quantization index to the value of the NN parameter. The mapping module 40 uses a mapping scheme 44 obtained by the decoder 10 from the data stream 14 to map the quantization index to the reconstruction level.
[0098] According to one embodiment, the decoder 10 is configured to dequantize or scale the reconstruction level 42 using, for example, the quantization step size, for example, the quantization parameter, so as to obtain reconstructed values of the NN parameters that can correspond to the values of the NN parameters encoded by the encoder 11 despite quantization losses. For example, the quantization index and the reconstruction level may be integers, and as a result, the mapping scheme is an integer-to-integer mapping.
[0099] The decoder 10 includes a mapping scheme analyzer 34 that checks whether the mapping scheme 44 satisfies predetermined criteria. The decoder 10 further includes a decoding module 30 configured to derive a quantization index 32 based on one or more syntax elements that may include predetermined syntax elements 22. Depending on whether the mapping scheme 44 satisfies predetermined criteria, the decoding module 30 derives predetermined syntax elements 22 from the data stream 14 or infers the state of predetermined syntax elements 22. For this purpose, the decoding module 30 may include an inference module 24 for providing the state of predetermined syntax elements 22. For example, if the mapping scheme 44 satisfies predetermined criteria, the decoder 10 may assume that the state is a predetermined state, i.e., the inference module 24 may provide a fixed state in that case. In other examples, if the mapping scheme 44 satisfies certain criteria, the state of a given syntax element 22 may be determined using further information, such as already decoded syntax elements, or properties of the mapping scheme 44 (for example, in the case of the sign flag described below, whether all indices in the codebook are non-positive or non-negative). Deriving a given syntax element 22 from the data stream 14 may be performed by the decoding module 30 using, for example, entropy decoding, e.g., context-adaptive binary arithmetic coding, or CABAC.
[0100] Figure 2 shows an encoder 11 according to one embodiment, which may be an encoder corresponding to a decoder 10, i.e., an encoder for encoding the data stream 14 in Figure 1. For this purpose, the encoder 11 can perform the inverse operation of the decoder 10. The encoder 10 includes a mapping module 41 for applying a reconstruction level to a mapping scheme 44 to obtain a quantization index 32. The mapping module 41 can apply the mapping scheme 44 in reverse, as the mapping module 40 in Figure 1. The encoder 10 further includes a coding module 31 for encoding the quantization index 32 into the data stream by using one or more syntax elements, including a predetermined syntax element 22, as described with respect to Figure 1, for example. If the mapping scheme 44 satisfies a predetermined criterion, the coding module 31 skips, suppresses, or refrains from encoding the predetermined syntax element 22 into the data stream 14. For example, the encoder 10 may include a mapping scheme analyzer 34 for checking a predetermined criterion on the mapping scheme 44. The encoder 10 may further include a mapping method acquisition unit 48 for acquiring a mapping method 44. For example, the mapping method acquisition unit 48 can acquire a mapping method 44 based on the entire range of reconstruction levels, e.g., the range of values in which the reconstruction levels are included, and / or the number of reconstruction levels, and / or the step size of the reconstruction levels. The encoder 11 can be configured to acquire the reconstruction levels of the NN parameters by quantizing or scaling the NN parameters, e.g., using a quantization step size, e.g., a quantization parameter.
[0101] Examples of mapping methods 44 that can be used with the decoder 10 in Figure 1 and the encoder 11 in Figure 2 are shown in Figures 3 to 7.
[0102] Figure 3 shows a mapping scheme 44 according to one embodiment, including an exemplary number or three quantization indices 32. In other words, the size of the mapping scheme 44, or codebook, is 3. Each quantization index 32 is associated with a corresponding reconstruction level 42. Note that the mapping between quantization indices 32 and reconstruction levels 42 is bijective. For example, each quantization index 32 has a position 33 in the sequence of quantization indices 32. Positions 33 may be indicated by position indices, for example indices 1, 2, and 3 in Figure 3, but instead, the numbering begins at index 0. According to this embodiment, the quantization indices 32 form a monotonic sequence of integers, and for example, the values of the quantization indices are monotonic with respect to their positions in the sequence.
[0103] According to one embodiment, the mapping scheme 44 may be signaled within the data stream 14 or characterized by indicating the size of the sequence and the position 33* of a predetermined quantization index 32*, for example, a quantization index having a value of 0 or any other predetermined value as shown in Figure 3. The size of the sequence may be indicated by the following codebook_size. For example, the size of the sequence can represent the number or count of quantization indices contained in the sequence, for example, 3 as shown in Figure 3. The position 33* may be indicated by the syntax element codebook_centre_offset or by the following variable cbZeroOffset.
[0104] According to one embodiment, the mapping scheme 44 may be further signaled by indicating the value of the reconstruction level 42* associated with a predetermined quantization index 32*, and, in this example, the distance 27 between adjacent reconstruction levels. The reconstruction level 42* may represent a predetermined reconstruction level, which may be indicated below by codebook_0_value. For this purpose, the reconstruction levels 42 may also form a monotonic sequence of numbers, which may optionally be integers.
[0105] For example, the mapping scheme 44 may conform to the one described in Section 2.1.2, and the mapping scheme 44 may be encoded into the data stream 14 as described in Section 2.1.2.
[0106] 3.1 Modified binarization scheme Subsections 3.1.1 and 3.1.2 describe two modifications of the binarization scheme based on the available information. Both methods may be used in combination whenever the requirements of each method are met. Section 3.1.3 describes special cases where it is not necessary to send bins.
[0107] 3.1.1 Skipping sign_flag sign_flag can represent a predetermined syntax element 22 as described with respect to Figures 1 and 2. sign_flag may also be a syntax element that indicates whether a value, such as a quantized value, or an index, such as a quantized index 32, is positive or negative. Optionally, sign_flag can indicate whether multiple values or indices are all positive or negative. In other words, a predetermined syntax element 22 can be a sign flag that indicates whether the quantized index 32 has a non-negative value, i.e., a positive value, or whether the quantized index 32 has a non-positive value, i.e., a negative value.
[0108] Whenever the set of indices to be encoded is not all negative or not all positive, each index, for example, quantized index 32, can be encoded as follows: In the first step, a binary syntax element sig_flag for index 32 is derived, which specifies whether the corresponding index 32 is equal to 0. Then, unlike the method described in 2.1.3.1, if sig_flag is equal to 1, the encoding of sign_flag is skipped, and instead the unary sequence of bins is encoded, followed by the encoding of a fixed-length sequence as follows:
[0109] The variable k is initialized with a non - negative integer, and X is initialized with 1 << k.
[0110] One or more syntax elements abs_level_greater_X indicating that the absolute value of the index is greater than X are encoded. When abs_level_greater_X is equal to 1, the variable k is updated (e.g., incremented by 1), then 1 << k is added to X, and further abs_level_greater_X is encoded. This procedure continues until abs_level_greater_X is equal to 0. Thereafter, a fixed - length code of length k is sufficient to complete the encoding of index 32. For example, the variable TIFF0007901433000085.tif728 can be encoded using k bits. Alternatively, the variable TIFF0007901433000086.tif710 is encoded using k bits and can be defined as TIFF0007901433000087.tif755. Alternatively, any other mapping of the variable TIFF0007901433000088.tif79 to a fixed - length code of k bits may be used.
[0111] 3.1.2 End of Absolute Level Encoding According to one embodiment, a predetermined syntax element 22 is a predetermined threshold flag indicating whether the absolute value of the quantization index is related to a threshold, that is, greater than X, for example, abs_level_greater_X. The threshold may be an integer value such as 1, 2, or 3.
[0112] The absolute level of index 32 is encoded using, for example, a single-term sequence of bins (abs_level_greater_X; for example, a sequence of threshold flags associated with different thresholds, i.e., X, where the thresholds may be monotonically increasing integers, e.g., a sequence of abs_level_greater_1, abs_level_greater_2, and abs_level_greater_3). From the perspective of decoder 10, this single-term sequence terminates when abs_level_greater_X decoded from the bitstream, i.e., datastream 14, is equal to 0. However, if the maximum absolute value is available to encoder 11 and decoder 10, this allows for early termination of the single-term sequence of all index 32 having an absolute value equal to the maximum absolute value. The binarization scheme for the encoded index 32 is as follows:
[0113] If the requirements of the method described in 3.1.1 are not met, the first step is to derive a binary syntax element sig_flag for index 32 that specifies whether the corresponding index 32 is equal to 0. If sig_flag is equal to 1, a further binary syntax element sign_flag is derived. The bin indicates whether the current index 32 is positive (e.g., bin=0) or negative (e.g., bin=1). Otherwise, if the requirements of the method described in 3.1.1 are met, sig_flag is encoded and the encoding of sign_flag is skipped. In this case, sign_flag may represent a given syntax element 22, the threshold flag may represent a further given syntax element, and the decoder 10 and / or encoder 11 may be configured to check whether the respective mapping schemes 44 meet further given criteria. In response to this check, the decoder 10 may be configured to decide whether to guess or decode further predetermined syntax elements, and the encoder 11 may be configured to decide whether to skip encoding or encode further predetermined syntax elements.
[0114] Next, the maximum absolute value Assuming that TIFF0007901433000089.tif74 is available, the following applies. If TIFF0007901433000090.tif74 is equal to 1, no further bins are encoded; otherwise, a single-term sequence of bins is encoded, and if necessary, followed by a fixed-length sequence as follows.
[0115] The variable k is initialized to a non-negative integer, and X is initialized to 1 << k.
[0116] One or more syntax elements abs_level_greater_X indicating that the absolute value of index 32 is greater than X are encoded. If abs_level_greater_X is equal to 1, the variable k is updated (e.g., incremented by 1), then 1 << k is added to X, and further abs_level_greater_X are encoded. This procedure continues until abs_level_greater_X is equal to 0 or X becomes TIFF0007901433000091.tif712 or more. For example, in particular, in an example where sig_flag is treated as a flag indicating whether the absolute value of q is greater than 0, a case X greater than M - 1 may occur for X == 0.
[0117] Next, after encoding abs_level_greater_X equal to 0, to complete the encoding of index 32, a fixed-length code of length k is sufficient. For example, the variable TIFF0007901433000092.tif728 can be encoded using k bits. Alternatively, the variable TIFF0007901433000093.tif710 is encoded using k bits TIFF0007901433000094.tif755 can be defined as. Alternatively, the variable to a fixed-length code of k bits Any other mapping for TIFF0007901433000095.tif79 may be used.
[0118] 3.1.3 Skipping Index Coding In special cases, none of the syntax elements need to be sent, and the integer index 32, and therefore the reconstruction level, can be entirely derived on the decoder side. This is given when the maximum absolute level of all sent indices is 0. The value is TIFF0007901433000096.tif714, which means that all indices sent in the tensor are 0.
[0119] 3.2 Derivation Concepts of sign_flag and Maximum Absolute Level The methods described in Section 3.1 require that the sign_flag value or maximum absolute value be available on both the encoder 11 and the decoder 10. Typically, this information is only available on the encoder 11 side, but this section provides concepts for how the decoder 10 can derive the information.
[0120] 3.2.1 Derivation of sign_flag value and threshold flag (abs_level_greater_X syntax element) from the codebook For example, when a codebook-based integer representation is applied, as in Section 2.1.2, the maximum absolute level and, optionally, the sign flag value can be derived by the decoder 10. Typically, the codebook is decoded from the bitstream 14 in the previous step. In some cases, the limits of the range, including all possible quantization indices, may be determined from the codebook.
[0121] In the example codebook according to Section 2.1.2 that can be decoded from bitstream 14, the decoded index depends on the codebook length (codebook_size) and the value of cbZeroOffset. The codebook may be a mapping scheme 44 as shown in one of Figures 3 to 7, and may be selected by the decoder 10 shown in Figure 1 and the encoder 11 shown in Figure 2.
[0122] a) Derivation of threshold flags: The threshold flag is an example of a predetermined syntax element 22 or another predetermined syntax element. The threshold flag indicates, for example, whether the absolute value of the quantization index 32 is greater than the threshold associated with the threshold flag. For example, the threshold is a positive value.
[0123] As described in Section 3.1.2, the parameter M can be used to terminate the encoding / decoding of the sequence of threshold flags, where M represents the maximum absolute value of the current index 32.
[0124] According to one embodiment, when codebook_size is equal to 1, the maximum absolute level TIFF0007901433000097.tif74 is derived to be 0 (as described in Section 3.1.3). In this case, decoder 10 may be configured to infer that the quantization index is 0.
[0125] Otherwise (if codebook_size is not equal to 1), the value of sign_flag for the current index 32 to be encoded is derived or decoded from bitstream 14. Then, the maximum absolute value M of the current index 32 can be derived as follows: TIFF0007901433000098.tif13125 This embodiment has the advantage that, in the example, even when the quantization indices are not symmetrically distributed around 0, M can still show exactly the maximum possible value M for both positive and negative quantization indices.
[0126] For example, there may be two cases in which a predetermined criterion is met and a predetermined threshold flag can be inferred by the decoder 10, or it can be omitted in encoding by the encoder 11. The predetermined criterion is met, for example, when the quantization index 32 to be reconstructed or encoded has a non-negative value and none of the quantization index 32 included in the mapping scheme 44 is greater than the threshold associated with the predetermined threshold flag. The predetermined criterion is also met, for example, when the quantization index 32 to be reconstructed or encoded has a non-positive value and none of the quantization index 32 included in the mapping scheme 44 is less than the negative threshold associated with the predetermined threshold flag. For example, the decoder 10 / encoder 11 may check the predetermined criterion depending on the sign of the quantization index 32. For example, the decoder 10 / encoder 11 can determine the predetermined threshold flag separately for the two cases of positive and negative quantization index 32.
[0127] For example, the decoder 10 / encoder 11 in Figures 1 and 2 can check one or more threshold flags, for example, it can sequentially check a sequence of threshold flags to check whether a predetermined criterion is met. The decoder 10 / encoder 11 can select a threshold flag associated with the smallest absolute threshold and satisfying a predetermined criterion as a predetermined threshold flag, for example, the threshold flag abs_level_greater_M. The decoder 10 / encoder 11 can select a threshold flag whose threshold is associated with the highest possible absolute value for the quantization index 32 according to the mapping scheme 44 and is selected as a predetermined threshold flag, for example, the threshold flag abs_level_greater_M. In other words, the decoder 10 / encoder 11 may be configured to select a predetermined threshold flag from one or more threshold flags so as to determine the maximum value M among the absolute values of the quantization index 32 included in the mapping scheme 44, and such that the threshold associated with the predetermined threshold flag is equal to the maximum value M among the absolute values of the quantization index 32. The maximum value M among the absolute values of the quantization index 32 may be determined based on the number of quantization index 32 included in the mapping scheme 44, for example, based on the count notification.
[0128] If a predetermined syntax element 22 is a predetermined threshold flag, the decoder 10 may be configured to infer that the predetermined threshold flag indicates that the absolute value of the quantization index 32 is less than or equal to a threshold associated with the predetermined threshold flag, and the encoder 11 may be configured to skip encoding the predetermined threshold flag into the data stream 14. In the inference, the decoder 10 may be configured to set the predetermined syntax element 22 to a predetermined state indicating that the quantization index 32 is less than or equal to a threshold. Based on this inference, the decoder 10 may be configured to set the value of the quantization index 32 to a threshold associated with the predetermined threshold flag, or to further decode the residual of the quantization index, for example, the difference between the quantization index 32 and the threshold associated with the predetermined threshold flag, from the data stream 14.
[0129] The quantization index 32 can be represented by two or more threshold flags, which can be sequentially decoded / encoded by deriving / encoding the first threshold flag among the two or more threshold flags. If the first threshold flag indicates that the absolute value of the quantization index 32 is greater than the threshold associated with the first threshold flag, for example, the value of the quantization index is adapted based on the threshold associated with the first threshold flag, and the sequential reading / encoding of the threshold flags continues. If the first threshold flag indicates that the absolute value of the quantization index is not greater than the threshold associated with the first threshold flag, the sequential derivation / encoding of the threshold flags is stopped, for example, the value of the quantization index is adapted based on the first threshold associated with the first threshold flag, and optionally, the decoding / encoding of the residual value of the quantization index 32 with respect to the data stream continues, adapting the value of the quantization index 32 based on the first threshold and the residual value. The decoding / encoding sequence of two or more threshold flags terminates, for example, when each threshold flag corresponds to a predetermined threshold flag, or when each threshold flag indicates that the absolute value of the quantization index is less than or equal to the threshold associated with each threshold flag. In the case of each threshold flag corresponding to a predetermined threshold flag, and when the mapping scheme 44 satisfies a predetermined criterion, the encoder 11 is configured to, for example, refrain from encoding the predetermined threshold flag into the data stream 14, skip encoding the threshold flag, and continue encoding, for example, the residual value of the quantization index into the data stream 14. In the case of each threshold flag corresponding to a predetermined threshold flag, and when the mapping scheme 44 satisfies a predetermined criterion, the decoder 10 is configured to terminate the sequence of deriving two or more threshold flags by, for example, inferring the state of the predetermined threshold flag, and optionally, the decoder 11 is configured to continue decoding, for example, the residual value of the quantization index 32 from the data stream 14. For example, two or more threshold flags form a sequence of threshold flags, each associated with its own threshold, and the threshold flags are monotonically arranged in the sequence with respect to their associated thresholds, for example, the thresholds are positive integers.The residual values can be represented, for example, in a binary representation with a fixed number of bins.
[0130] According to one embodiment, the decoder 10 may be configured to derive a sign flag for the quantization index 32 if the quantization index 32 is not zero, and to sequentially derive one or more threshold flags for the quantization index 32, including, for example, a predetermined threshold flag. Sequential derivation may be performed by deriving a first threshold flag among the threshold flags, continuing the sequential reading of threshold flags (for example, by fitting the value of the quantization index based on the threshold associated with the first threshold flag) if the first threshold flag indicates that the absolute value of the quantization index is greater than the threshold associated with the first threshold flag, and stopping the sequential derivation of threshold flags (and, for example, by fitting the value of the quantization index based on a first threshold associated with the first threshold flag) if the first threshold flag indicates that the absolute value of the quantization index is less than or equal to the threshold associated with the first threshold flag. The device is configured to perform the continuation of sequential derivation of threshold flags by deriving the next threshold flag among the threshold flags, and if the next threshold flag indicates that the absolute value of the quantization index is greater than the threshold associated with the next threshold flag, continuing the sequential derivation of threshold flags (for example, by fitting the value of the quantization index based on the threshold of the current threshold flag), and stopping the sequential derivation of threshold flags (and, for example, by fitting the value of the quantization index based on the first threshold associated with the first threshold flag) if the next threshold flag indicates that the absolute value of the quantization index is less than or equal to the threshold associated with the next threshold flag.
[0131] According to one embodiment, if the value of the quantization index is not 0, the encoder 11 may refrain from encoding the first threshold flag into the data stream and skip encoding the threshold flag (for example, continuing to encode the residual value of the quantization index (for example, the difference between the quantization index and the threshold associated with the predetermined threshold flag) into the data stream if the first threshold flag is a predetermined threshold flag, and if the value of the quantization index is not 0, the encoder 11 may be configured to encode the first threshold flag into the data stream, if the first threshold flag is not a predetermined threshold flag. If the absolute value of the quantization index is greater than the threshold associated with the first threshold flag, continue encoding the next threshold flag among one or more threshold flags. If the absolute value of the quantization index is less than or equal to the threshold associated with the first threshold flag, the coding of the threshold flag is skipped (and coding of the remaining values continues, for example), The device, If the next threshold flag is a predetermined threshold flag, then the encoding of the next threshold flag into the data stream is suppressed and the encoding of the threshold flag is stopped (and, for example, the remaining values continue to be encoded into the data stream), If the next threshold flag is not the specified threshold flag, encode the next threshold flag into the data stream, If the absolute value of the quantization index is greater than the threshold associated with the next threshold flag, continue encoding the next threshold flag among one or more threshold flags, If the absolute value of the quantization index is less than or equal to the threshold associated with the next threshold flag, then the encoding of the threshold flag is stopped (and, for example, the remaining value is continued to be encoded into the data stream), It is configured to perform the coding of the following threshold flags.
[0132] The techniques described herein under item a) also apply to the techniques described below under item c), where the term “predetermined syntax element” should therefore be understood as “further predetermined syntax element” and the term “predetermined criterion” should be understood as “further predetermined criterion.”
[0133] b) Derivation of the sign flag The sign flag is a predetermined syntax element 22 or an example of a further predetermined syntax element. The sign flag indicates, for example, whether the quantization index 32 is positive or negative.
[0134] In two cases, sign_flag can be derived in decoder 10 as follows, based on the notification of position 33* in the sequence of integers, where a quantization index 32* with value 0 is placed at position 33*.
[0135] Whenever cbZeroOffset is equal to 0, all decrypted indices are either 0 or positive. In this case, sign_flag is presumed to be 0.
[0136] Whenever cbZeroOffset is equal to codebook_size-1, all decrypted indices are either 0 or negative. In this case, sign_flag is presumed to be 1.
[0137] Referring to Figure 2, the encoder 11 may be configured to determine whether or not to encode the code flag based on position 33*. Furthermore, the encoder 11 may be configured to encode, for example, cbZeroOffset, i.e., notification of position 33*, into the data stream 14.
[0138] Referring to Figure 1, the decoder 10 can first check whether the quantization index 32 has a value of 0, for example by deriving a significance flag, for example by decoding the significance flag from the data stream 14, or by inferring the significance flag. For example, only if the quantization index 32 is not 0, a predetermined criterion is checked and the decoder 10 decides whether to infer or decode the sign flag from the data stream 14. After inferring or decoding the sign flag, the decoder 10 may be configured to continue deriving the quantization index 32 by deriving one or more threshold flags, each of which indicates whether the absolute value of the quantization index is greater than the threshold associated with the threshold flag. If the significance flag indicates that the value of the quantization index is 0, the decoder 10 may be configured to set the value of the quantization index to 0.
[0139] Referring to Figure 2, the encoder 11 may also be configured to check whether the quantization index 32 has a value of 0. Only if the quantization index 32 is not 0, a predetermined criterion is checked, and the encoder 11 decides whether to encode a code flag into the data stream 14. After this decision, the encoder 11 may be configured to continue encoding the quantization index 32 by encoding one or more threshold flags, each of which indicates whether the absolute value of the quantization index 32 is greater than a threshold associated with the threshold flag.
[0140] c) Derivation of threshold flag and sign flag Referring to Figures 1 and 2, the decoder 10 / encoder 11 is further configured to check whether the mapping scheme 44 satisfies further predetermined criteria. If the mapping scheme 44 satisfies further predetermined criteria, the decoder 10 may be configured to infer that further predetermined syntax elements have a predetermined state, for example, that the quantization index 32 is below a threshold, and the encoder 11 may be configured to skip encoding the further predetermined syntax elements into the data stream 14 so that the further predetermined syntax elements are inferred by the decoder 10. The further predetermined syntax elements are, for example, part of the representation of the quantization index 32, and for example, two or more syntax elements can represent the quantization index 32. If the mapping scheme 44 does not satisfy further predetermined criteria, the decoder 10 may be configured to derive or read further predetermined syntax elements from the data stream 10, and the encoder 11 may be configured to encode the further predetermined syntax elements into the data stream 14. A predetermined syntax element 22 is, for example, a sign flag, and a further predetermined syntax element is, for example, a predetermined threshold flag. A predetermined criterion checked in relation to a predetermined syntax element 22 is described under item b above, and a further predetermined criterion checked in relation to a further predetermined syntax element 22 is described under item a above.
[0141] 3.3 Preferred Embodiments In a preferred embodiment, the codebook according to Section 2.1.2 is decoded from the bitstream, where codebook_size is equal to 1 and cbZeroOffset is equal to 0. Then, sign_flag (presumably 0) and maximum absolute level are decoded. The derivation process in Section 3.2.1 is applied to TIFF0007901433000099.tif74 (which is derived to 0). Next, the encoding process in Section 3.1.3 is used, which skips the entire index coding. Referring to the decoder 10 in Figure 1 and the encoder 11 in Figure 2, in this embodiment, a predetermined criterion is met if the mapping scheme 44 includes exactly one quantization index 32 having a value of 0, and a predetermined syntax element 22 is a significance flag indicating whether the quantization index has a value of 0. The decoder 10 is configured to infer that the quantization index 32 is 0 if the mapping scheme 44 satisfies the predetermined criterion, and the encoder 11 is configured to suppress the coding of any syntax element, for example, any syntax element, into the data stream 14, where the syntax element is dedicated to signaling the quantization index 32. In this case, at least the predetermined syntax element 22, i.e., the significance flag, is not coded by the encoder 11 if the predetermined criterion is met. The decoder 10 may be configured to derive the quantization index 32 independently of the syntax elements of the data stream 14, which are dedicated to signaling the quantization index 32, for example, by deriving the quantization index without decoding such dedicated syntax elements from the data stream.
[0142] In the preferred embodiment shown in Figure 4, the codebook, i.e., mapping scheme 44, according to Section 2.1.2 is decoded / encoded to and from the bitstream 14, where codebook_size is equal to 2 and cbZeroOffset is equal to 0. In other words, the size of the monotonic sequence of integers of mapping scheme 44 is 2, and quantization index 32* having a value of 0 has the first position 33* in the sequence of quantization index 32. Next, sign_flag (presumably 0) and maximum absolute level The derivation process in Section 3.2.1 is applied to TIFF0007901433000100.tif74 (which is derived to 1). Then, the encoding process in Section 3.1.2, which includes skipping the sign_flag, is used.
[0143] In the preferred embodiment shown in Figure 5, the codebook, i.e., mapping scheme 44, according to Section 2.1.2 is decoded / encoded to and from the bitstream 14, where codebook_size is equal to 2 and cbZeroOffset is equal to 1. In other words, the size of the monotonic sequence of integers of mapping scheme 44 is 2, and quantization index 32* having a value of 0 has a second position 33* in the sequence of quantization index 32. Next, sign_flag (presumably 1) and maximum absolute level are determined. The derivation process in Section 3.2.1 is applied to TIFF0007901433000101.tif74 (which is derived to 1). Then, the encoding process in Section 3.1.2, which includes skipping the sign_flag, is used.
[0144] In a further preferred embodiment shown in Figure 3, the codebook, i.e., mapping scheme 44, according to Section 2.1.2, is decoded / encoded to and from the bitstream, where codebook_size is equal to 3 and cbZeroOffset is equal to 0. In other words, the size of the monotonic sequence of integers of mapping scheme 44 is 3, and quantization index 32* having a value of 0 has the first position 33* in the sequence of quantization index 32. Next, sign_flag (presumably 0) and maximum absolute level The derivation process in Section 3.2.1 is applied to TIFF0007901433000102.tif74 (which is derived to 2). Then, the encoding process in Section 3.1.2, which includes skipping the sign_flag, is used.
[0145] In a further preferred embodiment shown in Figure 6, the codebook, i.e., mapping scheme 44, according to Section 2.1.2 is decoded / encoded to and from the bitstream 14, where codebook_size is equal to 3 and cbZeroOffset is equal to 1. In other words, the size of the monotonic sequence of integers of mapping scheme 44 is 3, and quantization index 32* having a value of 0 has a second position 33* in the sequence of quantization index 32. Next, the maximum absolute level The derivation process in Section 3.2.1 is applied to TIFF0007901433000103.tif74 (which is derived to 1). Then, the encoding process in Section 3.1.2 is used, without skipping the sign_flag.
[0146] In the preferred embodiment shown in Figure 7, the codebook, i.e., mapping scheme 44, according to Section 2.1.2 is decoded / encoded to and from the bitstream 14, where codebook_size is equal to 3 and cbZeroOffset is equal to 2. In other words, the size of the monotonic sequence of integers of mapping scheme 44 is 3, and quantization index 32* having a value of 0 has the third position 33* in the sequence of quantization index 32. Next, sign_flag (presumably 1) and maximum absolute level The derivation process in Section 3.2.1 is applied to TIFF0007901433000104.tif74 (which is derived to 2). Then, the encoding process in Section 3.1.2, which includes skipping the sign_flag, is used.
[0147] The predetermined syntax element 22 may be a sign flag indicating the sign of the quantization index 32. The predetermined criterion is satisfied, for example, when the set of quantization indices included in the mapping scheme 44, for example, does not include both positive and negative quantization indices. This applies, for example, to the mapping scheme 44 shown in Figures 3 to 5 and Figure 7. The predetermined criterion is not satisfied, for example, when the set of quantization indices included in the mapping scheme 44 includes both positive and negative quantization indices. This applies, for example, to the mapping scheme 44 shown in Figure 6. For example, when the mapping scheme 44 of Figure 3 or Figure 4 is selected, the decoder 10 may be configured to infer that the predetermined syntax element 22 is not negative, i.e., positive. For example, when the mapping scheme 44 of Figure 5 or Figure 7 is selected, the decoder 10 may be configured to infer that the predetermined syntax element 22 is not positive, i.e., negative.
[0148] In a particularly preferred embodiment, the integer parameter / index is decoded (using the function int_param in Figure 8) as described below. The scheme in Figure 8 is a modified version of the scheme given in Section 10.2.1.5 of [2].
[0149] In Figure 8, if a codebook is not used, the values of codebookSize and cbZeroOffset are set to 0.
[0150] The variable `codebookSize` here is equal to the value of `codebook_size` used above.
[0151] The variable QuantParam can correspond to the reconstructed quantization index.
[0152] In the example in Figure 18, the quantization index is an integer. The variable maxNumNoRemMinus1 can represent the length of a sequence of threshold flags associated with a positive integer threshold having a monotonically increment of 1, e.g., the sequence of threshold flags associated with values 1, 2, 3, 4, ..., maxNumNoRemMinus1+1. For these threshold flags, there are no residuals in the case of purely integers, so the coding of residuals can be skipped.
[0153] The variable maxAbsVal may correspond to the variable M described in sections 2.1.2 and 3.2.1.
[0154] 4. Embodiment 2 of the Invention: History-Dependent Significant Encoding The present invention reduces the bitrate required to encode significance flags (sig_flag) as described in 2.1.3.1 for coefficients in an update layer as described in 2.2.1. Given a sig_flag for a coefficient q(x,y,k) at position (x,y) in an update layer having index k (where index k identifies a model update in a sequence of model updates), the basic idea is to model the probability of the sig_flag depending on the coefficient at position (x,y) occurring in a preceding update or a base layer having index l smaller than k. In contrast to the method described in 2.2.2.1, it should be noted that the present invention relates to a series of preceding coefficients, i.e., the coefficient history at (x,y).
[0155] 4.1 Overview Figure 9 shows a device 10, referred to as a decoder 10, for decoding NN parameters from a data stream 14, according to one embodiment. The NN parameters define a neural network. The decoder 10 receives a data stream 14 containing an encoded representation 122 of one of the neural network parameters, an update parameter 132. For example, the data stream 14 may contain a plurality of update parameters 20 arranged, for example, in a matrix representation, each update parameter providing, for example, update information for the corresponding NN parameter of the NN (which should be updated, for example, by the decoder 10). The decoder 10 includes an entropy decoding module 30 configured to decode the encoded representation 122 to obtain the update parameter 132. The decoder 10 further comprises an update module 50 configured to update the corresponding NN parameter using the update parameter 132, for example by adding the update parameter 132 to the NN parameter (e.g., combining the NN parameter and the update parameter) or by multiplication, or by using the update parameter 132 as input to control another operation for determining the updated NN parameter based on the current value of the NN parameter, or by replacing the NN parameter with the update parameter 132. In other words, the update may be incremental.
[0156] The decoder 10 further comprises a probabilistic model selector 140, which selects a probabilistic model 142 for entropy decoding of update parameter 132 according to previous update parameters (see 122' and 122''). For example, an update signaled in the data stream 14 may be one of a sequence of updates 26, each update containing its own set of update parameters, e.g., set of update parameters 20', 20''. Each of the set of update parameters may contain the corresponding update parameter 132 of the NN parameter referenced by update parameter 132. The update parameters may be signaled by coded representations 122', 122''.
[0157] Therefore, the probabilistic model selector 140 can obtain information based on previous update parameters 132', 132'', for example by storing them or by deriving historical parameters, for example, h, and can use this information to select a probabilistic model 142.
[0158] According to one embodiment, the decoder 10 can infer the characteristics of the previous update parameters 132', 132'' by analyzing the NN parameters, for example, by checking whether they are 0. According to this embodiment, the decoder can select a probabilistic model 142 depending on the NN parameters, and in this example, even independently of the previous update parameters 132', 132''. For example, the decoder 10 can check the value of the NN parameters and, if the NN parameters have a specific value, for example, 0, select a predetermined probabilistic model, for example, a special model. Otherwise, the decoder 10 may select a second probabilistic model, or it may select a probabilistic model 142 depending on the previous parameters.
[0159] Figure 10 shows a corresponding device 11, i.e., an encoder 11, for encoding update parameters for NN parameters according to one embodiment, and the decoder 10 performs the reverse calculation of the encoder 11. The encoder 11 includes an update derivator 51 for deriving update parameters 132.
[0160] According to one embodiment, the encoder 11 is configured to obtain an update parameter set 20 by reference and encode the update parameter set into a data stream 14, and the decoder 10 is configured to derive the update parameter set 20 from the data stream 14. The update parameter set 20 includes a plurality of update parameters for a plurality of NN parameters, and the decoder 10 is configured to update the NN parameters using the respective update parameters. For each of the update parameters, the decoder 10 / encoder 11 is configured to select a corresponding probabilistic model 142 depending on one or more previous update parameters of the respective NN parameter, for example, to use the selected probabilistic model 142 to entropy decode / encode each update parameter 132.
[0161] According to one embodiment, update parameter 132 and one or more previous update parameters 132', 132'' can be part of a sequence of update parameters for the NN parameters. Decoder 10 is configured to sequentially update the NN parameters based on the sequence of update parameters, for example, and encoder 11 retrieves update parameter 132, for example, based on update parameter 132, so that updating the NN parameters can be performed by decoder 10. For example, update parameter 132 allows decoder 10 to set or modify the NN parameters starting from the current values of the NN parameters. For example, encoder 11 is configured to sequentially update the NN parameters of the neural network based on the sequence of update parameters in order to make the NN available, for example, on the decoder 10 side. For example, encoder 11 is for retrieving update parameter 132 based on a sequentially updated neural network.
[0162] For both the decoder 10 and the encoder 11, the probability model 142 is, for example, a context model, or an adaptation model, or an adaptive context model. For example, the context of the update parameter 132 is part of the same set of update parameters (see 20), but is, for example, another parameter for another combination of input / output neurons, selected based on previously decoded, for example, directly adjacent update parameters (see, for example, 20a and 20b), and / or the probability is modeled according to the context. Alternatively, the context of the update parameter 132 is selected based on a sequence of previous update parameters 132', 132'' and / or the current NN parameters, and / or the probability is modeled according to the context.
[0163] Figure 11 shows the building blocks and data flow of the present invention. The target is the decoding of the encoded representation of q(x, y, k). The probability model selector 140 selects a probability model 142 m(x, y, k) based on the history h(x, y, k). The model is then used by the entropy decoder 30 to decode the sig_flag associated with q(x, y, k). After decoding other bins associated with q(x, y, k), the entropy decoder 30 reconstructs and provides the decoded value of q(x, y, k). Then, the history derivator derives a history h(x, y, k + 1) for the next update using q(x, y, k) and h(x, y, k).
[0164] 4.2 Decoding process The details of the decoding process are as follows.
[0165] 1) The history h(x, y, k) indicates whether any previous coefficient q(x, y, l) with l < k is non-zero, i.e., whether a significant coefficient was transmitted at the position (x, y) of any previous update layer (or base layer). For this purpose, in a preferred embodiment (E1), the history derivator iteratively derives h(x, y, k). This is intended as follows.
[0166] • Set h(x,y,0) to 0 before decoding the update of the first layer (i.e., the base layer). • To further update the layer, set h(x,y,k+1) to equal (h(x,y,k)||(q(x,y,k)!=0)) (where "||" represents the logical OR operator and "!=" represents the logical "not equal" operator). 2) If h(x,y,k) is equal to 1, the probability mode selector 140 selects a special probability model, such as a first probability model; otherwise (if h(x,y,k) is equal to 0), the probability mode selector 140 selects a default probability model. In a preferred embodiment, the special probability model provides, for example, a constant probability close to 0 or 0, or the minimum probability expressible by the entropy decoder 30. In a preferred embodiment, the default probability model is an adaptive context model selected by other means, such as those described in 2.1.3.2.
[0167] Referring to Figures 9 and 10, the decoder 10 / encoder 11 is configured to select a first probabilistic model for entropy decoding / coding of update parameter 132 (e.g., a special model, e.g., a constant probability of 0 or close to 0, or a constant model representing the minimum probability representable by the entropy decoder 30 / encoder 31) if, for example, the sequence of previous update parameters 132', 132'' does not satisfy a predetermined criterion, e.g., if not all of the previous update parameters in the sequence have a predetermined value, e.g., 0, e.g., if all of one or more previous update parameters indicate that the NN parameters will remain unchanged.
[0168] The decoder 10 / encoder 11 is configured to select a probabilistic model 142 for entropy decoding 30 / encoding 31 of the update parameter 132 according to historical data, for example, the historical data is derived by the decoder 10 / encoder 11 according to previous update parameters 132', 132''. The decoder 10 / encoder 11 is configured to entropy decoding 30 / encoding 31 of the update parameter 132 using the probabilistic model 142, for example, and to update the historical data according to the decoded update parameter 132.
[0169] The decoder 10 / encoder 11 is configured to update the history data, for example, according to previous update parameters. The history data can be updated, for example, by setting history parameters such as history derivator h. Alternatively, it can be done by storing the (previous) update parameters, for example, as shown in Section 4.3.3. Optionally, the history data is updated according to the history data (e.g., current and / or past history data), for example, the decoder 10 / encoder 11 can determine the values of the parameters included in the history data according to previous update parameters (see 132', 132'') and the parameter values before the parameter update (e.g., the current NN parameter values), and then replace the parameter values with the determined values.
[0170] According to one embodiment, the decoder 10 / encoder 11 is configured to select a first probabilistic model for entropy decoding / coding of update parameter 132, for example, a special model "0 model", if, for example, the sequence of previous update parameters 132', 132'' does not satisfy a predetermined criterion, for example, if not all previous update parameters 132', 132'' are equal to 0. Furthermore, the decoder 10 / encoder 11 is configured to update a history parameter, for example, the history parameter indicates whether the sequence of previous update parameters 132', 132'' satisfies a predetermined criterion, for example, the history parameter, for example, the history derivator h, according to the update parameter 132. If the history parameter has a first value, for example 1, indicating that a predetermined criterion is not met, or if the update parameter 132 does not have a predetermined value that is not equal to, for example 0, then the history parameter is set to the first value. If the update parameter 132 has a predetermined value, for example 0, and the history parameter has a second value, for example 0, indicating that a predetermined criterion has been met, then the history parameter is set to the second value. It is set by [this method].
[0171] Therefore, the historical data is updated using the current update parameters 132, thereby allowing the selection of an optimized possibility model 142 based on the updated historical data for subsequent update parameters of each NN parameter.
[0172] For example, the decoder 10 / encoder 11 is configured to select a first probabilistic model for entropy decoding / encoding of subsequent update parameters of the NN parameters if, for example, in the sequence of update parameters, the history parameter has a first value, for example, one, after the history parameter has been updated based on the update parameter 132. Optionally, the decoder 10 / encoder 11 is further configured to select a second probabilistic model, for example, a default model, for example, an adaptive model, for entropy decoding of subsequent update parameters of the NN parameters if, for example, the history parameter has a second value, for example, zero.
[0173] According to one embodiment, the decoder 10 / encoder 11 is configured to check whether one of the previous update parameters 132', 132'' has a predetermined value, for example, 0, by considering the significance flag of the previous update parameter. If one of the previous update parameters 132', 132'' does not have the predetermined value, the sequence of previous update parameters 132', 132'' does not satisfy the predetermined criterion. This check provides good notification of whether the update parameter 132 is 0 or not. The significance flag indicates whether the previous update parameter is 0 or not. The check can also be performed on each of the previous update parameters in the sequence of previous update parameters. For example, the decoder 10 / encoder 11 is configured to check whether the sequence of previous update parameters satisfies a predetermined criterion by considering the significance flag of each of the previous update parameters 132', 132'', where the significance flag of each previous update parameter indicates whether the respective update parameter is 0 or not. For example, the predetermined criterion is satisfied if, for all previous update parameters in the sequence of previous update parameters, the significance flag indicates that the respective previous update parameter is 0. This check may be further performed by considering the NN parameters, for example, by checking whether the NN parameters, such as the current NN parameters, are equal to 0.
[0174] The first probability model is, for example, a constant probability model, which represents / indicates, for example, a constant probability, for example, the probability of 0, or a probability of less than 0.1 or less than 0.05, or the minimum probability that can be represented by the entropy decoder 30 / encoder 31, or a predetermined probability that can be represented by the entropy decoder 30 / encoder 31, or the minimum probability that can be represented in entropy decoding / coding by the entropy decoder 10 / encoder 11 that performs entropy decoding / coding, for example, the minimum representable probability may depend on the configuration of the entropy decoder 10 / encoder 11, which is used to entropy decode / code the update parameter 132. The first probability model represents, for example, a constant probability of a predetermined syntax element, for example, a significance flag, having a predetermined state that indicates, for example, that the update parameter 132 is not 0. For example, the coded representation 122 of the update parameter 132 may be represented by one or more syntax elements that include a predetermined syntax element.
[0175] According to one embodiment, the decoder 10 / encoder 11 is configured to select a second probabilistic model, for example one from a set of probabilistic models, such as a default model, as the probabilistic model if the sequence of previous update parameters 132', 132'' satisfies a predetermined criterion. The second probabilistic model is, for example, a context model, or an adaptive model, or an adaptive context model.
[0176] 4.3 Other aspects / modified forms / embodiments 4.3.1 Enable / Disable For each layer (or for multiple sets of layers, or a subset of layers), the encoder signals a flag (or another syntax element) to enable or disable the history derivator and the probabilistic model selector 140. When they are disabled in a preferred embodiment, the entropy decoder 30 always uses the default probabilistic model.
[0177] A neural network includes multiple layers. Referring to Figures 1 and 2, the decoder 10 / encoder 11 may be configured to decode / encode each update parameter set for each layer of the multiple layers, including one or more update parameters of the NN parameters associated with each layer. The decoder 10 / encoder 11 may be configured to activate or deactivate the selection of a probabilistic model for entropy decoding / encoding the update parameters for each layer in response to notifications in the data stream 14, for example. The decoder 10 / encoder 11 may be configured to decode / encode notifications for each layer indicating whether the selection of a probabilistic model for each layer should be activated or deactivated. In other words, the decoder 10 / encoder 11 may be configured to derive / encode into the data stream notifications, such as syntax elements, e.g., flags, indicating whether the selection of a probabilistic model should be activated or deactivated for one or more update parameters of layers referenced by the syntax elements.
[0178] According to one embodiment, the decoder 10 / encoder 11 may be configured to use a predetermined probabilistic model 142, e.g., a default probabilistic model, e.g., a second probabilistic model, for entropy decoding / encoding of all update parameters of one of the layers when selection is deactivated for a layer. Alternatively, when selection is activated for a layer, the decoder 10 / encoder 11 may be configured to select a probabilistic model 142 for each of the update parameters of the layer, depending on one or more previous update parameters 132', 132'' for the NN parameters associated with each update parameter 132.
[0179] 4.3.2 Special Probability Models In other preferred embodiments, the special model, for example, the first probabilistic model, is a context model, or an adaptive model, or an adaptive context model.
[0180] In another preferred embodiment, the probability model selector 140 selects not only between a default probability model and a special probability model, but also from a set of probability models.
[0181] 4.3.3 Derivation of Alternative History In a modified form (V1) of the present invention, the history derivator stores the received coefficients (for example, information about previous update parameters 132', 132'') or their derived values such that h(x, y, k) represents a set of coefficients (or derived values). The probabilistic model selector 140 then uses h(x, y, k) to select a probabilistic model 142m(x, y, k) for example by logical and / or arithmetic operations (for example on previous update parameters 132', 132''), or by determining, for example, the number of non-zero coefficients, for example, a count (for example, coefficients that satisfy a predetermined criterion, for example, coefficients that have a predetermined value, or alternatively, coefficients that do not have a predetermined value such that in this case is non-zero), or by determining, for example, whether any of the coefficients are non-zero, or by determining whether any of the previous update parameters do not satisfy a predetermined criterion, or by comparing with a threshold (for example, comparing each of the previous update parameters 132', 132'' with a threshold), or for example based on the values or absolute values of the coefficients, or based on any combination of the means described previously.
[0182] A variant of the present invention (V2) operates similarly to variant V1, but discards coefficients from the set of coefficients, for example, depending on the index of the update to which they belong, or based on how many update layers have been received since they were added to the set of coefficients. The decoder 10 / encoder 11 is configured to store information about a limited number of previous update parameters 132', 132'' in historical data, and to discard information about the earliest update parameter (e.g., 132'') among the previous update parameters when the limit is reached. In other words, a sequence of previous update parameters can include a limited number of previous update parameters.
[0183] In a modified version of the present invention, the history derivator counts the number of previous coefficients q(x,y,k) that satisfy a certain condition (e.g., not equal to 0). The history data may include a parameter (see 132', 132'') indicating the count of previous update parameters that satisfy a predetermined criterion, e.g., a parameter related to a parameter having a predetermined value, or does not have a predetermined value. Alternatively, the history data may include a parameter (e.g., h(x,y,k), e.g., the number counted, e.g., the result of a binary determination) indicating whether a sequence of previous update parameters (see 132', 132'') satisfies a predetermined criterion, e.g., if all previous update parameters are equal to a predetermined value such as 0.
[0184] In a modified version of the present invention, the history derivator derives h(x,y,k) using an infinite impulse response (IIR) filter, for example, by providing an average of absolute values, or by any combination of h(x,y,k) and q(x,y,k) or values derived therefrom. The decoder 10 / encoder 11 is configured to determine the parameters of the history data, for example, by applying an infinite impulse response (IIR) filter to a plurality of previously updated parameters, for example, 132' and 132'', for example, by providing an average of absolute values.
[0185] 4.3.4 History Reset The history derivator resets h(x,y,k) to the initial history. In preferred embodiment E1, h(x,y,k) is set to 0. In modified forms V1 and V2, h(x,y,k) is set to an empty set. The history derivator resets the history if, for example, one or more of the following are true (for example, the decoder 10 / encoder 11 may be configured to reset the history data to a predetermined state):
[0186] - If explicitly signaled (for example, decoder 10 may be configured to derive a notification from data stream 14 indicating a reset of historical data. The notification is encoded in data stream 14 by encoder 11, for example, when the reset condition is met).
[0187] - If a flag is received indicating that the history derivator and the probabilistic model selector 140 are disabled (for example, the decoder 10 may be configured to derive a notification from the data stream 14 indicating the deactivation of the selection of probabilistic model 142. The encoder 11 may be configured to deactivate the selection of probabilistic model 142, reset the history data to a predetermined state, and encode a notification in the data stream 14 indicating the deactivation of the selection of probabilistic model 142 when the deactivation condition is met).
[0188] - After the decoder has received a certain number of model updates (for example, the count of one or more previous update parameters 132', 132'' received by decoder 10 / encoder 11 is greater than or equal to a predetermined count).
[0189] -After q(x,y,k) becomes equal to 0 with respect to a certain number of (e.g., consecutive) layer updates (e.g., the count of previous update parameters 132', 132'' that meet a predetermined criterion, e.g., have a predetermined value, or alternatively, do not have a predetermined value, falls below a predetermined threshold, e.g., update parameter 132 becomes equal to 0 with respect to a certain number of (e.g., consecutive) layer updates).
[0190] A reset condition is true, for example, if any of the conditions in the condition set are met, or if each of the conditions in a subset of the condition set is met.
[0191] 4.3.5 Others In a modified version of the present invention, the probabilistic model selector 140 selects a probabilistic model 142 for other bins that are not associated with sig_flag.
[0192] 4.4 Exemplary Embodiments Based on the ICNN Standard The following statements of this specification illustrate exemplary embodiments of the present invention in the working draft of the upcoming ICNN standard[8]. The following variables / syntax elements are defined by the following standard:
[0193] -QuantParam[j]: The currently decoded quantization parameter -sig_flag Significance flag -sign_flag sign flag -ctxInc is a variable that specifies the context model, i.e., the probabilistic model. 4.4.1 Specification Changes 4.4.1.1 Additional definitions Parameter identifier: A value that uniquely identifies a parameter within an incremental update, such that the same parameter in different incremental updates has the same parameter identifier.
[0194] NOTE: Parameters with the same parameter identifier will be in the same position within the same tensor in different incremental updates. This means they are located in the same place.
[0195] 4.4.1.2 Additional processing after decoding of the quantization parameter (QuantParam[j]) The variable curParaId is set to equal the parameter identifier of the parameter currently being decrypted. If the parameter whose parameter identifier QuantParam[i] is equal to curParaId has not been decrypted previously, the variable AnySigBeforeFlag[curParaId] is set to equal 0. Modify the variable AnySigBeforeFlag[curParaId] as follows:
[0196] AnySigBeforeFlag[curParaId]=AnySigBeforeFlag[ curParaId ]| |(QuantParam[ I ]!=0) 4.4.1.3 Changes to the derivation process of ctxInc for the syntax element sig_flag The inputs to this process are the sig_flag decoded before the current sig_flag, the state value stateId, the associated sign_flag if present, and, if present, the co-located parameter level (coLocParam) from the incremental update decoded before the current incremental update. If the sig_flag was not decoded before the current sig_flag, it is inferred to be 0. If the sign_flag associated with a previously decoded sig_flag was not decoded, it is inferred to be 0. If the co-located parameter level from the incremental update decoded before the current incremental update is not available, it is inferred to be 0. The co-located parameter level means the parameter level in the same tensor at the same position in the previously decoded incremental update.
[0197] The output of this process is the variable ctxInc.
[0198] The variable curParaId is set to be equal to the parameter identifier of the parameter currently being decoded.
[0199] The variable ctxInc is derived as follows: -If AnySigBeforeFlag[curParaId] is equal to 1, the following applies: Set -ctxInc to stateId+40 -Otherwise (AnySigBeforeFlag[curParaId] is equal to 0), the following applies: -If coLocParam is equal to 0, the following applies: -If sig_flag is equal to 0, set ctxInc to stateId * 3. - Otherwise, if sign_flag is equal to 0, set ctxInc to stateId*3+1. - Otherwise, set ctxInc to stateId*3+2. -If coLocParam is not equal to 0, the following applies: -If coLocParam is greater than 1 or less than -1, set ctxInc to stateId*2+24. - Otherwise, set ctxInc to stateId*2+25.
[0200] 4.4.2 Alternative specification changes Instead of 4.4.1.2, after decrypting sig_flag, the result will be as follows:
[0201] The variable curParaId is set to be equal to the parameter identifier of the sig_flag currently being decrypted. If the parameter whose parameter identifier is equal to curParaId has not been decrypted previously, the variable AnySigBeforeFlag[curParaId] is set to be equal to 0. Modify the variable AnySigBeforeFlag[curParaId] as follows:
[0202] AnySigBeforeFlag[curParaId]=AnySigBeforeFlag[ curParaId ]| |sig_flag) 5. Method based on the above principle Although several embodiments have been described in the context of apparatus, it is clear that these embodiments also represent descriptions of corresponding methods, where a block or device corresponds to a method step or a feature of a method step.
[0203] Figure 12 shows a method 200 for decoding neural network (NN) parameters that define a neural network from a data stream, which includes obtaining a mapping scheme 210 for mapping quantization indices to reconstruction levels from the data stream. Furthermore, method 200 includes checking whether the mapping scheme satisfies predetermined criteria 220 and reconstructing one of the NN parameters 230. Reconstruction of one of the NN parameters 230 is If the mapping method satisfies predetermined criteria, the state of a predetermined syntax element is inferred from the mapping method 232, If the mapping method does not meet the predetermined criteria, a predetermined syntax element is derived from the data stream 234, 236 In order to obtain the reconstruction level of the NN parameters, the quantization index derived using a given syntax element is applied to the mapping scheme, It will be carried out by [company name].
[0204] Method 200 is based on the same principles as described with respect to the decoder 10 in Section 3 above, and Method 200 may include method steps corresponding to the functions of the decoder 10, for example.
[0205] Figure 13 shows a method 300 for encoding neural network (NN) parameters that define a neural network into a data stream, the method 300 includes obtaining a mapping scheme 310 for mapping reconstruction levels to quantization indices, encoding the mapping scheme into a data stream 320, and encoding one of the NN parameters 330. Encoding one of the NN parameters 330 involves applying the reconstruction level of the NN parameter to the mapping scheme to obtain a quantization index, If the mapping method satisfies a predetermined criterion, skip the encoding of a predetermined syntax element into the data stream 334, wherein the predetermined syntax element is part of the representation of the quantization index, If the mapping method does not meet the predetermined criteria, the predetermined syntax elements are encoded into the data stream 336, It will be carried out by [company name].
[0206] Method 300 is based on the same principles as described with respect to the encoder 11 in Section 3 above, and Method 300 may include, for example, method steps corresponding to the functions of the encoder 11.
[0207] Figure 14 shows a method 400 for decoding neural network (NN) parameters that define a neural network from a data stream, the method 400 comprising receiving an update parameter 410 of the NN parameters and updating the NN parameters 420 by entropy decoding the update parameter from the data stream by selecting a probabilistic model for entropy decoding of the update parameter depending on the sequence of previous update parameters for the NN parameter and / or depending on the NN parameter.
[0208] Method 400 is based on the same principles as described with respect to the decoder 10 in Section 4 above, and Method 400 may include method steps corresponding to the functions of the decoder 10, for example.
[0209] Figure 15 shows a method 500 for encoding neural network (NN) parameters that define a neural network into a data stream, the method 500 comprising: obtaining an update parameter from among the NN parameters 510; and entropically encoding the update parameter into a data stream 520 by selecting a probabilistic model for entropy coding of the update parameter depending on the sequence of previous update parameters for the NN parameter and / or depending on the NN parameter.
[0210] Method 500 is based on the same principles as described with respect to the encoder 11 in Section 4 above, and Method 500 may include, for example, method steps corresponding to the functions of the encoder 11.
[0211] 6. Alternative forms of implementation This section describes alternative forms of implementation of the embodiments described in the previous section and in the claims.
[0212] While some aspects are described as features within the context of the apparatus, it is clear that such descriptions may also be considered descriptions of corresponding features of the method.
[0213] Some or all of the method steps may be performed by (or using) hardware devices such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such devices.
[0214] The encoded signals of the present invention can be stored in a digital storage medium and transmitted over a transmission medium such as the Internet, a wireless transmission medium, or a wired transmission medium. In other words, further embodiments provide a bitstream product including a bitstream according to any of the embodiments described herein, for example, a digital storage medium storing a video bitstream.
[0215] Depending on specific implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. Implementation may be carried out using a digital storage medium containing electronically readable control signals, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which may (or may) cooperate with a programmable computer system to perform the respective method. Therefore, the digital storage medium may be computer-readable.
[0216] Some embodiments of the present invention include a data carrier having an electronically readable control signal that can cooperate with a programmable computer system so that one of the methods described herein can be performed.
[0217] Generally, embodiments of the present invention can be implemented as a computer program product having program code, the program code operates to perform one of the methods when the computer program product is executed on a computer. The program code can be stored, for example, in a machine-readable carrier.
[0218] Other embodiments include a computer program stored in a machine-readable carrier for performing one of the methods described herein.
[0219] In other words, therefore, one embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer.
[0220] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) that records and includes a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is usually tangible and / or non-transitory.
[0221] Therefore, a further embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured to be transferred, for example, via a data communication connection, such as via the Internet.
[0222] A further embodiment includes processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0223] A further embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.
[0224] A further embodiment according to the present invention comprises an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0225] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the method herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods herein. Generally, the method is preferably performed by any hardware device.
[0226] The apparatus described herein may be implemented using hardware devices, or using a computer, or using a combination of hardware devices and a computer.
[0227] The methods described herein may be performed using hardware devices, or using a computer, or using a combination of hardware devices and a computer.
[0228] As can be seen in the detailed description above, various features are grouped together in the examples for the purpose of simplifying this disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed examples require more features than are explicitly stated in each claim. Rather, as reflected in the appended claims, the subject matter may consist of fewer features than all the features of a single disclosed example. Thus, the appended claims are incorporated into the detailed description herein, and each claim may stand alone as a separate example. While each claim may stand alone as a separate example, a dependent claim may refer to a specific combination of one or more other claims in the claim, but it should be noted that other examples may also include combinations of the subject matter of each of the other dependent claims, or combinations of each feature with other dependent or independent claims. Such combinations are proposed herein unless otherwise stated that a particular combination is not intended. Furthermore, it is intended that the features of a claim may be included in any other independent claim, even if that claim is not directly dependent on an independent claim.
[0229] The embodiments described above are merely illustrative of the principles of the present disclosure. Modifications and variations of the configurations and details described herein will be obvious to those skilled in the art. Therefore, it is intended to be limited only by the pending claims and not by any specific details presented as part of the description and explanation of the embodiments herein.
[0230] 7 References [1] S. Chetlur et al., "cuDNN:Efficient Primitives for Deep Learning," arXiv:1410.0759, 2014 [2] MPEG, "Text of ISO / IEC FDIS 15938-17 Compression of Neural Networks for Multimedia Content Description and Analysis", Document of ISO / IEC JTC1 / SC29 / WG11, w20331, OnLine, Apr.2021 [3] D. Marpe, H. Schwarz und T. Wiegand, "Context-Based Adaptive Binary Arithmetic Coding in the H.264 / AVC Video Compression Standard," IEEE transactions on circuits and systems for video technology, Vol. 13, No. 7, pp.620-636, July 2003. [4] H. Kirchhoffer, J. Stegemann, D. Marpe, H. Schwarz und T. Wiegand, "JVET-K0430-v3 - CE5-related:State-based probality estimator," in JVET, Ljubljana, 2018. [5] ITU - International Telecommunication Union, "ITU-T H.265 High efficiency video coding," Series H:Audiovisual and multimedia systems - Infrastructure of audiovisual services - Coding of moving video, April 2015. [6] B.Bross, J.Chen und S.Liu, "JVET-M1001-v6 - Versatile Video Coding (Draft 4)," in JVET, Marrakech, 2019. [7] S.Wiedemann et al., "DeepCABAC:A Universal Compression Algorithm for Deep Neural Networks," in IEEE Journal of Selected Topics in Signal Processing, vol. 14, no. 4, pp.700-714, May 2020, doi:10.1109 / JSTSP.2020.2969554. [8] "Working Draft 2 on Incremental Compression of Neural Networks" ISO / IEC / JTC29 / WG4, w20933, October 2021.
Claims
1. A device (10) for decoding neural network (NN) parameters that define a neural network from a data stream, From the aforementioned data stream (14), a mapping method (44) for mapping (40) the quantization index (32) to the reconstruction level (42) is obtained, (34) Checking whether the mapping method meets predetermined criteria, One of the aforementioned NN parameters, If the mapping method (44) satisfies the predetermined criteria, the state of a predetermined syntax element (22) is inferred from the mapping method (44), If the mapping method (44) does not satisfy the predetermined criteria, the predetermined syntax element (22) is derived from the data stream (14), To obtain the reconstruction level (42) of the NN parameters, the quantization index (32) derived using the predetermined syntax element (22) is applied to the mapping scheme (40), Reconstructing by, It is designed to do, The quantization index included in the mapping method (44) forms a monotonic sequence of integers, and the device takes the data stream, - The size of the sequence of the quantization index, - The position in the sequence of integers where a quantization index having a predetermined value is located, A device (10) configured to obtain information about one or more of the following.
2. The apparatus according to claim 1, wherein the mapping method (44) associates each of a finite set of quantization indices with one of a set of reconstruction levels.
3. The apparatus according to claim 1, wherein the quantization index is an integer.
4. The quantization index included in the mapping method (44) forms a monotonic sequence of integers, - The sequence size is 2, and the quantization index having a value of 0 occupies the first position in the sequence of quantization indexes, or - The sequence size is 2, and the quantization index having a value of 0 occupies the second position in the sequence of quantization indexes, or - The sequence size is 3, and the quantization index having a value of 0 occupies the first position in the sequence of quantization indices, or - The sequence size is 3, and the quantization index having a value of 0 occupies the second position in the sequence of quantization indexes, or - The sequence size is 3, and the quantization index having a value of 0 occupies the third position in the sequence of quantization indices. The apparatus according to claim 1.
5. The apparatus according to claim 1, wherein the predetermined syntax element (22) is a predetermined threshold flag that indicates whether the absolute value of the quantization index (32) is greater than a threshold associated with a predetermined threshold flag.
6. The NN parameters are reconstructed by deriving a sign flag indicating whether the quantization index (32) has a non-negative value or whether the quantization index (32) has a non-positive value. The aforementioned prescribed criteria If the quantization index (32) has a non-negative value, if none of the quantization indices included in the mapping scheme (44) are greater than the threshold associated with the predetermined threshold flag, and If the quantization index (32) has a non-positive value, then none of the quantization indices included in the mapping scheme (44) are less than the negative value of the threshold associated with the predetermined threshold flag. The apparatus according to claim 5, which is filled with
7. The apparatus according to claim 6, wherein, if the mapping method (44) satisfies the predetermined criteria, the predetermined threshold flag is configured to infer that the absolute value of the quantization index (32) is less than or equal to the threshold.
8. The mapping method (44) includes determining the maximum value among the absolute values of the quantization indices, Selecting one or more threshold flags such that the threshold associated with the predetermined threshold flag is equal to the maximum value among the absolute values of the quantization index, The apparatus according to claim 5, configured to do the following.
9. To derive a significance flag indicating whether or not the quantization index (32) has a value of 0, If the significance flag indicates that the value of the quantization index (32) is not 0, The sign flag is derived to indicate whether the quantization index (32) has a non-negative or non-positive value, Deriving the first threshold flag from one or more threshold flags, If the first threshold flag indicates that the absolute value of the quantization index (32) is greater than the threshold associated with the first threshold flag, the sequential derivation of the threshold flag is continued. If the first threshold flag indicates that the absolute value of the quantization index (32) is less than or equal to the threshold associated with the first threshold flag, the sequential derivation of the threshold flag is stopped. This involves deriving one or more threshold flags in sequence, each of which indicates whether the absolute value of the quantization index (32) is greater than the threshold associated with each of the threshold flags, The aforementioned device To derive the following threshold flag from among the aforementioned threshold flags, If the next threshold flag indicates that the absolute value of the quantization index (32) is greater than the threshold associated with the next threshold flag, the sequential derivation of the threshold flag is continued. If the next threshold flag indicates that the absolute value of the quantization index (32) is less than or equal to the threshold associated with the next threshold flag, the sequential derivation of the threshold flag is stopped. The system is configured to carry out the continuation of the sequential derivation of the threshold flag by Deriving one or more threshold flags in sequence, The derivation of the quantization index (32) is continued by decoding the remaining values of the quantization index (32) from the data stream, The apparatus according to claim 1, configured to derive a quantization index (32) for reconstructing the NN parameters by means of the above.
10. The apparatus according to claim 1, configured to dequantize the reconstruction level in order to obtain the reconstructed values of the NN parameters.
11. A device (11) for encoding neural network (NN) parameters that define a neural network into a data stream (14), Obtain a mapping method (44) for mapping reconstruction levels to quantization indices, The mapping method (44) is encoded into the data stream, The reconstruction level of the NN parameters is applied to the mapping method (44) in order to obtain the quantization index (32), If the mapping method (44) satisfies a predetermined criterion, the encoding of a predetermined syntax element (22) into the data stream is skipped, wherein the encoding of a predetermined syntax element (22) is part of the representation of the quantization index (32). If the mapping method (44) does not satisfy the predetermined criteria, the predetermined syntax element (22) is encoded into the data stream. This involves encoding one of the aforementioned NN parameters, It is configured to do so, The quantization index included in the mapping method (44) forms a monotonic sequence of integers, and the device takes the data stream, - The size of the sequence of the quantization index, - The position in the sequence of integers where a quantization index having a predetermined value is located, A device (11) configured to obtain information about one or more of the following.
12. The apparatus according to claim 11, configured to check whether the mapping method (44) meets predetermined criteria.
13. The apparatus according to claim 11, wherein the mapping method (44) associates each of a finite set of quantization indices with one of a set of reconstruction levels.
14. The apparatus according to claim 11, wherein the quantization index is an integer.
15. The quantization index included in the mapping method (44) forms a monotonic sequence of integers, - The sequence size is 2, and the quantization index having a value of 0 occupies the first position in the sequence of quantization indexes, or - The sequence size is 2, and the quantization index having a value of 0 occupies the second position in the sequence of quantization indexes, or - The sequence size is 3, and the quantization index having a value of 0 occupies the first position in the sequence of quantization indices, or - The sequence size is 3, and the quantization index having a value of 0 occupies the second position in the sequence of quantization indexes, or - The sequence size is 3, and the quantization index having a value of 0 occupies the third position in the sequence of quantization indices. The apparatus according to claim 11.
16. The apparatus according to claim 11, wherein the predetermined syntax element (22) is a predetermined threshold flag that indicates whether the absolute value of the quantization index (32) is greater than a threshold associated with a predetermined threshold flag.
17. If the quantization index (32) has a non-negative value, if none of the quantization indices included in the mapping scheme (44) are greater than the threshold associated with a predetermined threshold flag, and If the quantization index (32) has a non-positive value, then none of the quantization indices included in the mapping scheme (44) are less than the negative value of the threshold associated with the predetermined threshold flag. The apparatus according to claim 11, wherein the predetermined criteria are met.
18. The apparatus according to claim 17, wherein if the mapping method (44) satisfies the predetermined criteria, the encoding of the predetermined threshold flag into the data stream is skipped.
19. The apparatus according to claim 16, configured to select one or more threshold flags such that the threshold associated with the predetermined threshold flag is equal to the maximum value among the absolute values of the quantization indices included in the mapping method (44).
20. If the value of the quantization index (32) is not 0, If the first threshold flag among one or more threshold flags is the predetermined threshold flag, the encoding of the first threshold flag into the data stream is suppressed, and the encoding of the threshold flag is skipped. If the first threshold flag is not the predetermined threshold flag, the first threshold flag is encoded into the data stream, If the absolute value of the quantization index (32) is greater than the threshold associated with the first threshold flag, the encoding of the next threshold flag among the one or more threshold flags is continued, If the absolute value of the quantization index (32) is less than or equal to the threshold associated with the first threshold flag, the encoding of the threshold flag is skipped. It is configured to encode a quantization index (32) for encoding the NN parameters, The aforementioned device If the next threshold flag is the predetermined threshold flag, the encoding of the next threshold flag into the data stream is suppressed, and the encoding of the threshold flag is stopped. If the next threshold flag is not the predetermined threshold flag, the next threshold flag is encoded into the data stream, If the absolute value of the quantization index (32) is greater than the threshold associated with the next threshold flag, the encoding of the next threshold flag among the one or more threshold flags is continued, If the absolute value of the quantization index (32) is less than or equal to the threshold associated with the next threshold flag, the encoding of the threshold flag is stopped. The encoding of the following threshold flag is performed by The apparatus according to claim 16.
21. The apparatus according to claim 11, configured to quantize the NN parameters in order to obtain the reconstruction level of the NN parameters.
22. A method for decoding neural network (NN) parameters that define a neural network from a data stream, From the aforementioned data stream (14), a mapping method (44) for mapping (40) the quantization index (32) to the reconstruction level (42) is obtained, (34) Checking whether the mapping method meets predetermined criteria, One of the aforementioned NN parameters, If the mapping method (44) satisfies the predetermined criteria, the state of a predetermined syntax element (22) is inferred from the mapping method (44), If the mapping method (44) does not satisfy the predetermined criteria, the predetermined syntax element (22) is derived from the data stream (14), To obtain the reconstruction level (42) of the NN parameters, the quantization index (32) derived using the predetermined syntax element (22) is applied to the mapping scheme (40), Reconstructing by, Includes, The quantization index included in the mapping method (44) forms a monotonic sequence of integers, and the method takes the data stream, - The size of the sequence of the quantization index, - The position in the sequence of integers where a quantization index having a predetermined value is located, A method that includes obtaining information about one or more of the following.
23. A method for encoding neural network (NN) parameters that define a neural network into a data stream (14), Obtaining a mapping scheme for mapping reconstruction levels to quantization indices, The mapping method is encoded into the data stream, One of the aforementioned NN parameters, The reconstruction level of the NN parameters is applied to the mapping scheme in order to obtain a quantization index, If the mapping method satisfies predetermined criteria, the encoding of predetermined syntax elements into the data stream is skipped, wherein the encoding of predetermined syntax elements is part of the representation of the quantization index. If the mapping method does not satisfy the predetermined criteria, the predetermined syntax elements are encoded into the data stream. Encoding by and Includes, The quantization index included in the mapping scheme forms a monotonic sequence of integers, and the method takes the data stream from - The size of the sequence of the quantization index, - The position in the sequence of integers where a quantization index having a predetermined value is located, A method that includes obtaining information about one or more of the following.
24. A data stream (14) in which neural network (NN) parameters that define the neural network are internally encoded, Obtaining a mapping scheme for mapping reconstruction levels to quantization indices, The mapping method is encoded into the data stream, The reconstruction level of the NN parameters is applied to the mapping scheme in order to obtain a quantization index, If the mapping method satisfies predetermined criteria, the encoding of predetermined syntax elements into the data stream is skipped, wherein the encoding of predetermined syntax elements is part of the representation of the quantization index. If the mapping method does not satisfy the predetermined criteria, the predetermined syntax elements are encoded into the data stream. This involves encoding one of the aforementioned NN parameters, This is due to, The quantization index included in the mapping method (44) forms a monotonic sequence of integers, and the encoded data stream is derived from the data stream. - The size of the sequence of the quantization index, - The position in the sequence of integers where a quantization index having a predetermined value is located, A data stream (14) obtained by acquiring information about one or more of the following.
25. A computer program for carrying out the method described in claim 22 or 23, when executed on a computer or signal processor.
Citation Information
Patent Citations
Methods and apparatuses for compressing parameters of neural networks
WO2020188004A1
Concepts for coding neural networks parameters
WO2021123438A1