Method or apparatus for performing neural network-based processing with low complexity

By employing neural network-based video encoding and decoding with tensor product and addition layers and power-of-two quantization, the method addresses the challenges of reproducibility, complexity, and resource efficiency in existing techniques.

JP2025517698APending Publication Date: 2025-06-10INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024566696
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-18
Filing Date
2023-05-16
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing neural network-based video encoding and decoding techniques face challenges such as lack of reproducibility, high complexity, and significant memory and computational requirements.

Method used

The implementation of a neural network-based method that applies processing layers represented as tensor products and additions, with quantization using power-of-two scaling factors to reduce complexity and improve memory efficiency.

Benefits of technology

This approach enables fully reproducible neural network processing with reduced computational complexity and memory usage, optimizing video encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517698000001_ABST
    Figure 2025517698000001_ABST
Patent Text Reader

Abstract

At least a method and apparatus for efficiently encoding or decoding video are presented by applying neural network-based processing to a tensor of input data to generate a tensor of output data. For example, quantization of the tensor is limited to scaling by a power of two. For example, a layer of tensor products, a layer of bias addition, and activation are fused to reduce the number of operations and increase the bits available for representing values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - reference to Related Applications) This application claims the benefit of European Patent Application No. 22305731.6, filed on May 18, 2022, the entire content of which is incorporated herein by reference.

[0002] (Field of the Invention) At least one of the embodiments generally relates to a method or apparatus for video encoding or decoding, and in particular, to a method or apparatus for applying neural network - based processing to a tensor of input data to generate a tensor of output data with low complexity.

Background Art

[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction, including motion vector prediction, and transformation to exploit spatial and temporal redundancies within video content. Generally, intra - prediction or inter - prediction is used to take advantage of the correlation within or between frames, and then the difference between the original image and the predicted image, which often means the prediction error or prediction residual, is transformed, quantized, and entropy - coded. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to entropy coding, quantization, transformation, and prediction.

[0004] Neural network - based processing has recently been added to the high - compression techniques under consideration. The drawbacks of such neural network - based processing are the potential lack of reproducibility of the processing, the complexity of the processing (due to the number of operations or the nature of the operations themselves), and the huge amount of data to be stored. Therefore, it is desirable to provide an implementation of a neural network that enables a fully reproducible process and optimizes memory efficiency and computational power. Therefore, there is a need to improve this situation.

Summary of the Invention

[0005] According to the general aspects described in this specification, the drawbacks and shortcomings of the prior art are solved and addressed.

[0006] According to a first aspect, a method is provided. The method includes obtaining a tensor of input data representing a data sample and applying neural network-based processing to the tensor of input data to generate a tensor of output data. According to certain features, the neural network-based processing includes a plurality of processing layers, each processing layer generating an intermediate tensor. At least one processing layer is represented as a tensor product of the tensor of input data and a weight tensor, and at least one processing layer is represented as an addition of a bias tensor. Advantageously, any scaling factor of the quantized representation of tensors such as the tensor of input data, the weight tensor, the bias tensor, the intermediate tensor, and the tensor of output data uses a power of two.

[0007] According to another aspect, a method is provided. The method includes video decoding by applying neural network-based processing to a tensor of input data according to any of the disclosed embodiments to generate a tensor of output data, wherein the data samples of the input data tensor include at least samples of image blocks.

[0008] According to another aspect, a method is provided. The method includes video encoding by applying neural network-based processing to a tensor of input data according to any of the disclosed embodiments to generate a tensor of output data, wherein the data samples of the input data tensor include at least samples of image blocks.

[0009] According to another aspect, an apparatus is provided. The apparatus includes one or more processors, and the one or more processors are configured to implement a method for video decoding according to any of its variations. According to another aspect, an apparatus for video decoding includes means for applying neural network-based processing to a tensor of input data to generate a tensor of output data according to any of the disclosed embodiments.

[0010] According to another aspect, another apparatus is provided. The apparatus includes one or more processors, and the one or more processors are configured to implement a method for video encoding according to any of its variations. According to another aspect, an apparatus for video encoding includes means for applying neural network-based processing to a tensor of input data to generate a tensor of output data according to any of the disclosed embodiments.

[0011] According to another general aspect of at least one embodiment, a device is provided that includes an apparatus according to any of the embodiments related to decoding and at least one of (i) an antenna configured to receive a signal including video blocks, (ii) a band limiter configured to limit the received signal to a frequency band including video blocks, or (iii) a display configured to display an output representing the video blocks.

[0012] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided that includes data content generated according to any of the described encoding embodiments or variations.

[0013] According to another general aspect of at least one embodiment, a signal is provided that includes video data generated according to any of the described encoding embodiments or variations.

[0014] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variants.

[0015] According to another general aspect of at least one embodiment, when a program is executed by a computer, a computer program product is provided that includes instructions for causing the computer to execute any of the described embodiments or variants related to encoding / decoding.

[0016] The above and other aspects, features, and advantages of the general aspects will become apparent by reading on with reference to the accompanying drawings in the following detailed description of the exemplary embodiments.

Brief Description of the Drawings

[0017] In the drawings, examples of several embodiments are illustrated.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

[0018] Various embodiments relate to a video coding system, and in at least one embodiment, it is proposed to adapt video coding tools to low-complexity neural network processing. Different embodiments are proposed below that introduce changes to several tools to reduce the complexity of the codec when neural network processing is implemented in tools such as non-limiting examples of tool prediction or post-filtering. In particular, an encoding method, a decoding method, an encoding device, and a decoding device based on this principle are proposed.

[0019] Furthermore, while this aspect describes principles related to specific drafts of the VVC (Versatile Video Coding), or HEVC (High Efficiency Video Coding) specifications, or ECM (Enhanced Compression Model) reference software, it is not limited to VVC or HEVC or ECM, and can be applied, for example, to other standards and recommendations, as well as extensions of such standards and recommendations (including VVC, and HEVC, and ECM), whether existing or to be developed in the future. Unless otherwise indicated or technically excluded, the aspects described in this application can be used individually or in combination.

[0020] The acronyms used in this specification reflect the current state of video coding development and should therefore be considered as examples of naming that may be renamed at a later stage while still representing the same techniques.

[0021] FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments may be implemented. System 100 may be embodied as a device that includes various components described below and is configured to execute one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of System 100 may be embodied as a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing elements and encoder / decoder elements of System 100 are distributed across multiple ICs and / or discrete components. In various embodiments, System 100 is communicatively coupled to other systems or other electronic devices, for example, via a communication bus or dedicated input and / or output ports. In various embodiments, System 100 is configured to implement one or more of the aspects described in this application.

[0022] System 100 includes, for example, at least one processor 110 configured to execute instructions loaded internally to implement various aspects described in this application. Processor 110 may include built-in memory, an input / output interface, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. Storage device 140 may include, by way of non-limiting example, an internal storage device, a removable storage device, and / or a network-accessible storage device.

[0023] System 100 includes, for example, an encoder / decoder module 130 configured to process data to provide encoded video or decoded video. The encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, the device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 130 may be implemented as a separate element of the system 100 or may be incorporated within the processor 110 as a combination of hardware and software known to those skilled in the art.

[0024] To execute the various aspects described in this application, the program code loaded on the processor 110 or the encoder / decoder 130 may be stored in the storage device 140 and then loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of the various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or a portion of the decoded video, bitstream, matrix, variable, and intermediate or final results from the processing of equations, expressions, operations, and operation logic.

[0025] In some embodiments, the memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide a working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as a working memory for video coding operations and decoding operations such as HEVC or VVC.

[0026] Inputs to the elements of the system 100 may be provided through various input devices, as shown at block 105. Such input devices include, but are not limited to, (i) an RF section that receives, for example, an RF signal transmitted wirelessly by a broadcast station, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.

[0027] In various embodiments, the input device of block 105 has respective input processing elements known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) in certain embodiments, band-limiting again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel for example, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments may include one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF section may include a tuner for performing these various functions, including for example down-converting a received signal to a lower frequency (such as an intermediate frequency or near baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, down-converting, and filtering again to a desired frequency band an RF signal transmitted over a wired (e.g., cable) medium. In various embodiments, the order of the above (and other) elements is rearranged, some of these elements are omitted, and / or other elements performing similar or different functions are added. Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter for example. In various embodiments, the RF section includes an antenna.

[0028] In addition, the USB and / or HDMI terminals may each include an interface processor for connecting the system 100 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be performed, for example, within a separate input processing IC or within the processor 110 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within the processor 110. Demodulation, error correction, and the demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and an encoder / decoder 130 that operates in combination with memory and storage elements, to process the data streams necessary to present to the output devices.

[0029] The various elements of the system 100 may be provided within an integrated housing, in which the various elements are interconnected using an internal bus known in the art, including a suitable connection configuration 115, such as an I2C bus, wiring, and a printed circuit board, and may transmit data to each other.

[0030] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired medium and / or a wireless medium.

[0031] In various embodiments, data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of such embodiments are received via communication channel 190 and communication interface 150 adapted for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to an external network including the Internet to enable streaming applications and other over-the-top communications. In other embodiments, a set-top box that distributes data via the HDMI connection of input block 105 is used to provide the data streamed to system 100. In still other embodiments, the RF connection of input block 105 is used to provide the data streamed to system 100.

[0032] System 100 may provide an output signal to various output devices, including display 165, speaker 175, and other peripheral devices 185. Other peripheral devices 185 may include, in various examples of the embodiments, one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of system 100. In various embodiments, the control signal may be communicated between system 100 and display 165, speaker 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that enable control between devices with or without user intervention. The output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speaker 175 may be integrated into a single unit with other components of system 100 in an electronic device, such as a television, for example. In various embodiments, display interface 160 includes a display driver, such as a Timing Controller (T Con) chip, for example.

[0033] Alternatively, display 165 and speaker 175 may be separated from one or more of the other components, for example, if the RF section of input 105 is part of a separate set-top box. In various embodiments where display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0034] FIG. 2 illustrates an exemplary video encoder 200, such as a Versatile Video Coding (VVC) encoder. FIG. 2 may also illustrate an encoder with improvements made to the VVC standard, or an encoder that employs technology similar to VVC.

[0035] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "encoded" and "coded" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably. Usually, although not necessarily, the term "reconstructed" is used on the encoder side, while the term "decoded" is used on the decoder side.

[0036] Before being encoded, a video sequence undergoes pre-encoding processing (201), such as, for example, applying a color conversion to the input color picture (e.g., conversion from RGB4:4:4 to YCbCr4:2:0), or performing remapping of the input picture components to obtain a signal distribution that is more resistant to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the pre-processing and attached to the bitstream.

[0037] In encoder 200, a picture is encoded by encoder elements as follows. The picture to be encoded is partitioned (202) into units, such as coding units (CUs) for example, and processed. Each unit is encoded using either an intra mode or an inter mode, for example. When a unit is encoded in the intra mode, intra prediction (260) is performed. In the inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines (205) whether to use the intra mode or the inter mode to encode a unit, and indicates the intra or inter decision, for example, by a prediction mode flag. The prediction residual is calculated (210), for example, by subtracting the block predicted from the original image block.

[0038] Next, the prediction residual is transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying the transform process or quantization process.

[0039] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (240), inverse transformed (250), and the prediction residual is decoded. The decoded prediction residual and the predicted block are combined (255) to reconstruct the image block. A loop filter (265) is applied to the reconstructed picture, for example, to perform deblocking / Sample Adaptive Offset (SAO) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (280).

[0040] FIG. 3 illustrates a block diagram of an exemplary video decoder 300. As described below, in decoder 300, the bitstream is decoded by decoder elements. Video decoder 300 generally performs a decoding path that is inverse to the encoding path described in FIG. 2. Further, encoder 200 generally performs video decoding as part of video data encoding.

[0041] In particular, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. The picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture partitioning information (335). The transform coefficients are inverse quantized (340), inverse transformed (350), and the prediction residuals are decoded. The decoded prediction residuals and the predicted blocks are combined (355) to reconstruct the picture blocks. The predicted blocks can be obtained from intra prediction (360) or motion compensated prediction (i.e., inter prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed picture. The filtered picture is stored in the reference picture buffer (380).

[0042] The decoded picture can further undergo post-decoding processing (385), such as inverse color conversion (e.g., conversion from YcbCr 4:2:0 to RGB 4:4:4), or inverse remapping, which is the inverse process of the remapping process performed in the pre-encoding processing (201). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.

[0043] In recently considered video coding solutions, neural network-based processing has been proposed, for example, to provide a post-filtering stage or to provide block prediction.

[0044] Figure 4 illustrates a block-based pipeline for neural network processing in a video encoder / decoder in which various aspects of the embodiments can be implemented. A picture to be encoded (the original frame in Figure 4) is divided into a plurality of units (input blocks in Figure 4) and processed. The NN processing is applied to blocks of the picture where the picture data is supplied to the NN as input vectors, and the resulting processed blocks are output from the NN as output vectors and stored, for example, for additional encoding processing. Advantageously, the input data is not limited to picture samples and can convey any information / statistics associated with one or more blocks of the picture, as non-limiting examples, coding mode, quantization parameter, motion information, etc. Video decoding processing generally executes a decoding path opposite to the encoding path, so Figure 4 also illustrates the NN processing applied to blocks of the picture in the decoding process. In the context of video coding, strong constraints are required for processing including NN processing. - Inference should be fully reproducible, and thus all weights and operations should use integer operations. - Lower complexity is better, and thus it is desirable to limit the number of operations, limit complex operations (division, multiplication), and avoid some operations (e.g., square root, etc.). - Less memory usage is better, and thus it is desirable to have values quantized with a limited number of bits.

[0045] As shown in Figure 4, the NN processing includes multiple levels. Each level learns to transform its input data into a slightly more abstract and composite representation. In a video coding application, the raw input may be pixels / samples of a block, while the output is a processed block such as a predictor or a filtered block according to the above non-limiting examples. The output of a level uses the network representation. Inference refers to the process of supplying input data to the network and applying each layer to generate an output.

[0046] Here, three common methods used in general deep learning frameworks to meet the constraints of video coding are described.

[0047] According to the first embodiment, dynamic range quantization is used, and the weights w of the model are quantized to N bits (usually 8). The quantization is modeled using a scaling coefficient and a zero point (or offset) according to the following equation. W = clip(round(a × w + f)) Here, W is the quantized integer value of the floating-point weight w, a is the scaling coefficient, f is the zero point or offset, round() is a function that selects the nearest integer, clip() is a function that sets the value within the integer representation range, for example, [-128, 127] for 8 bits. The integer representation range is also referred to as the bit depth or representation type below.

[0048] However, in this embodiment, the weights are converted back to floating-point representation during inference, and the calculations are performed in floating-point.

[0049] According to the second embodiment, full integerization is used, and both the weights and the intermediate results are quantized and represented as integers. All operations use integer arithmetic. In this case, additional parameters are also defined to specify the scale and offset (or zero point) of the intermediate results (or tensors).

[0050] According to the third embodiment, quantization-aware training is performed. In addition to the types of representation and calculation, quantization constraints are also considered during the training itself. It enables directly considering the accuracy degradation of the parameters or tensors during training.

[0051] Figure 5 illustrates a block diagram of an embodiment of a layered neural network architecture in which various aspects of the embodiment can be implemented. The simple exemplary layered neural network NN of FIG. 5 includes three layers, namely a convolutional layer 510, a bias layer 520, and an activation layer 530 (ReLU here). However, the present principle is not limited to an NN having three layers and can be easily generalized to an NN modeled as one or more linear layers (matrix product and bias) together with one or more non-linear layers (activation functions such as ReLU, Gelu, sigmoid). FIG. 5 also shows the parameters a, f of the quantized NN model included in each layer of the NN. The parameter a represents a scaling coefficient, and the parameter b represents a zero point applied to any tensor of the quantized NN model that is a weight tensor, a bias tensor, and also applied to the input / output tensors of each layer X, Y, T. All parameters a and f are known in advance. FIG. 5 also shows the intermediate results or tensors Y, T of the quantized NN model. However, the implementation aspect of FIG. 5 still raises problems regarding complexity, for example, as detailed below.

[0052] First, by adding zero points, the number of operations performed during inference increases. For example, using a simple fully-connected layer that means a matrix product, the following equation holds.

[0053]

Number

[0054] Therefore, the integerized version is obtained as follows:

[0055]

number

[0056]

number

[0057] To perform all operations using only integer arithmetic, additional information regarding the rescaling of the result is required. The scaling factor is s t as shown. Advantageously, this scaling factor can be a power of 2 for execution using bit shift operations.

[0058]

Number

[0059] The clipping operation ensures that the result is included within the representation used for the intermediate result. One can notice that adding the calculation of the zero point term ΣX i . The bias term is also adapted to take into account the internal scaling s t and any potential offset of the result.

[0060] In practice, the above equation requires the following steps for calculation. - The cumulative product X i W ij is calculated. - Then, it is rescaled by the coefficient s t . - The scaling coefficient a t is calculated as

[0061]

Number

[0062] Second, in most methods, the bit depth of the weight and tensor representations is 8 bits because it targets common architectures such as CPUs, GPUs, or TPUs. However, in special hardware, the bit depth of both the weights and intermediate calculation results can have any bit depth.

[0063] Third, in most methods, the scaling factor is arbitrary and integer multiplication is required to calculate the output scaling factor. In general cases, division may also be required to adapt the scale of the layer output.

[0064] Fourth, the representation does not consider the nature of the operations in the model. For example, the output of the activation layer uses the same representation (scale, offset) regardless of the activation.

[0065] These problems are addressed and solved by the general aspects described herein that target representations that consider the following constraints. - Minimize the number of operations. - Simplify some operations, typically replacing multiplications and divisions with bit shift operations. - Take into account the nature of the operations. - Reduce the number of parameters representing the model.

[0066] In the following, this principle applies equally to matrix multiplications (dense layers, fully connected layers) or convolution-based layers. However, the matrix multiplication layer is described.

[0067] ​ According to the first embodiment, a low-complexity quantization is disclosed in which the scaling factor is a power of 2. In fact, to minimize complexity, the quantization is limited to scaling by a power of 2. This enables the use of bit shifts to perform the multiplication and division of the quantization. Moreover, since the quantization also includes the zero point of 0, no additional operations are performed on the quantization offset.

[0068] FIG. 6 illustrates a block diagram of an embodiment of a layered neural network architecture with low-complexity quantization. According to a specific variant of the first embodiment implemented in the exemplary NN of FIG. 5 having the same notation, the following is obtained.

[0069]

Number

[0070]

Number

[0071]

Number

[0072] In fact, the above equation requires the following steps for calculation. - Calculate the accumulated product X i W ij . - Bit shift by q y -(q x +q w ) and clip the result. The quantizer for this intermediate result is q y -(q x +q w ). - Divide the result by q b -(q y -(q x +qw ) Only perform bit shifting. - Bias

[0073]

Number

[0074] Advantageously, the number of parameters for controlling accuracy and bit depth is reduced. Assume that the bias layer drives the quantization of the input and output of the activation layer. All multiplication / division operations for quantization are advantageously replaced by shifts (multiplication / division by powers of 2). No additional operations are required to achieve the zero point.

[0075] According to a specific variant in which the convolutional / matrix multiplication layer and the bias layer are fused, the number of operations can be further reduced.

[0076] Figure 7 illustrates a block diagram of an embodiment of a layered neural network architecture with low-complexity quantization in a fused convolutional and bias layer. Thus, the scaling factors of the intermediate tensors Y and T are equal, q y = q b It is.

[0077] Assuming a fused convolutional / matrix multiplication and bias layer, the above equation can be further simplified.

[0078]

Number

[0079] In practice, the steps for calculating the value of tensor T are as follows. - Use an intermediate variable H to accumulate the sum ΣX i W ij Accumulate. - H’ = H ≫ ((q x + qw ) - q b ) Use to shift H so that variable H is quantized at q b and becomes quantized at q - Bias

[0080]

Number

[0081] One skilled in the art will understand that a right shift with a negative value is considered equivalent to a left shift

[0082] According to a particular variation, the processing of the sum of partial products is split. This variation is particularly suitable for input tensors with a very large bit depth. In fact, when the input tensor has a very large bit depth, there is a possibility of overflowing the underlying type in intermediate calculations

[0083] According to a non - limiting example of this variation, the processing

[0084]

Number

[0085]

Number

[0086] The steps of the processing can be described as follows - Only use the indices i within Ω1 to accumulate the sum ΣX i W ij into the intermediate variable H1, and accumulate the sum ΣX iW ij is accumulated.

[0087]

Number

[0088] Advantageously, using the same principle, the accumulation is divided into N stages to avoid overflow.

[0089] According to the second embodiment, the activation operation is also fused inside the convolution / matrix multiplication. FIG. 8 illustrates a block diagram of an embodiment of a layered neural network architecture with a fused activation operation. In fact, to further optimize the quantization, the activation operation is fused inside the convolution / matrix multiplication. According to a non-limiting variant, the activation layer is, for example, ReLU. In the context of a neural network, the rectifier or ReLU (Rectified Linear Unit) activation function is an activation function defined as the positive part of its argument x. f(x) = max(0, x)

[0090] Therefore, in such a case, the output is known to be positive. Advantageously, the underlying representation of the output is transformed to avoid the sign bit. FIG. 8 shows an example of such a process using the associated bit depth. - The input tensor X is assumed to be positive (thus requiring no bits for sign). This is the case for model inputs as well as for intermediate tensors after activation when ReLU is used. - The weights used a base type of 16 bits (15 bits + 1 sign bit). - The intermediate result Y uses a base type of 32 bits (31 bits + 1 sign bit). - They are added to a 17-bit bias (16 bits + 1 sign bit). - The activation uses 16 bits (no sign bit) to clip and shift the result to obtain the intermediate tensor.

[0091] This example shows that since the output of the activation layer is positive, a 1-bit sign is avoided for the intermediate results (between each set of convolution + bias + activation layers). The same principle applies to other types of layers (such as fully-connected layers) where the activation outputs only positive values.

[0092] According to the third embodiment, training for quantization recognition is disclosed, and the training stage also generates the quantization parameter q for each layer. Several techniques are possible to obtain the quantization parameter q for each layer. - Offline quantization: Each parameter q of the weights is found offline by checking the results of a small representative dataset. - Training for quantization recognition is performed: Each layer is replaced by a quantized version of the weights.

[0093] Figure 9 illustrates a block diagram of an embodiment of the transformation of a layered neural network architecture for performing training for quantization recognition. The model at the top of Figure 9 is replaced by the model at the bottom. Thus, for each weight, a quantization layer Q and an inverse quantization layer Q -1is inserted and both layers use the q parameter. All calculations are still performed with floating-point numerical values. Since the quantization operation is not differentiable, a substitute is used for quantization, typically using, for example, STE (Straight-through estimator), uniform noise, quantization function approximation, etc. Then, the output of the multiplication or convolution is also quantized / dequantized using the same method.

[0094] Further embodiments and information FIG. 10 illustrates a general decoding method (300) according to a general aspect of at least one embodiment. The block diagram of FIG. 10 partially represents a decoder or modules of a decoding method implemented, for example, in the exemplary decoder of FIG. 3.

[0095] The method includes applying neural network-based processing (1020) to a tensor of input data to generate a tensor of output data, where the input data includes at least samples of image blocks. Advantageously, the inference of the neural network-based processing uses any of the disclosed features to reduce the complexity of the neural network-based processing. Then, the NN-processed blocks (output data) are encoded (1020) according to any of the variant forms described herein.

[0096] FIG. 11 illustrates a general encoding method (200) according to a general aspect of at least one embodiment. The block diagram of FIG. 11 partially represents a module of an encoder or a module of an encoding method implemented, for example, in the exemplary encoder of FIG. 2. The method includes applying neural network-based processing (1120) to a tensor of input data to generate a tensor of output data, where the input data includes at least samples of image blocks. Advantageously, the inference of the neural network-based processing uses any of the disclosed features to reduce the complexity of the neural network-based processing. The NN-processed blocks (output data) are then encoded (1120) according to any of the variant forms described herein.

[0097] Various methods are described herein, and each of the methods includes one or more steps or acts for achieving the described method. The order and / or use of specific steps and / or acts may be modified or combined, provided that a particular order of the steps or acts is not required for proper operation of the method. Additionally, terms such as "first", "second", etc. may be used in various embodiments to modify elements, components, steps, acts, etc., for example, as "first decoding" and "second decoding". The use of such terms does not imply an ordering with respect to the modified acts, unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding and may occur, for example, before, during, or overlapping with the second decoding.

[0098] Using the various methods and other aspects described in this application, as non-limiting examples, module components of video encoder 200 and video decoder 300 as shown in FIGS. 2 and 3, and modules such as intra prediction modules (202, 260, 335, 360) can be modified. Further, this aspect is not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, and extensions of any such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application can be used individually or in combination.

[0099] In this application, various numerical values are used. The specific values are for illustrative purposes only, and the described aspects are not limited to these specific values.

[0100] Various implementations involve decoding. As used in this application, "decoding" can include, for example, all or part of the processing performed on the received encoded sequence to generate a final output suitable for display. In various embodiments, such processing typically includes one or more of the processes performed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to the broader decoding process as a whole will become apparent based on the context of the specific description and is considered to be fully understood by those skilled in the art.

[0101] Various implementations involve encoding. Similar to the above considerations regarding "decoding", "encoding" as used in this application can include all or part of the processing performed on the input video sequence to generate an encoded bitstream.

[0102] Note that the syntactic elements used in this specification are for descriptive purposes. Therefore, they do not exclude the use of other syntactic element names.

[0103] The implementations and aspects described in this specification can be implemented as various information such as, for example, syntax that can be transmitted or stored. This information can be packaged or arranged in various ways, including, for example, general ways in video standards such as putting the information into an SPS, PPS, NAL unit, header (e.g., NAL unit header, or slice header), or SEI message. Other ways are also available, including, for example, general ways in system-level or application-level standards such as putting the information into one or more of the following. ● SDP (Session Description Protocol), for example, a format for describing a multimedia communication session for the purpose of session announcement and session invitation, as described in an RFC and used in conjunction with RTP (Real-time Transport Protocol) transmission, ● DASH MPD (Media Presentation Description) descriptor, for example, used in DASH and transmitted via HTTP, which is associated with a representation or set of representations to provide additional characteristics to the content representation, ● RTP header extension, for example, used during RTP streaming, ● ISO-based media file format that uses boxes, which are object-oriented building blocks defined by a unique type identifier and length and are also known as "atoms" in some specifications and are used, for example, in OMAF, ● HLS (HTTP live Streaming) manifest transmitted via HTTP. The manifest can be associated with a version or set of versions of the content, for example, to provide characteristics of the version or set of versions.

[0104] The implementations and aspects described in this specification can be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if considered only in the context of a single form of implementation (e.g., considered only as a method), the implementation of the features considered can also be implemented in other forms (e.g., an apparatus or a program). The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. The method can be implemented, for example, in an apparatus such as a processor that refers to general processing devices including a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor can also include, for example, a communication device such as a computer, a mobile phone, a portable / personal digital assistant (PDA), and other devices that facilitate the communication of information between the end user.

[0105] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", and other variations thereof, mean that the specific features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment. Thus, the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" that appear at various places throughout this application, and other variations, do not necessarily all refer to the same embodiment.

[0106] In addition, this application may refer to "determining" various information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0107] Furthermore, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.

[0108] In addition, this application may refer to "receiving" various information. Receiving is intended to be a broad term, similar to "accessing". Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory). Further, "receiving" typically involves being involved in some way during an operation, such as while storing information, processing information, transmitting information, moving information, copying information, deleting information, computing information, determining information, predicting information, or estimating information.

[0109] For example, in the case of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of any of the following " / ", "and / or", and "at least one of" is intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those skilled in the art in the relevant art fields, this can be extended for any number of listed items.

[0110] Also, as used herein, the term "signaling" specifically means indicating something to the corresponding decoder. For example, in certain embodiments, the encoder signals a quantization matrix for dequantization. Thus, in one embodiment, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, the encoder can send specific parameters to the decoder (explicit signaling) so that the decoder can use the same specific parameters. In contrast, if the decoder already has other parameters along with that specific parameter, signaling that does not involve sending (implicit signaling) can be used to simply enable the decoder to know and select that specific parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It will be understood that signaling can be accomplished in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above relates to the verb form of the term "signal", although the term "signal" may also be used as a noun herein.

[0111] As will be apparent to those skilled in the art, the implementation may generate various signals formatted to carry information that can be stored or transmitted, for example. The information may include, for example, instructions for executing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be analog information or digital information, for example. The signal may be transmitted by various different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

[0112] Some embodiments will be described. The features of these embodiments may be provided singly or in any combination across various claim categories and types. Further, the embodiments can include one or more of the following features, devices, or aspects, singly or in any combination, across various claim categories and types. ● Adapting neural network inference in a decoder and / or encoder. ● A method, a process, an apparatus, a medium storing instructions, a medium storing data or a signal, according to any of the described embodiments. ● A TV, a set-top box, a mobile phone, a tablet, or other electronic device that executes neural network processing adapted to low complexity according to any of the described embodiments. ● A TV, a set-top box, a mobile phone, a tablet, or other electronic device that executes neural network processing adapted to low complexity according to any of the described embodiments and displays the resulting image (e.g., using a monitor, a screen, or other type of display). A television, set-top box, mobile phone, tablet, or other electronic device that selects a channel (e.g., using a tuner) to receive a signal including a symbolized image and performs neural network processing adapted to low complexity according to any of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device that wirelessly receives a signal including a symbolized image (e.g., using an antenna) and performs neural network processing adapted to low complexity according to any of the described embodiments.

Claims

1. A method implemented by a computer, comprising: obtaining a tensor (X) of input data representing a data sample; applying neural network-based processing to the tensor of the input data to generate a tensor (Z) of output data, wherein the neural network-based processing includes a plurality of processing layers (501, 502, 503), each processing layer generates an intermediate tensor (Y, T), at least one processing layer is represented as a tensor product of the tensor of the input data and a weight tensor (W), and at least one processing layer is represented as an addition of a bias tensor (B); A method implemented by a computer, wherein a scaling coefficient of a quantized representation of the tensor of the input data, a scaling coefficient of a quantized representation of the weight tensor, a scaling coefficient of a quantized representation of the bias tensor, a scaling coefficient of a quantized representation of the intermediate tensor, and a scaling coefficient of a quantized representation of the tensor of the output data use a power of 2.

2. An apparatus comprising a memory and one or more processors, wherein the one or more processors are configured to: obtain a tensor (X) of input data representing a data sample; apply neural network-based processing to the tensor of the input data to generate a tensor (Z) of output data; wherein the neural network-based processing includes a plurality of processing layers, each processing layer generates an intermediate tensor (Y, T), at least one processing layer is represented as a tensor product of the tensor of the input data and a weight tensor (W), and at least one processing layer is represented as an addition of a bias (B) tensor; An apparatus, wherein a scaling coefficient of a quantized representation of the tensor of the input data, a scaling coefficient of a quantized representation of the weight tensor, a scaling coefficient of a quantized representation of the bias tensor, a scaling coefficient of a quantized representation of the intermediate tensor, and a scaling coefficient of a quantized representation of the tensor of the output data use a power of 2.

3. The method according to claim 1 or the apparatus according to claim 2, wherein the quantized representation of the tensor is obtained by a shift according to a power of 2 of the scaling coefficient.

4. The method according to claim 1 or 3, or the apparatus according to claim 2 or 3, wherein the offset parameter of the quantization representation of the input data tensor, the offset parameter of the quantization representation of the weight tensor, the offset parameter of the quantization representation of the bias tensor, the offset parameter of the quantization representation of the intermediate tensor, and the offset parameter of the quantization representation of the output data tensor are equal to zero.

5. The method according to any one of claims 1, 3 or 4, or the apparatus according to any one of claims 2 to 4, wherein the at least one processing layer representing the addition of the bias tensor is fused with the at least one processing layer representing the tensor product.

6. The intermediate tensor T, which is the result of the fused tensor product and bias tensor addition, The sum of partial products (ΣX i W ij ) of the quantized representation of the input tensor and the quantized representation of the weight tensor is accumulated in an intermediate variable, and using the scaling coefficient of the quantization representation of the input data tensor, the scaling coefficient of the quantization representation of the weight tensor, and the scaling coefficient of the quantization representation of the intermediate tensor to shift the intermediate variable; The method according to claim 5, or the apparatus according to claim 5, wherein the result of the sum of the bias tensors is clipped to the shifted intermediate variable to obtain the intermediate tensor at the bit depth of the quantization representation of the intermediate tensor.

7. The sum (ΣX i W ij ) of the quantization representation of the tensor of the input data and the quantization representation of the weight tensor, accumulating the same, uses at least two intermediate variables to avoid overflow, the method according to claim 6 or the apparatus according to claim 6.

8. The method according to any one of claims 5 to 7, or the apparatus according to any one of claims 5 to 7, wherein the at least one processing layer includes an activation layer fused with the at least one processing layer representing the fused tensor product and bias tensor addition.

9. The input data tensor (X) is positive and represented without a bit sign, The output data tensor (Z) is positive and represented without a bit sign, The method according to claim 8, or the apparatus according to claim 8, wherein the activation layer clips and shifts the at least one processing layer representing the tensor product using addition of a bias to generate an intermediate tensor at the bit depth of the input data tensor.

10. A method implemented by a computer, the method including decrypting an image block, the decrypting including applying neural network-based processing to a tensor of input data according to any one of claims 1, 3 to 9 to generate a tensor of output data, the data sample including at least a sample of the image block, a method implemented by a computer.

11. The method according to claim 10, wherein the data sample further includes other information related to the image block.

12. A method implemented by a computer, the method including encrypting an image block, the encrypting including applying neural network-based processing to a tensor of input data according to any one of claims 1, 3 to 9 to generate a tensor of output data, the data sample including at least a sample of the image block, a method implemented by a computer.

13. The method according to claim 12, wherein the data sample further includes other information related to the image block.

14. An apparatus comprising a memory and one or more processors, the one or more processors being configured to decrypt an image block by applying neural network-based processing to a tensor of input data according to any one of claims 1, 3 to 9 to generate a tensor of output data, the data sample including at least a sample of the image block.

15. An apparatus comprising a memory and one or more processors, the one or more processors being configured to encrypt an image block by applying neural network-based processing to a tensor of input data according to any one of claims 1, 3 to 9 to generate a tensor of output data, the data sample including at least a sample of the image block.

16. A non-transitory program storage device readable by a computer, tangibly embodying a program of instructions executable by the computer for performing the method according to any one of claims 1, 3 to 9.

17. A method of training implemented by a computer, Obtain a tensor (X) of input data representing a sample of an image block, Apply a neural network-based training process to the tensor of the input data to generate a tensor (Z) of output data representing a sample of a compressed image block, The neural network-based training process includes a plurality of processing layers, each processing layer generates an intermediate tensor, at least one processing layer is represented as a tensor product of the tensor of the input data and a weight tensor (W), and at least one processing layer is represented as an addition of a bias tensor, The quantization processing layer (Q) and the inverse quantization processing layer (Q -1 ) generate a quantized representation of the weight tensor used in the neural network-based training process, The quantization processing layer (Q) and the inverse quantization layer (Q -1 ) generate a quantized representation of the bias tensor used in the neural network-based training process, The quantization processing layer (Q) and the inverse quantization layer (Q -1 ) generate a quantized representation of the intermediate tensor used in the neural network-based training process, The neural network-based training process further generates scaling factors of the quantization representation of the tensor of the input data, scaling factors of the quantization representation of the weight tensor (q w ), scaling factors of the quantization representation of the bias tensor (q b ), scaling factors of the quantization representation of the intermediate tensor (q b ), and scaling factors of the quantization representation of the tensor of the output data using powers of two, a training method implemented by a computer.

18. An apparatus comprising a memory and one or more processors, the one or more processors Obtain a tensor (X) of input data representing a sample of an image block, Apply a neural network-based training process to the tensor of the input data to generate a tensor (Z) of output data representing a sample of a compressed image block, The neural network-based training process includes a plurality of processing layers, each processing layer generates an intermediate tensor, at least one processing layer is represented as a tensor product of the tensor of the input data and a weight tensor (W), and at least one processing layer is represented as an addition of a bias tensor, The quantization processing layer (Q) and the inverse quantization processing layer (Q -1 ) generate a quantized representation of the weight tensor used in the neural network-based training process, The quantization processing layer (Q) and the inverse quantization layer (Q -1 ) generate a quantized representation of the bias tensor used in the neural network-based training process, The quantization processing layer (Q) and the inverse quantization layer (Q -1 ) generate a quantized representation of the intermediate tensor used in the neural network-based training process, The apparatus, wherein the neural network-based training process further generates, using powers of two, a scaling factor of a quantized representation of a tensor of input data, a scaling factor of a quantized representation of the weight tensor (q w ), a scaling factor of a quantized representation of the bias tensor (q b ), a scaling factor of a quantized representation of the intermediate tensor (q b ), and a scaling factor of a quantized representation of a tensor of the output data.

19. A trained machine learning model trained according to the method of claim 17.