Decryption method, apparatus, and device
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-08-14
AI Technical Summary
【0009】 以上の技術案から分かるように、本発明の実施例は、現在の画像ブロックに対応するビットストリームから現在の画像ブロックに対応する目標特徴を復号し、目標特徴に基づいて目標復号ネットワークの第1入力特徴を決定し、目標特徴値量子化ビット幅及び目標特徴値量子化ハイパーパラメータに基づいて第1入力特徴を第2入力特徴に変換し、目標復号ネットワークの整数型重みに基づいて第2入力特徴を処理して目標復号ネットワークの出力特徴を得、該出力特徴に基づいて現在の画像ブロックに対応する再構成画像ブロックを決定することにより、エンドツーエンドのビデオ画像圧縮方法を提供し、復号ネットワークに基づいてビデオ画像の復号を実現することができ、符号化効率及び復号効率を向上させる目的を達成する。目標重み量子化ビット幅及び目標重み量子化ハイパーパラメータを用いて固定小数点化重みの復号ネットワークを構築し、目標特徴値量子化ビット幅及び目標特徴値量子化ハイパーパラメータを用いて固定小数点化の入力特徴を生成し、適応復号加速を実現し、復号品質を保証した上で、復号計算量を削減する。
Smart Images

Figure 2026527476000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the technical field of encoding and decoding, and more particularly to decoding methods, apparatus and devices thereof. [Background technology]
[0002] To save space, all video images are encoded before transmission, and complete video encoding may include processes such as prediction, transformation, quantization, entropy coding, and filtering. Regarding the prediction process, it may include intra-frame prediction and inter-frame prediction. Inter-frame prediction effectively eliminates the temporal redundancy of the video by utilizing the temporal correlation of the video to predict the current pixel using pixels from adjacent encoded images. Intra-frame prediction eliminates the spatial redundancy of the video by utilizing the spatial correlation of the video to predict the current pixel using pixels from encoded blocks of the image in the current frame.
[0003] With the rapid development of deep learning technology, it has achieved success in many high-level computer vision problems, such as image classification and target detection. Deep learning technology is also gradually being applied in the fields of coding and decoding, meaning that it is now possible to code and decode images using neural networks. [Overview of the project] [Means for solving the problem]
[0004] In view of this, the present invention provides a decoding method, apparatus, and device that improve decoding performance and reduce the complexity of decoding.
[0005] The present invention The steps include decoding a target feature corresponding to the current image block from a bitstream corresponding to the current image block, The steps include determining the first input feature of the target decoding network based on the aforementioned target features, The steps include obtaining the target feature value quantization bit width and target feature value quantization hyperparameter of the target decoding network, and converting the first input feature of the target decoding network to the second input feature of the target decoding network based on the target feature value quantization bit width and target feature value quantization hyperparameter, A step of processing the second input features of the target decoding network based on the integer weights of the target decoding network to obtain the output features of the target decoding network, wherein the integer weights of the target decoding network are determined based on the target weight quantization bit width and target weight quantization hyperparameters of the target decoding network. The present invention provides a decoding method that includes the step of determining a reconstructed image block corresponding to the current image block based on the output features of the target decoding network.
[0006] The present invention A decoding module for decoding a target feature corresponding to the current image block from a bitstream corresponding to the current image block, A decision module for determining the first input feature of the target decoding network based on the aforementioned target features, A processing module for obtaining the target feature value quantization bit width and target feature value quantization hyperparameter of the target decoding network, converting the first input features of the target decoding network to second input features of the target decoding network based on the target feature value quantization bit width and target feature value quantization hyperparameter, processing the second input features of the target decoding network based on the integer weights of the target decoding network, and obtaining the output features of the target decoding network, wherein the integer weights of the target decoding network are determined based on the target weight quantization bit width and target weight quantization hyperparameter of the target decoding network, The decision module provides a decoding device used to determine the reconstructed image block corresponding to the current image block based on the output features of the target decoding network.
[0007] The present invention A decoding device comprising a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor provides a decoding device used to execute the above decoding method by executing machine-executable instructions.
[0008] The present invention The present invention provides a machine-readable storage medium in which multiple computer instructions are stored, wherein the above-described decoding method is performed when the computer instructions are executed by a processor.
[0009] As can be seen from the above technical proposals, embodiments of the present invention provide an end-to-end video image compression method by decoding target features corresponding to the current image block from a bitstream corresponding to the current image block, determining a first input feature of the target decoding network based on the target features, converting the first input feature to a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameters, processing the second input feature based on the integer weights of the target decoding network to obtain the output feature of the target decoding network, and determining the reconstructed image block corresponding to the current image block based on the output feature, thereby achieving the objective of improving encoding efficiency and decoding efficiency. A decoding network with fixed-point weights is constructed using the target weight quantization bit width and target weight quantization hyperparameters, and fixed-point input features are generated using the target feature value quantization bit width and target feature value quantization hyperparameters, thereby achieving adaptive decoding acceleration and reducing the decoding computation amount while guaranteeing decoding quality. [Brief explanation of the drawing]
[0010] [Figure 1] This is a schematic diagram of a 3D feature matrix according to one embodiment of the present invention. [Figure 2]This is a flowchart of a decoding method according to one embodiment of the present invention. [Figure 3] This is a schematic diagram of the encoding process according to one embodiment of the present invention. [Figure 4] This is a schematic diagram of the decoding process according to one embodiment of the present invention. [Figure 5A] This is a schematic diagram of a decryption network according to one embodiment of the present invention. [Figure 5B] This is a schematic diagram of a decryption network according to one embodiment of the present invention. [Figure 6A] This is a schematic diagram of a decryption network according to one embodiment of the present invention. [Figure 6B] This is a schematic diagram of a decryption network according to one embodiment of the present invention. [Figure 7] This is a hardware structure diagram of a decoding device according to one embodiment of the present invention. [Modes for carrying out the invention]
[0011] The terms used in the embodiments of this invention are merely for the purpose of describing specific embodiments and are not intended to limit the invention. The singular forms “one,” “the said,” and “the” used in the embodiments and claims of this invention are also intended to include the plural form unless the context clearly indicates otherwise. Furthermore, it should be understood that the term “and / or” used in this invention means including any or all possible combinations of one or more related enumerated items. The embodiments of this invention may use terms such as first, second, third, etc. to describe various types of information, but it should be understood that this information is not limited to these terms. These terms are used only to distinguish the same type of information. For example, as long as it does not deviate from the scope of the embodiments of this invention, depending on the context, first information may be called second information, and similarly, second information may be called first information. Furthermore, the word “…case” used herein may be interpreted as “…and,” “…when,” or “in response to a decision.”
[0012] Embodiments of the present invention provide a decoding method and may relate to the following technical terms.
[0013] JPEG (Joint Photographic Experts Group): JPEG is a standard for compressing continuous-tone still images. Its file extension can be .jpg or .jpeg, and it is a common image file format. JPEG employs a joint coding scheme of predictive coding (e.g., DPCM, Differential pulse-code modulation), discrete cosine transform (DCT), and entropy coding to remove redundant image and color data. It is a lossy compression format, capable of compressing images into a small memory space, but it does cause some damage to the image data. In particular, if the compression ratio is too high, the quality of the decompressed image will deteriorate, so it is not advisable to use excessively high JPEG compression when pursuing high-quality images.
[0014] JPEG-AI (Joint Photographic Experts Group Artificial Intelligence): The scope of JPEG-AI is to create a learning-based image coding standard that provides a single-stream, compact, compressed region representation, significantly improving compression efficiency compared to commonly used image coding standards at the same subjective quality, and effectively enhancing performance in image processing and computer vision tasks. JPEG-AI supports a wide range of applications, including cloud storage, vision management, autonomous driving for automobiles and equipment, image acquisition storage and management, real-time management of vision data, and media distribution. The goal of JPEG-AI is to design coding and decoding solutions that significantly improve compression efficiency at the same subjective quality, providing effective compressed region processing for machine learning-based image processing and computer vision tasks. JPEG-AI needs to achieve hardware and software-friendly coding and decoding, supporting 8-bit and 10-bit depths, and performing efficient coding and progressive decoding of images using text and graphics.
[0015] Entropy coding: Entropy coding is a method of coding in which no information is lost during the coding process, according to the principle of entropy. Information entropy is the average amount of information (degree of uncertainty) of the information source. Entropy coding schemes may include, but are not limited to, Shannon coding, Huffman coding, and arithmetic coding.
[0016] Neural Network (NN): A neural network refers to an artificial neural network, which is a computational model composed of a large number of nodes (called neurons) connected to one another. In a neural network, neuron processing units can represent different objects, such as features, alphabets, concepts, or several meaningful abstract modes. There are three types of processing units in a neural network: input units, output units, and hidden units. Input units receive external signals and data, output units realize the output of the processing results, and hidden units are units that are between the input and output units and cannot be observed from outside the system. The connection weights between neurons reflect the strength of the connections between units, and the representation and processing of information are reflected in the connection relationships of the processing units. A neural network is an unprogrammed, brain-like information processing method, and the essence of a neural network is to acquire parallel and distributed information processing capabilities through the transformation and dynamic actions of the neural network, mimicking the information processing capabilities of the human brain and nervous system to different degrees and levels. In the field of video processing, commonly used neural networks may include, but are not limited to, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and fully connected networks.
[0017] Convolutional Neural Networks (CNNs): Convolutional neural networks are feedforward neural networks and one of the representative network structures in deep learning techniques. The artificial neurons in a convolutional neural network can respond to peripheral units within a certain coverage area, exhibiting excellent performance in large-scale image processing. The basic structure of a convolutional neural network may include two layers: one is a feature extraction layer (also called a convolutional layer), where the input of each neuron is connected to the local receptive field of the previous layer, and the local features are extracted. Once the local features are extracted, their positional relationship with other features is also determined. The other is a feature mapping layer (also called an activation layer), where each computational layer of the neural network consists of multiple feature mappings, each feature mapping is a plane, and the weights of all neurons in the plane are equal. The feature mapping structure may use functions such as the Sigmoid function, ReLU function, Leaky-ReLU function, PReLU function, and GDN function as activation functions for the convolutional network. Furthermore, because neurons on a single mapping plane share weights, the number of free parameters in the network decreases.
[0018] For example, one advantage of convolutional neural networks compared to image processing algorithms is that they can avoid complex pre-processing processes for images (such as extracting artificial features), directly inputting original images and performing end-to-end learning. Another advantage of convolutional neural networks compared to general neural networks is that while general neural networks employ a fully connected architecture, meaning all neurons from the input layer to the hidden layer are connected, resulting in a huge number of parameters and making network training time-consuming and difficult, convolutional neural networks avoid this difficulty through methods such as local connectivity and weight sharing.
[0019] For example, convolutional neural networks can perform large-scale image processing. Convolutional neural networks typically include convolutional layers, pooling layers, and fully connected layers, and have made significant progress in fields such as image classification, target detection, and semantic segmentation.
[0020] Deconvolution: Also known as transposed convolution, the deconvolution layer and convolutional layer operate similarly. The main difference is that the deconvolution layer uses padding to make the output larger than the input (though they may be the same). A stride of 1 indicates that the output size is equal to the input size. A stride of N indicates that the width of the output features is N times the width of the input features, and the height of the output features is N times the height of the input features.
[0021] Model Fixed-Point / Model Quantization: Model fixed-point quantization, also known as model quantization, is a commonly used neural network model acceleration algorithm. By converting the model's float-type parameters and input / output feature values to fixed-bit values, such as 8 bits or 4 bits, it reduces the computational cost, data bandwidth, and memory space of the network model, allowing the network model to be applied more quickly and efficiently even on hardware with limited memory and computing power. Model fixed-point quantization includes fixed-point quantization of the model's weight values and fixed-point quantization of feature values. The lower the number of bits used for fixed-point quantization, the more pronounced the acceleration on hardware becomes, but the performance decreases accordingly.
[0022] Image Encoding and Decoding: The purpose of image encoding and decoding technology is to encode and compress images to reduce the transmission and storage costs of image data, while simultaneously allowing them to be decoded and restored to their original image content. Higher encoding and compression ratios result in lower transmission and storage consumption, but also make image restoration significantly more difficult. As artificial intelligence has achieved success in various fields, it has also been applied to the field of image encoding and decoding, achieving lower compression ratios and superior image restoration effects.
[0023] Feature: The feature according to the present invention may be a 3D feature matrix of type C*W*H. Figure 1 is a schematic diagram of the 3D feature matrix, where C represents the number of channels, H represents the feature height, and W represents the feature width. The 3D feature matrix may be the input to a neural network or the output of a neural network.
[0024] The Rate-Distortion Optimization principle: Encoding efficiency is evaluated using two metrics: bitrate and PSNR (Peak Signal to Noise Ratio). A smaller bitstream results in greater compression, and a higher PSNR results in better reconstructed image quality. When selecting a mode, the discriminant is essentially a comprehensive evaluation of both. For example, the cost corresponding to a mode is J(mode) = D + λ*R, where D represents distortion, which can usually be evaluated using the SSE (Sum of the Squared Errors) metric, where SSE is the mean square sum of the differences between the reconstructed image block and the source image. To consider the cost, the SAD metric may also be used, where SAD is the sum of the absolute differences between the reconstructed image block and the source image, λ is the Lagrangian multiplier, and R is the actual number of bits required to encode the image block in that mode, including the total number of bits required to encode mode information, motion information, residuals, etc. When selecting a mode, comparing and determining the encoding mode using the rate distortion principle usually guarantees optimal encoding performance.
[0025] While neural network-based encoding and decoding methods have shown great performance potential in related technologies, they still suffer from challenges such as poor decoding performance and high complexity.
[0026] For each module on the encoding side, a great many encoding tools have been proposed, and each tool usually has many modes. Often, different encoding tools yield optimal encoding performance for different video sequences. Therefore, in the encoding process, Rate-Distortion Optimize (RDO) is typically used to compare the encoding performance of different tools or modes and select the optimal mode. After determining the optimal tool or mode, the decision information is transmitted by encoding mark information into the bitstream. Such a method can adaptively select the optimal mode combination for different content and obtain optimal encoding performance. On the decoding side, the relevant mode information can be obtained by directly analyzing the mark information, resulting in less complexity.
[0027] The decoding method in the embodiments of the present invention will be described in detail below with reference to several specific examples.
[0028] Example 1: An embodiment of the present invention provides a decoding method, which may include the following steps, as shown in Figure 2.
[0029] Step 201: Decode the target feature corresponding to the current image block from the bitstream corresponding to the current image block.
[0030] Step 202: Determine the first input feature of the target decoding network based on the target feature.
[0031] Step 203: Obtain the target feature value quantization bit width and target feature value quantization hyperparameter of the target decoding network, and convert the first input feature to a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter.
[0032] Step 204: The second input features are processed based on the integer type weights of the target decoding network to obtain the output features of the target decoding network, where the integer type weights are determined based on the target weight quantization bit width and target weight quantization hyperparameters of the target decoding network.
[0033] Step 205: Based on the output features of the target decoding network, determine the reconstructed image block corresponding to the current image block.
[0034] Exemplary, a target decoding network includes at least one target network layer, which may be a network layer employing integer weights, and for each target network layer in the target decoding network, a first input feature of the target network layer is converted to a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter of the target network layer, the second input feature is processed based on the integer weights of the target network layer to obtain the output feature of the target network layer, the integer weights of the target network layer are determined based on the target weight quantization bit width and target weight quantization hyperparameter of the target network layer, the target feature value quantization bit widths of different target network layers may be the same or different, the target weight quantization bit widths of different target network layers may all be fixed quantization bit widths, and the target weight quantization hyperparameters of different target network layers may be the same or different.
[0035] Exemplarily, based on the target eigenvalue quantization bit width and the target eigenvalue quantization hyperparameter of the target network layer, the step of converting the first input feature of the target network layer into the second input feature is: quant = clip(round(data × 2 param ), -2 bw-1 , 2 bw-1 -1), or quant = clip(round(data ÷ param), -2 bw-1 , 2 bw-1 -1) is used to convert the first input feature of the target network layer into the second input feature, but not limited thereto. quant represents the second input feature, data represents the first input feature, param represents the target eigenvalue quantization hyperparameter, bw represents the target eigenvalue quantization bit width, clip represents the clipping function, and round represents rounding.
[0036] By the operation of the clipping function clip, when round(data × 2 param ) is less than -2 bw-1 , quant becomes -2 bw-1 ; when round(data × 2 param ) is greater than 2 bw-1 -1, quant becomes 2 bw-1 -1; when round(data × 2 param ) is between -2 bw-1 and 2 bw-1 -1, quant becomes round(data × 2 param ).
[0037] By the operation of the clipping function clip, when round(data ÷ param) is less than -2 bw-1 , quant becomes -2 bw-1 ; when round(data ÷ param) is greater than 2 bw-1 -1, quant becomes 2 bw-1 -1; when round(data ÷ param) is between -2 bw-1 and 2 bw-1 -1, quant becomes round(data ÷ param).
[0038] In subsequent embodiments, the meaning of the clip operation is similar, so it will not be explained again in the subsequent processes.
[0039] Exemplary, a sample decoding network may include multiple network layers employing floating-point type weights, and the multiple network layers may include a target network layer. For each target network layer, a candidate list of weight quantization hyperparameters corresponding to the target network layer is obtained, and the candidate list of weight quantization hyperparameters may include multiple candidate weight quantization hyperparameters. For each candidate weight quantization hyperparameter, the floating-point weights of the target network layer are simulated quantized using a fixed quantization bit width and the candidate weight quantization hyperparameter to obtain candidate simulation weights. Based on these candidate simulation weights, the floating-point input features of the target network layer are processed, and the target network... The floating-point output features of the network layer are obtained, the quantization error corresponding to the candidate weight quantization hyperparameter is determined based on the floating-point output features of the target network layer, the candidate weight quantization hyperparameter corresponding to the minimum quantization error is determined as the target weight quantization hyperparameter based on the quantization error corresponding to each candidate weight quantization hyperparameter, the fixed quantization bit width is determined as the target weight quantization bit width, the floating-point weights of the target network layer are quantized to fixed-point values based on the target weight quantization bit width and the target weight quantization hyperparameter to obtain integer weights of the target network layer, and the target decoding network is generated based on the integer weights of each target network layer.
[0040] Exemplary steps include, but are not limited to, obtaining a list of candidate weight quantization hyperparameters corresponding to the target network layer, which includes: determining a maximum weight value based on all floating-point weights of the target network layer, where each floating-point weight includes multiple weight values, and the maximum weight value is the maximum value among all weight values of all floating-point weights; generating initial weight quantization hyperparameters corresponding to the target network layer based on the maximum weight value; and constructing a list of candidate weight quantization hyperparameters based on the initial weight quantization hyperparameters.
[0041] For example, the step of generating initial weight quantization hyperparameters corresponding to the target network layer based on the maximum weight value is param=bw-1-ceil(log2(max)) or param=(max) / 2 bw-1 The process may include, but is not limited to, generating initial weight quantization hyperparameters using , where param represents the initial weight quantization hyperparameters corresponding to the target network layer, bw represents the fixed quantization bit width, max represents the maximum weight value, and ceil represents the rounding up operation.
[0042] For example, the step of obtaining candidate simulation weights by simulating quantizing the floating-point weights of the target network layer using a fixed quantization bit width and the candidate weight quantization hyperparameter is quant=2 -param ×clip(round(data×2 param ),-2 bw-1 ,2 bw-1 -1), or quant=param×clip(round(data÷param),-2 bw-1 ,2 bw-1The process may include the step of obtaining candidate simulation weights by simulating quantizing the floating-point weights of the target network layer using (-1), where quant represents the candidate simulation weights, param represents the candidate weight quantization hyperparameters, data represents the floating-point weights of the target network layer, bw represents the fixed quantization bit width, clip represents the clipping function, and round represents rounding.
[0043] Exemplary, the step of determining a quantization error corresponding to the candidate weight quantization hyperparameters based on the floating-point output features of the target network layer may include, but are not limited to, the steps of determining a sample output feature based on the floating-point output features of the target network layer, wherein the sample output feature may be an output feature of a reference network layer in a sample decoding network, and the reference network layer is any network layer located after the target network layer; and determining a quantization error corresponding to the candidate weight quantization hyperparameters based on the sample output feature and the reference output feature. The method for obtaining the reference output feature may include, but are not limited to, decoding a sample feature from a sample bitstream, determining a floating-point input feature corresponding to a sample decoding network based on the sample feature, and determining a reference output feature which may be an output feature of a reference network layer in a sample decoding network based on the floating-point input feature.
[0044] Exemplary, a sample decoding network may include multiple network layers employing floating-point weights, each including a target network layer, for each target network layer, a list of candidate feature quantization hyperparameters and a set of quantization bit widths corresponding to the target network layer are obtained, the list of candidate feature quantization hyperparameters includes multiple candidate feature quantization hyperparameters, and the set of quantization bit widths includes at least two quantization bit widths, traversing the at least two quantization bit widths from the set of quantization bit widths, and for each candidate feature quantization hyperparameter, simulating the floating-point input features of the target network layer using the current quantization bit width and the candidate feature quantization hyperparameter. The system performs quantization to obtain candidate input features, processes these candidate input features based on the pseudoweights of the target network layer, obtains floating-point output features of the target network layer, determines the quantization error corresponding to the candidate feature value quantization hyperparameter based on the floating-point output features of the target network layer, determines the candidate feature value quantization hyperparameter corresponding to the minimum quantization error based on the quantization error corresponding to each candidate feature value quantization hyperparameter in the current quantization bit width, determines the candidate feature value quantization hyperparameter corresponding to the minimum quantization error as the target feature value quantization hyperparameter if the minimum quantization error is smaller than a preset threshold, determines the current quantization bit width as the target feature value quantization bit width, and records the target feature value quantization hyperparameter and target feature value quantization bit width of the target network layer.
[0045] For example, based on the quantization error corresponding to each candidate feature value quantization hyperparameter, if the minimum quantization error is greater than or equal to a preset threshold, it is determined whether the current quantization bit width is the last quantization bit width in the quantization bit width set. If it is the last quantization bit width, the candidate feature value quantization hyperparameter corresponding to the minimum quantization error is determined as the target feature value quantization hyperparameter, the current quantization bit width is determined as the target feature value quantization bit width, and the target feature value quantization hyperparameter and target feature value quantization bit width of the target network layer are recorded. If it is not the last quantization bit width, the next quantization bit width after the current quantization bit width is set as the new current quantization bit width, and the process returns to the step of simulating quantization of the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter to obtain candidate input features.
[0046] Exemplary, the step of obtaining a list of candidate feature value quantization hyperparameters corresponding to the target network layer includes: determining a maximum feature value based on all floating-point input features of the target network layer, where each floating-point input feature includes multiple feature values, and the maximum feature value is the maximum value among all feature values of all floating-point input features; generating initial feature value quantization hyperparameters corresponding to the target network layer based on the maximum feature value; and constructing a list of candidate feature value quantization hyperparameters based on the initial feature value quantization hyperparameters.
[0047] For example, the step of generating initial feature value quantization hyperparameters corresponding to the target network layer based on the maximum feature value is param=bw-1-ceil(log2(max)) or param=(max) / 2 bw-1 The process may include, but is not limited to, generating initial feature value quantization hyperparameters using , where param represents the initial feature value quantization hyperparameters corresponding to the target network layer, bw represents the current quantization bit width, max represents the maximum feature value, and ceil represents the rounding up operation.
[0048] For example, the step of obtaining candidate input features by simulating quantization of the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter is quant=2 -param ×clip(round(data×2 param ),-2 bw-1 ,2 bw-1 -1), or quant=param×clip(round(data÷param),-2 bw-1 ,2 bw-1 The process may include, but is not limited to, the step of obtaining candidate input features by simulating quantization of the floating-point input features of the target network layer using -1), where quant represents the candidate input feature, param represents the candidate feature value quantization hyperparameter, data represents the floating-point input feature, bw represents the current quantization bit width, clip represents the clipping function, and round represents rounding.
[0049] For example, before processing the candidate input features based on the pseudoweights of the target network layer and obtaining the floating-point output features of the target network layer, the floating-point weights of the target network layer may be simulated quantized using the target weight quantization hyperparameters and target weight quantization bit width of the target network layer to obtain the pseudoweights of the target network layer.
[0050] Exemplary, the step of determining a quantization error corresponding to the candidate feature value quantization hyperparameter based on the floating-point output features of the target network layer may include, but are not limited to, the steps of determining a sample output feature based on the floating-point output features of the target network layer, wherein the sample output feature may be an output feature of a reference network layer in a sample decoding network, and the reference network layer is any network layer located after the target network layer; and determining a quantization error corresponding to the candidate feature value quantization hyperparameter based on the sample output feature and the reference output feature. The method for obtaining the reference output feature may include, but are not limited to, decoding a sample feature from a sample bitstream, determining a floating-point input feature corresponding to a sample decoding network based on the sample feature, and determining a reference output feature which may be an output feature of a reference network layer in a sample decoding network based on the floating-point input feature.
[0051] For illustrative purposes, the above execution order is merely illustrative for the sake of clarity, and in actual applications, the order of execution between steps may be changed, and this order is not limiting. Furthermore, in other embodiments, the steps of the corresponding method may not necessarily be performed in the order shown and described herein, and the method may include more or fewer steps than those described herein. Also, a single step described herein may be broken down into multiple steps in other embodiments, and multiple steps described herein may be combined into a single step in other embodiments.
[0052] As can be seen from the above technical proposals, embodiments of the present invention provide an end-to-end video image compression method by decoding target features corresponding to the current image block from a bitstream corresponding to the current image block, determining a first input feature of the target decoding network based on the target features, converting the first input feature to a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameters, processing the second input feature based on the integer weights of the target decoding network to obtain the output feature of the target decoding network, and determining the reconstructed image block corresponding to the current image block based on the output feature, thereby achieving the objective of improving encoding efficiency and decoding efficiency. A decoding network with fixed-point weights is constructed using the target weight quantization bit width and target weight quantization hyperparameters, and fixed-point input features are generated using the target feature value quantization bit width and target feature value quantization hyperparameters, thereby achieving adaptive decoding acceleration and reducing the decoding computation amount while guaranteeing decoding quality.
[0053] Example 2: The encoding process is shown in Figure 3, which is merely an example and not limiting.
[0054] The encoding side may, after obtaining the current image block x (which may be the original image block x, i.e., the input image block), perform an analytical transformation on the current image block x using an analytical transformation network (i.e., a neural network) to obtain the image features y corresponding to the current image block x. Here, performing a feature transformation on the current image block x using an analytical transformation network means transforming the current image block x into image features y in the latent domain, so that all subsequent processes can operate in the latent domain.
[0055] The image may be divided into one image block or into multiple image blocks. If the image is divided into one image block, the current image block x may be the current image, that is, the encoding and decoding processes for the image block may be applied directly to the image.
[0056] The encoding side obtains image features y, then performs a coefficient hyperparameter feature transformation on image features y to obtain coefficient hyperparameter features z, and for example, inputs image features y into a hyperparameter coding network (i.e., a neural network), and the hyperparameter coding network performs a coefficient hyperparameter feature transformation on image features y to obtain coefficient hyperparameter features z. The hyperparameter coding network may be a trained neural network, and the training process of this hyperparameter coding network is not limited as long as it can perform a coefficient hyperparameter feature transformation on image features y. Here, after the image features y of the latent domain pass through the hyperparameter coding network, hyperprior latent information z is obtained.
[0057] After obtaining the coefficient hyperparameter feature z, the encoding side may quantize the coefficient hyperparameter feature z to obtain the corresponding hyperparameter quantization feature; that is, the Q operation in Figure 3 is a quantization process. After obtaining the hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z, the hyperparameter quantization feature is encoded to obtain Bitstream#1 corresponding to the current image block (i.e., the first bitstream, and the first bitstream is the main bitstream corresponding to the current image block); that is, the AE operation in Figure 3 represents an encoding process such as an entropy encoding process. Alternatively, after obtaining the coefficient hyperparameter feature z, the encoding side may directly encode the coefficient hyperparameter feature z to obtain Bitstream#1 corresponding to the current image block without performing a quantization process. The hyperparameter quantization feature or coefficient hyperparameter feature z included in Bitstream#1 is mainly used to obtain the mean value and the probability distribution parameters of the probability distribution model.
[0058] After obtaining Bitstream#1 corresponding to the current image block, the encoding side may send Bitstream#1 corresponding to the current image block to the decoding side. The decoding side's processing process for Bitstream#1 corresponding to the current image block will be described in subsequent embodiments.
[0059] After obtaining Bitstream#1 corresponding to the current image block, the encoding side may decode Bitstream#1 to obtain the hyperparameter quantization feature, i.e., AD in Figure 3 represents the decoding process, and then the hyperparameter quantization feature is inversely quantized to obtain the coefficient hyperparameter feature z_hat, which may be the same as or different from the coefficient hyperparameter feature z, and the IQ operation in Figure 3 is the inverse quantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the encoding side may decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat without performing the inverse quantization process of the coefficient hyperparameter feature z_hat.
[0060] The encoding process for Bitstream#1 may employ a fixed probability density model encoding method, and the decoding process for Bitstream#1 may employ a fixed probability density model decoding method; however, the encoding and decoding processes are not limited to these methods.
[0061] After obtaining the coefficient hyperparameter feature z_hat, the encoding side may perform a context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the image feature y_hat of the previous image block (the process for determining the image feature y_hat will be described in subsequent examples) to obtain a predicted value mu (i.e., mean mu) corresponding to the current image block. For example, the coefficient hyperparameter feature z_hat and the image feature y_hat may be input to a mean prediction network, and the mean prediction network may determine the predicted value mu based on the coefficient hyperparameter feature z_hat and the image feature y_hat. This prediction process is not limited. Here, for the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded image feature y_hat. These two are combined and input to obtain a more accurate predicted value mu. The predicted value mu is used to obtain a residual by calculating the difference from the original feature, and the residual and the decoded feature are added together to obtain the reconstructed y.
[0062] Note that the mean prediction network is a selectable neural network; that is, it is not necessary to have a mean prediction network, meaning that it is not necessary to determine the predicted value mu by the mean prediction network. The dashed box in Figure 3 indicates that the mean prediction network is selectable.
[0063] The encoding side can obtain image features y and then determine residual features r based on image features y and predicted values mu. For example, residual features r are defined as the difference between image features y and predicted values mu. Subsequently, feature processing is performed on residual features r to obtain image features s. This feature processing process is not limited to any specific method. In this case, an average value prediction network is required, and the predicted values mu are provided by this network. Alternatively, the encoding side may obtain image features s after obtaining image features y, and this feature processing process is not limited to any specific method. In this case, an average value prediction network is not required, and the dashed box indicates that the residual process is a selectable process.
[0064] After obtaining the image feature s, the encoding side may quantize the image feature s to obtain the corresponding image quantized feature; that is, the Q operation in Figure 3 is a quantization process. After obtaining the corresponding image quantized feature, the encoding side may encode the image quantized feature to obtain Bitstream #2 (i.e., the second bitstream) corresponding to the current image block; that is, the AE operation in Figure 3 represents an encoding process such as an entropy encoding process. Alternatively, the encoding side may directly encode the image feature s to obtain Bitstream #2 corresponding to the current image block without performing a quantization process of the image feature s.
[0065] After obtaining Bitstream#2 corresponding to the current image block, the encoding side may send Bitstream#2 corresponding to the current image block to the decoding side. The decoding side's processing process for Bitstream#2 corresponding to the current image block will be described in subsequent embodiments.
[0066] After obtaining Bitstream#2 corresponding to the current image block, the encoding side may decode Bitstream#2 to obtain image quantization features, i.e., AD in Figure 3 represents the decoding process, and then the encoding side may dequantize the image quantization features to obtain image features s', image features s' may be the same as or different from image features s, and the IQ operation in Figure 3 is the dequantization process. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoding side may decode Bitstream#2 to obtain image features s' without performing the dequantization process of image quantization features.
[0067] The encoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on image features s'. This feature reconstruction process is not limited to any specific method; any feature reconstruction method may be used, resulting in a residual feature r_hat, which may be the same as or different from the residual feature r. After obtaining the residual feature r_hat, the encoding side determines the image feature y_hat based on the residual feature r_hat and the predicted value mu. The image feature y_hat may be the same as or different from the image feature y. For example, the sum of the residual feature r_hat and the predicted value mu is taken as the image feature y_hat. In this case, it is necessary to implement an mean prediction network, which provides the predicted value mu. Alternatively, the encoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on image features s' to obtain the image feature y_hat. The image feature y_hat may be the same as or different from the image feature y. In this case, it is not necessary to implement an mean prediction network, and the dashed box indicates that the residual process is a selectable process.
[0068] The encoding side may, after obtaining the image feature y_hat, perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat can be input to a composite transformation network, and the composite transformation network can perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is complete.
[0069] In one possible embodiment, when the encoding side encodes image quantization features or image features s to obtain Bitstream#2 corresponding to the current image block, the encoding side first needs to determine a probability distribution model, and then encodes the image quantization features or image features s based on the probability distribution model. Similarly, when decoding Bitstream#2, the encoding side first needs to determine a probability distribution model, and then decodes Bitstream#2 based on the probability distribution model.
[0070] To obtain a probability distribution model, as shown in Figure 3, the encoding side may, after obtaining the coefficient hyperparameter feature z_hat, perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. Alternatively, the coefficient hyperparameter feature z_hat may be input to a probabilistic hyperparameter decoding network, and the probabilistic hyperparameter decoding network may perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. Or, after obtaining the probability distribution parameter p, a probability distribution model may be generated based on the probability distribution parameter p. Here, the probabilistic hyperparameter decoding network may be a trained neural network, and the training process of this probabilistic hyperparameter decoding network is not limited; it is sufficient that an inverse coefficient hyperparameter feature transformation can be performed on the coefficient hyperparameter feature z_hat.
[0071] In one possible embodiment, the encoding-side processing process may be performed by a deep learning model or a neural network model to realize an end-to-end image compression and encoding process, but is not limited to this encoding process.
[0072] Example 3: The decoding process is shown in Figure 4, which is merely an example and not limiting.
[0073] After obtaining Bitstream#1 corresponding to the current image block, the decoding side may decode Bitstream#1 to obtain the hyperparameter quantization feature, i.e., AD in Figure 4 represents the decoding process, and then the hyperparameter quantization feature is inversely quantized to obtain the coefficient hyperparameter feature z_hat, which may be the same as or different from the coefficient hyperparameter feature z, and the IQ operation in Figure 4 is the inverse quantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the decoding side may decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat without performing the inverse quantization process of the coefficient hyperparameter feature z_hat.
[0074] For the decoding process of Bitstream#1, a fixed probability density model decoding method may be used, but is not limited to this method.
[0075] The image may be divided into one image block or into multiple image blocks. If the image is divided into one image block, the current image block x may be the current image; that is, the decoding process for the image block may be applied directly to the image.
[0076] After obtaining the coefficient hyperparameter feature z_hat, the decoding side may perform a context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the image feature y_hat of the previous image block (the process for determining the image feature y_hat will be described in subsequent examples) to obtain a predicted value mu (i.e., mean mu) corresponding to the current image block. For example, the coefficient hyperparameter feature z_hat and the image feature y_hat may be input to a mean prediction network, and the mean prediction network may determine the predicted value mu based on the coefficient hyperparameter feature z_hat and the image feature y_hat. This prediction process is not limited. Here, for the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded image feature y_hat, and both are combined and input to obtain a more accurate predicted value mu.
[0077] Note that the mean prediction network is a selectable neural network; that is, it is not necessary to have a mean prediction network, meaning that it is not necessary to determine the predicted value mu by the mean prediction network. The dashed box in Figure 4 indicates that the mean prediction network is selectable.
[0078] After obtaining Bitstream#2 corresponding to the current image block, the decoding side may decode Bitstream#2 to obtain image quantization features, i.e., AD in Figure 4 represents the decoding process, and then the decoding side may dequantize the image quantization features to obtain image features s', image features s' may be the same as or different from image features s, and the IQ operation in Figure 4 is the dequantization process. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding side may decode Bitstream#2 to obtain image features s' without performing the dequantization process of image quantization features.
[0079] The decoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on image features s' to obtain residual features r_hat, where residual features r_hat may be the same as or different from residual features r. After obtaining residual features r_hat, the decoding side determines image features y_hat based on residual features r_hat and predicted values mu, where image features y_hat may be the same as or different from image features y, for example, the sum of residual features r_hat and predicted values mu is taken as image features y_hat. In this case, it is necessary to set up an mean prediction network, which provides the predicted values mu. Alternatively, the decoding side may, after obtaining image features s', perform feature reconstruction on image features s' to obtain image features y_hat, where image features y_hat may be the same as or different from image features y. In this case, it is not necessary to set up an mean prediction network, and the dashed box indicates that the residual process is a selectable process.
[0080] The decoding side may, after obtaining the image feature y_hat, perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat may be input to a composite transformation network, and the composite transformation network may perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is complete.
[0081] In one possible embodiment, when decoding Bitstream#2, the decoding side first needs to determine a probability distribution model, and then decodes Bitstream#2 based on this probability distribution model. To obtain the probability distribution model, as shown in Figure 4, the decoding side obtains the coefficient hyperparameter feature z_hat, then performs an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. For example, the coefficient hyperparameter feature z_hat may be input to a probabilistic hyperparameter decoding network, and the probabilistic hyperparameter decoding network may perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. Alternatively, after obtaining the probability distribution parameter p, a probability distribution model may be generated based on the probability distribution parameter p. Here, the probabilistic hyperparameter decoding network may be a trained neural network, and the training process of this probabilistic hyperparameter decoding network is not limited; it is sufficient that the probability distribution parameter p can be obtained by performing an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat.
[0082] In one possible embodiment, the decoding process may be performed by a deep learning model or a neural network model to realize an end-to-end image compression and encoding process, and is not limited to this decoding process.
[0083] Example 4: In Examples 2 and 3, the encoding side includes an analytical transform network, a hyperparameter coding network, a mean prediction network, a probabilistic hyperparameter decoding network, and a synthetic transform network, while the decoding side includes a mean prediction network, a probabilistic hyperparameter decoding network, and a synthetic transform network. Any of these networks may be neural networks. All neural networks on the decoding side may be used as the decoding network, or some of the neural networks on the decoding side may be used as the decoding network; this is not limited to these cases. For example, all of the networks among the mean prediction network, the probabilistic hyperparameter decoding network, and the synthetic transform network may be used as the decoding network, or some of the networks among the mean prediction network, the probabilistic hyperparameter decoding network, and the synthetic transform network may be used as the decoding network. If other neural networks are included on the decoding side, those other neural networks may be used as the decoding network; this is not limited to these cases.
[0084] In one possible embodiment, we will describe an example where all neural networks on the decoding side are the decoding network, and as shown in Figure 5A, the decoding network may include a probabilistic hyperparameter decoding network and a synthetic transformation network. As shown in Figure 5B, the decoding network may include a mean prediction network, a probabilistic hyperparameter decoding network and a synthetic transformation network. In Figures 5A and 5B, the decoding network may not include any other networks and may include only the synthetic transformation network, but is not limited to this.
[0085] In one possible embodiment, the encoding or decoding side may decode a target feature corresponding to the current image block from the bitstreams corresponding to the current image block (e.g., a first bitstream and a second bitstream), determine the input features of the decoding network based on the target features, process the input features based on the decoding network, obtain the output features of the decoding network, and then determine the reconstructed image block corresponding to the current image block based on the output features, and this process will be described below.
[0086] As shown in Figure 5A, the encoding or decoding side may, after obtaining Bitstream#1 (i.e., the first bitstream) corresponding to the current image block, decode Bitstream#1 to obtain the hyperparameter quantization feature (i.e., the target feature of the current image block), and then dequantize the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat (i.e., determine the input features of the probabilistic hyperparameter decoding network based on this target feature). Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the encoding or decoding side may decode Bitstream#1 without performing the dequantization process of the coefficient hyperparameter feature z_hat to obtain the coefficient hyperparameter feature z_hat (i.e., the target feature of the current image block, which is directly used as the input feature of the probabilistic hyperparameter decoding network).
[0087] After obtaining the coefficient hyperparameter feature z_hat (i.e., the input feature of the probabilistic hyperparameter decoding network), the coefficient hyperparameter feature z_hat is input to the probabilistic hyperparameter decoding network (i.e., the decoding network), and the probabilistic hyperparameter decoding network performs an inverse transformation of the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p (i.e., the output feature of the probabilistic hyperparameter decoding network) corresponding to the current image block. After obtaining the probability distribution parameter p, a probability distribution model may be generated based on the probability distribution parameter p.
[0088] The encoding or decoding side may, after obtaining Bitstream#2 (i.e., the second bitstream) corresponding to the current image block, decode Bitstream#2 to obtain image quantization features (i.e., the target features of the current image block), dequantize the image quantization features to obtain image features s', and perform feature reconstruction on image features s' to obtain image features y_hat (i.e., determine the input features of the composite transformation network based on the target features). Alternatively, the encoding or decoding side may, after obtaining Bitstream#2 corresponding to the current image block, decode Bitstream#2 to obtain image features s' (i.e., the target features of the current image block) without performing the dequantization process of image quantization features, and perform feature reconstruction on image features s' to obtain image features y_hat (i.e., determine the input features of the composite transformation network based on the target features). The encoding or decoding side may decode Bitstream#2 based on a probability distribution model when decoding Bitstream#2, and this process will not be explained further.
[0089] The encoding or decoding side obtains the image feature y_hat, then inputs the image feature y_hat into the composite transformation network (i.e., the decoding network), and the composite transformation network performs a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x (i.e., the output feature of the composite transformation network). As can be seen from the above, a composite transformation is performed on the image feature y_hat based on the composite transformation network to obtain the reconstructed image block x_hat corresponding to the current image block x.
[0090] After obtaining the output features of the composite transformation network, the output features of the composite transformation network may be used as the reconstructed image block x_hat corresponding to the current image block x; that is, the reconstructed image block x_hat corresponding to the current image block x is determined based on these output features.
[0091] As shown in Figure 5B, the encoding or decoding side may, after obtaining Bitstream#1 corresponding to the current image block, decode Bitstream#1 to obtain the hyperparameter quantization feature (i.e., the target feature of the current image block), and then dequantize the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat (i.e., determine the input features of the probabilistic hyperparameter decoding network based on this target feature). Alternatively, after obtaining Bitstream#1 corresponding to the current image block, Bitstream#1 may be decoded to obtain the coefficient hyperparameter feature z_hat (i.e., the target feature of the current image block, which is directly used as the input feature of the probabilistic hyperparameter decoding network).
[0092] After obtaining the coefficient hyperparameter feature z_hat (i.e., the input feature of the probabilistic hyperparameter decoding network), the coefficient hyperparameter feature z_hat is input to the probabilistic hyperparameter decoding network (i.e., the decoding network), and the probabilistic hyperparameter decoding network performs an inverse transformation of the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p (i.e., the output feature of the probabilistic hyperparameter decoding network) corresponding to the current image block. After obtaining the probability distribution parameter p, a probability distribution model may be generated based on the probability distribution parameter p.
[0093] After obtaining the coefficient hyperparameter feature z_hat (i.e., the input feature of the mean prediction network), the coefficient hyperparameter feature z_hat and the image feature y_hat from the previous image block are input to the mean prediction network (i.e., the decoding network). The mean prediction network determines the predicted value mu (i.e., the output feature of the mean prediction network) based on the coefficient hyperparameter feature z_hat and the image feature y_hat. In other words, the mean prediction network makes a context-based prediction and obtains the predicted value mu (i.e., mean mu) corresponding to the current image block.
[0094] The encoding or decoding side obtains Bitstream#2 corresponding to the current image block, decodes Bitstream#2 to obtain image quantization features (i.e., target features of the current image block), dequantizes the image quantization features to obtain image features s', performs feature reconstruction (i.e., residual reconstruction, which is the reverse process of residual processing) on image features s' to obtain residual features r_hat, determines image features y_hat based on residual features r_hat and predicted values mu, for example, the sum of residual features r_hat and predicted values mu may be used as image features y_hat (i.e., the input features of the composite transformation network are determined based on the target features).
[0095] Alternatively, the encoding or decoding side obtains Bitstream#2 corresponding to the current image block, decodes Bitstream#2 to obtain image features s' (i.e., the target features of the current image block), performs feature reconstruction on image features s' to obtain residual features r_hat, and determines image features y_hat based on residual features r_hat and predicted values mu. For example, the sum of residual features r_hat and predicted values mu may be used as image features y_hat (i.e., the input features of the composite transformation network are determined based on the target features).
[0096] The encoding or decoding side obtains the image feature y_hat, then inputs the image feature y_hat into the composite transformation network (i.e., the decoding network), and the composite transformation network performs a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x (i.e., the output feature of the composite transformation network). As can be seen from the above, a composite transformation is performed on the image feature y_hat based on the composite transformation network to obtain the reconstructed image block x_hat corresponding to the current image block x.
[0097] After obtaining the output features of the composite transformation network, the output features of the composite transformation network may be used as the reconstructed image block x_hat corresponding to the current image block x; that is, the reconstructed image block x_hat corresponding to the current image block x is determined based on these output features.
[0098] Example 5: In relation to Example 4, the decoding network may be a floating-point decoding network, that is, the decoding network may employ floating-point weights (i.e., floating-point parameters), that is, all weights of the decoding network are of type float, the input feature values of the decoding network are of type float, and the output feature values of the decoding network are of type float. For example, for each network layer of the decoding network, the network layer may employ floating-point weights, that is, all weights of the network layer are of type float, the input feature values of the network layer are of type float, and the output feature values of the network layer are of type float.
[0099] However, in order to implement a decoding network for floating-point weights, the decoding side needs to store weights of type float, and storing float-type weights requires a large amount of memory resources. Also, if the decoding side processes float-type feature values using float-type weights, it requires a large amount of computational resources, resulting in slow processing speed for the decoding side.
[0100] In response to the above findings, this embodiment provides an efficient inference method for decoding after image coding, and allows processing of image features using a fixed-point decoding network. The decoding network in Embodiment 4 is a fixed-point decoding network, meaning that the decoding network may use fixed-point weights (i.e., fixed-point parameters), where all weights of the decoding network are fixed-bit weights, the input feature values of the decoding network are fixed-bit feature values, and the output feature values of the decoding network are fixed-bit feature values. For example, the network layer of the decoding network may employ fixed-point weights, where the weights of the network layer are fixed-bit weights, the input feature values of the network layer are fixed-bit feature values, and the output feature values of the network layer are fixed-bit feature values.
[0101] Clearly, in order to implement a decoding network with fixed-point weights, the decoding side needs to store fixed-bit weights (e.g., 8 bits, 16 bits, etc.). However, storing fixed-bit weights only requires a small amount of memory resources, thus saving memory resources. Furthermore, when the decoding side processes fixed-bit feature values using fixed-bit weights, it only requires a small amount of computational resources, thus saving computational resources. Moreover, when the decoding side processes fixed-bit feature values using fixed-bit weights, it can improve processing speed and achieve adaptive acceleration of decoding.
[0102] This embodiment is a training-free method that can be quickly integrated into any floating-point decoding network to obtain a fixed-point weight decoding network. The fixed-point weight decoding network can be inferred directly on the hardware platform. When obtaining the fixed-point weight decoding network, the sensitivity of each network layer to quantization is first evaluated to determine an appropriate quantization bit width for each network layer, thereby ensuring the quantization compression ratio and quantization performance of the decoding network.
[0103] For example, quantizing floating-point parameters of a decoding network to fixed-point parameters can reduce the memory consumption of the decoding network, and quantizing floating-point feature values of a decoding network to fixed-point feature values can reduce the bandwidth during forward computation. Furthermore, converting floating-point operations in a decoding network to fixed-point operations can reduce the computational complexity of the decoding network, and fixed-point computation also contributes to decoding consistency between different devices. However, quantization of the decoding network itself is achieved by reducing the precision of the data type and computation, which affects the performance of the decoding network, and this is unacceptable in image coding and decoding. In response to the above findings, when a neural network-based decoder decodes a bitstream, the sensitivity to quantization differs among the network layers of the neural network. That is, some network layers can maintain quantization performance with low bits, while some network layers require high bits to guarantee quantization accuracy. Considering that this relates to the parameters and feature distribution learned by the neural network, this embodiment proposes an efficient compression and inference method to achieve the maximum possible compressed bit width while guaranteeing performance after compression. A decoding network for fixed-point weights is constructed using the target weight quantization bit width and target weight quantization hyperparameters. Fixed-point input features are generated using the target feature value quantization bit width and target feature value quantization hyperparameters. Adaptive decoding acceleration is achieved, reducing the decoding computation while guaranteeing decoding quality.
[0104] Example 6: As shown in Figure 6A, three modules are added to the decoding network shown in Figure 5A (e.g., a probabilistic hyperparameter decoding network, a synthetic transformation network): a bitstream extraction module, a quantization analysis module, and a decoding network quantization module. As shown in Figure 6B, three modules are added to the decoding network shown in Figure 5B (e.g., a mean prediction network, a probabilistic hyperparameter decoding network, a synthetic transformation network): a bitstream extraction module, a quantization analysis module, and a decoding network quantization module.
[0105] For example, after the decoding network has been trained (the decoding network is obtained using floating-point weight training), the floating-point weights of the decoding network are fixed. To facilitate distinction, the decoding network that employs floating-point weights is called the sample decoding network. The sample decoding network may contain multiple network layers that employ floating-point weights, and these multiple network layers may contain a target network layer, which is a network layer that needs to employ integer weights (i.e., the floating-point weights of the target network layer need to be converted to integer weights). If the sample decoding network contains network layer 1 and network layer 2, and network layer 1 needs to employ integer weights, and network layer 2 does not need to employ integer weights, then network layer 1 is the target network layer, and network layer 2 is not the target network layer.
[0106] In one possible embodiment, all network layers in the sample decoding network may be the target network layer, some network layers in the sample decoding network may be the target network layer, or all convolutional layers in the sample decoding network may be the target network layer. These are just some examples and are not limited thereto. Any one or more network layers in the sample decoding network may be the target network layer, or all network layers may be the target network layer.
[0107] For example, the following steps may be taken to transform a sample decoding network into a target decoding network with fixed-point weights.
[0108] In step S11, the bitstream extraction module acquires a sample bitstream and inputs the sample bitstream into the quantization analysis module.
[0109] For example, images can be collected from several real-world scenes (e.g., application scenes for image acquisition devices), such as 100 frames or 200 frames. The image of each frame is processed using the flow shown in Figure 3 to obtain a first bitstream (Bitstream #1) and a second bitstream (Bitstream #2) corresponding to the image of that frame. The first and second bitstreams can be called sample bitstreams, or the first bitstream can be called a sample bitstream, or the second bitstream can be called a sample bitstream. For the sake of clarity, in the subsequent process, we will use the example of referring to the first and second bitstreams as sample bitstreams. The bitstream extraction module can obtain the sample bitstreams corresponding to the image of each frame and input these sample bitstreams into the quantization analysis module.
[0110] In step S12, the quantization analysis module receives the sample bitstream corresponding to the image of each frame and stores the sample bitstream corresponding to the image of each frame.
[0111] Step S13, the quantization analysis module decodes sample features from the sample bitstream, determines floating-point input features corresponding to the sample decoding network based on the sample features, determines reference output features corresponding to the sample features based on the floating-point input features, the reference output features may be output features of the reference network layer in the sample decoding network. Here, the reference network layer may be any network layer in the sample decoding network, for example, the last network layer in the sample decoding network, or the second to last network layer in the sample decoding network, or the last target network layer in the sample decoding network, or any network layer located after the last target network layer in the sample decoding network, and the location of this reference network layer is not limited.
[0112] For example, for each sample bitstream, the sample bitstream can be decoded to obtain sample features. This process refers to the process for obtaining target features in Example 4, and is omitted here. After obtaining the sample features, a floating-point input feature corresponding to the sample decoding network is determined based on the sample features. This process refers to the process for determining the input features of the decoding network based on target features in Example 4, and is omitted here. After obtaining the floating-point input feature corresponding to the sample decoding network, the floating-point input feature is input to the sample decoding network, passes through each network layer of the sample decoding network sequentially, and after passing through the reference network layer of the sample decoding network, the output feature of the reference network layer is used as the reference output feature corresponding to the sample features.
[0113] Clearly, for each sample bitstream, a reference output feature corresponding to that sample bitstream can be obtained. Based on this, the quantization analysis module may store the average value of the reference output features corresponding to all sample bitstreams, the maximum value of the reference output features corresponding to all sample bitstreams, the minimum value of the reference output features corresponding to all sample bitstreams, and, as an example, the average value of the reference output features corresponding to all sample bitstreams is stored as the reference output feature.
[0114] In step S14, for each target network layer in the sample decoding network, the quantization analysis module obtains a list of candidate weight quantization hyperparameters corresponding to the target network layer, and the list of candidate weight quantization hyperparameters may include multiple candidate weight quantization hyperparameters.
[0115] First, the quantization analysis module determines the maximum weight value based on all floating-point weights of the target network layer, where each floating-point weight contains multiple weight values, and the maximum weight value is the maximum value among all the weight values of all floating-point weights. For example, the target network layer may contain multiple floating-point weights, each floating-point weight may contain multiple weight values, and for each floating-point weight, the maximum value may be selected from all the weight values corresponding to that floating-point weight, and then the maximum weight value may be selected from the maximum values corresponding to all floating-point weights. For example, if the target network layer is network layer l, the maximum weight value is max l =max(abs(w 0_l ),abs(w 1_l ),…,abs(w n_l It can also be written as )) and abs(w 0_l ) represents the maximum value corresponding to the first floating-point weight of network layer l, ..., max l This represents the maximum weight value for all floating-point weights.
[0116] The target network layer may contain multiple channels. If all channels correspond to the same target weight quantization hyperparameter, the maximum weight value is the maximum value among all floating-point weights of all channels. If each channel corresponds to a target weight quantization hyperparameter individually, each channel corresponds to one maximum weight value; that is, the maximum weight value is the maximum value among all floating-point weights of that channel. For the sake of explanation, the following explanation will use the example that all channels correspond to the same target weight quantization hyperparameter. The implementation methods for each channel corresponding to a target weight quantization hyperparameter are similar, and the target weight quantization hyperparameter for each channel can be determined using the same method. The determination method is the same as the determination method for the target weight quantization hyperparameter for all channels.
[0117] Next, the quantization analysis module generates initial weight quantization hyperparameters corresponding to the target network layer based on the maximum weight value. For example, the initial weight quantization hyperparameters corresponding to the target network layer may be generated using param=bw-1-ceil(log2(max)), where param represents the initial weight quantization hyperparameters corresponding to the target network layer, bw represents the fixed quantization bit width, max represents the maximum weight value, and ceil represents the rounding up operation.
[0118] For example, the initial weight quantization hyperparameter corresponding to the target network layer is param, where param=get_param(bw,max), and bw represents the quantization bit width corresponding to the target network layer. The quantization bit width corresponding to the target network layer may be a fixed quantization bit width, and the fixed quantization bit width may be set empirically, for example, 4 bits, 8 bits, 16 bits, etc. We will explain using the example that the fixed quantization bit width is 8 bits. param=get_param(bw,max) represents the initial weight quantization hyperparameter calculated based on the data range max under the quantization bit width bw. Since the calculation method of the initial weight quantization hyperparameter differs depending on the quantization algorithm, here is an example of an initial weight quantization hyperparameter param=bw-1-ceil(log2(max)).
[0119] Next, the quantization analysis module constructs a list of candidate weight quantization hyperparameters based on the initial weight quantization hyperparameters, and this list of candidate weight quantization hyperparameters may include multiple candidate weight quantization hyperparameters, i.e., multiple candidate weight quantization hyperparameters corresponding to the target network layer.
[0120] For example, in the weight values of a neural network, there are usually outliers (anomalous weight values, and the range of values for anomalous weight values is usually particularly large), and considering that if these outliers are too large, they will affect the quantization accuracy, a list of candidate weight quantization hyperparameters can be constructed based on the initial weight quantization hyperparameters. For example, one example of such a list of candidate weight quantization hyperparameters is param l_list =[param l -2,param l -1,param l ,param l +1,param lIt may also be +2], and the above is merely an example of a list of candidate weight quantization hyperparameters, and does not limit the candidate weight quantization hyperparameters within this list. As an example, if the target network layer is network layer l, param l represents the initial weight quantization hyperparameter corresponding to the network layer l, and param l -2, param l -1, param l , param l +1, param l +2 are the five candidate weight quantization hyperparameters corresponding to the network layer l.
[0121] In step S15, for each candidate weight quantization hyperparameter, the quantization analysis module simulates quantizing the floating-point weights of the target network layer using a fixed quantization bit width and the candidate weight quantization hyperparameter to obtain candidate simulated weights.
[0122] For illustrative purposes, the quantization bit width corresponding to the target network layer may be a fixed quantization bit width, and this fixed quantization bit width may be set based on experience, for example, 4 bits, 8 bits, or 16 bits. We will explain this using the example of a fixed quantization bit width of 8 bits.
[0123] For example, one could sequentially traverse each candidate weight quantization hyperparameter in the list of weight quantization hyperparameter candidates, and for the currently traversed candidate weight quantization hyperparameter, simulate quantize the floating-point weights of the target network layer using a fixed quantization bit width and the candidate weight quantization hyperparameter to obtain the candidate simulation weights corresponding to the candidate weight quantization hyperparameter. For example, the following formula could be used to simulate quantize the floating-point weights of the target network layer to obtain the candidate simulation weights. quant=2 -param ×clip(round(data×2 param ),-2bw-1 ,2 bw-1 -1)
[0124] quant represents the candidate simulation weight, param represents the candidate weight quantization hyperparameter, data represents the floating-point weight of the target network layer, bw represents the fixed quantization bit width, clip represents the clipping function, and round represents the rounding operation. Since different quantization algorithms use different simulation quantization methods, the above formula is merely one example of simulation quantization.
[0125] In summary, for each candidate weight quantization hyperparameter, substituting the candidate weight quantization hyperparameter param into the above formula yields the candidate simulation weight quant corresponding to that candidate weight quantization hyperparameter.
[0126] In step S16, after obtaining candidate simulation weights, the quantization analysis module processes the floating-point input features of the target network layer based on the candidate simulation weights to obtain the floating-point output features of the target network layer. Based on the floating-point output features of the target network layer, the quantization analysis module determines the quantization error corresponding to the candidate weight quantization hyperparameter.
[0127] First, the floating-point input features of the target network layer can be obtained. For example, if the target network layer is the first network layer of the sample decoding network, the input features of the sample decoding network are the floating-point input features of the target network layer. Alternatively, if the target network layer is not the first network layer of the sample decoding network, the output features of the network layer immediately preceding the target network layer are used as the floating-point input features of the target network layer. For the network layer preceding the target network layer (denoted as network layer A), if network layer A is not a network layer that requires integer weights, network layer A directly processes its input features to obtain floating-point output features, and these floating-point output features are used as the floating-point input features of the target network layer. If network layer A is a network layer that requires integer weights, after obtaining the target weight quantization hyperparameter and target weight quantization bit width of network layer A, the floating-point weights of network layer A are simulated quantized using the target weight quantization hyperparameter and target weight quantization bit width to obtain pseudoweights of network layer A. The floating-point input features of network layer A are then processed using these pseudoweights to obtain floating-point output features of network layer A, and these floating-point output features can be used as the floating-point input features of the target network layer.
[0128] Next, the floating-point input features of the target network layer are processed based on the candidate simulation weights to obtain the floating-point output features of the target network layer. For example, if the target network layer is a convolutional layer, a convolution operation is performed on the floating-point input features of the target network layer based on the candidate simulation weights to obtain the floating-point output features of the target network layer. If the target network layer is a pooling layer, a pooling operation is performed on the floating-point input features of the target network layer based on the candidate simulation weights to obtain the floating-point output features of the target network layer. The above is merely an example and is not limited to this processing method; it is related to the functions supported by the target network layer.
[0129] Next, the quantization error corresponding to the candidate weight quantization hyperparameter is determined based on the floating-point output features of the target network layer. For example, the sample output features are determined based on the floating-point output features of the target network layer, and these sample output features may also be the output features of the reference network layer in the sample decoding network, where the reference network layer is any network layer located after the target network layer.
[0130] For example, the floating-point output features of the target network layer are used as input features for the next network layer (which also uses floating-point weights), the next network layer processes the input features to obtain the floating-point output features of the next network layer, and the same process is carried out up to the reference network layer in the sample decoding network, the reference network layer processes the input features to obtain the output features of the reference network layer, and the output features of the reference network layer are used as sample output features.
[0131] Next, based on the sample output features and the reference output features, the quantization error corresponding to the candidate weight quantization hyperparameters is determined. For example, the quantization error can be determined using an error loss function based on the sample output features and the reference output features, and this error loss function may include, but is not limited to, loss functions such as MSE, cosine similarity, and KL divergence.
[0132] Here, the sample output feature may be the average value of the sample output features corresponding to all sample bitstreams, the reference output feature may be the average value of the reference output features corresponding to all sample bitstreams, or the sample output feature may be the maximum value of the sample output features corresponding to all sample bitstreams, the reference output feature may be the maximum value of the reference output features corresponding to all sample bitstreams, or the sample output feature may be the minimum value of the sample output features corresponding to all sample bitstreams, the reference output feature may be the minimum value of the reference output features corresponding to all sample bitstreams, and this is not limited to these.
[0133] In step S17, based on the quantization error corresponding to each candidate weight quantization hyperparameter, the quantization analysis module determines the candidate weight quantization hyperparameter corresponding to the minimum quantization error as the target weight quantization hyperparameter corresponding to the target network layer, and determines the fixed quantization bit width as the target weight quantization bit width corresponding to the target network layer. Up to this point, the quantization analysis module can obtain the target weight quantization hyperparameter and the target weight quantization bit width corresponding to the target network layer.
[0134] Step S18, the quantization analysis module simulates quantizing the floating-point weights of the target network layer using the target weight quantization hyperparameters and the target weight quantization bit width of the target network layer to obtain pseudoweights of the target network layer.
[0135] For example, after obtaining the target weight quantization hyperparameters and target weight quantization bit width, the floating-point weights of the target network layer may be simulated quantized using the target weight quantization hyperparameters and target weight quantization bit width to obtain pseudoweights of the target network layer, that is, the floating-point weights of the target network layer are replaced with the pseudoweights of the target network layer. For the floating-point input features of the target network layer, the target network layer processes the floating-point input features of the target network layer using these pseudoweights to obtain floating-point output features of the target network layer, and these floating-point output features are used as floating-point input features of subsequent network layers.
[0136] For example, based on the target weight quantization hyperparameter and the target weight quantization bit width, pseudoweights may be obtained by simulating quantization of floating-point weights using the following formula. quant=2 -param ×clip(round(data×2 param ),-2 bw-1 ,2 bw-1 -1)
[0137] Here, quant represents the pseudoweights of the target network layer, param represents the target weight quantization hyperparameters of the target network layer, data represents the floating-point weights of the target network layer, bw represents the target weight quantization bit width of the target network layer, clip represents the clipping function, and round represents the rounding operation. Since different quantization algorithms use different simulation quantization methods, the above equation is merely one example of simulation quantization.
[0138] For example, for each target network layer, after obtaining the target weight quantization hyperparameter and target weight quantization bit width of the target network layer, the target feature value quantization bit width and target feature value quantization hyperparameter may be obtained. For example, the target feature value quantization bit width and target feature value quantization hyperparameter may be obtained using the following steps.
[0139] In step S19, for each target network layer in the sample decoding network, the quantization analysis module obtains a list of candidate feature value quantization hyperparameters and a set of quantization bit widths corresponding to the target network layer, wherein the list of candidate feature value quantization hyperparameters may include multiple candidate feature value quantization hyperparameters, and the set of quantization bit widths may include at least two quantization bit widths.
[0140] First, the quantization analysis module may determine the maximum feature value based on all floating-point input features of the target network layer, where each floating-point input feature contains multiple feature values, and the maximum feature value is the maximum value among all feature values of all floating-point input features. For example, the target network layer may correspond to multiple floating-point input features (for example, 100 sample bitstreams correspond to 100 floating-point input features), where each floating-point input feature contains multiple feature values, and for each floating-point input feature, the maximum value may be selected from all feature values corresponding to that floating-point input feature, and then the maximum feature value may be selected from the maximum values corresponding to all floating-point input features. For example, if the target network layer is network layer l, the maximum feature value is max l =max(abs(f 0_l ),abs(f 1_l ),…,abs(f n_l It can also be written as )) and abs(f 0_l ) represents the maximum value corresponding to the first floating-point input feature of network layer l, ..., max l This represents the maximum feature value across all floating-point input features.
[0141] Next, the quantization analysis module generates initial feature value quantization hyperparameters corresponding to the target network layer based on the maximum feature value. For example, the initial feature value quantization hyperparameters corresponding to the target network layer may be generated using param=bw-1-ceil(log2(max)), where param represents the initial feature value quantization hyperparameters corresponding to the target network layer, bw represents the current quantization bit width, max represents the maximum feature value, and ceil represents the rounding up operation.
[0142] For example, the initial feature value quantization hyperparameter corresponding to the target network layer is param, where param = get_param(bw, max), where bw represents the quantization bit width corresponding to the target network layer, and the quantization bit width corresponding to the target network layer is the current quantization bit width (i.e., the quantization bit width currently traversed from the quantization bit width set). param = get_param(bw, max) represents the initial feature value quantization hyperparameter calculated based on the data range max under the quantization bit width bw. Since the calculation method of the initial feature value quantization hyperparameter differs for different quantization algorithms, this is just one example of the initial feature value quantization hyperparameter, and is not the only one that is applicable.
[0143] Next, the quantization analysis module constructs a list of candidate feature value quantization hyperparameters based on the initial feature value quantization hyperparameters, and this list of candidate feature value quantization hyperparameters may include multiple candidate feature value quantization hyperparameters, i.e., multiple candidate feature value quantization hyperparameters corresponding to the target network layer. For example, considering that in the feature values of a neural network, outliers (anomalous feature values, whose value range is usually particularly large) usually exist, and that if these outliers are too large, it will affect the quantization accuracy, a list of candidate feature value quantization hyperparameters can be constructed based on the initial feature value quantization hyperparameters, for example, an example of such a list of candidate feature value quantization hyperparameters is param l_list=[param l -2,param l -1,param l ,param l +1,param l It may also be +2], and the above is merely an example of a list of candidate feature value quantization hyperparameters, and is not limited to the candidate feature value quantization hyperparameters in this list of candidate feature value quantization hyperparameters. As an example, if the target network layer is network layer l, param l represents the initial feature value quantization hyperparameter corresponding to network layer l, and the list of candidate feature value quantization hyperparameters contains five candidate feature value quantization hyperparameters corresponding to network layer l.
[0144] The quantization analysis module may generate a set of quantization bit widths, which may include at least two quantization bit widths, for example, two quantization bit widths such as 8-bit, 16-bit, or 8-bit, 32-bit, or 16-bit, 32-bit, or 4-bit, 8-bit. In another example, the set of quantization bit widths may include three quantization bit widths, for example, 8-bit, 16-bit, 32-bit, or 4-bit, 8-bit, 16-bit, or 8-bit, 16-bit, 24-bit. In yet another example, the set of quantization bit widths may include four quantization bit widths. The above are just some examples of quantization bit width sets and are not limited thereto.
[0145] In step S20, the quantization analysis module traverses the quantization bit widths in the set of quantization bit widths, based on increasing order of quantization bit widths. For each candidate feature value quantization hyperparameter, the quantization analysis module simulates quantizing the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter to obtain candidate input features.
[0146] For example, based on the quantization bit widths in ascending order, the quantization analysis module first sets the first quantization bit width as the current quantization bit width. For instance, if the quantization bit width set includes two quantization bit widths, such as 8-bit and 16-bit, the quantization analysis module first sets the 8-bit quantization bit width as the current quantization bit width and then performs step S20 based on the current quantization bit width.
[0147] For example, one may sequentially traverse each candidate feature quantization hyperparameter in the list of candidate feature quantization hyperparameters, and for the currently traversed candidate feature quantization hyperparameter, simulate quantization of the floating-point input features of the target network layer using the current quantization bit width and the candidate feature quantization hyperparameter to obtain the candidate input features corresponding to the candidate feature quantization hyperparameter. For example, quant=2 -param ×clip(round(data×2 param ),-2 bw-1 ,2 bw-1 The floating-point input features of the target network layer may be simulated quantized using (-1) to obtain the candidate input features, where quant represents the candidate input feature, param represents the candidate feature value quantization hyperparameter, data represents the floating-point input feature, bw represents the current quantization bit width, clip represents the clipping function, and round represents the rounding operation. Since different quantization algorithms use different simulation quantization methods, the above formula is merely one example of simulation quantization.
[0148] In summary, for each candidate feature value quantization hyperparameter, substituting the candidate feature value quantization hyperparameter param into the above formula yields the candidate input feature quant corresponding to that candidate feature value quantization hyperparameter.
[0149] In step S21, after obtaining candidate input features, the quantization analysis module processes the candidate input features of the target network layer based on the pseudoweights of the target network layer to obtain floating-point output features of the target network layer. Based on the floating-point output features of the target network layer, the quantization analysis module determines the quantization error corresponding to the candidate feature value quantization hyperparameter.
[0150] First, the floating-point input features of the target network layer (the details of these floating-point input features can be found in step S16) are simulated quantized using the current quantization bit width and candidate feature value quantization hyperparameters to obtain candidate input features corresponding to these candidate feature value quantization hyperparameters, and these candidate input features are used as input features for the target network layer.
[0151] Next, the candidate input features are processed based on the pseudoweights of the target network layer to obtain the floating-point output features of the target network layer. For example, if the target network layer is a convolutional layer, a convolution operation is performed on the candidate input features based on the pseudoweights to obtain the floating-point output features of the target network layer. If the target network layer is a pooling layer, a pooling operation is performed on the candidate input features based on the pseudoweights to obtain the floating-point output features of the target network layer. The above is merely an example and does not limit the processing method, but relates to the functions supported by the target network layer. After obtaining the target weight quantization hyperparameters and target weight quantization bit width corresponding to the target network layer, the candidate input features are processed based on the pseudoweights of the target network layer to obtain the floating-point output features in order to simulate quantize the floating-point weights of the target network layer and obtain pseudoweights.
[0152] Next, based on the floating-point output features of the target network layer, the quantization error corresponding to the candidate feature value quantization hyperparameter is determined. For example, the sample output features are determined based on the floating-point output features of the target network layer, and these sample output features may also be the output features of the reference network layer in the sample decoding network, where the reference network layer is any network layer located after the target network layer.
[0153] For example, the floating-point output features of the target network layer are used as input features for the next network layer (which also uses floating-point weights), the next network layer processes the input features to obtain the floating-point output features of the next network layer, and the same process is carried out up to the reference network layer in the sample decoding network, the reference network layer processes the input features to obtain the output features of the reference network layer, and the output features of the reference network layer are used as sample output features.
[0154] Next, based on the sample output features and the reference output features, the quantization error corresponding to the candidate feature value quantization hyperparameter is determined. For example, the quantization error can be determined using an error loss function based on the sample output features and the reference output features, and this error loss function may include, but is not limited to, loss functions such as MSE, cosine similarity, and KL divergence.
[0155] Here, the sample output feature may be the average value of the sample output features corresponding to all sample bitstreams, the reference output feature may be the average value of the reference output features corresponding to all sample bitstreams, or the sample output feature may be the maximum value of the sample output features corresponding to all sample bitstreams, the reference output feature may be the maximum value of the reference output features corresponding to all sample bitstreams, or the sample output feature may be the minimum value of the sample output features corresponding to all sample bitstreams, the reference output feature may be the minimum value of the reference output features corresponding to all sample bitstreams, and this is not limited to these.
[0156] In step S22, based on the quantization error corresponding to each candidate feature value quantization hyperparameter, the quantization analysis module determines whether the minimum quantization error is smaller than a preset threshold (which may be set based on experience). If it is smaller than the threshold, step S23 is executed; otherwise, step S24 is executed.
[0157] In step S23, the quantization analysis module determines a candidate feature value quantization hyperparameter corresponding to the minimum quantization error as the target feature value quantization hyperparameter corresponding to the target network layer, determines the current quantization bit width as the target feature value quantization bit width corresponding to the target network layer, and records the target feature value quantization hyperparameter and target feature value quantization bit width of the target network layer.
[0158] Up to this point, the quantization analysis module can obtain the target feature value quantization hyperparameters and target feature value quantization bit width corresponding to the target network layer, and record the target feature value quantization hyperparameters and target feature value quantization bit width of the target network layer.
[0159] Step S24, the quantization analysis module determines whether the current quantization bit width is the last quantization bit width in the quantization bit width set.
[0160] If it is the last quantization bit width, the candidate feature value quantization hyperparameter corresponding to the minimum quantization error is determined as the target feature value quantization hyperparameter corresponding to the target network layer, the current quantization bit width is determined as the target feature value quantization bit width corresponding to the target network layer, and the target feature value quantization hyperparameter and target feature value quantization bit width of the target network layer are recorded. This process can be referred to in step S23.
[0161] If it is not the last quantization bit width, the next quantization bit width after the current quantization bit width is set as the new current quantization bit width, and the process returns to step S20. Based on the current quantization bit width, for each candidate feature value quantization hyperparameter, the quantization analysis module simulates quantizing the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter to obtain candidate input features.
[0162] For example, if the current quantization bit width is set to 8 bits in the first step, and the minimum quantization error is greater than or equal to a preset threshold, then the current quantization bit width is set to 16 bits in the second step, and the above steps are repeated, and so on.
[0163] For example, after recording the target feature value quantization hyperparameters and target feature value quantization bit width of the target network layer, the quantization analysis module obtains the target weight quantization hyperparameters, target weight quantization bit width, target feature value quantization hyperparameters, and target feature value quantization bit width corresponding to the target network layer. The above steps are repeated for the next target network layer of the target network layer until the processing process for that target network layer is completed and the target weight quantization hyperparameters, target weight quantization bit width, target feature value quantization hyperparameters, and target feature value quantization bit width corresponding to each target network layer are obtained.
[0164] In step S25, after obtaining the target weight quantization hyperparameters, target weight quantization bit width, target feature value quantization hyperparameters, and target feature value quantization bit width corresponding to each target network layer, the decoding network quantization module generates a target decoding network based on the sample decoding network. For example, for each target network layer in the sample decoding network, the decoding network quantization module quantizes the floating-point weights of the target network layer into fixed-point weights based on the target weight quantization bit width and target weight quantization hyperparameters corresponding to the target network layer to obtain integer weights for the target network layer. After obtaining the integer weights for each target network layer, the decoding network quantization module can generate a target decoding network based on the integer weights for each target network layer.
[0165] For example, for each target network layer, the quantization analysis module may output the target weight quantization bit width and target weight quantization hyperparameters corresponding to the target network layer to the decoding network quantization module, and the decoding network quantization module may obtain integer weights of the target network layer by quantizing the floating-point weights of the target network layer to fixed-point quantization based on the target weight quantization bit width and target weight quantization hyperparameters corresponding to the target network layer, and this is not limited to the fixed-point quantization process of floating-point weights.
[0166] For each target network layer, the quantization analysis module may output the target feature value quantization hyperparameter and target feature value quantization bit width corresponding to the target network layer to the decoding network quantization module, which records the target feature value quantization hyperparameter and target feature value quantization bit width of the target network layer. For example, the decoding network quantization module may record the target feature value quantization hyperparameter and target feature value quantization bit width in the configuration information of the target network layer, and when the target network layer performs the relevant processing, it can read the target feature value quantization hyperparameter and target feature value quantization bit width from the configuration information. The target feature value quantization hyperparameter and target feature value quantization bit width may be recorded in other memory areas as long as the target network layer can obtain the target feature value quantization hyperparameter and target feature value quantization bit width.
[0167] After obtaining the integer weights for each target network layer, the floating-point weights of the target network layer can be replaced with the integer weights, and in this way, a target decoding network can be constructed with all the replaced target network layers. At this point, the target decoding network can be obtained and introduced into either the encoding or decoding side. For example, the target decoding network may include a probabilistic hyperparameter decoding network and a composite transformation network, or the target decoding network may include a mean prediction network, a probabilistic hyperparameter decoding network and a composite transformation network. By introducing these target decoding networks into the encoding side, when the encoding side performs processing using Example 2, the processing will be performed based on the target decoding network. Also, by introducing these target decoding networks into the decoding side, when the decoding side performs processing using Example 3, the processing will be performed based on the target decoding network.
[0168] Example 7: As shown in Figure 6A, three modules—a bitstream extraction module, a quantization analysis module, and a decoding network quantization module—are added to the decoding network shown in Figure 5A. As shown in Figure 6B, three modules—a bitstream extraction module, a quantization analysis module, and a decoding network quantization module—are added to the decoding network shown in Figure 5B. The sample decoding network includes multiple network layers that employ floating-point weights, and these multiple network layers may include a target network layer, which is a network layer that must employ integer weights (i.e., the floating-point weights of the target network layer must be converted to integer weights). For example, all network layers in the sample decoding network may be the target network layer, some network layers in the sample decoding network may be the target network layer, or all convolutional layers in the sample decoding network may be the target network layer.
[0169] For example, the following steps may be taken to transform a sample decoding network into a target decoding network with fixed-point weights.
[0170] In step S31, the bitstream extraction module acquires a sample bitstream and inputs the sample bitstream into the quantization analysis module.
[0171] In step S32, the quantization analysis module receives the sample bitstream corresponding to the image of each frame and stores the sample bitstream corresponding to the image of each frame.
[0172] In step S33, the quantization analysis module decodes sample features from the sample bitstream, determines floating-point input features corresponding to the sample decoding network based on the sample features, determines reference output features corresponding to the sample features based on the floating-point input features, and the reference output features may be output features of the reference network layer in the sample decoding network. Here, the reference network layer may be any network layer in the sample decoding network, for example, the last network layer in the sample decoding network, or the second to last network layer in the sample decoding network, or the last target network layer in the sample decoding network, or any network layer located after the last target network layer in the sample decoding network, and the location of this reference network layer is not limited.
[0173] For illustrative purposes, steps S31 to S33 are similar to steps S11 to S13 and will not be explained again here.
[0174] In step S34, for each target network layer in the sample decoding network, the quantization analysis module obtains a list of candidate weight quantization hyperparameters corresponding to the target network layer, and the list of candidate weight quantization hyperparameters may include multiple candidate weight quantization hyperparameters.
[0175] First, the quantization analysis module determines the maximum weight value based on all floating-point weights of the target network layer, where each floating-point weight may contain multiple weight values, and the maximum weight value may be the maximum value among all weight values of all floating-point weights. For example, the target network layer may contain multiple floating-point weights, each floating-point weight may contain multiple weight values, and for each floating-point weight, the maximum value may be selected from all weight values corresponding to that floating-point weight, and then the maximum weight value may be selected from the maximum values corresponding to all floating-point weights.
[0176] Next, the quantization analysis module generates initial weight quantization hyperparameters corresponding to the target network layer based on the maximum weight value. For example, param = (max) / 2 bw-1 The initial weight quantization hyperparameters corresponding to the target network layer are generated using the following: param represents the initial weight quantization hyperparameters corresponding to the target network layer, bw represents the fixed quantization bit width, and max represents the maximum weight value. The initial weight quantization hyperparameters corresponding to the target network layer are param, where param=get_param(bw,max), where bw represents the quantization bit width corresponding to the target network layer, and the quantization bit width corresponding to the target network layer may be a fixed quantization bit width, for example, 4 bits, 8 bits, 16 bits, etc. param=get_param(bw,max) represents the initial weight quantization hyperparameters calculated based on the data range max under the quantization bit width bw. Since the calculation method of the initial weight quantization hyperparameters differs depending on the quantization algorithm, an example of the initial weight quantization hyperparameters is shown here.
[0177] Next, the quantization analysis module constructs a list of candidate weight quantization hyperparameters based on the initial weight quantization hyperparameters, and this list of candidate weight quantization hyperparameters may include multiple candidate weight quantization hyperparameters, i.e., multiple candidate weight quantization hyperparameters corresponding to the target network layer. For example, considering that outliers (anomalous weight values) usually exist in the weight values of a neural network, a list of candidate weight quantization hyperparameters can be constructed based on the initial weight quantization hyperparameters, and for example, an example of such a list of candidate weight quantization hyperparameters is:
number
[0178] The above is merely an example of a list of candidate weight quantization hyperparameters, and does not limit the list to candidate weight quantization hyperparameters. As an example, if the target network layer is network layer l, then param l The initial weight quantization hyperparameters corresponding to network layer l are shown, and the list of candidate weight quantization hyperparameters shows seven candidate weight quantization hyperparameters corresponding to network layer l.
[0179] In step S35, for each candidate weight quantization hyperparameter, the quantization analysis module simulates quantizing the floating-point weights of the target network layer using a fixed quantization bit width and the candidate weight quantization hyperparameter to obtain candidate simulated weights.
[0180] For example, the quantization bit width corresponding to the target network layer may be a fixed quantization bit width, and the fixed quantization bit width may be set empirically, but is not limited thereto. Each candidate weight quantization hyperparameter in the list of candidate weight quantization hyperparameters may be traversed sequentially, and for the currently traversed candidate weight quantization hyperparameter, the floating-point weights of the target network layer may be simulated quantized using the fixed quantization bit width and the candidate weight quantization hyperparameter to obtain the candidate simulation weights corresponding to the candidate weight quantization hyperparameter. For example, the floating-point weights of the target network layer may be simulated quantized using the following formula to obtain the candidate simulation weights. quant=param×clip(round(data÷param),-2 bw-1 ,2 bw-1 -1)
[0181] quant represents the candidate simulation weight, param represents the candidate weight quantization hyperparameter, data represents the floating-point weight of the target network layer, bw represents the fixed quantization bit width, clip represents the clipping function, and round represents the rounding operation. Since different quantization algorithms use different simulation quantization methods, the above formula is merely one example of simulation quantization.
[0182] In summary, for each candidate weight quantization hyperparameter, substituting the candidate weight quantization hyperparameter param into the above formula yields the candidate simulation weight quant corresponding to that candidate weight quantization hyperparameter.
[0183] In step S36, after obtaining candidate simulation weights, the quantization analysis module processes the floating-point input features of the target network layer based on the candidate simulation weights to obtain the floating-point output features of the target network layer. Based on the floating-point output features of the target network layer, the quantization analysis module determines the quantization error corresponding to the candidate weight quantization hyperparameters.
[0184] In step S37, based on the quantization error corresponding to each candidate weight quantization hyperparameter, the quantization analysis module determines the candidate weight quantization hyperparameter corresponding to the minimum quantization error as the target weight quantization hyperparameter corresponding to the target network layer, and determines the fixed quantization bit width as the target weight quantization bit width corresponding to the target network layer. Up to this point, the quantization analysis module can obtain the target weight quantization hyperparameter and the target weight quantization bit width corresponding to the target network layer.
[0185] For illustrative purposes, steps S36 to S37 are similar to steps S16 to S17 and will not be repeated here.
[0186] Step S38, the quantization analysis module simulates quantizing the floating-point weights of the target network layer using the target weight quantization hyperparameters and the target weight quantization bit width of the target network layer to obtain pseudoweights of the target network layer.
[0187] For example, after obtaining the target weight quantization hyperparameters and target weight quantization bit width, the floating-point weights of the target network layer may be simulated and quantized using the target weight quantization hyperparameters and target weight quantization bit width to obtain pseudoweights for the target network layer. For the floating-point input features of the target network layer, the target network layer processes the floating-point input features of the target network layer using these pseudoweights to obtain floating-point output features of the target network layer, and these floating-point output features are used as floating-point input features for subsequent network layers. For example, pseudoweights may be obtained by simulating and quantizing floating-point weights using the following formula. quant=param×clip(round(data÷param),-2 bw-1 ,2 bw-1 -1)
[0188] Here, quant represents the pseudoweights of the target network layer, param represents the target weight quantization hyperparameters of the target network layer, data represents the floating-point weights of the target network layer, bw represents the target weight quantization bit width of the target network layer, clip represents the clipping function, and round represents the rounding operation. Since different quantization algorithms use different simulation quantization methods, the above equation is merely one example of simulation quantization.
[0189] For example, for each target network layer, after obtaining the target weight quantization hyperparameter and target weight quantization bit width of the target network layer, the target feature value quantization bit width and target feature value quantization hyperparameter may be obtained. For example, the target feature value quantization bit width and target feature value quantization hyperparameter may be obtained using the following steps.
[0190] In step S39, for each target network layer in the sample decoding network, the quantization analysis module obtains a list of candidate feature value quantization hyperparameters and a set of quantization bit widths corresponding to the target network layer, wherein the list of candidate feature value quantization hyperparameters may include multiple candidate feature value quantization hyperparameters, and the set of quantization bit widths may include at least two quantization bit widths.
[0191] First, the quantization analysis module may determine the maximum feature value based on all floating-point input features of the target network layer, where each floating-point input feature contains multiple feature values, and the maximum feature value is the maximum value among all feature values of all floating-point input features. For example, the target network layer may support multiple floating-point input features, each floating-point input feature containing multiple feature values, and for each floating-point input feature, the maximum value may be selected from all feature values corresponding to that floating-point input feature, and then the maximum feature value may be selected from the maximum values corresponding to all floating-point input features.
[0192] Next, the quantization analysis module generates initial feature value quantization hyperparameters corresponding to the target network layer based on the maximum feature value. For example, param = (max) / 2 bw-1The initial feature value quantization hyperparameters corresponding to the target network layer are generated using the following method: param represents the initial feature value quantization hyperparameter, bw represents the fixed quantization bit width, and max represents the maximum feature value. The initial feature value quantization hyperparameter is param, where param=get_param(bw,max), where bw represents the quantization bit width corresponding to the target network layer, and the quantization bit width corresponding to the target network layer is the current quantization bit width (i.e., the quantization bit width currently traversed from the quantization bit width set). param=get_param(bw,max) represents the initial feature value quantization hyperparameter calculated based on the data range max under the quantization bit width bw. Since the calculation method of the initial feature value quantization hyperparameter differs depending on the quantization algorithm, only an example of the initial feature value quantization hyperparameter is shown here, and it is not limited to this example.
[0193] Next, the quantization analysis module may construct a list of candidate feature value quantization hyperparameters based on the initial feature value quantization hyperparameters, and this list of candidate feature value quantization hyperparameters may include multiple candidate feature value quantization hyperparameters, i.e., multiple candidate feature value quantization hyperparameters corresponding to the target network layer. For example, considering that in the feature values of a neural network, outliers (anomalous feature values, whose value range is usually particularly large) usually exist, and that if these outliers are too large, it affects the quantization accuracy, the quantization analysis module can construct a list of candidate feature value quantization hyperparameters based on the initial feature value quantization hyperparameters, for example, an example of such a list of candidate feature value quantization hyperparameters is:
number
[0194] The above is merely an example of a list of candidate feature value quantization hyperparameters, and is not limited to this list.
[0195] The quantization analysis module may generate a set of quantization bit widths, which may include at least two quantization bit widths.
[0196] In step S40, the quantization analysis module traverses the quantization bit widths in the set of quantization bit widths, based on increasing order of quantization bit widths. For each candidate feature value quantization hyperparameter, the quantization analysis module simulates quantization of the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter to obtain candidate input features.
[0197] For example, based on the order of increasing quantization bit widths, the quantization analysis module may first set the first quantization bit width as the current quantization bit width. It may then sequentially traverse each candidate feature value quantization hyperparameter in the feature value quantization hyperparameter candidate list, and for the currently traversed candidate feature value quantization hyperparameter, simulate quantization of the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter to obtain the candidate input features corresponding to the candidate feature value quantization hyperparameter. For example, quant = param × clip(round(data ÷ param), -2 bw-1 ,2 bw-1 The candidate input features may be obtained by simulating quantization of the floating-point input features of the target network layer using (-1).
[0198] quant represents the candidate input feature, param represents the candidate feature value quantization hyperparameter, data represents the floating-point input feature, bw represents the current quantization bit width, clip represents the clipping function, and round represents the rounding operation. Since different quantization algorithms use different simulation quantization methods, the above formula is merely one example of simulation quantization.
[0199] In summary, for each candidate feature value quantization hyperparameter, substituting the candidate feature value quantization hyperparameter param into the above formula yields the candidate input feature quant corresponding to that candidate feature value quantization hyperparameter.
[0200] In step S41, after obtaining candidate input features, the quantization analysis module processes the candidate input features of the target network layer based on the pseudoweights of the target network layer to obtain floating-point output features of the target network layer. Based on the floating-point output features of the target network layer, the quantization analysis module determines the quantization error corresponding to the candidate feature value quantization hyperparameter.
[0201] In step S42, based on the quantization error corresponding to each candidate feature value quantization hyperparameter, the quantization analysis module determines whether the minimum quantization error is smaller than a preset threshold (which may be set based on experience). If it is smaller than the threshold, step S43 is performed; otherwise, step S44 is performed.
[0202] In step S43, the quantization analysis module determines a candidate feature value quantization hyperparameter corresponding to the minimum quantization error as the target feature value quantization hyperparameter corresponding to the target network layer, determines the current quantization bit width as the target feature value quantization bit width corresponding to the target network layer, and records the target feature value quantization hyperparameter and target feature value quantization bit width of the target network layer.
[0203] Step S44, the quantization analysis module determines whether the current quantization bit width is the last quantization bit width in the quantization bit width set.
[0204] If it is the last quantization bit width, the candidate feature value quantization hyperparameter corresponding to the minimum quantization error is determined as the target feature value quantization hyperparameter corresponding to the target network layer, the current quantization bit width is determined as the target feature value quantization bit width corresponding to the target network layer, and the target feature value quantization hyperparameter and target feature value quantization bit width of the target network layer are recorded. This process can be referred to in step S43.
[0205] If it is not the last quantization bit width, the next quantization bit width after the current quantization bit width is set as the new current quantization bit width, and the process returns to step S40. Based on the current quantization bit width, for each candidate feature value quantization hyperparameter, the quantization analysis module simulates quantizing the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter to obtain candidate input features.
[0206] In step S45, after obtaining the target weight quantization hyperparameters, target weight quantization bit width, target feature value quantization hyperparameters, and target feature value quantization bit width corresponding to each target network layer, the decoding network quantization module generates a target decoding network based on the sample decoding network. For example, for each target network layer in the sample decoding network, the decoding network quantization module obtains integer weights for the target network layer by quantizing the floating-point weights of the target network layer into fixed-point weights based on the target weight quantization bit width and target weight quantization hyperparameters corresponding to the target network layer. After obtaining the integer weights for each target network layer, the decoding network quantization module can generate a target decoding network based on the integer weights of each target network layer.
[0207] For illustrative purposes, steps S41 to S45 are similar to steps S21 to S25 and will not be explained again here.
[0208] After obtaining the integer weights for each target network layer, the floating-point weights of the target network layer can be replaced with the integer weights, and in this way, a target decoding network can be constructed with all the replaced target network layers. At this point, the target decoding network can be obtained and introduced into either the encoding or decoding side. For example, the target decoding network may include a probabilistic hyperparameter decoding network and a composite transformation network, or the target decoding network may include a mean prediction network, a probabilistic hyperparameter decoding network and a composite transformation network. By introducing these target decoding networks into the encoding side, when the encoding side performs processing using Example 2, the processing will be performed based on the target decoding network. Also, by introducing these target decoding networks into the decoding side, when the decoding side performs processing using Example 3, the processing will be performed based on the target decoding network.
[0209] Example 8: In a neural network-based decoder, when decoding a bitstream, the sensitivity to quantization differs among the network layers of the neural network. That is, some network layers can maintain quantization performance even with low bit counts, while some network layers require high bit counts to guarantee quantization accuracy. This relates to the parameters and feature distribution learned by the neural network. Therefore, when using more user-friendly PTQ quantization, in order to achieve maximum compression of the model's bit width, guarantee the performance of the compressed model, and ensure that the compressed model is acceptable on most devices, this embodiment provides an efficient decoder compression and inference method that mainly includes the following steps: 1. Collect images from a real-world scene (application scene for an image acquisition device) of several frames (e.g., 100-200 frames). 2. Extract and store the bitstreams (i.e., the first bitstream and the second bitstream) that will be input to the encoding network and transmitted to the decoding network. 3. Transmit the stored bitstreams to a quantization analysis module and obtain appropriate quantization bit widths and quantization parameters (e.g., quantization hyperparameters) for each network layer of the decoding network. 4. Based on the quantization bit width and quantization parameters of each network layer obtained by the quantization analysis module, each network layer of the decoding network is quantized to obtain the final quantization model, that is, each network layer of the sample decoding network is quantized to obtain the target decoding network.
[0210] Based on the decoding network shown in Figure 5A or Figure 5B, a bitstream extraction module, a quantization analysis module, and a decoding network quantization module can be added. The functions of the bitstream extraction module, quantization analysis module, and decoding network quantization module will be described below.
[0211] Bitstream Extraction Module: This module collects images from a real-world scene (an application scene for an image acquisition device) for several frames (e.g., 100-200 frames), inputs the images into an encoding network for forward computation, extracts the output of the encoding network (i.e., bitstreams to be transmitted to the decoding network, i.e., the first bitstream and the second bitstream), and stores the bitstream corresponding to the image of each frame.
[0212] Quantization Analysis Module: The bitstream information extracted by the bitstream extraction module is entered into the quantization analysis module. The role of the quantization analysis module is to analyze the sensitivity of each network layer of the decoding network to quantization, guarantee the decoding performance, and then obtain the quantization bit width to be allocated to each network layer of the decoding network, that is, to obtain the quantization bit width for each network layer of the decoding network.
[0213] The quantization analysis module stores the forward calculation results of the decoding network. For example, after the bitstream enters the decoding network, it first performs a forward decoding of all floating-point data types once, and the input features (f) of each network layer of the decoding network in floating-point are calculated. i_l (where i represents the i-th bitstream and l represents the l-th network layer), and the output features of the last layer (f i_o Here, i represents the i-th bitstream, and the output feature of the last layer is the reference output feature of the above embodiment.
[0214] The quantization analysis module further supports performing simulation quantization operations on the weight values and input feature values of each network layer of the decoding network, and can perform simulation quantization calculations, and is used to simulate the quantization inference of the decoding network. Here, in the decoding network, the weight values of network layers (such as convolutional layers) are 8-bit and not sensitive to quantization, while the input feature values of network layers (such as convolutional layers) are more sensitive to quantization. Therefore, performance loss can be satisfied by quantizing some network layers using 8-bit, but it may be difficult to satisfy the quantization performance loss in some network layers. Thus, the quantization analysis module mainly analyzes the quantization of the input feature values of the network layers (such as convolutional layers) of the decoding network, and selects appropriate quantization bit widths and quantization hyperparameters for the inputs of each network layer of the decoding network.
[0215] The quantization analysis module initializes the quantization hyperparameters for the input feature values of each network layer of the decoding network. The initialization of the quantization hyperparameters depends on the data range of the floating-point features. The data range is obtained from the statistics of the input features (where f is the input feature of the l-th network layer, i represents the i-th bitstream, and l represents the l-th network layer), that is, max i_l = max(abs(f l ), abs(f 0_l ), …, abs(f 1_l ), …, abs(f n_l )). The initialized quantization hyperparameters are obtained by the formula param l = get_param(bw l , max l ), where bw l`param` represents the bit width of the input feature of the l-th layer, initialized to 8, and represents the quantization hyperparameter calculated based on the data range max under the bw quantization bit width. Since different quantization algorithms use different methods to calculate quantization hyperparameters, here we show one example: param = bw - 1 - ceil(log2(max)). Considering that in neural network feature values, there are usually outliers (outliers with particularly large value ranges), and that these outliers can affect quantization accuracy if they are too large, a list of candidate hyperparameters, i.e., param, is created based on the initialized quantization hyperparameter. l_list =[param l -2,param l -1,param l ,param l +1,param l Set [+2]. Another example of calculating the quantization hyperparameter is param=(max) / 2 bw-1 This may also be the case, and the list of candidate hyperparameters set based on the initialization quantization hyperparameters is,
number
[0216] The quantization analysis module analyzes the quantization of the decoding network layer by layer. The analysis method involves performing simulated quantization (for example, 8-bit quantization initially) on the input feature values for each layer, inputting the simulated quantized features into the network for forward calculation, saving the output of the decoding network, and storing the previously saved floating-point output features f i_o An error analysis is performed, and if the error is greater than a certain threshold, the quantization bit width of that layer is set to 16 bits. The calculation formula for simulation quantization is data q = quant(data f It may also be ,bw,param), and since the calculation formula for simulation quantization differs depending on the quantization algorithm, here is an example of a simulation quantization method quant(data f ,bw,param)=2-param ×clip(round(data f ×2 param ),-2 bw-1 ,2 bw-1 -1) is shown, and clip represents the clipping function. Also, regarding the simulation quantization process, another example of a simulation quantization method is quant(data f ,bw,param)=param×clip(round(data f ÷param),-2 bw-1 ,2 bw-1 -1) is also acceptable.
[0217] In one possible embodiment, the quantization hyperparameters and quantization bit width of the feature values may be calculated using the following steps.
[0218] Initialize the quantization bit width, for example, by setting the initial value of the quantization bit width to 8 bits.
[0219] The quantization hyperparameters are initialized, for example, by determining their initial values using param=bw-1-ceil(log2(max)). The initial values of the quantization hyperparameters are the initial feature value quantization hyperparameters of the above example.
[0220] The quantization error is initialized, for example, by setting a predetermined threshold for the quantization error. This predetermined threshold may be set based on experience.
[0221] Based on the initial values of the quantization hyperparameters, param l_list =[param l -2,param l -1,param l ,param l +1,param l Obtain a list of candidate quantization hyperparameters, such as by using [+2] to obtain the list of candidate quantization hyperparameters.
[0222] Based on the quantization bit width and quantization hyperparameters, quant(data f , bw, param) = 2 -param × clip(round(data f × 2 param ), -2 bw-1 , 2 bw-1 - 1) is used to perform simulation quantization, etc., and the input feature values are simulated and quantized.
[0223] The feature values after simulation quantization are input into the decoding network for forward calculation.
[0224] The quantization error between the output of the decoding network after simulation quantization and the output of the floating - point decoding network is calculated.
[0225] When the quantization error is smaller than a predetermined threshold, the quantization bit width of the output feature values is 8 bits, and the quantization hyperparameters of the output feature values are the quantization hyperparameters corresponding to the minimum quantization error (that is, a certain quantization hyperparameter in the quantization hyperparameter candidate list).
[0226] When the quantization error is greater than or equal to the predetermined threshold, the quantization bit width is set to 16 bits.
[0227] The initial value of the quantization hyperparameters is reacquired. That is, when obtaining the initial value of the quantization hyperparameters using the above formula, the quantization bit width bw is changed from 8 bits to 16 bits, and the initial value of the quantization hyperparameters is obtained on the premise of 16 bits.
[0228] The quantization hyperparameter candidate list is reacquired, such as obtaining the quantization hyperparameter candidate list based on the initial value of the quantization hyperparameters.
[0229] The input feature values are simulated and quantized, such as performing simulation quantization using the above formula.
[0230] The feature values after simulation quantization are input into the decoding network, and a forward calculation is performed.
[0231] The quantization error between the decoded network output after simulation quantization and the floating-point decoded network output is calculated.
[0232] The quantization bit width of the output feature value is 16 bits, and the quantization hyperparameter of the output feature value is the quantization hyperparameter corresponding to the minimum quantization error.
[0233] In the above process, when calculating the quantization error, an error loss function can be used to calculate the quantization error. The calculation method for the error loss function may include, but is not limited to, MSE, cosine similarity, KL divergence, etc. Through the calculation in the above process, the quantization bit width bw of the input to each network layer of the decoding network is obtained. l and quantization hyperparameter param l You can obtain this.
[0234] Decoding Network Quantization Module: The quantization bit width bw of the input to each network layer output from the quantization analysis module. l and quantization hyperparameter param l Based on the fixed 8-bit quantization bit width of the weight values of each network layer, the decoding network quantization module can quantize the decoding network. For example, by performing 8-bit quantization on the weight parameters of the decoding network, the memory space of the model can be significantly compressed. At the same time, the decoding network quantization module marks the input of each network layer with bit width information (8-bit or 16-bit) and writes the corresponding quantization hyperparameters to the computation process of the network layer (usually expressed as a shift), thereby further compressing the computational complexity and throughput bandwidth of the model while guaranteeing the decoding quality of the decoding network.
[0235] Example 9: After obtaining a target decoding network, the target decoding network can be introduced to either the encoding side or the decoding side. For example, the target decoding network includes a probabilistic hyperparameter decoding network and a composite transformation network, or the target decoding network includes a mean prediction network, a probabilistic hyperparameter decoding network and a composite transformation network. By introducing these target decoding networks to the encoding side, when the encoding side performs processing using Example 2, the processing will be performed based on the target decoding network. By introducing these target decoding networks to the decoding side, when the decoding side performs processing using Example 3, the processing will be performed based on the target decoding network.
[0236] In one possible embodiment, the encoding or decoding side may decode a target feature corresponding to the current image block from the bitstreams corresponding to the current image block (e.g., a first bitstream and a second bitstream), determine the input features of the decoding network based on the target features, process the input features based on the decoding network, obtain the output features of the decoding network, and then determine the reconstructed image block corresponding to the current image block based on the output features, and this process will be described below.
[0237] Exemplary, the encoding or decoding side decodes the target feature corresponding to the current image block from the bitstream corresponding to the current image block and determines the first input feature of the target decoding network based on the target feature. The target feature value quantization bit width and target feature value quantization hyperparameters of the target decoding network are obtained (for example, obtained from the target decoding network configuration information), and the first input feature is transformed into a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameters. The second input feature is processed based on the integer weights of the target decoding network (the integer weights may be determined based on the target weight quantization bit width and target weight quantization hyperparameters) to obtain the output feature of the target decoding network. Based on the output feature of the target decoding network, the reconstructed image block corresponding to the current image block is determined.
[0238] Exemplary, a target decoding network may include at least one target network layer, the target network layer being a network layer employing integer weights, and for each target network layer in the target decoding network, the first input feature of the target network layer may be converted to a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter of the target network layer, the second input feature may be processed based on the integer weights of the target network layer (the integer weights of the target network layer are determined based on the target weight quantization bit width and target weight quantization hyperparameter of the target network layer), and the output feature of the target network layer may be obtained. Different target network layers may have the same or different target feature value quantization bit widths, different target network layers may have the same or different target feature value quantization hyperparameters, different target network layers may have fixed quantization bit widths, and different target network layers may have the same or different target weight quantization hyperparameters.
[0239] For example, the step of transforming the first input feature of the target network layer into a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter of the target network layer is given by quant=clip(round(data×2 param ),-2 bw-1 ,2 bw-1 -1), or quant=clip(round(data÷param),-2 bw-1 ,2 bw-1 -1) The process includes, but is not limited to, the step of converting the first input feature of the target network layer to the second input feature, where quant represents the second input feature, data represents the first input feature, param represents the target feature value quantization hyperparameter, bw represents the target feature value quantization bit width, clip represents the clipping function, and round represents rounding.
[0240] As shown in Figure 5A, the encoding or decoding side obtains Bitstream#1 (i.e., the first bitstream) corresponding to the current image block, then decodes Bitstream#1 to obtain the hyperparameter quantization feature (i.e., the target feature of the current image block), and dequantizes the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat (i.e., the first input feature of the probabilistic hyperparameter decoding network is determined based on the target feature). Alternatively, Bitstream#1 may be decoded to obtain the coefficient hyperparameter feature z_hat (i.e., the target feature of the current image block, which may be directly used as the first input feature of the probabilistic hyperparameter decoding network).
[0241] After obtaining the coefficient hyperparameter feature z_hat (i.e., the first input feature of the probabilistic hyperparameter decoding network), the coefficient hyperparameter feature z_hat is input to the probabilistic hyperparameter decoding network (i.e., the target decoding network). The probabilistic hyperparameter decoding network may convert the floating-point type coefficient hyperparameter feature z_hat (i.e., the first input feature) to an integer type coefficient hyperparameter feature z_hat (i.e., the second input feature) using the target feature value quantization bit width and the target feature value quantization hyperparameter, and then perform an inverse coefficient hyperparameter feature transformation on the integer type coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p (i.e., the output feature of the probabilistic hyperparameter decoding network) corresponding to the current image block. After obtaining the probability distribution parameter p, a probability distribution model may be generated based on the probability distribution parameter p.
[0242] For example, for each target network layer in a probabilistic hyperparameter decoding network, after obtaining a floating-point first input feature, the target network layer can convert the floating-point first input feature to an integer second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter of the target network layer. Furthermore, since the target network layer employs integer weights, the integer second input feature can be processed based on the integer weights of the target network layer to obtain the output feature of the target network layer. Here, the output feature of the target network layer can be the input feature of the next network layer, and the output feature of the target network layer may also be a floating-point feature, and so on.
[0243] As shown in Figure 5A, the encoding or decoding side obtains Bitstream#2 (i.e., the second bitstream) corresponding to the current image block, decodes Bitstream#2 to obtain the image quantization feature (i.e., the target feature of the current image block), dequantizes the image quantization feature to obtain the image feature s', and performs feature reconstruction on the image feature s' to obtain the image feature y_hat (i.e., determines the first input feature of the composite transformation network based on the target feature). Alternatively, Bitstream#2 may be decoded to obtain the image feature s' (i.e., the target feature of the current image block), and feature reconstruction on the image feature s' to obtain the image feature y_hat (i.e., determines the first input feature of the composite transformation network based on the target feature). When decoding Bitstream#2, the encoding or decoding side may decode Bitstream#2 based on a probability distribution model, and this process will not be explained.
[0244] After obtaining the image feature y_hat (i.e., the first input feature of the composite transformation network), the image feature y_hat is input to the composite transformation network (i.e., the target decoding network). The composite transformation network may convert the floating-point image feature y_hat (i.e., the first input feature) to an integer type image feature y_hat (i.e., the second input feature) using the target feature value quantization bit width and target feature value quantization hyperparameters, and then perform a composite transformation on the integer type image feature y_hat to obtain the reconstructed image block x_hat (i.e., the output feature of the composite transformation network) corresponding to the current image block x. As can be seen from the above, a composite transformation can be performed on the image feature y_hat based on the composite transformation network to obtain the reconstructed image block x_hat corresponding to the current image block x.
[0245] For example, for each target network layer in a composite transformation network, after obtaining a floating-point first input feature, the target network layer can convert the floating-point first input feature to an integer second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter of the target network layer. Furthermore, since the target network layer employs integer weights, it can process the integer second input feature based on the integer weights of the target network layer to obtain the output feature of the target network layer. Here, the output feature of the target network layer can be the input feature of the next network layer, and the output feature of the target network layer may also be a floating-point feature, and so on.
[0246] After obtaining the output features of the composite transformation network, the output features of the composite transformation network may be used as the reconstructed image block x_hat corresponding to the current image block x; that is, the reconstructed image block x_hat corresponding to the current image block x is determined based on these output features.
[0247] As shown in Figure 5B, the encoding or decoding side may, after obtaining Bitstream #1 corresponding to the current image block, decode Bitstream #1 to obtain a hyperparameter quantization feature (i.e., the target feature of the current image block), and then dequantize the hyperparameter quantization feature to obtain a coefficient hyperparameter feature z_hat (i.e., the first input feature of the probabilistic hyperparameter decoding network is determined based on this target feature). Alternatively, Bitstream #1 may be decoded to obtain a coefficient hyperparameter feature z_hat (i.e., the target feature of the current image block, which is directly used as the first input feature of the probabilistic hyperparameter decoding network).
[0248] After obtaining the coefficient hyperparameter feature z_hat (i.e., the first input feature of the probabilistic hyperparameter decoding network), the coefficient hyperparameter feature z_hat is input to the probabilistic hyperparameter decoding network (i.e., the target decoding network). The probabilistic hyperparameter decoding network may convert the floating-point type coefficient hyperparameter feature z_hat (i.e., the first input feature) to an integer type coefficient hyperparameter feature z_hat (i.e., the second input feature) using the target feature value quantization bit width and the target feature value quantization hyperparameter, and then perform an inverse coefficient hyperparameter feature transformation on the integer type coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p (i.e., the output feature of the probabilistic hyperparameter decoding network) corresponding to the current image block.
[0249] For example, for each target network layer in a probabilistic hyperparameter decoding network, the target network layer obtains a floating-point first input feature, and then converts the floating-point first input feature to an integer second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter of the target network layer. The target network layer employs integer weights and processes the integer second input feature based on the integer weights of the target network layer to obtain the output feature of the target network layer.
[0250] After obtaining the coefficient hyperparameter feature z_hat (i.e., the first input feature of the mean prediction network), the coefficient hyperparameter feature z_hat and the image feature y_hat of the previous image block are input into the mean prediction network (i.e., the target decoding network). The mean prediction network uses the target feature value quantization bit width and the target feature value quantization hyperparameter to convert the floating-point coefficient hyperparameter feature z_hat (the first input feature) into an integer coefficient hyperparameter feature z_hat (the second input feature), and uses the target feature value quantization bit width and the target feature value quantization hyperparameter to convert the floating-point image feature y_hat (the first input feature) into an integer image feature y_hat (the second input feature). The mean prediction network makes a context-based prediction based on the integer coefficient hyperparameter feature z_hat and the integer image feature y_hat, and obtains the predicted value mu corresponding to the current image block (i.e., the output feature of the mean prediction network).
[0251] For example, for each target network layer in the mean prediction network, after obtaining the floating-point first input feature, the target network layer converts the floating-point first input feature into an integer second input feature based on the target feature value quantization bit width and the target feature value quantization hyperparameter of this target network layer. The target network layer adopts integer weights, processes the integer second input feature based on the integer weights of the target network layer, and obtains the output feature of the target network layer.
[0252] The encoding or decoding side obtains Bitstream#2 corresponding to the current image block, decodes Bitstream#2 to obtain image quantization features (i.e., target features of the current image block), dequantizes the image quantization features to obtain image features s', performs feature reconstruction (i.e., residual reconstruction) on image features s' to obtain residual features r_hat, determines image features y_hat based on residual features r_hat and predicted values mu, for example, setting the sum of residual features r_hat and predicted values mu as image features y_hat (i.e., determining the first input feature of the composite transformation network based on the target feature). Alternatively, Bitstream#2 is decoded to obtain image features s' (i.e., the target features of the current image block), feature reconstruction is performed on image features s' to obtain residual features r_hat, and image features y_hat are determined based on residual features r_hat and predicted values mu. For example, the sum of residual features r_hat and predicted values mu is taken as image features y_hat (i.e., the first input feature of the composite transformation network is determined based on this target feature).
[0253] After obtaining the image feature y_hat (i.e., the first input feature of the composite transformation network), the image feature y_hat is input to the composite transformation network (i.e., the target decoding network). The composite transformation network may convert the floating-point image feature y_hat (i.e., the first input feature) to an integer type image feature y_hat (i.e., the second input feature) using the target feature value quantization bit width and target feature value quantization hyperparameters, and then perform a composite transformation on the integer type image feature y_hat to obtain the reconstructed image block x_hat (i.e., the output feature of the composite transformation network) corresponding to the current image block x. As can be seen from the above, a composite transformation can be performed on the image feature y_hat based on the composite transformation network to obtain the reconstructed image block x_hat corresponding to the current image block x.
[0254] For example, for each target network layer in a composite transformation network, after obtaining a floating-point first input feature, the target network layer can convert the floating-point first input feature to an integer second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter of the target network layer. Furthermore, since the target network layer employs integer weights, it can process the integer second input feature based on the integer weights of the target network layer to obtain the output feature of the target network layer. Here, the output feature of the target network layer can be the input feature of the next network layer, and the output feature of the target network layer may also be a floating-point feature, and so on.
[0255] After obtaining the output features of the composite transformation network, the output features of the composite transformation network may be used as the reconstructed image block x_hat corresponding to the current image block x; that is, the reconstructed image block x_hat corresponding to the current image block x is determined based on these output features.
[0256] As can be seen from the above technical proposals, embodiments of the present invention provide an end-to-end video image compression method that can perform video image decoding based on a decoding network, thereby achieving the objective of improving coding efficiency and decoding efficiency. A decoding network of fixed-point weights is constructed using the target weight quantization bit width and target weight quantization hyperparameters, and fixed-point input features are generated using the target feature value quantization bit width and target feature value quantization hyperparameters, thereby achieving adaptive decoding acceleration and reducing decoding computation while guaranteeing decoding quality. A compression method for decoding networks in a neural network-based image coding and decoding framework is provided, and the quantization bit width and quantization hyperparameters of each layer are determined by analyzing the quantization loss of each layer when the decoding network performs decoding calculations of the input bitstream, and the decoding network is quantized based on the obtained quantization bit width and quantization hyperparameters, thereby guaranteeing the decoding quality of the decoding network and significantly compressing the memory space, computational complexity, model throughput and bandwidth of the decoding network, thereby enabling massive decoding networks to operate on devices with limited computing resources. Changing the decoding network's inference from floating-point to fixed-point calculations contributes to the decoding consistency of the decoding network across different devices.
[0257] Exemplary examples, each of the above embodiments may be implemented individually or in combination. For example, each of the embodiments from Embodiments 1 to 9 may be implemented individually, and at least two of the embodiments from Embodiments 1 to 9 may be implemented in combination.
[0258] For example, in each of the above embodiments, the contents of the encoding side may be applied to the decoding side, that is, processed by the decoding side in the same manner, and the contents of the decoding side may be applied to the encoding side, that is, processed by the encoding side in the same manner.
[0259] Based on the same concept as described above, embodiments of the present invention further provide a decoding device which is applied to the decoding side and includes a memory configured to store video data and a decoder configured to carry out the decoding method in embodiments 1 to 9 above, i.e., the decoding side processing process.
[0260] For example, in one possible embodiment, the decoder is The steps include decoding a target feature corresponding to the current image block from a bitstream corresponding to the current image block, The steps include determining the first input feature of the target decoding network based on the aforementioned target features, The steps include obtaining the target feature value quantization bit width and target feature value quantization hyperparameter of the target decoding network, and converting the first input feature to a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter, A step of processing the second input features based on integer weights of the target decoding network to obtain output features of the target decoding network, wherein the integer weights are determined based on the target weight quantization bit width and the target weight quantization hyperparameters. The system is configured to perform the steps of determining a reconstructed image block corresponding to the current image block based on the output characteristics of the target decoding network.
[0261] Based on the same concept as described above, a schematic diagram of the hardware architecture of a decoding device (also called a video decoder) provided by an embodiment of the present invention may be shown in detail in Figure 7. The decoding device includes a processor 701 and a machine-readable storage medium 702, the machine-readable storage medium 702 stores machine-executable instructions that can be executed by the processor 701, and the processor 701 is used to execute the machine-executable instructions and carry out the decoding methods of embodiments 1 to 9 of the present invention.
[0262] For example, in one possible embodiment, the decoding device is The steps include decoding a target feature corresponding to the current image block from a bitstream corresponding to the current image block, The steps include determining the first input feature of the target decoding network based on the aforementioned target features, The steps include obtaining the target feature value quantization bit width and target feature value quantization hyperparameter of the target decoding network, and converting the first input feature to a second input feature based on the target feature value quantization bit width and target feature value quantization hyperparameter, A step of processing the second input features based on integer weights of the target decoding network to obtain output features of the target decoding network, wherein the integer weights are determined based on the target weight quantization bit width and the target weight quantization hyperparameters. This is used to carry out the step of determining the reconstructed image block corresponding to the current image block based on the output characteristics of the target decoding network.
[0263] Based on the same concept as described above, embodiments of the present invention provide an electronic device. The electronic device includes a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions and carry out the decoding methods of embodiments 1 to 9 of the present invention.
[0264] Based on the same concept as described above, embodiments of the present invention further provide a machine-readable storage medium in which several computer instructions are stored, and when the computer instructions are executed by a processor, the decoding methods disclosed in the above embodiments of the present invention, such as the decoding methods in each of the above embodiments, can be performed.
[0265] Based on the same concept as described above, embodiments of the present invention further provide a computer application, and when the computer application is executed by a processor, the decoding method disclosed in the above examples of the present invention can be performed.
[0266] Embodiments of the present invention further provide a decoding device applicable to the decoding side, the decoding device including: a decoding module for decoding a target feature corresponding to the current image block from a bitstream corresponding to the current image block; a determination module for determining a first input feature of a target decoding network based on the target feature; and a processing module for obtaining the target feature value quantization bit width and target feature value quantization hyperparameter of the target decoding network, converting the first input feature of the target decoding network to a second input feature of the target decoding network based on the target feature value quantization bit width and target feature value quantization hyperparameter, processing the second input feature of the target decoding network based on integer weights of the target decoding network, and obtaining output features of the target decoding network, wherein the integer weights of the target decoding network are determined based on the target weight quantization bit width and target weight quantization hyperparameter of the target decoding network, the determination module is used to determine a reconstructed image block corresponding to the current image block based on the output features of the target decoding network.
[0267] Exemplary, the target decoding network includes at least one target network layer, the at least one target network layer employing integer weights, and for each target network layer in the target decoding network, the processing module is further used to convert a first input feature of the target network layer to a second input feature of the target network layer based on the target feature value quantization bit width and target feature value quantization hyperparameter of the target network layer, to process the second input feature of the target network layer based on the integer weights of the target network layer, and to obtain an output feature of the target network layer, wherein the integer weights of the target network layer are determined based on the target weight quantization bit width and target weight quantization hyperparameter of the target network layer, the target feature value quantization bit widths of different target network layers may be the same or different, the target feature value quantization hyperparameters of different target network layers may be the same or different, the target weight quantization bit widths of different target network layers may all be fixed quantization bit widths, and the target weight quantization hyperparameters of different target network layers may be the same or different.
[0268] For example, when the processing module converts the first input feature of the target network layer to the second input feature of the target network layer based on the target feature value quantization bit width and target feature value quantization hyperparameters of the target network layer, specifically, quant = clip(round(data × 2 param ),-2 bw-1 ,2 bw-1 -1), or quant=clip(round(data÷param),-2 bw-1 ,2 bw-1 -1) is used to convert the first input feature to the second input feature, where quant represents the second input feature of the target network layer, data represents the first input feature of the target network layer, param represents the target feature value quantization hyperparameter of the target network layer, bw represents the target feature value quantization bit width of the target network layer, clip represents the clipping function, and round represents rounding.
[0269] Exemplary, the sample decoding network includes a plurality of network layers employing floating-point weights, the plurality of network layers include at least one target network layer of the sample decoding network, and for each target network layer of the at least one target network layer of the sample decoding network, the processing module further obtains a list of weight quantization hyperparameter candidates corresponding to the target network layer, the list of weight quantization hyperparameter candidates includes a plurality of candidate weight quantization hyperparameters, and for each candidate weight quantization hyperparameter, the floating-point weights of the target network layer are simulated quantized using a fixed quantization bit width and the candidate weight quantization hyperparameter to obtain candidate simulation weights, and the floating-point input characteristics of the target network layer are calculated based on the candidate simulation weights. The system processes the characteristics to obtain floating-point output features of the target network layer, determines the quantization error corresponding to the candidate weight quantization hyperparameter based on the floating-point output features of the target network layer, determines the candidate weight quantization hyperparameter corresponding to the minimum quantization error as the target weight quantization hyperparameter of the target network layer based on the quantization error corresponding to each candidate weight quantization hyperparameter, determines the fixed quantization bit width as the target weight quantization bit width of the target network layer, and obtains integer weights of the target network layer by fixed-point quantization of the floating-point weights of the target network layer based on the target weight quantization bit width and the target weight quantization hyperparameter of the target network layer, which are used to generate the target decoding network based on the integer weights of each target network layer of the sample decoding network.
[0270] Exemplary, when the processing module obtains a list of candidate weight quantization hyperparameters corresponding to the target network layer, it specifically determines a maximum weight value based on all floating-point weights of the target network layer, where each floating-point weight includes multiple weight values, and the maximum weight value is the maximum value among all weight values of all floating-point weights. Based on the maximum weight value, it generates initial weight quantization hyperparameters corresponding to the target network layer, and is used to construct the list of candidate weight quantization hyperparameters based on the initial weight quantization hyperparameters.
[0271] For example, when the processing module generates initial weight quantization hyperparameters corresponding to the target network layer based on the maximum weight value, specifically, param=bw-1-ceil(log2(max)) or param=(max) / 2 bw-1 The initial weight quantization hyperparameters are generated using the following: param represents the initial weight quantization hyperparameters corresponding to the target network layer, bw represents the fixed quantization bit width, max represents the maximum weight value, and ceil represents the rounding up operation.
[0272] For example, when the processing module obtains candidate simulation weights by simulating quantizing the floating-point weights of the target network layer using a fixed quantization bit width and the candidate weight quantization hyperparameter, specifically, quant=2 -param ×clip(round(data×2 param ),-2 bw-1 ,2 bw-1 -1), or quant=param×clip(round(data÷param),-2 bw-1 ,2 bw-1-1) is used to obtain candidate simulation weights by simulation quantizing the floating-point weights, where quant represents the candidate simulation weight, param represents the candidate weight quantization hyperparameter, data represents the floating-point weight of the target network layer, bw represents the fixed quantization bit width, clip represents the clipping function, and round represents rounding.
[0273] Exemplary, when the processing module determines the quantization error corresponding to the candidate weight quantization hyperparameter based on the floating-point output features of the target network layer, it specifically determines a sample output feature based on the floating-point output features of the target network layer, the sample output feature being the output feature of a reference network layer in a sample decoding network, the reference network layer being any network layer located after the target network layer, and is used to determine the quantization error corresponding to the candidate weight quantization hyperparameter based on the sample output feature and the reference output feature, the method for obtaining the reference output feature includes decoding a sample feature from a sample bitstream, determining a floating-point input feature corresponding to the sample decoding network based on the sample feature, and determining the reference output feature, which is the output feature of the reference network layer in the sample decoding network, based on the floating-point input feature.
[0274] Exemplary, the sample decoding network includes a plurality of target network layers employing floating-point weights, the plurality of network layers each include at least one target network layer of the sample decoding network, and for each target network layer of the at least one target network layer of the sample decoding network, the processing module further obtains a feature value quantization hyperparameter candidate list and a quantization bit width set corresponding to the target network layer, the feature value quantization hyperparameter candidate list includes a plurality of candidate feature value quantization hyperparameters, the quantization bit width set includes at least two quantization bit widths, traverses the at least two quantization bit widths from the quantization bit width set, and for each candidate feature value quantization hyperparameter, obtains the current quantization bit width and the candidate The floating-point input features of the target network layer are simulated quantized using feature value quantization hyperparameters to obtain candidate input features, these candidate input features are processed based on pseudoweights of the target network layer to obtain floating-point output features of the target network layer, the quantization error corresponding to the candidate feature value quantization hyperparameter is determined based on the floating-point output features of the target network layer, the quantization error corresponding to the candidate feature value quantization hyperparameter is determined based on the quantization error corresponding to each candidate feature value quantization hyperparameter in the current quantization bit width if the minimum quantization error is smaller than a preset threshold, the candidate feature value quantization hyperparameter corresponding to the minimum quantization error is determined as the target feature value quantization hyperparameter of the target network layer, and the current quantization bit width is used to determine the target feature value quantization bit width of the target network layer.The processing module further determines, based on the quantization error corresponding to each candidate feature value quantization hyperparameter in the current quantization bit width, whether the current quantization bit width is the last quantization bit width in the quantization bit width set if the minimum quantization error is greater than or equal to a preset threshold; if the current quantization bit width is the last quantization bit width in the quantization bit width set, it determines the candidate feature value quantization hyperparameter corresponding to the minimum quantization error as the target feature value quantization hyperparameter of the target network layer, and uses the current quantization bit width to determine the target feature value quantization bit width of the target network layer.
[0275] Exemplary, when the processing module obtains a list of candidate feature value quantization hyperparameters corresponding to a target network layer, it specifically determines the maximum feature value based on all floating-point input features of the target network layer, where each floating-point input feature contains multiple feature values, and the maximum feature value is the maximum value among all feature values of all floating-point input features. Based on the maximum feature value, it generates initial feature value quantization hyperparameters corresponding to the target network layer, and is used to construct the list of candidate feature value quantization hyperparameters based on the initial feature value quantization hyperparameters. When the processing module generates initial feature value quantization hyperparameters corresponding to the target network layer based on the maximum feature value, it specifically uses get_param=bw-1-ceil(log2(max)) or get_param=(max) / 2 bw-1 The initial feature value quantization hyperparameters are generated using the following: get_param represents the initial feature value quantization hyperparameters corresponding to the target network layer, bw represents the current quantization bit width, max represents the maximum feature value, and ceil represents the rounding up operation.
[0276] The processing module obtains candidate input features by simulating quantization of the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter, specifically, quant=2 -param ×clip(round(data×2 param ),-2 bw-1 ,2 bw-1 -1), or quant=param×clip(round(data÷param),-2 bw-1 ,2 bw-1 -1) is used to obtain candidate input features by simulating quantization of floating-point input features, where quant represents the candidate input feature, param represents the candidate feature value quantization hyperparameter, data represents the floating-point input feature, bw represents the current quantization bit width, clip represents the clipping function, and round represents rounding.
[0277] For example, the processing module is further used to simulate quantize the floating-point weights of the target network layer using the target weight quantization hyperparameters and the target weight quantization bit width of the target network layer to obtain pseudoweights of the target network layer.
[0278] Exemplary, when the processing module determines a quantization error corresponding to the candidate feature value quantization hyperparameter based on the floating-point output features of the target network layer, it specifically determines a sample output feature based on the floating-point output features of the target network layer, the sample output feature being the output feature of a reference network layer in a sample decoding network, the reference network layer being any network layer located after the target network layer, and is used to determine the quantization error corresponding to the candidate feature value quantization hyperparameter based on the sample output feature and the reference output feature, the method for obtaining the reference output feature includes decoding a sample feature from a sample bitstream, determining a floating-point input feature corresponding to the sample decoding network based on the sample feature, and determining the reference output feature, which is the output feature of the reference network layer in the sample decoding network, based on the floating-point input feature.
[0279] Those skilled in the art will understand that embodiments of the present invention may be provided as methods, systems, or computer program products. The present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Embodiments of the present invention may take the form of computer program products implemented on one or more computer-compatible storage media (including, but not limited to, magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-compatible program code. The above are merely embodiments of the present invention and do not limit the present invention.
[0280] To those skilled in the art, the present invention is subject to various modifications and changes. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A decoding method, The steps include decoding a target feature corresponding to the current image block from a bitstream corresponding to the current image block, The steps include determining the first input feature of the target decoding network based on the aforementioned target features, The steps include obtaining the target feature value quantization bit width and target feature value quantization hyperparameter of the target decoding network, and converting the first input feature of the target decoding network to the second input feature of the target decoding network based on the target feature value quantization bit width and target feature value quantization hyperparameter, A step of processing the second input features of the target decoding network based on the integer weights of the target decoding network to obtain the output features of the target decoding network, wherein the integer weights of the target decoding network are determined based on the target weight quantization bit width and target weight quantization hyperparameters of the target decoding network. The step includes determining a reconstructed image block corresponding to the current image block based on the output characteristics of the target decoding network, A method characterized by the following:
2. The aforementioned target decoding network includes at least one target network layer, the at least one target network layer employs integer weights, and for each target network layer in the aforementioned target decoding network, Based on the target feature value quantization bit width and target feature value quantization hyperparameters of the target network layer, the first input feature of the target network layer is converted to the second input feature of the target network layer. The second input features of the target network layer are processed based on the integer weights of the target network layer to obtain the output features of the target network layer, the integer weights of the target network layer are determined based on the target weight quantization bit width and target weight quantization hyperparameters of the target network layer, The target feature value quantization bit widths of different target network layers are the same or different. The target feature value quantization hyperparameters for different target network layers may be the same or different. The target weight quantization bit widths for different target network layers are all fixed quantization bit widths. The target weight quantization hyperparameters of different target network layers are either the same or different. The method according to feature 1.
3. The step of converting the first input feature of the target network layer to the second input feature of the target network layer based on the target feature value quantization bit width and target feature value quantization hyperparameters of the target network layer is: quant=clip(round(data×2 param ), -2 bw-1 ,2 bw-1 -1), or, quant=clip(round(data÷param), -2 bw-1 ,2 bw-1 -1) The step of converting the first input feature of the target network layer to the second input feature of the target network layer using the method, quant represents the second input feature of the target network layer, `data` represents the first input feature of the target network layer, param represents the target feature value quantization hyperparameter of the target network layer, bw represents the target feature value quantization bit width of the target network layer, `clip` represents the clipping function, "Round" indicates rounding to the nearest whole number. The method according to feature 2.
4. The sample decoding network includes a plurality of network layers employing floating-point weights, the plurality of network layers include at least one target network layer of the sample decoding network, and for each target network layer of the at least one target network layer of the sample decoding network, A list of candidate weight quantization hyperparameters corresponding to the target network layer is obtained, and the list of candidate weight quantization hyperparameters includes multiple candidate weight quantization hyperparameters. For each candidate weight quantization hyperparameter, Using a fixed quantization bit width and the candidate weight quantization hyperparameters, the floating-point weights of the target network layer are simulated to obtain candidate simulation weights. Based on the candidate simulation weights, the floating-point input features of the target network layer are processed to obtain the floating-point output features of the target network layer. Based on the floating-point output features of the target network layer, the quantization error corresponding to the candidate weight quantization hyperparameter is determined. Based on the quantization error corresponding to each candidate weight quantization hyperparameter, the candidate weight quantization hyperparameter corresponding to the minimum quantization error is determined as the target weight quantization hyperparameter of the target network layer, and the fixed quantization bit width is determined as the target weight quantization bit width of the target network layer. Based on the target weight quantization bit width and target weight quantization hyperparameter of the target network layer, the floating-point weights of the target network layer are quantized to fixed-point values to obtain integer weights of the target network layer. The target decoding network is generated based on the integer weights of each target network layer in the sample decoding network. The method according to any one of claims 1 to 3, characterized by the features described above.
5. The step of obtaining a list of candidate weight quantization hyperparameters corresponding to the target network layer is: A step of determining a maximum weight value based on all floating-point weights of the target network layer, wherein each floating-point weight includes a plurality of weight values, and the maximum weight value is the maximum value among all the weight values of all floating-point weights. The steps include generating initial weight quantization hyperparameters corresponding to the target network layer based on the aforementioned maximum weight value, The step of constructing a list of candidate weight quantization hyperparameters based on the initial weight quantization hyperparameters includes: The method according to feature 4.
6. The step of generating initial weight quantization hyperparameters corresponding to the target network layer based on the aforementioned maximum weight value is: param = bw - 1 - ceil (log2(max)), or, param=(max) / 2 bw-1 The step includes generating the initial weight quantization hyperparameters using the above method, param represents the initial weight quantization hyperparameter corresponding to the target network layer, bw represents the fixed quantization bit width, max represents the maximum weight value, ceil represents the rounding up operation. The method according to specification 5.
7. The step of obtaining candidate simulation weights by simulating quantization of the floating-point weights of the target network layer using a fixed quantization bit width and the candidate weight quantization hyperparameter is: quant = 2 -param * clip(round(data * 2 param ), -2 bw-1 , 2 bw-1 -1), or, quant=param×clip(round(data÷param), -2 bw-1 ,2 bw-1 -1) The process includes the step of obtaining candidate simulation weights by simulating quantizing the floating-point weights of the target network layer using the method described above. quant represents the candidate simulation weights, param represents the candidate weight quantization hyperparameter, `data` represents the floating-point weights of the target network layer. bw represents the fixed quantization bit width, `clip` represents the clipping function, "Round" indicates rounding to the nearest whole number. The method according to feature 4.
8. The step of determining the quantization error corresponding to the candidate weight quantization hyperparameter based on the floating-point output features of the target network layer is: A step of determining sample output features based on the floating-point output features of the target network layer, wherein the sample output features are the output features of a reference network layer in a sample decoding network, and the reference network layer is any network layer located after the target network layer. The steps include determining the quantization error corresponding to the candidate weight quantization hyperparameter based on the sample output features and reference output features, The method for obtaining the reference output features includes decoding sample features from a sample bitstream, determining floating-point input features corresponding to the sample decoding network based on the sample features, and determining the reference output features, which are the output features of the reference network layer in the sample decoding network, based on the floating-point input features. The method according to feature 4.
9. The sample decoding network includes a plurality of network layers employing floating-point weights, the plurality of network layers include at least one target network layer of the sample decoding network, and for each target network layer of the at least one target network layer of the sample decoding network, A list of candidate feature value quantization hyperparameters and a set of quantization bit widths corresponding to the target network layer are obtained, wherein the list of candidate feature value quantization hyperparameters includes multiple candidate feature value quantization hyperparameters, and the set of quantization bit widths includes at least two quantization bit widths. Traverse at least two quantization bit widths from the aforementioned quantization bit width set, For each candidate feature value quantization hyperparameter, Using the current quantization bit width and the candidate feature value quantization hyperparameters, the floating-point input features of the target network layer are simulated quantized to obtain candidate input features. The candidate input features are processed based on the pseudoweights of the target network layer, and the floating-point output features of the target network layer are obtained. Based on the floating-point output features of the target network layer, the quantization error corresponding to the candidate feature value quantization hyperparameter is determined. Based on the quantization error corresponding to each candidate feature value quantization hyperparameter in the current quantization bit width, if the minimum quantization error is smaller than a preset threshold, the candidate feature value quantization hyperparameter corresponding to the minimum quantization error is determined as the target feature value quantization hyperparameter of the target network layer, and the current quantization bit width is determined as the target feature value quantization bit width of the target network layer. The method according to any one of claims 1 to 3, characterized by the features described above.
10. The steps include determining whether the current quantization bit width is the last quantization bit width in the quantization bit width set, based on the quantization error corresponding to each candidate feature value quantization hyperparameter in the current quantization bit width, if the minimum quantization error is greater than or equal to a preset threshold, The process further includes the steps of determining a candidate feature value quantization hyperparameter corresponding to the minimum quantization error as the target feature value quantization hyperparameter of the target network layer, and determining the current quantization bit width as the target feature value quantization bit width of the target network layer, if the current quantization bit width is the last quantization bit width in the set of quantization bit widths. The method according to feature 9.
11. The step of obtaining a list of candidate feature value quantization hyperparameters corresponding to the target network layer is: A step of determining the maximum feature value based on all floating-point input features of the target network layer, wherein each floating-point input feature includes multiple feature values, and the maximum feature value is the maximum value among all feature values of all floating-point input features. The steps include generating initial feature value quantization hyperparameters corresponding to the target network layer based on the aforementioned maximum feature value, The steps include constructing a list of candidate feature value quantization hyperparameters based on the initial feature value quantization hyperparameters, The method according to feature 9.
12. The step of generating initial feature value quantization hyperparameters corresponding to the target network layer based on the aforementioned maximum feature value is: param = bw - 1 - ceil (log2(max)), or, param=(max) / 2 bw-1 The step includes generating the initial feature value quantization hyperparameters using the above, param represents the initial feature value quantization hyperparameter corresponding to the target network layer, bw represents the current quantization bit width, max represents the maximum feature value, ceil represents the rounding up operation. The method according to the present invention, characterized by the features described in the present invention.
13. The step of obtaining candidate input features by simulating quantization of the floating-point input features of the target network layer using the current quantization bit width and the candidate feature value quantization hyperparameter is: quant = 2 -param ×clip(round(data×2 param ), -2 bw-1 ,2 bw-1 -1), or, quant=param×clip(round(data÷param), -2 bw-1 ,2 bw-1 -1) The process includes the step of obtaining candidate input features by simulating quantization of the floating-point input features of the target network layer using the method described above. quant represents the candidate input feature, param represents the candidate feature value quantization hyperparameter, `data` represents the floating-point input feature, bw represents the current quantization bit width, `clip` represents the clipping function, "Round" indicates rounding to the nearest whole number. The method according to feature 9.
14. The process further includes the step of obtaining pseudoweights of the target network layer by simulating quantization of the floating-point weights of the target network layer using the target weight quantization hyperparameters and the target weight quantization bit width of the target network layer. The method according to feature 9.
15. The step of determining the quantization error corresponding to the candidate feature value quantization hyperparameter based on the floating-point output features of the target network layer is: A step of determining sample output features based on the floating-point output features of the target network layer, wherein the sample output features are the output features of a reference network layer in a sample decoding network, and the reference network layer is any network layer located after the target network layer. The step includes determining a quantization error corresponding to the candidate feature value quantization hyperparameter based on the sample output features and reference output features, The method for obtaining the reference output features includes decoding sample features from a sample bitstream, determining floating-point input features corresponding to the sample decoding network based on the sample features, and determining the reference output features, which are the output features of the reference network layer in the sample decoding network, based on the floating-point input features. The method according to feature 9.
16. A decoding module for decoding a target feature corresponding to the current image block from a bitstream corresponding to the current image block, A decision module for determining the first input feature of the target decoding network based on the aforementioned target features, A processing module for obtaining the target feature value quantization bit width and target feature value quantization hyperparameter of the target decoding network, converting the first input features of the target decoding network to second input features of the target decoding network based on the target feature value quantization bit width and target feature value quantization hyperparameter, processing the second input features of the target decoding network based on the integer weights of the target decoding network, and obtaining the output features of the target decoding network, wherein the integer weights of the target decoding network are determined based on the target weight quantization bit width and target weight quantization hyperparameter of the target decoding network, The decision module is used to determine the reconstructed image block corresponding to the current image block based on the output features of the target decoding network. A decoding device characterized by the following features.
17. A decoding device comprising a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor is used to execute the machine-executable instructions and carry out the method according to any one of claims 1 to 15. A decoding device characterized by the following features.
18. A machine-readable storage medium storing a plurality of computer instructions, wherein when the computer instructions are executed by a processor, the method according to any one of claims 1 to 15 is performed. A machine-readable storage medium characterized by the following features.