Feature map encoding and decoding method and apparatus
The feature map encoding and decoding method improves image compression by using peak probability-based selection to enhance accuracy and efficiency, addressing the challenge of balancing quality and efficiency in modern multimedia applications.
Patent Information
- Application Number
- JP2024516958
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-25
- Filing Date
- 2022-09-08
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2042-09-08
AI Technical Summary
Existing image compression technologies struggle to balance image quality with compression efficiency, particularly in modern multimedia applications.
A feature map encoding and decoding method that utilizes probability estimation to determine feature elements based on peak probabilities, selectively performing entropy encoding, thereby improving accuracy and reducing complexity.
Enhances the accuracy and efficiency of data decoding and encoding by accurately selecting feature elements based on peak probabilities, reducing the need for unnecessary entropy coding and lowering computational complexity.
Smart Images

Figure 0007799046000024 
Figure 0007799046000025 
Figure 0007799046000026
Abstract
Description
[Technical Field]
[0001] This application claims priority to Chinese Patent Application No. 202111101920.9, entitled "FEATURE MAP ENCODING AND DECODING METHOD AND APPARATUS," filed with the State Intellectual Property Office of China on September 18, 2021, which is incorporated herein by reference in its entirety. This application claims priority to Chinese Patent Application No. 202210300566.0, entitled "FEATURE MAP ENCODING AND DECODING METHOD AND APPARATUS," filed with the State Intellectual Property Office of China on March 25, 2022, which is incorporated herein by reference in its entirety.
[0002] TECHNICAL FIELD Embodiments of the present application relate to the field of artificial intelligence (AI)-based audio / video or image compression technology, and in particular to a feature map encoding and decoding method and apparatus. [Background technology]
[0003] Image compression is a technique that uses image data characteristics, such as spatial redundancy, visual redundancy, and statistical redundancy, to represent the original image pixel matrix in a lossy or lossless manner using fewer bits to implement efficient transmission and storage of image information. Image compression is classified as lossless compression or lossy compression. Lossless compression causes no loss of image detail, while lossy compression achieves a large compression ratio at the expense of reducing image quality to a certain extent. Lossy image compression algorithms typically use many techniques to remove redundant information in image data. For example, quantization techniques are used to remove spatial redundancy caused by correlation between adjacent pixels in an image and visual redundancy determined by the human visual system's perception. Entropy coding and transform techniques are used to remove statistical redundancy in image data. Mature lossy image compression standards such as JPEG and BPG After decades of research and optimization by those skilled in the art of traditional lossy image compression techniques was formed.
[0004] However, image compression technology ,If it is not possible to guarantee the image compression quality while improving the ,compression efficiency, Image compression techniques cannot meet the increasing data requirements of modern multimedia applications. Summary of the Invention [Means for solving the problem]
[0005] The present application reduces the encoding and decoding complexity. While To improve encoding and decoding performance 、 Feature map encoding and decoding methods and A device is provided.
[0006] According to a first aspect, the present application provides a feature map decoding method, which includes: obtaining a bitstream of a feature map to be decoded, the bitstream including a plurality of feature elements; obtaining a first probability estimation result corresponding to each of the plurality of feature elements based on the bitstream of the feature map to be decoded, the first probability estimation result including a first peak probability; and selecting a first feature element from the plurality of feature elements based on a first threshold and the first peak probability corresponding to each feature element. of Set and second feature element of Determining the set and the first feature element of Set and 2nd of and obtaining a decoded feature map based on the set of feature elements.
[0007] Compared with the method for determining a first feature element and a second feature element from a plurality of feature elements based on a first threshold value and a corresponding probability that the numerical value of each feature element is a fixed value, in the present application, the method for determining a first feature element and a second feature element based on a first threshold value and a peak probability corresponding to each feature element is more accurate, thereby improving the accuracy of the obtained decoded feature map and improving data decoding performance.
[0008] In a possible implementation, the first probability estimation result is a Gaussian distribution and the first peak probability is the mean probability of the Gaussian distribution.
[0009] Alternatively, the first probability estimation result is a Gaussian mixture distribution. The Gaussian mixture distribution includes a plurality of Gaussian distributions. The first peak probability is the maximum value in the average probabilities of the Gaussian distributions, or the first peak probability is calculated based on the average probabilities of the Gaussian distributions and the weights of the Gaussian distributions in the Gaussian mixture distribution.
[0010] In a possible implementation, the decoded feature map value is of The numerical values of all primary features in the set and the secondary features of and the numerical values of all second feature elements in the set.
[0011] In a possible implementation, the first feature of The set is either empty or contains the second characteristic element of The set is the empty set.
[0012] In a possible implementation, the first probability estimation result further includes a feature value corresponding to the first peak probability. Furthermore, entropy decoding may be performed on the first feature based on the first probability estimation result corresponding to the first feature to obtain a numerical value of the first feature. The numerical value of the second feature is obtained based on the feature value corresponding to the first peak probability of the second feature. In this possible implementation, compared to assigning a fixed value to the value of the uncoded feature (i.e., the second feature), the present application assigns a feature value corresponding to the first peak probability of the second feature to the value of the uncoded feature (i.e., the second feature), thereby obtaining a decoded value. Features The accuracy of the numerical value of the second feature element in the map value is improved, and the data decoding performance is improved.
[0013] In a possible implementation, a first feature is selected from the plurality of features based on a first threshold and a first peak probability corresponding to each feature. of Set and second feature element of Before determining the set, the first threshold value may be further obtained based on the bitstream of the feature map to be decoded. In this possible implementation, compared with the method in which the first threshold value is an empirical preset value, the feature map to be decoded corresponds to the first threshold value of the feature map to be decoded, and the variability and flexibility of the first threshold value are increased, thereby reducing the difference between the replacement value and the true value of the uncoded feature element (i.e., the second feature element), and improving the accuracy of the decoded feature map.
[0014] In a possible implementation, the first peak probability of the first feature is less than or equal to a first threshold, and the first peak probability of the second feature is greater than the first threshold.
[0015] In a possible implementation, the first probability estimation result has a Gaussian distribution. The first probability estimation result further includes a first probability variance value. In this case, the first probability variance value of the first feature element is equal to or greater than a first threshold, and the first probability variance value of the second feature element is less than the first threshold. In this possible implementation, when the probability estimation result has a Gaussian distribution, the time complexity of determining the first feature element and the second feature element based on the probability variance value is lower than the time complexity of determining the first feature element and the second feature element based on the peak probability, thereby improving data decoding speed.
[0016] In one possible implementation, side information corresponding to the feature map to be decoded is obtained based on the bitstream of the feature map to be decoded, and a first probability estimation result corresponding to each feature element is obtained based on the side information.
[0017] In one possible implementation, side information corresponding to the feature map to be decoded is obtained based on the bitstream of the feature map to be decoded. A first probability estimation result for each feature element is calculated based on the side information and the first context information: revenge The first context information is estimated for each feature element in the feature map to be encoded.、 Characteristic elements Corresponding to, The feature elements are within a predetermined range in the feature map to be decoded. In this possible implementation, the probability estimation result of each feature element is obtained based on side information and context information, thereby improving the accuracy of the probability estimation result and improving the encoding and decoding performance.
[0018] According to a second aspect, the present application provides a feature map encoding method, the method including: obtaining a first feature map to be encoded, the first feature map including a plurality of features; determining a first probability estimation result for each of the plurality of features based on the first feature map to be encoded, the first probability estimation result including a first peak probability; determining whether a feature is a first feature based on the first peak probability of each feature in the first feature map to be encoded; and performing entropy encoding on the first feature only if the feature is the first feature.
[0019] According to the method of the second aspect, whether entropy coding needs to be performed on each feature element in the feature map to be coded is determined, thereby skipping the coding process of some feature elements in the feature map to be coded, significantly reducing the amount of elements for entropy coding and the complexity of entropy coding. In addition, compared with determining whether a feature element needs to be coded based on a probability corresponding to a fixed value in the probability estimation result corresponding to each feature element, the reliability of the determination result (whether entropy coding needs to be performed on a feature element) is improved based on the probability peak of each feature element, and the coding process of more feature elements is skipped, thereby further improving the coding speed and coding performance.
[0020] In a possible implementation, the first probability estimation result is a Gaussian distribution and the first peak probability is the mean probability of the Gaussian distribution.
[0021] Alternatively, the first probability estimation result is a Gaussian mixture distribution. The Gaussian mixture distribution includes a plurality of Gaussian distributions. The first peak probability is the maximum value in the average probabilities of the Gaussian distributions, or the first peak probability is calculated based on the average probabilities of the Gaussian distributions and the weights of the Gaussian distributions in the Gaussian mixture distribution.
[0022] In a possible implementation, for each feature in the first feature map to be coded, it is determined whether the feature is a first feature based on a first threshold and a first peak probability of the feature.
[0023] In a possible implementation, a second probability estimation result for each of the plurality of features is determined based on the first feature map to be coded, the second probability estimation result including a second peak probability. A third feature element is selected from the plurality of features based on the second probability estimation result for each feature element. of The first threshold is the third feature element. of The first threshold is determined based on the second peak probabilities of all features in the set. Entropy coding is performed for the first threshold. In this possible implementation, the first threshold for the feature map to be coded may be determined for the feature map to be coded based on the features of the feature map to be coded, so that the first threshold has a better fit to the feature map to be coded, thereby improving the reliability of the decision result (i.e., whether entropy coding needs to be performed for the feature) determined based on the first threshold and the first peak probabilities of the feature.
[0024] In a possible implementation, the first threshold is a third characteristic element. of The second peak probability is the maximum of the second peak probabilities corresponding to the features in the set.
[0025] In a possible implementation, the first peak probability of the first feature is less than or equal to a first threshold.
[0026] In a possible implementation, the second probability estimation result is a Gaussian distribution, and the second probability estimation result further includes a second probability variance value. of The first probability estimation result may be a Gaussian distribution, and the first probability estimation result may further include a first probability variance value. The first probability variance value of the first feature may be greater than or equal to a first threshold. In this possible implementation, when the probability estimation result is a Gaussian distribution, the time complexity of determining the first feature based on the probability variance value is lower than the time complexity of determining the first feature based on the peak probability, thereby improving the data encoding speed.
[0027] In a possible implementation, the second probability estimation result further includes a feature value corresponding to the second peak probability. of The set is determined from the plurality of features based on a preset error, a numerical value of each feature, and a feature value corresponding to a second peak probability for each feature.
[0028] In a possible implementation, the third feature of The features in the set are
number
number
[0029] In a possible implementation, the first probability estimation result is the same as the second probability estimation result. In this case, side information of the first feature map to be coded is obtained based on the first feature map to be coded. Probability estimation is performed on the side information to obtain the first probability estimation result for each feature element.
[0030] In a possible implementation, the first probability estimation result is different from the second probability estimation result, in which side information of the first feature map to be coded and second context information of each feature element are obtained based on the first feature map to be coded. The second context information is 、 Characteristic elements Corresponding to, The feature elements are within a predetermined range in the first feature map to be coded. A second probability estimation result for each feature element is obtained based on the side information and the second context information.
[0031] In a possible implementation, side information of the first feature map to be coded is obtained based on the first feature map to be coded. For any feature element in the first feature map to be coded, a first probability estimation result of the feature element is determined based on the first context information and the side information. The first probability estimation result further includes a feature value corresponding to a first probability peak. The first context information 、 Characteristic elements Corresponding to, The feature element is a feature element within a predetermined range in the second feature map to be coded. The values of the second feature map to be coded include a numerical value of the first feature element and a feature value corresponding to a first peak probability of the second feature element. The second feature element is a feature element other than the first feature element in the first feature map to be coded. In this way, the probability estimation result of each feature element is obtained by referring to the side information and the context information, thereby improving the accuracy of the probability estimation result of each feature element compared to a method in which the probability estimation result of each feature element is obtained based on the side information alone.
[0032] In a possible implementation, the entropy coding results of all first feature elements are written into the coded bitstream.
[0033] According to a third aspect, the present application provides a method for manufacturing a semiconductor device comprising: an acquisition module configured to acquire a bitstream of a feature map to be decoded, the feature map to be decoded including a plurality of feature elements, and to acquire a first probability estimation result corresponding to each of the plurality of feature elements based on the bitstream of the feature map to be decoded, the first probability estimation result including a first peak probability; a decoding module for selecting a first feature from the plurality of features based on a first threshold and a first peak probability corresponding to each feature; of Set and second feature element of Determine the set of first characteristic elements of Set and second feature element of a decoding module configured to obtain a decoded feature map based on the set.
[0034] For further implementation functions of the acquisition module and the decoding module, please refer to the first aspect or any one of the implementation forms of the first aspect, and the details will not be described again in this specification.
[0035] According to a fourth aspect, the present application provides: an acquisition module for acquiring a first encoding target feature map, the first encoding target feature map being configured to include a plurality of feature elements; and an encoding module configured to determine a first probability estimation result for each of a plurality of features based on a first feature map to be encoded, the first probability estimation result including a first peak probability, determine whether a feature is a first feature based on the first peak probability of each feature in the first feature map to be encoded, and perform entropy encoding on the first feature only if the feature is the first feature.
[0036] For further implementation functions of the acquisition module and the encoding module, please refer to the second aspect or any one of the implementation forms of the second aspect, and the details will not be described again in this specification.
[0037] According to a fifth aspect, the present application provides a decoder, the decoder including a processing circuit and configured to determine a method according to any one of the first aspect and the implementation forms of the first aspect.
[0038] According to a sixth aspect, the present application provides an encoder, the encoder including a processing circuit and configured to determine a method according to any one of the second aspect and the implementation forms of the second aspect.
[0039] According to a seventh aspect, the present application provides a computer program product including program code, which, when determined by a computer or processor, determines a method according to any one of the first aspect and implementations of the first aspect, or a method according to any one of the second aspect and implementations of the second aspect.
[0040] According to an eighth aspect, the present application provides a decoder including one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing a program determined by the processors, the program, when determined by the processors, enabling the decoder to determine a method according to the first aspect and any one of the implementation forms of the first aspect.
[0041] According to a ninth aspect, the present application provides an encoder including one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing a program determined by the processors, the program, when determined by the processors, enabling the encoder to determine a method according to any one of the second aspect and implementation forms of the second aspect.
[0042] According to a tenth aspect, the present application provides a non-transitory computer-readable storage medium including a program code, which, when determined by a computer device, determines a method according to any one of the first aspect and the implementation forms of the first aspect, or a method according to any one of the second aspect and the implementation forms of the second aspect.
[0043] According to an eleventh aspect, application The present invention relates to a decoding device. The decoding device has a function for implementing the behavior according to any one of the embodiments of the first aspect or the method of the first aspect. The function may be implemented by hardware or by hardware that determines corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions.
[0044] According to a twelfth aspect, application The present invention relates to an encoding device. The encoding device has a function for implementing the behavior according to any one of the embodiments of the second aspect or the method of the second aspect. The function may be implemented by hardware or by hardware that determines the corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions. [Brief explanation of the drawings]
[0045] [Figure 1] 1 is a schematic diagram of the architecture of a data coding system according to an embodiment of the present application; [Figure 4a] FIG. 10 is a schematic diagram of the input and output results of the probability estimation module 103 according to an embodiment of the present application. [Figure 2a] FIG. 1 is a schematic diagram of the output result of the probability estimation module 103 according to an embodiment of the present application. [Figure 2b] FIG. 1 is a schematic diagram of a probability estimation result according to an embodiment of the present application; [Figure 3] 1 is a schematic flowchart of a feature map encoding method according to an embodiment of the present application; [Figure 4a] FIG. 10 is a schematic diagram of the input and output results of the probability estimation module 103 according to an embodiment of the present application. [Figure 4b] 1 is a schematic diagram of the structure of a probability estimation network according to an embodiment of the present application; [Figure 4c] 1 is a schematic flowchart of a method for determining a first threshold value according to an embodiment of the present application; [Figure 5] 1 is a schematic flowchart of a feature map decoding method according to an embodiment of the present application; [Figure 6a] 1 is a schematic flowchart of another feature map encoding method according to an embodiment of the present application; [Figure 6b] FIG. 10 is a schematic diagram of input and output results of another probability estimation module 103 according to an embodiment of the present application. [Figure 7a] 1 is a schematic flowchart of another feature map decoding method according to an embodiment of the present application; [Figure 7b] FIG. 1 is a schematic diagram of experimental results of a compression performance comparison test according to an embodiment of the present application. [Figure 7c] FIG. 10 is a schematic diagram of experimental results of another compression performance comparison test according to an embodiment of the present application. [Figure 8] 1 is a schematic diagram of the structure of a feature map encoding device according to an embodiment of the present application; [Figure 9] 1 is a schematic diagram of the structure of a feature map decoding device according to an embodiment of the present application; [Figure 10] 1 is a schematic diagram of the structure of a computing device according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION
[0046] The following clearly and completely describes the technical solutions in the embodiments of the present application with reference to the accompanying drawings. It is obvious that the described embodiments are only some embodiments of the present application, but not all embodiments.
[0047] It should be noted that in the specification of this application and the accompanying drawings, terms such as "first," "second," etc. are intended to distinguish between different objects or between different processes of the same object, but are not used to describe a particular order of objects. Additionally, the terms "including," "having," or any other variations thereof in the description of this application are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include other unlisted steps or units, or may optionally include other inherent steps or units of the process, method, product, or device. It should be noted that in the embodiments of this application, words such as "an example," "for example," etc. are used to denote providing an example, illustration, or explanation. Any embodiment or design scheme described as an "example" or "for example" in the embodiments of this application should not be described as preferred or having more advantages than another embodiment or design scheme. Specifically, the use of the words "example" and "for example" is intended to present related concepts in a particular way. In the embodiments of the present application, "A and / or B" represents two meanings: A and B, and A or B. A and / or B, and / or C represent any one of A, B, and C, or any two of A, B, and C, or A, B, and C. The technical solutions of the present application are described below with reference to the accompanying drawings.
[0048] The feature map decoding method and the feature map encoding method provided in the embodiments of the present application can be used in data coding fields (including audio coding fields, video coding fields, and image coding fields). Specifically, the feature map decoding method and the feature map encoding method can be used in the scenarios of album management, human-computer interaction, audio compression or transmission, video compression or transmission, image compression or transmission, and data compression or transmission. For ease of understanding, it should be noted that the embodiments of the present application are only described by using an example in which the feature map decoding method and the feature map encoding method are used in the image coding field, and this cannot be considered as a limitation on the methods provided in the present application.
[0049] Specifically, an example is used in which the feature map encoding method and the feature map decoding method are used in an end-to-end image feature map encoding and decoding system. The end-to-end image feature map encoding and decoding system includes two parts: image encoding and image decoding. Image encoding is determined on the source side and typically involves processing the original video image (e.g., by compressing it) to reduce the amount of data required to represent the video image (for more efficient storage and / or transmission). Image decoding is determined on the destination side and typically involves the reverse process of the encoder to reconstruct the image. In the end-to-end image feature map encoding and decoding system, according to the feature map decoding method and the feature map encoding method provided in this application, it can be determined whether entropy coding needs to be performed on each feature element in the feature map to be encoded, thereby skipping the encoding process of some features, reducing the amount of elements for entropy coding, and reducing the complexity of entropy coding. In addition, based on the probability peak of each feature element, the reliability of the decision result (whether entropy coding needs to be performed on the feature element) is improved, thereby improving image compression performance.
[0050] The embodiments of the present application relate to large-scale applications of neural networks, and therefore, for ease of understanding, the following will first explain terms and concepts related to neural networks in the embodiments of the present application.
[0051] 1. Entropy Coding Entropy coding is a coding process that does not lose information according to the entropy principle. Entropy coding uses an entropy coding algorithm or solution in quantization coefficients or other syntax elements to obtain coded data that can be output by an output terminal in the form of a coded bitstream, etc., so that a decoder or the like can receive and use parameters used for decoding. The coded bitstream can be transmitted to a decoder or stored in a memory for later transmission or retrieval by the decoder. The entropy coding algorithm or solution may include, but is not limited to, a variable length coding (VLC) solution, a context adaptive VLC solution (CALVC), an arithmetic coding scheme, a binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique.
[0052] 2. Neural Networks A neural network may include neurons, where x sand the intercept of 1 can be used as input. The output of the operational unit can be shown as equation (1):
number
[0053] s=1, 2,..., or n, where n is a natural number greater than 1, and W s x s where a is the weight of the neuron, and b is the bias of the neuron. f is the activation function of the neuron, which is used to introduce nonlinear features into the neural network and convert the input signal at the neuron into an output signal. The signal output from the activation function can serve as the input of the next convolutional layer. The activation function may be a sigmoid function. A neural network is a network formed by connecting many single neurons to each other. Specifically, the output of a neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract features of the local receptive field. The local receptive field may be a region containing multiple neurons.
[0054] 3. Deep neural network (DNN) DNN, also known as multi-layer neural network, can be understood as a neural network with multiple hidden layers. DNN is divided based on the location of different layers, and as a result, the neural network within DNN can be classified into three types: input layer, hidden layer, and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the intermediate layer is a hidden layer. The layers are fully connected. Specifically, any neuron in the i-th layer is always connected to any neuron in the i+1-th layer.
[0055] Although DNNs may seem complex, the work of each layer is not.
number
number
[0056] In conclusion, the coefficient from the kth neuron in the (L-1)th layer to the jth neuron in the Lth layer is
number
[0057] Note that the input layer does not have the parameter W. In a deep neural network, the more hidden layers, the more capable the network is of explaining complex cases in the real world. Theoretically, a model with more parameters has higher complexity and greater "capacity," which indicates that the model can complete more complex learning tasks. The process of training a deep neural network is the process of learning weight matrices, and the ultimate goal of training is to obtain the weight matrices of all layers of the trained deep neural network (weight matrices consisting of vectors W of multiple layers).
[0058] 4. Convolutional neural networks al network, CNN) A CNN is a deep neural network with a convolutional structure. A convolutional network includes a feature extractor, which includes a convolutional layer and a subsampling layer. The feature extractor can be considered a filter. The convolution process can be considered as performing convolution by using a trainable filter and an input image or a convolutional feature map. A convolutional layer is a neuron layer in a convolutional network that performs convolutional processing on an input signal. In the convolutional layer portion of a convolutional network, a neuron may only be connected to some of the neurons in adjacent layers. A convolutional layer usually includes some feature planes, and each feature plane may include some neurons in a rectangular arrangement. Neural units in the same feature plane share weights, and the shared weights herein are convolution kernels. Weight sharing can be understood as the position-independent nature of the image information extraction method. The principle implied herein is that the statistics of one part of an image are the same as those of other parts. This means that image information learned in one part can be utilized in other parts. Therefore, the same image information obtained by training can be used for all positions on the image. Multiple convolution kernels can be used in the same convolution layer to extract different image information. Generally, the more convolution kernels there are, the richer the image information reflected in the convolution operation.
[0059] The convolution kernels may be initialized in the form of a random-size matrix. During the process of training the convolution network, the convolution kernels can acquire appropriate weights through learning. In addition, a direct benefit of weight sharing is that it reduces the connections between layers of the convolution network, thereby reducing the risk of overfitting.
[0060] 5. Recurrent neural networks (RNN) In the real world, many elements are ordered and interconnected. To enable machines to have human-like memory capabilities, RNNs are developed to perform inference from context.
[0061] RNNs process sequence data. Specifically, the current output of a sequence is related to the previous output. In other words, the output of an RNN depends on the current input information and historical memory information. A specific representation is that the network memorizes previous information and applies it to the calculation of the current output. Specifically, nodes in the hidden layer are connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. Theoretically, RNNs can process sequence data of any length. Training an RNN is similar to training a conventional CNN or DNN. The backpropagation algorithm is also used, but the difference is that when an RNN is extended, the RNN parameters (such as W) are shared. This differs from the conventional neural network described in the previous example. In addition, when using the gradient descent algorithm, the output at each step depends not only on the network state at the current step but also on the network state at some previous step. The training algorithm is called the backpropagation through time (BPTT) algorithm.
[0062] 6. Loss Function In the process of training a deep neural network, the current network prediction value may be compared with the actual expected target value, so that the output of the deep neural network is expected to be as close as possible to the actual expected predicted value. Then, the weight vector of each layer of the neural network is updated based on the difference between the prediction value and the target value. (Indeed, there is usually an initialization process before the first update, specifically, parameters are preconfigured for all layers of the deep neural network.) For example, if the network prediction value is large, the weight vector is adjusted to reduce the prediction value, and adjustments are continuously performed until the deep neural network can predict the actual expected target value or a value very close to the actual expected target value. Therefore, "how to obtain the difference between the prediction value and the target value through comparison" needs to be predefined. This is the loss function or objective function. The loss function and objective function are important equations that measure the difference between the prediction value and the target value. The loss function is used as an example. A higher output value (loss) of the loss function indicates a larger difference. Therefore, training a deep neural network is a process of minimizing the loss as much as possible.
[0063] 7. Backpropagation Algorithm The convolutional network may correct the parameter values of the initial super-resolution model during the training process using the back propagation (BP) algorithm, resulting in a smaller error loss in reconstructing the super-resolution model. Specifically, the input signal is forward-transferred until an error loss occurs at the output, and the parameters of the initial super-resolution model are updated based on the back propagation error loss information to converge the error loss. The back propagation algorithm is an error loss-centered back propagation algorithm intended to obtain parameters such as the weight matrix of the optimal super-resolution model.
[0064] 8. Generative Adversarial Networks A generative adversarial network (GAN) is a deep learning model. The model includes at least two modules: one is a generative model and the other is a discriminative model. The two modules are used to learn from each other through games to generate better outputs. Both the generative model and the discriminative model may be neural networks, specifically, deep neural networks or convolutional neural networks. The basic principle of GAN is as follows: Using a GAN for generating pictures as an example, it is assumed that there are two networks, namely, G (Generator) and D (Discriminator). G is a network for generating pictures. G receives random noise z and generates a picture by using the noise, where the picture is denoted as G(z). D is a discriminator network used to determine whether the picture is "real." The input parameter of D is x, where x represents the picture, and the output D(x) represents the probability that x is a real picture. A value of 1 for D(x) indicates that the picture is 100% realistic. A value of 0 for D(x) indicates that the picture cannot be realistic. In the process of training a generative adversarial network, the goal of the generative network G is to generate pictures that are as realistic as possible to deceive the discriminator network D, and the goal of the discriminator network D is to distinguish between pictures generated by G and real pictures as much as possible. In this way, a dynamic "game" process, specifically an "adversary" in the "generative adversarial network," exists between G and D. The final game result, in an ideal state, is that G may generate an image G(z) that is difficult to distinguish from a real image, and D has difficulty determining whether the image generated by G is real. Specifically, D(G(z)) = 0.5.In this way, a good generative model G can be obtained and used to generate pictures.
[0065] 9. Pixel Values The pixel values of an image may be red-green-blue (RGB) color values. The pixel values may be long integers that represent colors. For example, the pixel value may be 256*Red+100*Green+76 * Blue, where Blue represents the blue component, Green represents the green component, and Red represents the red component. For each color component, a smaller value indicates lower brightness, and a larger value indicates higher brightness. For grayscale images, the pixel value can be a grayscale value.
[0066] The following describes the system architecture provided in the embodiment of the present application. coding The system architecture is shown below. coding The system architecture includes a data acquisition module 101 , a feature extraction module 102 , a probability estimation module 103 , a data encoding module 104 , a data decoding module 105 , a data reconstruction module 106 , and a display module 107 .
[0067] The data capture module 101 is configured to capture an original image, e.g., any kind of image capture device for capturing real-world images, and / or any type of image generation device, e.g., computer graphics for generating computer-animated images. Processing Unit, or any type of other device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). Data capture module 101 may also be any type of memory or storage for storing images.
[0068] The feature extraction module 102 is configured to receive an original image from the data ingestion module 101, preprocess the original image, and further extract a feature map (i.e., a feature map to be encoded) from the preprocessed image through a feature extraction network. The feature map (i.e., a feature map to be encoded) includes multiple feature elements. Specifically, the preprocessing of the original image includes, but is not limited to, cropping, color format conversion (e.g., conversion from RGB to YcbCr), color correction, noise removal, normalization, etc. The feature extraction network may be one or a variation of a neural network, a DNN, a CNN, or an RNN. The specific form of the feature extraction network is not specifically limited herein. Optionally, the feature extraction module 102 is further configured to perform rounding on the feature map (i.e., a feature map to be encoded), for example, through scalar quantization or vector quantization. The feature map includes multiple feature elements, and the value of the feature map should be learned to include the numerical values of all the feature elements. Optionally, the feature extraction module 102 further includes a side information extraction network. Specifically, in addition to outputting the feature map output by the feature extraction network, the feature extraction module 102 further outputs side information of the feature map extracted through the side information extraction network. The side information extraction network may be one or a variation of a neural network, a DNN, a CNN, or an RNN. The specific form of the feature extraction network is not specifically limited herein.
[0069] The probability estimation module 103 estimates the probability of a value corresponding to each of multiple feature elements in a feature map (i.e., a feature map to be coded). For example, the feature map to be coded includes m feature elements, where m is a positive integer. As shown in FIG. 2a, the probability estimation module 103 outputs a probability estimation result for each of the m feature elements. For example, the probability estimation result of the feature elements can be shown in FIG. 2b. The horizontal coordinate in FIG. 2b is the possible numerical value of the feature element (also referred to as the possible value of the feature element). The vertical coordinate indicates the likelihood of each possible numerical value (also referred to as the possible value of the feature element). For example, point P indicates that the probability that the value of the feature element is [a-0.5, a+0.5] is p.
[0070] The data encoding module 104 is configured to perform entropy encoding based on the feature map (i.e., the feature map to be encoded) from the feature extraction module 102 and the probability estimation results of each feature element from the probability estimation module 103, to generate an encoded bitstream (also referred to in this specification as a bitstream of the feature map to be decoded).
[0071] The data decoding module 105 is configured to receive the encoded bitstream from the data encoding module 104, and further perform entropy decoding based on the encoded bitstream and the probability estimation result of each feature element from the probability estimation module 103 to obtain a decoded feature map (or understood as a decoded feature map value).
[0072] The data reconstruction module 106 is configured to perform post-processing on the decoded image feature map from the data decoding module 105 and perform image reconstruction on the post-processed decoded image feature map via an image reconstruction network to obtain a decoded image. The post-processing operations include, but are not limited to, color format conversion (e.g., YcbCr to RGB conversion), color correction, cropping, resampling, etc. The image reconstruction network may be one or a variant of a neural network, a DNN, a CNN, or an RNN. The specific form of the feature extraction network is not specifically limited herein.
[0073] The display module 107 is configured to display the decoded images from the data reconstruction module 106 to display the images to a user, viewer, etc. The display module 107 may be configured to display the reconstructed audio or the reconstructed images. of The player may be or include any type of player or display, for example, an integrated or external display. For example, the display may be a liquid crystal display (LCD), an organic light emitting diode (OLED), or a - The display may include a light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any class of other display.
[0074] data coding Note that the architecture of the system can be the functional modules of the device. coding The architecture of the system can alternatively be coding It may be a system, i.e., a data codingThe system architecture includes two devices: a source device and a destination device. The source device may include a data capture module 101, a feature extraction module 102, a probability estimation module 103, and a data encoding module 104. The destination device may include a data decoding module 105, a data reconstruction module 106, and a display module 107. Scheme 1 is configured such that the source device provides an encoded bitstream to the destination device. So, The source device can transmit the encoded bitstream to the destination device via a communication interface. The communication interface can be a direct communication link between the source device and the destination device, e.g., a direct wired or wireless connection, or through any type of network, e.g., a wired network, a wireless network, any combination thereof, any type of private network and public network, or any combination thereof. Method 2 in which the source device is configured to provide the encoded bitstream to the destination device So, The source device may store the encoded bitstream in a storage device, and the destination device may retrieve the encoded bitstream from the storage device.
[0075] It should be noted that the feature map encoding method referred to in this application may be mainly implemented by the probability estimation module 103 and the data encoding module 104 in Figure 1. The feature map decoding method referred to in this application may be mainly implemented by the probability estimation module 103 and the data decoding module 105 in Figure 1.
[0076] In one example, the feature map encoding method provided in the present application is implemented by an encoding device, which may mainly include the probability estimation module 103 and the data encoding module 104 in FIG. 1. For the feature map encoding method provided in the present application, the encoding device performs the following steps, namely, step 11 to step 14: Implemented obtain.
[0077] Step 11: The encoding device obtains a first encoding target feature map, where the first encoding target feature map includes a plurality of feature elements.
[0078] Step 12: The probability estimation module 103 in the encoding device determines a first probability estimation result for each of the plurality of feature elements based on the first feature map to be encoded, and the first probability estimation result includes a first peak probability.
[0079] Step 13: The encoding device determines whether the feature is the first feature based on the first peak probability of each feature in the first feature map to be encoded.
[0080] Step 14: The data encoding module 104 in the encoding device performs entropy encoding on the first feature only if the feature is the first feature.
[0081] In another example, the feature map decoding method provided in the present application may be implemented by a decoding device. implementation 1. In the feature map decoding method provided in the present application, the decoding device may include the following steps: step 21 to step 24.
[0082] Step 21: The decoding device obtains a bitstream of a feature map to be decoded, where the feature map to be decoded includes a plurality of feature elements.
[0083] Step 22: The probability estimation module 103 in the decoding device obtains a first probability estimation result corresponding to each of the plurality of feature elements based on the bitstream of the feature map to be decoded, and the first probability estimation result includes a first peak probability.
[0084] Step 23: The decoding device selects a first feature element from the plurality of features based on a first threshold value and a first peak probability corresponding to each feature element. of Set and second feature element of Determine the set.
[0085] Step 24: The data decoding module 105 in the decoding device detects the first characteristic element of Set and second feature element of Based on the set, a decoded feature map is obtained.
[0086]
[0023] In the following, specific implementation forms of the feature map decoding method and the feature map encoding method provided in the present application will be described in detail with reference to the accompanying drawings. In the following, the schematic diagram of the implementation procedure at the encoder side shown in Figure 3 and the schematic diagram of the implementation procedure at the decoder side shown in Figure 5 can be regarded as a schematic flowchart of the feature map encoding and decoding method. The schematic diagram of the implementation procedure at the encoder side shown in Figure 6a and the schematic diagram of the implementation procedure at the decoder side shown in Figure 7a can be regarded as a schematic flow diagram of the feature map encoding and decoding method.
[0087] Encoder Side: Figure 3 is a schematic flowchart of a feature map encoding method according to an embodiment of the present application. The procedure of the feature map encoding method includes S301 to S306.
[0088] S301: Obtain a first encoding target feature map, where the first encoding target feature map includes a plurality of feature elements.
[0089] After feature extraction is performed on the original data, a feature map y to be encoded is obtained. Then, the feature map y to be encoded is quantized, i.e., the floating-point feature values are rounded to obtain integer feature values, and the quantized feature map y to be encoded is obtained.
number
number
number
[0090] S302: Based on the first encoding target feature map, side information of the first encoding target feature map is obtained.
[0091] The side information may be understood as a feature map obtained through further feature extraction on the feature map to be coded, where the amount of feature elements contained in the side information is less than the amount of feature elements in the feature map to be coded.
[0092] In a possible implementation, the side information of the first encoding target feature map may be obtained through a side information extraction network. The side information extraction network may use an RNN, a CNN, a variant of an RNN, a variant of a CNN, or another deep neural network (or a variant of another deep neural network). This is not specifically limited in this application.
[0093] S303: Obtain a first probability estimation result for each feature element based on the side information, where the first probability estimation result includes a first peak probability.
[0094] As shown in FIG. 4a, the side information is used as input to the probability estimation module 103 of FIG. 1, and the output from the probability estimation module 103 is a first probability estimation result for each feature element. The probability estimation module 103 may be a probability estimation network, which may use an RNN, a CNN, a variant of an RNN, a variant of a CNN, or another deep neural network (or a variant of another deep neural network). FIG. 4b is a schematic diagram of the structure of the probability estimation network. In FIG. 4b, the probability estimation network is a convolutional network, which includes five network layers: three convolutional layers and two nonlinear activation layers. The probability estimation module 103 may alternatively be implemented according to a conventional non-network probability estimation method. Probability estimation methods include, but are not limited to, statistical methods such as maximum likelihood estimation, maximum a posteriori estimation, and maximum likelihood estimation.
[0095] Any feature element in the first encoding target feature map
number
number
number
number
[0096] In a possible implementation, the first probability estimation result is a Gaussian distribution, and the first peak probability is the average probability of the Gaussian distribution. For example, the first probability estimation result is a Gaussian distribution as shown in FIG. 2b, and the first peak is the average probability of the Gaussian distribution, i.e., the probability p corresponding to the average value a.
[0097] In another possible implementation, the first probability estimation result is a Gaussian mixture distribution. The Gaussian mixture distribution includes a plurality of Gaussian distributions. In other words, the Gaussian mixture distribution can be obtained by multiplying a Gaussian distribution by the weight of the Gaussian distribution through weighting. In a possible case, the first peak probability is the maximum value in the average probability of the Gaussian distribution. Alternatively, in another possible case, the first peak probability is calculated based on the average probability of the Gaussian distribution and the weight of the Gaussian distribution in the Gaussian mixture distribution.
[0098] For example, the first probability estimation result is a Gaussian mixture distribution, and the Gaussian mixture distribution is obtained by weighting Gaussian distribution 1, Gaussian distribution 2, and Gaussian distribution 3. The weight of Gaussian distribution 1 is w1, the weight of Gaussian distribution 2 is w2, and the weight of Gaussian distribution 3 is w3. The average probability of Gaussian distribution 1 is p1. The average probability of Gaussian distribution 2 is p2. The average probability of Gaussian distribution 3 is p3, where p1>p2>p3. When the first peak probability is the maximum value in the average probabilities of the Gaussian distributions, the first peak probability is the maximum value of the average probabilities of the Gaussian distributions (i.e., the average probability of Gaussian distribution 1 is p1). When the first peak probability is calculated based on the average probabilities of the Gaussian distributions and the weights of the Gaussian distributions in the Gaussian mixture distribution, the first peak probability is shown in Equation (2). First peak probability = p1 × w1 + p2 × w2 + p3 × w3 (2)
[0099] If the first probability estimation result is a Gaussian mixture distribution, RugaIt should be learned that the weights corresponding to the Gaussian mixture distributions can be obtained and output through a probability estimation network (e.g., probability estimation module 103). In other words, when obtaining the first probability estimation result (i.e., Gaussian mixture distribution) of each feature element, the probability estimation network can estimate the weights corresponding to the Gaussian mixture distributions. Ruga We also obtain the weights corresponding to the Usian distribution.
[0100] S304: First probability of each feature element Estimate A first threshold is determined based on the result.
[0101] In a possible implementation, the third feature of The set is determined from the plurality of feature elements in the first feature map to be encoded based on the first probability estimation result of each feature element in the first feature map to be encoded. Further, the first threshold is set to a value corresponding to a third feature element. of The probability is determined based on the first probability estimates of all features in the set.
[0102] In other words, the process of determining the first threshold value can be divided into two steps. Specifically, step S401 and A schematic flow chart of determining the first threshold, including S402, is shown in FIG. 4c.
[0103] S401: Selecting a third feature element from a plurality of feature elements included in a first encoding target feature map of Determine the set.
[0104] Third characteristic element of The set is determined from the plurality of feature elements in the first feature map to be encoded based on a first probability estimation result for each feature element in the first feature map to be encoded. of The set may be understood as a set of features for determining the first threshold.
[0105] In a possible implementation, the third feature ofThe set may be determined from a plurality of features based on a preset error, a numerical value of each feature in the first feature map to be coded, and a feature value corresponding to a first peak probability of each feature. 1 The feature value corresponding to the peak probability of is the possible value (or possible numerical value) of the feature element corresponding to the first peak probability in the first probability estimation result of the feature element, for example, the horizontal coordinate numerical value a of point P in Figure 2b. The preset error value may be understood as the allowable error in the feature map encoding method, and may be determined based on experience or according to an algorithm.
[0106] Specifically, the feature elements in the determined third feature element set have the characteristics shown in equation (3).
number
[0107]
number
number
number
[0108] For example, the plurality of feature elements included in the first feature map to be coded are feature element 1, feature element 2, feature element 3, feature element 4, and feature element 5. The first probability estimation result of each feature element in the plurality of feature elements in the first feature map to be coded is obtained through a probability estimation module. In this case, based on the preset error e, the value of each feature element, and the first peak probability of the first probability estimation result corresponding to each feature element (hereinafter referred to as the first peak probability of the feature element for short), a feature element that satisfies Equation (3) is selected from feature element 1, feature element 2, feature element 3, feature element 4, and feature element 5, and is selected as a third feature element. of Form a set. Feature 1 satisfies equation (3) if the absolute difference between the value of feature 1 and the feature value of the first peak probability corresponding to feature 1 is greater than TH_2. Feature 2 satisfies equation (3) if the absolute difference between the value of feature 2 and the feature value of the first peak probability corresponding to feature 2 is greater than TH_2. Feature 3 does not satisfy equation (3) if the absolute difference between the value of feature 3 and the feature value of the first peak probability corresponding to feature 3 is less than TH_2. Feature 4 does not satisfy equation (3) if the absolute difference between the value of feature 4 and the feature value of the first peak probability corresponding to feature 4 is equal to TH_2. Feature 5 satisfies equation (3) if the absolute difference between the value of feature 5 and the feature value of the first peak probability corresponding to feature 5 is greater than TH_2. In conclusion, the feature 1, the feature 2, and the feature 5 are determined to be the third feature from the feature 1, the feature 2, the feature 3, the feature 4, and the feature 5, and the third feature of Form a set.
[0109] S402: Third characteristic element of A first threshold is determined based on the first probability estimates of all features in the set.
[0110] The first threshold is the third feature element ofThe probability distribution is determined based on the form of a first probability estimation result of the features in the set, the form of the first probability estimation result including a Gaussian distribution or another form of probability distribution (including but not limited to a Laplace distribution or a Gaussian mixture distribution).
[0111] Below is the first probability Estimate The method for determining the first threshold based on the type of result will now be described in detail.
[0112] Method 1: The first threshold is the third feature element of The maximum first peak probability among the first peak probabilities corresponding to the features in the set.
[0113] In this way, the first probability Estimate It should be learned that the form of the result may be a Gaussian distribution or another form of probability distribution (including but not limited to a Laplace distribution or a mixture of Gaussians).
[0114] For example, feature 1, feature 2, and feature 5 are combined into a third feature of If the first peak probability of feature 1 is 70%, the first peak probability of feature 2 is 65%, and the first peak probability of feature 5 is 75%, then the third feature is determined to form the set. of The largest first peak probability corresponding to the features in the set (ie, the first peak probability of feature 5 is 75%) is determined to be the first threshold.
[0115] Method 2: The first probability estimation result is a Gaussian distribution, and the first probability estimation result further includes a first probability variance value. The first threshold is a third feature element. of The smallest first probability variance value among the first probability variance values corresponding to the features in the set.
[0116] It should be learned that the mathematical characteristics of the Gaussian distribution can be summarized as follows: in a Gaussian distribution, a larger first probability variance value indicates a smaller first peak probability. In addition, when the first probability estimation result is a Gaussian distribution, the speed of obtaining the first probability variance value from the first probability estimation result is faster than the speed of obtaining the first peak probability from the first probability estimation result. It can be learned that when the first probability estimation result is a Gaussian distribution, the efficiency of determining the first threshold based on the first probability variance value can be higher than the efficiency of determining the first threshold based on the first peak probability.
[0117] For example, feature 1, feature 2, and feature 5 are combined into a third feature of If the first probability variance value σ of feature 1 is 0.6, the first probability variance value σ of feature 2 is 0.7, and the first probability variance value σ of feature 5 is 0.5, then the third feature is determined to form the set. of The smallest first probability variance value σ corresponding to the features in the set (ie, the probability variance value 0.5 for feature 5) is determined to be the first threshold value.
[0118] Since the first threshold is determined based on the feature elements in the first feature map to be coded, i.e., it should be known that the first threshold corresponds to the first feature map to be coded, entropy coding may be performed on the first threshold to facilitate data decoding, and the result of the entropy coding is written into the coded bitstream of the first feature map to be coded.
[0119] S305: Determine whether the feature is a first feature based on the first threshold and the first probability estimation result of each feature.
[0120] For each of the plurality of features in the first feature map to be coded, whether the feature is a first feature can be determined based on a first threshold and a first probability estimation result of the feature. It can be learned that the first threshold is an important determining condition for determining whether the feature is a first feature. Below, a method for determining whether the feature is a first feature will be specifically considered based on a specific method for determining the first threshold.
[0121] Method 1: The first threshold is the third feature element of The maximum first peak probability corresponding to the feature element in the set 1 When the peak probability of the first feature element is equal to or less than the first threshold value, the first feature element determined based on the first threshold value satisfies the following condition: the first peak probability of the first feature element is equal to or less than the first threshold value.
[0122] For example, the plurality of feature elements included in the first encoding target feature map are feature element 1, feature element 2, feature element 3, feature element 4, and feature element 5. Feature element 1, feature element 2, and feature element 5 are included in the third feature element of Forming a set, the third characteristic element of Based on the set, a first threshold is determined to be 75%. In this case, if the first peak probability of feature 1 is 70% and less than the first threshold, the first peak probability of feature 2 is 65% and less than the first threshold, the first peak probability of feature 3 is 80% and greater than the first threshold, the first peak probability of feature 4 is 60% and less than the first threshold, and the first peak probability of feature 5 is 75% and equal to the first threshold. In conclusion, feature 1, feature 2, feature 4, and feature 5 are determined to be first features.
[0123] Method 2: The first threshold is the third feature element of When the first probability variance value is the smallest among the first probability variance values corresponding to the feature elements in the set, the first feature element determined based on the first threshold satisfies the condition that the first probability variance value of the first feature element is greater than or equal to the first threshold.
[0124] For example, the plurality of feature elements included in the first encoding target feature map are feature element 1, feature element 2, feature element 3, feature element 4, and feature element 5. Feature element 1, feature element 2, and feature element 5 are included in the third feature element of Forming a set, the third characteristic element of Based on the set, the first threshold is determined to be 0.5. In this case, if the first peak probability of feature 1 is 0.6 and is greater than the first threshold, the first peak probability of feature 2 is 0.7 and is greater than the first threshold, the first peak probability of feature 3 is 0.4 and is less than the first threshold, the first peak probability of feature 4 is 0.75 and is greater than the first threshold, and the first peak probability of feature 5 is 0.5 and is equal to the first threshold. In conclusion, feature 1, feature 2, feature 4, and feature 5 are determined to be first features.
[0125] S306: Only if the feature is the first feature, entropy coding is performed on the first feature.
[0126] Each feature element in the first feature map to be coded is determined, and it is determined whether the feature element is a first feature element. If the feature element is a first feature element, the first feature element is coded, and the coding result of the first feature element is written into the coded bitstream. In other words, it can be understood that entropy coding is performed on all first feature elements in the feature map, and the entropy coding results of all first feature elements are written into the coded bitstream.
[0127] For example, the feature elements included in the first feature map to be coded are feature element 1, feature element 2, feature element 3, feature element 4, and feature element 5. Feature element 1, feature element 2, feature element 4, and feature element 5 are determined to be the first feature elements. In this case, entropy coding is performed on the feature elements 3, but is performed on feature 1, feature 2, feature 4, and feature 5, and the entropy coding results of all first features are written into the coded bitstream.
[0128] It should be noted that if the determination result for each feature in S305 is that the feature is not the first feature, then entropy coding is not performed on any of the features. If the determination result for each feature in S305 is that the feature is the first feature, then entropy coding is performed on each feature, and the entropy coding result for each feature is written into the coded bitstream.
[0129] In a possible implementation, entropy coding may be further performed on the side information of the first feature map to be coded, and the entropy coding result of the side information may be written into the bitstream. Alternatively, the side information of the first feature map to be coded may be transmitted to the decoder side to facilitate subsequent data decoding.
[0130] Decoder side: FIG. 5 shows a feature map according to one embodiment of the present application. Decryption 1 is a schematic flow chart of the method. Map Decryption The method procedure includes steps S501 to S504.
[0131] S501: Obtain a bitstream of a feature map to be decoded, where the feature map to be decoded includes a plurality of feature elements.
[0132] revenge The bitstream of the feature map to be decoded can be understood as the encoded bitstream obtained in S306. The feature map to be decoded is a feature map obtained after data decoding is performed on the bitstream. The feature map to be decoded includes a plurality of feature elements. The plurality of feature elements includes a first feature element of Set and second feature element of The first feature element is divided into two parts: of The set is the set of features on which entropy coding is performed in the feature map coding stage of FIG. 3. of The set is a set of features for which no entropy coding is performed in the feature map coding stage of FIG.
[0133] In a possible implementation, the first feature of The set is either empty or contains the second characteristic element of The set is the empty set. of The set is an empty set, i.e., entropy coding is not performed on any of the features in the feature map coding stage of FIG. 3. of The set is an empty set, i.e., in the feature map encoding stage of FIG. 3, entropy encoding is performed on each feature element.
[0134] S502: Obtain a first probability estimation result corresponding to each of a plurality of feature elements based on the bitstream of the feature map to be decoded, where the first probability estimation result includes a first peak probability.
[0135] Entropy decoding is performed on the bitstream of the feature map to be decoded. Further, a first probability estimation result corresponding to each of the plurality of feature elements may be obtained based on the entropy decoding result. The first probability estimation result includes a first peak probability.
[0136] In one possible implementation, side information corresponding to the feature map to be decoded is obtained based on the bitstream of the feature map to be decoded, and a first probability estimation result corresponding to each feature element is obtained based on the side information.
[0137] Specifically, the bitstream of the feature map to be decoded includes the entropy coding result of the side information. Therefore, entropy decoding may be performed on the bitstream of the feature map to be decoded, and the obtained entropy decoding result includes the side information corresponding to the feature map to be decoded. Furthermore, as shown in Figure 4a, the side information is used as an input to the probability estimation module 103 of Figure 1, and the output from the probability estimation module 103 is of A feature element in the set and a second feature element of The first probability estimate for the feature set.
[0138] For example, see Figure 2b for the first probability estimation result of the feature.
number
[0139] The probability estimation module 103 may be a probability estimation network, which may use an RNN, a CNN, a variant of an RNN, a variant of a CNN, or another deep neural network (or a variant of another deep neural network). FIG. 4b is a schematic diagram of the structure of the probability estimation network. In FIG. 4b, the probability estimation network is a convolutional network, which includes five network layers: three convolutional layers and two nonlinear activation layers. The probability estimation module 103 may alternatively be implemented according to a non-network conventional probability estimation method. Probability estimation methods include, but are not limited to, statistical methods such as maximum likelihood estimation, maximum a posteriori estimation, and maximum likelihood estimation.
[0140] S503: Selecting a first feature element from the plurality of feature elements based on a first threshold value and a first peak probability corresponding to each feature element. of Set and second feature element of Determine the set.
[0141] First characteristic element of Set and second feature element of The set is determined from the plurality of feature elements in the feature map to be decoded based on a numerical relationship between a first threshold and a first peak probability corresponding to each feature element. The first threshold may be determined through negotiation between a device corresponding to the feature map encoding method and a device corresponding to the feature map decoding method, or may be set based on empirical values. Alternatively, the first threshold may be obtained based on the bitstream of the feature map to be decoded.
[0142] Specifically, the first threshold is the third characteristic element set in the method 1 in S402. of The first peak probability may be the largest in the set. In this case, for each feature in the feature map to be decoded, if the first peak probability of the feature is greater than the first threshold, the feature is classified as a second feature (i.e., ofAlternatively, if the first peak probability of the feature is less than or equal to the first threshold, the feature is determined to be a first feature (i.e., a first feature in the set). of The element is determined to be a feature element in the set.
[0143] For example, the first threshold is 75%, and the multiple feature elements of the feature map to be decoded are feature 1, feature 2, feature 3, feature 4, and feature 5. The first peak probability of feature 1 is 70%, which is less than the first threshold, the first peak probability of feature 2 is 65%, which is less than the first threshold, the first peak probability of feature 3 is 80%, which is greater than the first threshold, the first peak probability of feature 4 is 60%, which is less than the first threshold, and the first peak probability of feature 5 is 75%, which is equal to the first threshold. As a result, feature 1, feature 2, feature 4, and feature 5 are determined to be first feature elements. As a result, feature 1, feature 2, feature 4, and feature 5 are determined to be first feature elements. of The feature 3 is determined to be a feature in the set, and the second feature of It is determined to be a feature element in the set.
[0144] In some cases, the first probability estimation result has a Gaussian distribution, and the first probability estimation result further includes a first probability variance value. In this case, S 50 Optional implementation of item 3 selects a first feature element from the plurality of features based on a first threshold value and a first probability variance value for each feature element. of Set and second feature element of Specifically, the first threshold is determined by the third characteristic element set in the method 2 in S402. of Further, for each feature element in the feature map to be decoded, if the first probability variance value of the feature element is less than the first threshold, the feature element is classified as a second feature element (i.e., the second feature element ofIf the first probability variance value of the feature is equal to or greater than the first threshold, the feature is determined to be a first feature (i.e., a first feature of The element is determined to be a feature element in the set.
[0145] For example, the first threshold is 0.5, and the multiple features included in the first feature map to be coded are feature 1, feature 2, feature 3, feature 4, and feature 5. The first peak probability of feature 1 is 0.6, which is greater than the first threshold; the first peak probability of feature 2 is 0.7, which is greater than the first threshold; the first peak probability of feature 3 is 0.4, which is less than the first threshold; the first peak probability of feature 4 is 0.75, which is greater than the first threshold; and the first peak probability of feature 5 is 0.5, which is equal to the first threshold. In conclusion, feature 1, feature 2, feature 4, and feature 5 are the first feature of The feature 3 is determined to be a feature in the set, and the second feature of It is determined to be a feature element in the set.
[0146] S504: First characteristic element of Set and second feature element of Based on the set, a decoded feature map is obtained.
[0147] In other words, the value of the decoded feature map is of The numeric value of each feature in the set and the second feature of The probability estimate is obtained based on the first probability estimate of each feature in the set.
[0148] In a possible implementation, entropy decoding is performed on the first probability estimation result corresponding to the first feature element, (first feature element ofA numerical value of a first feature element (understood as a general term for features in a set) is obtained. The first probability estimation result includes a first peak probability and a feature value corresponding to the first peak probability. Furthermore, a numerical value of a second feature element (understood as a general term for features in a set) is obtained based on the feature value corresponding to the first peak probability of the second feature element. of (This is understood as a general term for the features in the set). In other words, the first feature of To get the values of all features in the set, use the first feature of It can be understood that entropy decoding is performed on the first probability estimation results corresponding to all features in the set. of The numerical values of all the feature elements in the set are obtained based on the feature values corresponding to the first peak probabilities of all the feature elements in the second feature element, and the entropy decoding is performed on the second feature element. of It need not be performed for every feature in the set.
[0149] For example, data decoding is performed on the feature map to be decoded, i.e., the numerical value of each feature element is obtained. The feature elements in the feature map to be decoded are feature element 1, feature element 2, feature element 3, feature element 4, and feature element 5. Feature element 1, feature element 2, feature element 4, and feature element 5 are the first feature element. of The feature 3 is determined to be a feature in the set, and the second feature of 1 to obtain a value of feature 1, a value of feature 2, a value of feature 4, and a value of feature 5. The feature value corresponding to the first peak probability in the first probability estimation result of feature 3 is determined to be the value of feature 3 in the feature map to be decoded. In this way, the value of feature 1, the value of feature 2, the value of feature 3, the value of feature 4, and the value of feature 5 are combined into a value of the feature map to be decoded.
[0150] First characteristic element of It should be noted that if the set is empty (i.e., entropy coding is not performed on any of the features), the value of the decoded feature map can be obtained based on the first probability estimation result of each feature (herein, the feature value corresponding to the first peak probability in the first probability estimation result is referred to). of If the set is empty (i.e., entropy coding is performed on each feature element), entropy decoding is performed on the first probability estimation results corresponding to each feature element to obtain decoded feature map values.
[0151] 3 for determining whether the entropy coding process should be skipped for a feature based on the peak probability of the probability estimation result corresponding to the feature, compared with determining whether coding should be performed for the feature based on the probability corresponding to a fixed value in the probability estimation result corresponding to each feature, can improve the reliability of the determination result (whether entropy coding should be performed for the feature), significantly reduce the number of elements for performing entropy coding, and reduce the complexity of entropy coding. In addition, as shown in FIG. 5, the reliability of using the feature value of the first probability peak of the feature not subjected to entropy coding (i.e., the second feature) as the numerical value of the second feature to form the value of the feature map to be decoded is better than that of the prior art, which replaces the numerical value of the second feature with a fixed value to form the value of the feature map to be decoded, thereby further improving the data decoding accuracy and performance of the data encoding and decoding method.
[0152] Encoder Side: Figure 6a is a schematic flowchart of another feature map encoding method according to an embodiment of the present application. The steps of the feature map encoding method include: S601 to S607.
[0153] S601: Obtain a first encoding target feature map, where the first encoding target feature map includes a plurality of feature elements.
[0154] For a specific implementation of S601, please refer to the description of the specific implementation of S301, and the details will not be described again in this specification.
[0155] S602: Based on the first encoding target feature map, side information of the first encoding target feature map and second context information of each feature element are obtained.
[0156] For a specific implementation of obtaining the side information of the first encoding target feature map, please refer to the description of the specific implementation of S302, and the details will not be described again in this specification.
[0157] The second context information may be obtained from the first feature map to be encoded via a network module, which may be an RNN or a modified RNN network. The second context information may be understood as a feature element of the feature element (or a value of the feature element) within a predetermined range in the first feature map to be encoded.
[0158] S603: Obtain a second probability estimation result for each feature element based on the side information and the second context information.
[0159] As shown in Figure 6b, the side information and the second context information are used as inputs to the probability estimation module 103 of Figure 1, and the output from the probability estimation module 103 is the second probability estimation result of each feature element. For a specific description of the probability estimation module 103, please refer to S303. The form of the second probability estimation result includes a Gaussian distribution or another form of probability distribution (including, but not limited to, a Laplace distribution or a Gaussian mixture distribution). The second probability of the feature element Estimate A schematic diagram of the results is shown in Fig. 2b. Estimate The results are the same as in the schematic diagram, and the details will not be explained again here.
[0160] S604: Second probability of each feature element Estimate A first threshold is determined based on the result.
[0161] In one possible implementation, a third feature element is selected from the plurality of feature elements in the first feature map to be encoded based on the second probability estimation result of each feature element in the first feature map to be encoded. of Furthermore, the first threshold is set to a third characteristic element. of The probability of the third feature is determined based on the second probability estimate of all features in the set. of The specific method for determining the first threshold based on the second probability estimation result of each feature in the set is shown in FIG. 4c. of Please refer to the specific manner of determining the first threshold value based on the first probability estimation result of each feature element in the set, the details of which will not be described again in this specification.
[0162] S605: Determine a first probability estimation result for each feature element in the first feature map to be coded based on the side information of the feature element and the first context information.
[0163] The first context information is a feature element corresponding to the feature element and within a predetermined range in the second feature map to be encoded; The values of the second feature map to be encoded include a numerical value of the first feature element and a feature value corresponding to a first peak probability of the second feature element, where the second feature element is a feature element other than the first feature element in the first feature map to be encoded. It should be understood that the amount of feature elements included in the first feature map to be encoded is the same as the amount of feature elements included in the second feature map to be encoded, and the values of the first feature map to be encoded are different from the values of the second feature map to be encoded. The second feature map to be encoded can be understood as a feature map obtained after the first feature map to be encoded is decoded (i.e., a feature map to be decoded in this application). The first context information describes the relationship between the feature elements in the second feature map to be encoded, and the second context information describes the relationship between the feature elements in the first feature map to be encoded.
[0164] For example, the features included in the first feature map to be coded are feature 1, feature 2, feature 3, ..., and feature m. After the first threshold is obtained based on the specific description scheme of S604, alternative probability estimation and entropy coding are performed on feature 1, feature 2, feature 3, feature 4, and feature 5. That is, it can be understood that probability estimation and entropy coding are performed on feature 1 first. Because feature 1 is the first feature to be entropy coded, the first context information for feature 1 is empty. In this case, to obtain a first probability estimation result corresponding to feature 1, only probability estimation needs to be performed on feature 1 based on the side information. Furthermore, whether feature 1 is the first feature is determined based on the first probability estimation result and the first threshold. Only if feature 1 is the first feature is entropy coding performed on feature 1, and the value of feature 1 in the second feature map to be coded is determined. Next, for feature 2, a first probability estimation result for feature 2 is estimated based on the side information and the first context information (which in this case can be understood as the numerical value of the first feature in the second feature map to be coded). Furthermore, whether feature 2 is the first feature is determined based on the first probability estimation result and a first threshold. Only if feature 2 is the first feature is entropy coding performed on feature 2 and the numerical value of feature 2 in the second feature map to be coded determined. Next, for feature 3, a first probability estimation result for feature 3 is estimated based on the side information and the first context information (which in this case can be understood as the numerical value of the first feature in the second feature map to be coded and the numerical value of the second feature in the second feature map to be coded). Further, whether feature 3 is the first feature is determined based on the first probability estimation result and the first threshold, and only if feature 3 is the first feature is entropy coded and the value of feature 3 in the second feature map to be coded is determined. The rest can be inferred by analogy until the probabilities of all features in the first feature map to be coded are estimated.
[0165] S606: Determine whether the feature is a first feature based on the first probability estimation result of the feature and a first threshold.
[0166] S607: Only if the feature is the first feature, entropy coding is performed on the first feature.
[0167] For specific implementations of S606 and S607, please refer to the description of specific implementations of S305 and S306, and the details will not be described again in this specification.
[0168] For any feature element in the feature map, the probability estimation result for determining whether the feature element is a first feature element (i.e., a feature element requiring entropy coding) is denoted as the first probability estimation result of the feature element, and the probability for determining the first threshold is Estimate It should be understood that the result is shown as the second probability estimation result. In the feature map encoding method shown in Figure 6a, the first probability estimation result of the feature element is different from the second probability estimation result of the feature element. However, in the feature map encoding method shown in Figure 3, since no context features are introduced for probability estimation, the first probability estimation result of the feature element is the same as the second probability estimation result of the feature element.
[0169] Decoder side: Figure 7a is a schematic flowchart of a feature map decoding method according to an embodiment of the present application. The steps of the feature map decoding method include: S701 to S706.
[0170] S701: Obtain a bitstream of a feature map to be decoded, where the feature map to be decoded includes a plurality of feature elements.
[0171] For a specific implementation of S701, please refer to the description of the specific implementation of S501, and the details will not be described again in this specification.
[0172] S702: Obtain side information corresponding to the feature map to be decoded based on the bitstream of the feature map to be decoded.
[0173] In one possible implementation, side information corresponding to the feature map to be decoded is obtained based on the bitstream of the feature map to be decoded, and a first probability estimation result corresponding to each feature element is obtained based on the side information.
[0174] Specifically, the bitstream of the feature map to be decoded includes the entropy encoding result of the side information, so that entropy decoding can be performed on the bitstream of the feature map to be decoded, and the obtained entropy decoding result includes the side information corresponding to the feature map to be decoded.
[0175] S703: Estimate a first probability estimation result for each feature based on the side information and the first context information of the feature.
[0176] The first context information is 、 Characteristic elements Corresponding to, The feature elements are within a predetermined range in the feature map to be decoded (i.e., the second feature map to be encoded in S605). In this case, it should be noted that probability estimation and entropy decoding are performed sequentially and alternately on the feature elements in the feature map to be decoded.
[0177] For example, the features in the feature map to be decoded are feature 1, feature 2, feature 3, ..., and feature m. First, probability estimation and entropy decoding are performed on feature 1. Because feature 1 is the first feature to be entropy decoded, the first context information for feature 1 is empty. In this case, to obtain a first probability estimation result corresponding to feature 1, only probability estimation needs to be performed on feature 1 based on the side information. Furthermore, it is determined (or is determined) that feature 1 is the first feature or the second feature, and the numerical value of feature 1 in the feature map to be decoded is determined based on the determination result. Next, for feature 2, the first probability estimation result of feature 2 is estimated based on the side information and the first context information (which in this case can be understood as the numerical value of the first feature element in the feature map to be decoded). Furthermore, it is determined (or is determined) whether feature 2 is the first feature or the second feature. The numerical value of feature 2 in the feature map to be decoded is determined based on the determination result. Then, for feature 3, a first probability estimation result of feature 3 is estimated based on the side information and the first context information (which in this case can be understood as the numerical value of the first feature in the feature map to be decoded and the numerical value of the second feature in the feature map to be decoded). Also, feature 3 is determined to be the first feature or the second feature. The numerical value of feature 3 in the feature map to be decoded is determined based on the determination result. The rest can be inferred by analogy until the probabilities of all features are estimated.
[0178] S704: Determine whether the feature is the first feature or the second feature based on the first probability estimation result of the feature and the first threshold.
[0179] For a specific implementation of S704, please refer to the description of the specific implementation of S503, and the details will not be described again in this specification.
[0180] S705: When the feature element is a first feature element, perform entropy decoding based on the first probability estimation result of the first feature element and the bit stream of the feature map to be decoded to obtain the numerical value of the first feature element.
[0181] If the feature determination result is that the feature is the first feature, entropy decoding is performed on the first feature based on the first probability estimation result of the first feature to obtain a numerical value of the first feature in the decoded feature map, where the numerical value of the first feature in the decoded feature map is the same as the numerical value of the first feature in the feature map to be encoded.
[0182] S706: When the feature element is the second feature element, obtain the numerical value of the second feature element based on the first probability estimation result of the second feature element.
[0183] If the determination result for a feature is that the feature is a second feature, the feature value corresponding to the first peak probability of the second feature is determined to be the numerical value of the second feature. In other words, entropy decoding does not need to be performed on the second feature, and the numerical value of the second feature in the decoded feature map may be the same as or different from the numerical value of the second feature in the feature map to be encoded. The value of the decoded feature map is determined based on both the numerical values of all second features and the numerical values of all first features to obtain the decoded feature map.
[0184] Compared with the feature map encoding method provided in Fig. 3, in the feature map encoding method provided in Fig. 6a, probability estimation is performed with reference to context information, thereby improving the accuracy of the probability estimation result corresponding to each feature element, increasing the amount of features skipped in the encoding process, and further improving data encoding efficiency. Compared with the feature map decoding method provided in Fig. 5, in the feature map decoding method provided in Fig. 7a, probability estimation is performed with reference to context information, thereby improving the accuracy of the probability estimation result corresponding to each feature element, improving the reliability of features not entropy coded in the feature map to be decoded (i.e., the second feature element), and improving data decoding performance.
[0185] The applicant calls the feature map encoding and decoding method without skipped coding (i.e., when entropy coding is performed on a feature map to be encoded, the entropy coding process is performed on all feature elements in the feature map to be encoded) the baseline method, and conducts a comparative experiment between the feature map encoding and decoding method provided in FIGS. 6a and 7a (referred to as the feature map encoding and decoding method with dynamic peak-based skipping) and a method for feature map encoding with skipped feature elements based on probabilities corresponding to fixed values in the probability estimation results corresponding to each feature element (referred to as the feature map encoding and decoding method with fixed peak-based skipping).
[0186] For the results of the comparative experiments, see Table 1. Compared with the baseline method, the feature map decoding method using fixed peak-based skipping reduces the amount of data to obtain the same image quality by 0.11%, and our solution reduces the amount of data to obtain the same image quality by 1%.
[0187] [Table 1]
[0188] It can be learned that when the decoded image quality is guaranteed, the technical method provided in this application can reduce a larger amount of data and improve data compression performance (including but not limited to compression rate).
[0189] The present applicant further conducted a comparative experiment between the feature map encoding and decoding methods provided in FIGS. 6A and 7A and the feature map encoding and decoding method that skips based on fixed peaks. The comparative experimental results are shown in FIGS. 7B and 7C. In FIG. 7B, the vertical axis can be understood as the image quality of the reconstructed image, and the horizontal axis is the image compression rate. Generally, as the image compression rate increases, the image quality of the reconstructed picture improves. From FIG. 7B, it can be seen that the curve of the feature map encoding and decoding method that skips based on dynamic peaks (i.e., marked as dynamic peaks in FIG. 7B) almost overlaps with the curve of the feature map encoding method that skips based on fixed peaks (i.e., marked as fixed peaks in FIG. 7B). In other words, when the reconstructed picture quality (i.e., the numerical values of the vertical coordinates are the same), the feature map encoding and decoding method that skips based on dynamic peaks (i.e., marked as dynamic peaks in FIG. 7B) is slightly better than the feature map encoding method that skips based on fixed peaks (i.e., marked as fixed peaks in FIG. 7B). In Figure 7c, the vertical axis is Skippable Characteristic elements ratio where the horizontal axis is the video compression ratio. Typically, as the video compression ratio increases, the number of skippable features decreases. ratiogradually decreases. From FIG. 7c, it can be seen that the curve of the feature map encoding and decoding method that skips based on dynamic peaks (i.e., marked as dynamic peaks in FIG. 7c) is above the curve of the feature map encoding method that skips based on fixed peaks (i.e., marked as fixed peaks in FIG. 7c). In other words, when the image compression ratio (i.e., the numerical values of the horizontal coordinates are the same), the feature map encoding and decoding method that skips based on dynamic peaks (i.e., marked as dynamic peaks in FIG. 7c) skips more features than the feature map encoding method that skips based on fixed peaks (i.e., marked as fixed peaks in FIG. 7c).
[0190] 8 is a schematic diagram of the structure of a feature map encoding device according to the present application. The feature map encoding device may be an integration of the probability estimation module 103 and the data encoding module 104 of FIG. 1. The device: The encoding system includes an acquisition module 80 configured to acquire a first feature map to be encoded, the first feature map to be encoded including a plurality of feature elements; and an encoding module 81 configured to determine a first probability estimation result for each of the plurality of feature elements based on the first feature map to be encoded, the first probability estimation result including a first peak probability, determine whether a feature element is a first feature element based on the first peak probability of each feature element in the first feature map to be encoded, and perform entropy encoding on the first feature element only if the feature element is the first feature element.
[0191] In a possible implementation, the first probability estimation result is a Gaussian distribution and the first peak probability is the mean probability of the Gaussian distribution.
[0192] Alternatively, the first probability estimation result is a Gaussian mixture distribution. The Gaussian mixture distribution includes a plurality of Gaussian distributions. The first peak probability is the maximum value in the average probabilities of the Gaussian distributions, or the first peak probability is calculated based on the average probabilities of the Gaussian distributions and the weights of the Gaussian distributions in the Gaussian mixture distribution.
[0193] In a possible implementation, the encoding module 81 is specifically configured to determine whether the feature is a first feature based on a first threshold and a first peak probability of the feature.
[0194] In a possible implementation, the encoding module 81 determines a second probability estimation result for each of the plurality of features based on the first feature map to be encoded, the second probability estimation result including a second peak probability, and selects a third feature from the plurality of features based on the second probability estimation result for each feature. of Determine the set and the third characteristic element of The method is further configured to determine a first threshold based on the second peak probabilities of all features in the set, and to perform entropy coding on the first threshold.
[0195] In a possible implementation, the first threshold is a third characteristic element. of The second peak probability is the maximum of the second peak probabilities corresponding to the features in the set.
[0196] In a possible implementation, the first peak probability of the first feature is less than or equal to a first threshold.
[0197] In a possible implementation, the second probability estimation result is a Gaussian distribution, and the second probability estimation result further includes a second probability variance value. of The second probability variance value is the smallest of the second probability variance values corresponding to the feature elements in the set.
[0198] In one possible implementation, the first probability estimation result is a Gaussian distribution, and the first probability estimation result further includes a first probability variance value, and the first probability variance value of the first feature element is equal to or greater than a first threshold value.
[0199] In a possible implementation, the second probability estimation result further includes a feature value corresponding to the second peak probability. The encoding module 81 selects a third feature element from the plurality of features based on the preset error, the value of each feature element, and the feature value corresponding to the second peak probability of each feature element. of The device is specifically configured to determine a set of
[0200] In a possible implementation, the third feature of The features in the set are
number
number
[0201] In a possible implementation, the first probability estimation result is the same as the second probability estimation result. The encoding module 81 is specifically configured to obtain side information of the first feature map to be encoded based on the first feature map to be encoded, and perform probability estimation on the side information to obtain a first probability estimation result for each feature element.
[0202] In a possible implementation, the first probability estimation result is different from the second probability estimation result. The encoding module 81 obtains side information of the first feature map to be encoded and second context information of each feature element based on the first feature map to be encoded, and the second context information is 、 Characteristic elements Corresponding to,The feature elements are within a predetermined region range in the first feature map to be coded, and the feature elements are specifically configured to obtain a second probability estimation result for each feature element based on the side information and the second context information.
[0203] In a possible implementation, the encoding module 81 is specifically configured to obtain side information of the first feature map to be encoded based on the first feature map to be encoded, and for any feature element in the first feature map to be encoded, determine a first probability estimation result of the feature element based on the first context information and the side information. The first probability estimation result further includes a feature value corresponding to a first probability peak. The first context information is 、 Characteristic elements Corresponding to, The second feature map to be encoded includes a feature element that is within a predetermined range in the second feature map to be encoded. The value of the second feature map to be encoded includes a numerical value of the first feature element and a feature value corresponding to a first peak probability of the second feature element. The second feature element is a feature element other than the first feature element in the first feature map to be encoded.
[0204] In a possible implementation, the encoding module 81 is further configured to write the entropy encoding results of all first feature elements into an encoded bitstream.
[0205] 9 is a schematic diagram of the structure of a feature map decoding device according to the present application. The feature map decoding device may be an integration of the probability estimation module 103 and the data decoding module 105 in FIG. 1. The feature map decoding device includes: an acquisition module 90 configured to acquire a bitstream of a feature map to be decoded, the feature map to be decoded including a plurality of feature elements, and to acquire a first probability estimation result corresponding to each of the plurality of feature elements based on the bitstream of the feature map to be decoded, the first probability estimation result including a first peak probability; A decoding module 91 for selecting a first feature from the plurality of features based on a first threshold and a first peak probability corresponding to each feature. of Set and second feature element of Determine the set of first characteristic elements of Set and second feature element of and a decoding module 91 configured to obtain a feature map to be decoded based on the set.
[0206] In a possible implementation, the first probability estimation result is a Gaussian distribution and the first peak probability is the mean probability of the Gaussian distribution.
[0207] Alternatively, the first probability estimation result is a Gaussian mixture distribution. The Gaussian mixture distribution includes a plurality of Gaussian distributions. The first peak probability is the maximum value in the average probabilities of the Gaussian distributions, or the first peak probability is calculated based on the average probabilities of the Gaussian distributions and the weights of the Gaussian distributions in the Gaussian mixture distribution.
[0208] In a possible implementation, the value of the feature map to be decoded is a first feature element of The numerical values of all primary features in the set and the secondary features of and the numerical values of all second feature elements in the set.
[0209] In a possible implementation, the first feature of The set is either empty or contains the second characteristic element of The set is the empty set.
[0210] In a possible implementation, the first probability estimation result further includes a feature value corresponding to the first peak probability. The decoding module 91 is further configured to: perform entropy decoding on the first feature based on the first probability estimation result corresponding to the first feature to obtain a numerical value of the first feature, and to obtain a numerical value of the second feature based on the feature value corresponding to the first peak probability of the second feature.
[0211] In a possible implementation, the decoding module 91 is further configured to obtain the first threshold value based on the bitstream of the feature map to be decoded.
[0212] In a possible implementation, the first peak probability of the first feature is less than or equal to a first threshold, and the first peak probability of the second feature is greater than the first threshold.
[0213] In one possible implementation, the first probability estimation result is a Gaussian distribution. The first probability estimation result further includes a first probability variance value. The first probability variance value of the first feature element is equal to or greater than a first threshold, and the first probability variance value of the second feature element is less than the first threshold.
[0214] In a possible implementation, the acquisition module 90 is further configured to acquire side information corresponding to the feature map to be decoded based on the bitstream of the feature map to be decoded, and to acquire a first probability estimation result corresponding to each feature element based on the side information.
[0215] In a possible implementation, the decoding module 91 obtains side information corresponding to the feature map to be decoded based on the bitstream of the feature map to be decoded, and based on the side information and the first context information: revenge The first context information is further configured to estimate a first probability estimate of each feature element for each feature element in the feature map to be coded. 、 Characteristic elements Corresponding to, A feature element that is within a preset range in the feature map to be decoded.
[0216] 10 is a schematic diagram of a hardware structure of a feature map encoding device or a feature map decoding device according to an embodiment of the present application. The device shown in FIG. 10 (which may specifically be a computer device 1000) includes a memory 1001, a processor 1002, a communication interface 1003, and a bus 1004. The memory 1001, the processor 1002, and the communication interface 1003 are communicatively connected to each other via the bus 1004.
[0217] The memory 1001 is a read-only memory (Read- The memory 1001 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1001 may store a program. When the program stored in the memory 1001 is executed by the processor 1002, steps of the feature map encoding method provided in the embodiments of the present application are performed, or steps of the feature map decoding method provided in the embodiments of the present application are performed.
[0218] The processor 1002 may be a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or a - The device may be an ASIC (Application Specific Integrated Circuit), a graphics processing unit (GPU), or one or more ICs, configured to execute associated programs to implement the functions that need to be performed by the units of the feature map encoding device or feature map decoding device in the embodiments of the present application, or to perform the steps of the feature map encoding method in the method embodiments of the present application, or to perform the steps of the feature map decoding method provided in the embodiments of the present application.
[0219] Alternatively, the processor 1002 may be an integrated circuit chip and have signal processing capabilities. In one implementation process, the steps of the feature map encoding method or the feature map decoding method in the present application may be completed through instructions in the form of hardware integrated logic circuits or software in the processor 1002. The processor 1002 may be a general-purpose processor, a digital signal processor (DSP), or a similar processor. or, DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, which can implement or perform the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps in the methods disclosed with reference to the embodiments of the present application may be implemented by hardware. coding may be performed and completed directly by a processor, or coding The implementation and completion may be achieved by using a combination of hardware and software modules within the processor. The software modules may be located in a storage medium established in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is located within the memory 1001. The processor 1002 reads information within the memory 1001 and, in combination with the hardware of the processor 1002, completes the functions required to be performed by the units included in the feature map encoding device or the feature map decoding device in the embodiments of the present application, or implements the feature map encoding method or the feature map decoding method in the method embodiments of the present application.
[0220] The communication interface 1003 uses a transceiver device, such as, but not limited to, a transceiver, to implement communications between the computing device 1000 and another device or communications network.
[0221] The bus 1004 may include a path for transmitting information between components of the computer device 1000 (eg, the memory 1001, the processor 1002, and the communication interface 1003).
[0222] It should be understood that in the feature map encoding apparatus of Figure 8, the acquisition module 80 corresponds to the communication interface 1003 in the computing device 1000, and the encoding module 81 corresponds to the processor 1002 in the computing device 1000. Alternatively, in the feature map decoding apparatus of Figure 9, the acquisition module 90 corresponds to the communication interface 1003 in the computing device 1000, and the decoding module 91 corresponds to the processor 1002 in the computing device 1000.
[0223] It should be noted that for the functions of the functional units in the computer device 1000 described in this embodiment of the present application, please refer to the description of the relevant steps in the above method embodiment, and the details will not be described again in this specification.
[0224] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, which, when executed by a processor, can implement some or all of the steps recorded in any one of the above-described method embodiments and the functions of any functional modules shown in FIG.
[0225] An embodiment of the present application further provides a computer program product. When the computer program product runs on a computer or a processor, the computer or the processor is enabled to perform one or more steps of any one of the above-mentioned methods. When the above-mentioned modules in the device are implemented in the form of software functional units and sold or used as independent products, the modules can be stored in a computer-readable storage medium.
[0226] In the above-mentioned embodiments, the descriptions in the embodiments have their own focus. For parts not described in detail in one embodiment, please refer to the related descriptions in other embodiments. It should be understood that the sequence numbers of the above-mentioned processes do not mean the execution sequence in various embodiments of the present application. The execution order of the processes should be determined according to the functions and internal logic of the processes, and should not be interpreted as any limitation on the implementation process of the embodiments of the present application.
[0227] Those skilled in the art will appreciate that the functions described with reference to the various illustrative logical blocks, modules, and algorithm steps disclosed and described herein may be implemented by hardware, software, firmware, or any combination thereof. If implemented by software, the functions described with reference to the illustrative logical blocks, modules, and steps may be stored on or transmitted via a computer-readable medium as one or more instructions or code and determined by a hardware-based processing unit. Computer-readable media may include computer-readable storage media corresponding to tangible media, such as data storage media, or any communication medium that facilitates the transfer of a computer program from one place to another (e.g., according to a communication protocol). In this manner, computer-readable media may generally correspond to (1) non-transitory tangible computer-readable storage media or (2) a communication medium, such as a signal or carrier. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described herein. A computer program product may include computer-readable media.
[0228] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can store the required program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the instructions are transmitted from a website, server, or another remote source over coaxial cable, fiber optic, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, or microwave, the coaxial cable, fiber optic, twisted pair, DSL, or wireless technologies such as infrared, radio, or microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other transitory media, but actually refer to non-transitory tangible media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), and Blu-ray discs. Disks typically reproduce data magnetically, while discs reproduce data optically by using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0229] The instructions may be determined by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or equivalent integrated circuits or discrete logic circuits. Accordingly, the term "processor" as used herein may refer to the foregoing structure or any other structure that may be applied to an implementation of the techniques described herein. Furthermore, in some aspects, the functionality described with reference to the exemplary logic blocks, modules, and steps described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Furthermore, the techniques may be implemented entirely in one or more circuits or logic elements.
[0230] The technology in this application may be implemented in a variety of apparatuses or devices, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Although various components, modules, or units have been described in this application to emphasize functional aspects of an apparatus configured to determine the disclosed technology, those components, modules, or units do not necessarily require realization by different hardware units. In practice, as described above, the various units may be combined into a codec hardware unit in combination with appropriate software and / or firmware, or may be provided by interoperable hardware units (including one or more processors as described above).
[0231] The foregoing description is merely an exemplary specific implementation of the present application and is not intended to limit the scope of protection of the present application. Any variations or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application shall fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims. [Explanation of symbols]
[0232] 1. Characteristic elements 2. Characteristic elements 3. Characteristic elements 4. Characteristic elements 5. Characteristic elements 80 Acquisition Module 81 Encoding Module 90 Acquisition Module 91 Decryption Module 101 Data Acquisition Module 102 Feature Extraction Module 103 Probability Estimation Module 104 Data Encoding Module 105 Data Decryption Module 106 Data Reconstruction Module 107 Display Module 1000 computer devices 1001 memory 1002 processor 1003 Communication Interface 1004 Bus
Claims
1. 1. A feature map decoding method, the method comprising: obtaining a bitstream of a feature map to be decoded, the feature map to be decoded comprising a plurality of feature elements; obtaining a first probability estimation result corresponding to each of the plurality of feature elements based on the bitstream of the feature map to be decoded, the first probability estimation result including a first peak probability; determining a first set of features and a second set of features from the plurality of features based on a first threshold and the first peak probability corresponding to each feature; and obtaining a decoded feature map based on the first set of features and the second set of features.
2. 2. The method of claim 1, wherein the first probability estimation result is a Gaussian distribution and the first peak probability is an average probability of the Gaussian distribution, or the first probability estimation result is a Gaussian mixture distribution, the Gaussian mixture distribution includes a plurality of Gaussian distributions and the first peak probability is a maximum value among the average probabilities of the Gaussian distributions, or the first peak probability is calculated based on the average probability of the Gaussian distribution and weights of the Gaussian distributions in the Gaussian mixture distribution.
3. 3. The method of claim 1, wherein the decoded feature map values include the numerical values of all first features in the first set of features and the numerical values of all second features in the second set of features.
4. The method of claim 3 , wherein the first set of features is an empty set or the second set of features is an empty set.
5. The first probability estimation result further includes a feature value corresponding to the first peak probability, and the method further comprises: performing entropy decoding on the first feature based on a first probability estimation result corresponding to the first feature to obtain the numerical value of the first feature; The method of claim 3 , further comprising: obtaining the numerical value of the second feature based on a feature value corresponding to a first peak probability of the second feature.
6. Prior to the step of determining a first set of features and a second set of features from the plurality of features based on a first threshold and the first peak probability corresponding to each feature, the method further comprises: The method of claim 1 or 2, further comprising: obtaining the first threshold value based on the bitstream of the feature map to be decoded.
7. The method of claim 1 or 2, wherein a first peak probability of the first feature is less than or equal to the first threshold, and a first peak probability of the second feature is greater than the first threshold.
8. The step of obtaining a first probability estimation result corresponding to each of the plurality of feature elements based on the bitstream of the feature map to be decoded includes: obtaining side information corresponding to the feature map to be decoded based on the bitstream of the feature map to be decoded; and obtaining the first probability estimation result corresponding to each feature element based on the side information.
9. The step of obtaining a first probability estimation result corresponding to each of the plurality of feature elements based on the bitstream of the feature map to be decoded includes: obtaining side information corresponding to the feature map to be decoded based on the bitstream of the feature map to be decoded; and estimating the first probability estimate for each feature element in the feature map to be decoded based on the side information and first context information, wherein the first context information is a feature element of the feature element and is a feature element within a preset range in the feature map to be decoded.
10. 1. A feature map encoding method, the method comprising: obtaining a first encoding target feature map, the first encoding target feature map including a plurality of feature elements; determining a first probability estimate for each of the plurality of features based on the first feature map to be coded, the first probability estimate including a first peak probability; determining whether each feature in the first feature map to be coded is a first feature based on the first peak probability of the feature; performing entropy coding on the first feature element only if the feature element is the first feature element.
11. The step of determining whether each feature element in the first feature map to be coded is a first feature element based on the first peak probability of the feature element includes: The method of claim 10 , comprising determining whether the feature is the first feature based on a first threshold and the first peak probability of the feature.
12. The method comprises: determining a second probability estimate for each of the plurality of features based on the first feature map, the second probability estimate including a second peak probability; determining a third set of features from the plurality of features based on the second probability estimate for each feature; determining the first threshold based on second peak probabilities of all features in the third set of features; The method of claim 11 , further comprising: performing entropy coding on the first threshold value.
13. The method of claim 12 , wherein the first threshold is a maximum second peak probability among the second peak probabilities corresponding to the features in the third set of features.
14. The method of claim 13 , wherein a first peak probability of the first feature is less than or equal to the first threshold.
15. the second probability estimation result further includes a feature value corresponding to the second peak probability, and the step of determining a third set of feature elements from the plurality of feature elements based on the second probability estimation result for each feature element includes:
13. The method of claim 12, further comprising determining the third set of features from the plurality of features based on a preset error, a numerical value of each feature, and the feature value corresponding to the second peak probability for each feature.
16. The features in the third set of features are: [Equation 1] It has the characteristics of [Equation 2] 16. The method of claim 15, wherein p(x, y, i) is the numeric value of the feature, p(x, y, i) is the feature value corresponding to the second peak probability of the feature, and TH_2 is the preset error.
17. the first probability estimation result is the same as the second probability estimation result, and the step of determining the first probability estimation result for each of the plurality of feature elements based on the first encoding target feature map includes: obtaining side information of the first encoding target feature map based on the first encoding target feature map; and performing probability estimation on the side information to obtain the first probability estimation result for each feature.
18. the first probability estimation result is different from the second probability estimation result, and the step of determining the second probability estimation result for each of the plurality of feature elements based on the first encoding target feature map comprises: obtaining side information of the first feature map to be coded and second context information of each feature element based on the first feature map to be coded, wherein the second context information is a feature element of the feature element and is a feature element within a preset region range in the first feature map to be coded; and obtaining the second probability estimate for each feature based on the side information and the second context information.
19. determining a first probability estimate for each of the plurality of feature elements based on the first encoding target feature map, obtaining the side information of the first encoding target feature map based on the first encoding target feature map; and determining, for any feature element in the first feature map to be coded, a first probability estimation result for the feature element based on first context information and the side information, the first probability estimation result further including a feature value corresponding to the first peak probability, the first context information being a feature element corresponding to the feature element and within a predetermined range in a second feature map to be coded, values of the second feature map to be coded including a numerical value of the first feature element and a feature value corresponding to the first peak probability of a second feature element, the second feature element being a feature element other than the first feature element in the first feature map to be coded.
20. The method comprises: The method according to claim 10 or 11, further comprising the step of writing entropy coding results of all the first features into a coded bitstream.
21. A feature map decoding device, comprising: an acquisition module configured to acquire a bitstream of a feature map to be decoded, the feature map to be decoded including a plurality of feature elements, and to acquire a first probability estimation result corresponding to each of the plurality of feature elements based on the bitstream of the feature map to be decoded, the first probability estimation result including a first peak probability; a decoding module configured to determine a first set of feature elements and a second set of feature elements from the plurality of feature elements based on a first threshold and the first peak probability corresponding to each feature element, and to obtain the feature map to be decoded based on the first set of feature elements and the second set of feature elements.
22. 1. A feature map encoding device, comprising: an acquisition module configured to acquire a first encoding target feature map, the first encoding target feature map including a plurality of feature elements; and an encoding module configured to determine a first probability estimation result for each of the plurality of features based on the first feature map to be encoded, the first probability estimation result including a first peak probability, determine whether each feature in the first feature map to be encoded is a first feature based on the first peak probability of the feature, and perform entropy encoding on the first feature only if the feature is the first feature.
23. A decoder comprising processing circuitry configured to perform the method according to claim 1 or 2.
24. An encoder comprising a processing circuit configured to perform the method according to claim 10 or 11.
25. 12. A computer program comprising program code which, when determined by a computer or processor, determines the method according to claim 1 or 2 or the method according to claim 10 or 11.
26. A decoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing a program determined by the processor, the program, when determined by the processor, enabling the decoder to perform the method of claim 1 or 2.
27. 1. An encoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing a program determined by the processor, the program, when determined by the processor, enabling the encoder to perform the method of claim 10 or 11.
28. A data processor comprising processing circuitry configured to perform the method of claim 1 or 2 or configured to perform the method of claim 10 or 11.
29. A non-transitory computer-readable storage medium containing program code, the program code being determined by a computer device to implement the method of claim 1 or 2 or the method of claim 10 or 11.
Citation Information
Patent Citations
Image encoding device, probability model generating apparatus, and image compression system
JP2020191631A
Methods And Apparatuses For Learned Image Compression
US20200160565A1