Feature data encoding method and apparatus
By dynamically determining entropy coding based on probability estimation, the method reduces encoding/decoding complexity and maintains compression efficiency in picture and audio data processing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2026-02-20
- Publication Date
- 2026-04-23
AI Technical Summary
Existing picture and audio compression technologies face challenges in achieving high compression efficiency with minimal quality loss, particularly due to the complexity of entropy coding processes in deep learning-based methods.
A method and apparatus that dynamically determine whether to perform entropy coding on feature elements based on probability estimation results, skipping encoding/decoding for elements that do not meet predetermined conditions, thereby reducing the complexity of encoding and decoding processes.
Significantly reduces the complexity of entropy coding/decoding by selectively applying these processes only to feature elements that meet specific probability criteria, maintaining compression efficiency while minimizing computational overhead.
Smart Images

Figure 2026069706000001_ABST
Abstract
Description
[Technical Field]
[0001] Embodiments of the present invention relate to the field of picture or audio compression technology based on artificial intelligence (AI), and more particularly to methods and apparatus for feature data encoding and decoding. [Background technology]
[0002] Picture or audio encoding and decoding (abbreviated as encoding and decoding) are widely used in digital picture or audio applications such as digital television broadcasting, transmission of pictures or audio over the internet and mobile networks, real-time conversation applications such as video or voice chat and video or audio conferencing, DVD and Blu-ray discs, picture or audio content capture and editing systems, and secure applications for camcorders. A video contains multiple frames of a picture. Therefore, a picture in this application may be a single picture or a picture within a video.
[0003] Even short videos require a considerable amount of video data to depict, which can pose a challenge when data is to be streamed or transmitted over networks with limited bandwidth. Therefore, picture (or audio) data is generally compressed before being transmitted over modern telecommunication networks. The size of the picture (or audio) data can also be a problem when it is stored on a storage device, as memory resources can be limited. Picture (or audio) compression devices often use source-side software and / or hardware to encode the picture (or audio) data before transmission or storage. This reduces the amount of data required to represent the digital picture (or audio). The compressed data is then received at the destination by a picture (or audio) decompression device. Due to limited network resources and the ever-increasing demand for higher picture (or audio) quality, improved compression and decompression techniques that improve the compression ratio with little or no sacrifice in picture (or audio) quality are desirable.
[0004] In recent years, deep learning has gained popularity in the field of picture (or audio) coding and decoding. For example, Google has hosted the CLIC (Challenge on Learned Image Compression) competition at the CVPR (IEEE Conference on Computer Vision and Pattern Recognition) for several consecutive years. CLIC focuses on using deep neural networks to improve picture compression efficiency. A picture challenge category was also added to CLIC2020. Based on the performance evaluation of the solutions in that competition, the overall compression efficiency of current picture coding and decoding solutions based on deep learning techniques is equivalent to that of the latest generation video-picture coding and decoding standard, VVC (Versatile Video Coding), and has a unique advantage in improving the quality perceived by the user.
[0005] The VVC video standard was completed in June 2020. This standard includes almost all technical algorithms that can significantly improve compression efficiency. Therefore, it is difficult to make rapid technological breakthroughs by continuing to research new compression coding algorithms along conventional signal processing paths. Unlike conventional picture algorithms that optimize picture compression modules through manual design, end-to-end AI picture compression is optimized as a whole. Therefore, AI picture compression has a higher compression effect. The variational autoencoder (VAE) method is the mainstream technical solution for current AI picture lossy compression technology. In the current mainstream technical solution, a picture feature map is obtained for the picture to be encoded by using an encoder network, and entropy coding is further performed on the picture feature map. However, the entropy coding process is overly complex. [Overview of the Initiative] [Means for solving the problem]
[0006] This application provides a method and apparatus for encoding and decoding feature data in order to reduce the complexity of encoding and decoding without affecting the performance of encoding and decoding.
[0007] According to the first aspect, A step of obtaining feature data to be encoded, wherein the feature data to be encoded includes multiple feature elements, and each of the multiple feature elements includes a first feature element. A step to obtain the probability estimation result of the first feature element, The steps include determining whether to perform entropy coding on the first feature element based on the probability estimation result of the first feature element, The step of performing entropy coding on the first feature element only when it is determined that entropy coding should be performed on the first feature element. A feature data encoding method is provided, which includes [the following].
[0008] Feature data may include picture feature maps, audio feature variables, or picture feature maps and audio feature variables, and may be one-dimensional, two-dimensional, or multi-dimensional data output by an encoder network, with each data being a feature element. It should be noted that the meanings of feature point and feature element are the same in this application.
[0009] Specifically, the first feature element is any feature element of the feature data to be encoded.
[0010] In one possibility, the probability estimation process to obtain the probability estimation result of the first feature element may be carried out by using a probability estimation network. In another possibility, the probability estimation process may use conventional non-network probability estimation methods to perform probability estimation on the feature data.
[0011] It should be noted that when only side information is used as input for probability estimation, the probability estimation results for feature elements may be output in parallel. When the input for probability estimation includes context information, the probability estimation results for feature elements must be output in serial order. Side information is feature information that is further extracted by inputting feature data into a neural network, and the number of feature elements contained in side information is less than the number of feature elements in the feature data. Optionally, the side information of the feature data may be encoded into a bitstream.
[0012] In the event that the first feature element of the feature data does not satisfy a predetermined condition, entropy coding does not need to be performed on the first feature element of the feature data.
[0013] Specifically, if the current first feature element is the P-th feature element of the feature data, then the determination of the P-th feature element is completed, and entropy coding is performed or not performed based on the determination result. After that, the determination of the (P+1)-th feature element of the feature data is started, and the entropy coding process is performed or not performed based on the determination result. P is a positive integer and is less than M, where M is the total number of feature elements in the feature data. For example, if it is determined that entropy coding is not necessary for the second feature element, then entropy coding for the second feature element is skipped.
[0014] In the aforementioned technical solution, whether or not entropy coding needs to be performed is determined for each feature element to be coded, thereby allowing the entropy coding process to be skipped for some feature elements, significantly reducing the number of elements that need to be coded. In this way, the complexity of entropy coding can be reduced.
[0015] In possible implementations, deciding whether to perform entropy coding on a first feature element involves deciding that entropy coding should be performed on the first feature element if the probability estimation result of the first feature element satisfies a predetermined condition, or deciding that entropy coding is not necessary for the first feature element if the probability estimation result of the first feature element does not satisfy the predetermined condition.
[0016] In a possible implementation, when the probability estimation result of the first feature element is the probability value that the value of the first feature element is k, the pre-set condition is that the probability value that the value of the first feature element is k is less than or equal to a first threshold, where k is an integer.
[0017] k is a value within the range of possible values for the first feature element. For example, the range of values for the first feature element could be [-255, 255]. k may be set to 0, in which case entropy coding is performed on the first feature element whose probability value is 0.5 or less. Entropy coding is not performed on the first feature element whose probability value is greater than 0.5.
[0018] In a possible implementation, the probability value of the first feature element being k is the maximum probability value among all possible values of the first feature element.
[0019] The first threshold selected for a bitstream encoded at a low bitrate is smaller than the first threshold selected for a bitstream encoded at a high bitrate. The specific bitrate is related to the picture resolution and picture content. For example, a publicly available Kodak dataset is used. A bitrate lower than 0.5 bpp is considered low bitrate; otherwise, the bitrate is considered high bitrate.
[0020] In the case of a specific bitrate, the first threshold may be configured based on the actual requirements, but is not limited thereto.
[0021] In the foregoing technical solution, the complexity of entropy encoding can be flexibly reduced by flexibly setting a flexible first threshold based on requirements.
[0022] In a possible implementation, the probability estimation result of the first feature element includes the first parameter and the second parameter of the probability distribution of the first feature element.
[0023] When the probability distribution is a Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean value of the Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the variance of the Gaussian distribution of the first feature element. Alternatively, when the probability distribution is a Laplace distribution, the first parameter of the probability distribution of the first feature element is the location parameter of the Laplace distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the scale parameter of the Laplace distribution of the first feature element. The preset conditions are as follows, that is, the absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to the second threshold, the second parameter of the probability distribution of the first feature element is greater than or equal to the third threshold, or the sum of the second parameter of the probability distribution of the first feature element, the absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to the fourth threshold can be any one of them.
[0024] When the probability distribution is a mixture Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean value of the mixture Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the variance of the mixture Gaussian distribution of the first feature element. The preset conditions are as follows, that is, the sum of the sum of the absolute value of the difference between any variance of the mixture Gaussian distribution of the first feature element and the value k of the first feature element and all the mean values of the mixture Gaussian distribution of the first feature element is greater than or equal to the fifth threshold, The difference between any mean value of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than or equal to the sixth threshold, or The variance of any Gaussian mixture distribution of the first feature element is greater than or equal to the seventh threshold. It could be any one of the following.
[0025] When the probability distribution is an asymmetric Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean value of the asymmetric Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the first and second variances of the asymmetric Gaussian distribution of the first feature element. The pre-set conditions are as follows: The absolute value of the difference between the mean of the asymmetric Gaussian distribution of the first feature element and the value k of the first feature element is greater than or equal to the eighth threshold. The first variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to the ninth threshold, or The second variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to the tenth threshold. It could be any one of the following.
[0026] When the probability distribution of the first feature element is a Gaussian mixture, the range of the determined value for the first feature element is determined. Multiple mean values of the probability distribution of the first feature element do not fall within the range of the determined value for the first feature element.
[0027] When the probability distribution of the first feature element is a Gaussian distribution, the range of the determined value for the first feature element is determined. The mean of the probability distribution of the first feature element is not within the range of the determined value for the first feature element.
[0028] When the probability distribution of the first feature element is a Gaussian distribution, the range of determined values for the first feature element is determined, and this range of determined values includes multiple possible values for the first feature element. The absolute value of the difference between the mean parameter of the Gaussian distribution of the first feature element and each value within the range of determined values for the first feature element is greater than or equal to the 11th threshold, or the variance of the probability distribution of the first feature element is greater than or equal to the 12th threshold.
[0029] The value of the first feature element is not within the range of the determined values for the first feature element.
[0030] The probability value corresponding to the value of the first feature element is less than or equal to the 13th threshold.
[0031] In a possible implementation, the method further includes the steps of: constructing a list of threshold candidates for a first threshold; adding a first threshold to the list of threshold candidates for the first threshold, which has an index number corresponding to the first threshold; and writing the index number of the first threshold to an encoded bitstream, the length of the list of threshold candidates for the first threshold may be set to T, where T is an integer greater than or equal to 1. It can be understood that other thresholds may be constructed in a manner such as constructing the list of threshold candidates for the first threshold, which have a corresponding index number written to an encoded bitstream.
[0032] Specifically, the index number can be written to the bitstream and stored in the sequence header, picture header, slice header, or SEI (supplemental enhancement information), and then sent to the decoder. Alternatively, other methods may be used, and are not limited thereto. The method of constructing the candidate list is not limited.
[0033] Alternatively, decision information can be obtained by inputting the probability estimation results into the generative network. The generative network may be a convolutional network and may contain multiple network layers. Any network layer may be a convolutional layer, a normalization layer, a nonlinear activation layer, or similar.
[0034] In possible implementations, the probability estimation results of the feature data are input to the generative network to obtain decision information for the first feature element. This decision information indicates whether to perform entropy coding on the first feature element.
[0035] In possible implementations, the decision information for feature data is a decision map, which may also be called a decision map. The decision map is preferably a binary map, which may also be called a binary map. The decision information values for feature elements in the binary map are usually 0 or 1. Therefore, when the value corresponding to the position of the first feature element in the decision map is a pre-set value, entropy coding needs to be performed on the first feature element. When the value corresponding to the position of the first feature element in the decision map is not a pre-set value, entropy coding does not need to be performed on the first feature element.
[0036] In possible implementations, the determination information for a feature element in the feature data is a pre-set value. This pre-set value is typically 1. Therefore, when the determination information is a pre-set value, entropy coding must be performed on the first feature element. When the determination information is not a pre-set value, entropy coding is not required for the first feature element. The determination information can be an identifier or the value of an identifier. The decision of whether to perform entropy coding on the first feature element depends on whether the identifier or the value of the identifier is a pre-set value. When the identifier or the value of the identifier is a pre-set value, entropy coding must be performed on the first feature element. When the identifier or the value of the identifier is not a pre-set value, entropy coding is not required for the first feature element. Alternatively, the set of determination information for a feature element in the feature data may be a floating-point number. In other words, the value may be something other than 0 and 1. In this case, a pre-set value may be set. When the value of the determination information for the first feature element is greater than or equal to the pre-set value, it is determined that entropy coding should be performed on the first feature element. When the value of the determination information for the first feature element is smaller than a predetermined value, it is determined that entropy coding does not need to be performed on the first feature element.
[0037] In a possible implementation, the method further includes the steps of obtaining feature data by passing the picture to be encoded through an encoder network, obtaining feature data by rounding the picture to be encoded after it has passed through the encoder network, or obtaining feature data by quantizing and rounding the picture to be encoded after it has passed through the encoder network.
[0038] The encoder network may use an autoencoder structure. The encoder network may be a convolutional neural network. The encoder network may contain multiple subnets, each subnet containing one or more convolutional layers. The network structures between subnets may be the same or different.
[0039] The picture that will be encoded can be the original picture or the residual picture.
[0040] It should be understood that the picture to be encoded may be in RGB format, or in other representation formats such as YUV or RAW. Preprocessing operations may be performed on the picture to be encoded before it is input to the encoder network. These preprocessing operations may include operations such as conversion, block splitting, filtering, and pruning.
[0041] It should be understood that multiple pictures or picture blocks to be encoded are acceptable input into an encoder and decoder network to obtain feature data, either within the same timestamp or at the same moment.
[0042] According to the second aspect, A step of obtaining a bitstream of feature data to be decoded, Steps include: the feature data to be decoded contains multiple feature elements, and each of the multiple feature elements contains a first feature element; A step to obtain the probability estimation result of the first feature element, The steps include determining whether to perform entropy decoding on the first feature element based on the probability estimation result of the first feature element, The step of performing entropy decoding on the first feature element only when it is determined that entropy decoding needs to be performed on the first feature element. A feature data decoding method is provided, which includes the following.
[0043] The first feature element can be understood as any feature element of the feature data to be decoded. After all feature elements of the feature data to be decoded have been determined, and entropy decoding is performed or not performed based on the determination results, the decoded feature data is obtained.
[0044] The decoded feature data may be one-dimensional, two-dimensional, or multi-dimensional, and each data point is a feature element. It should be noted that the meanings of feature points and feature elements are the same in this application.
[0045] Specifically, the first feature element is any feature element of the feature data to be decoded.
[0046] In one possibility, the probability estimation process to obtain the probability estimation result of the first feature element may be carried out by using a probability estimation network. In another possibility, the probability estimation process may use conventional non-network probability estimation methods to perform probability estimation on the feature data.
[0047] It should be noted that when only side information is used as input for probability estimation, the probability estimation results for feature elements may be output in parallel. When the input for probability estimation includes context information, the probability estimation results for feature elements must be output in serial order. The number of feature elements included in the side information is less than the number of feature elements in the feature data.
[0048] Possibly, the bitstream contains side information, which needs to be decoded during the process of decoding the bitstream.
[0049] Specifically, the process for determining each feature element of the feature data includes determining conditions and deciding whether to perform entropy decoding based on the result of the condition determination.
[0050] Possibly, entropy decoding can be performed by using neural networks.
[0051] In another possibility, entropy decoding can be performed through conventional entropy decoding.
[0052] Specifically, if the current first feature element is the P-th feature element of the feature data, then after the determination of the P-th feature element is complete and entropy decoding is performed or not performed based on the determination result, the determination of the (P+1)-th feature element of the feature data is started, and the entropy decoding process is performed or not performed based on the determination result. P is a positive integer, and P is less than M, where M is the total number of feature elements in the feature data. For example, if it is determined that entropy decoding is not necessary for the second feature element, then entropy decoding for the second feature element is skipped.
[0053] In the aforementioned technical solution, whether or not entropy decoding needs to be performed is determined for each feature element to be decoded, thereby allowing the entropy decoding process for some feature elements to be skipped and significantly reducing the number of elements for which entropy decoding needs to be performed. In this way, the complexity of entropy decoding can be reduced.
[0054] In possible implementations, deciding whether to perform entropy decoding on a first feature element of the feature data involves deciding that entropy decoding should be performed on the first feature element if the probability estimation result of the first feature element of the feature data satisfies a predetermined condition, or deciding that entropy decoding is not necessary for the first feature element if the probability estimation result of the first feature element does not satisfy the predetermined condition, and setting the feature value of the first feature element to k, where k is an integer.
[0055] In a possible implementation, when the probability estimation result of the first feature element is the probability value that the value of the first feature element is k, the pre-set condition is that the probability value that the value of the first feature element is k is less than or equal to a first threshold, where k is an integer.
[0056] In terms of probability, if a predetermined condition is not met, the first feature element is set to k. For example, the range of values for the first feature element can be [-255, 255]. k may also be set to 0, in which case entropy coding is performed on the first feature element whose probability value is 0.5 or less. Entropy coding is not performed on the first feature element whose probability value is greater than 0.5.
[0057] In another possibility, when the predetermined conditions are not met, the value of the first feature element is determined by using a list.
[0058] In another possibility, if the predetermined conditions are not met, the first feature element is set to a fixed integer value.
[0059] k is a value within the range of possible values for the first feature element.
[0060] In terms of probability, k is the value corresponding to the maximum probability within the range of all possible values for the first feature element.
[0061] The first threshold selected for a decoded bitstream at a low bitrate is smaller than the first threshold selected for a decoded bitstream at a high bitrate. The specific bitrate is related to the picture resolution and picture content. For example, a publicly available Kodak dataset is used. A bitrate lower than 0.5 bpp is considered low bitrate; otherwise, the bitrate is considered high bitrate.
[0062] In the case of a specific bitrate, the first threshold may be configured based on the actual requirements, but is not limited thereto.
[0063] In the aforementioned technical solutions, the complexity of entropy decoding can be flexibly reduced by flexibly setting a flexible first threshold based on the requirements.
[0064] In possible implementations, the probability estimation result for the first feature element includes the first and second parameters of the probability distribution of the first feature element.
[0065] When the probability distribution is a Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean of the Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the variance of the Gaussian distribution of the first feature element. Alternatively, when the probability distribution is a Laplace distribution, the first parameter of the probability distribution of the first feature element is the position parameter of the Laplace distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the scale parameter of the Laplace distribution of the first feature element. The pre-set conditions are as follows: The absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to the second threshold. The second parameter of the first feature element is greater than or equal to the third threshold, or The sum of the second parameter of the probability distribution of the first feature element and the absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to the fourth threshold. It could be any one of the following.
[0066] When the probability distribution is a Gaussian mixture, the first parameter of the probability distribution of the first feature element is the mean of the Gaussian mixture of the first feature element, and the second parameter of the probability distribution of the first feature element is the variance of the Gaussian mixture of the first feature element. The pre-set conditions are as follows: The sum of the arbitrary variance of the Gaussian mixture distribution of the first feature element and the sum of the absolute differences between all the mean values of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than or equal to the fifth threshold. The difference between any mean of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than the sixth threshold, or The variance of any Gaussian mixture distribution of the first feature element is greater than or equal to the seventh threshold. It could be any one of the following.
[0067] When the probability distribution is an asymmetric Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean value of the asymmetric Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the first and second variances of the asymmetric Gaussian distribution of the first feature element. The pre-set conditions are as follows: The absolute value of the difference between the mean parameter of the asymmetric Gaussian distribution of the first feature element and the value k of the first feature element is greater than the eighth threshold. The first variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to the ninth threshold, or The second variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to the tenth threshold. It could be any one of the following.
[0068] When the probability distribution of the first feature element is a Gaussian mixture, the range of the determined value for the first feature element is determined. Multiple mean values of the probability distribution of the first feature element do not fall within the range of the determined value for the first feature element.
[0069] When the probability distribution of the first feature element is a Gaussian distribution, the range of the determined value for the first feature element is determined. The mean of the probability distribution of the first feature element is not within the range of the determined value for the first feature element.
[0070] When the probability distribution of the first feature element is a Gaussian distribution, the range of determined values for the first feature element is determined, and this range of determined values includes multiple possible values for the first feature element. The absolute value of the difference between the mean parameter of the Gaussian distribution of the first feature element and each value within the range of determined values for the first feature element is greater than or equal to the 11th threshold, or the variance of the probability distribution of the first feature element is greater than or equal to the 12th threshold.
[0071] The value k of the first feature element is not within the range of the determined values for the first feature element.
[0072] The probability value corresponding to the value k of the first feature element is less than or equal to the 13th threshold.
[0073] In possible implementations, a list of candidate thresholds for a first threshold is constructed, the index number of the candidate threshold for the first threshold is obtained by decoding the bitstream, and the value at the position in the candidate threshold for the first threshold that corresponds to the index number of the first threshold is used as the value of the first threshold. The length of the candidate threshold for the first threshold may be set to T, where T is an integer greater than or equal to 1. It can be understood that any other thresholds may be constructed in a manner such as constructing the candidate threshold for the first threshold. The index number corresponding to the threshold may be obtained through decoding, and the value in the constructed list is selected as the threshold based on the index number.
[0074] Alternatively, decision information can be obtained by inputting the probability estimation results into the generative network. The generative network may be a convolutional network and may contain multiple network layers. Any network layer may be a convolutional layer, a normalization layer, a nonlinear activation layer, or similar.
[0075] In possible implementations, the probability estimation results of the feature data are input to the generative network to obtain decision information for the first feature element. This decision information indicates whether to perform entropy decoding on the first feature element.
[0076] In possible implementations, the decision information for feature elements in feature data is a decision map, which may also be called a decision map. The decision map is preferably a binary map, which may also be called a binary map. The values of the decision information for feature elements in the binary map are usually 0 or 1. Therefore, when the value corresponding to the position of the first feature element in the decision map is a pre-set value, entropy decoding needs to be performed on the first feature element. When the value corresponding to the position of the first feature element in the decision map is not a pre-set value, entropy decoding does not need to be performed on the first feature element.
[0077] Alternatively, the set of determination information for feature elements in the feature data may be a floating-point number. In other words, the value may be a value other than 0 and 1. In this case, a pre-set value may be set. When the value of the determination information for the first feature element is greater than or equal to the pre-set value, it is determined that entropy decoding should be performed on the first feature element. When the value of the determination information for the first feature element is less than the pre-set value, it is determined that entropy decoding does not need to be performed on the first feature element.
[0078] In possible implementations, the feature data passes through a decoder network to obtain a reconstructed picture.
[0079] In another possible implementation, machine-ready task data is obtained by passing feature data through a decoder network. Specifically, machine-ready task data is obtained by passing feature data through a machine-ready task module, which includes a target recognition network, a classification network, or a semantic segmentation network.
[0080] According to the third aspect, An acquisition module configured to acquire feature data to be encoded, wherein the feature data to be encoded includes multiple feature elements, the multiple feature elements include a first feature element, and the acquisition module is configured to acquire the probability estimation result of the first feature element. An encoding module configured to determine whether to perform entropy coding on a first feature element based on the probability estimation result of the first feature element, and to perform entropy coding on the first feature element only when it is determined that entropy coding should be performed on the first feature element. A feature data encoding device is provided, which includes [the following].
[0081] For further implementations of the acquisition and encoding modules, please refer to either the first embodiment or the implementation of the first embodiment. Further details are not provided here.
[0082] According to the fourth aspect, An acquisition module configured to acquire a bitstream of feature data to be decoded, wherein the feature data to be decoded includes multiple feature elements, the multiple feature elements include a first feature element, and the acquisition module is configured to acquire the probability estimation result of the first feature element. A decoding module configured to determine whether to perform entropy decoding on a first feature element based on the probability estimation result of the first feature element, and to perform entropy decoding on the first feature element only when it is determined that entropy decoding should be performed on the first feature element. A feature data decoding device is provided, which includes [the necessary components].
[0083] For further implementation details of the acquisition and decryption modules, please refer to either the second embodiment or the implementation of the second embodiment. Further details are not provided here.
[0084] According to a fifth aspect, the application provides an encoder including a processing circuit configured to determine a method according to either the first aspect or one of the first aspects.
[0085] According to a sixth aspect, the application provides a decoder including a processing circuit configured to determine a method according to the second aspect and one of the second aspect.
[0086] According to the seventh aspect, the application provides a computer program product including program code. When the program code is determined in a computer or processor, the program code is used to determine a method according to the first aspect and one of the first aspects, and a method according to the second aspect and one of the second aspects.
[0087] According to the eighth aspect, the application provides an encoder comprising one or more processors and a non-temporary computer-readable storage medium coupled to the processors and storing a program determined by the processors. When the program is determined by the processors, the decoder is enabled to determine a method according to either the first aspect or one of the first aspects.
[0088] According to the ninth aspect, the application provides a decoder comprising one or more processors and a non-temporary computer-readable storage medium coupled to the processors and storing a program determined by the processors. When the program is determined by the processors, the encoder is enabled to determine a method according to either the second aspect or one of the second aspect.
[0089] According to a tenth aspect, the application provides a non-temporary computer-readable storage medium containing program code. When the program code is determined by a computer device, the program code is used to determine a method according to the first aspect and one of the first aspects, and a method according to the second aspect and one of the second aspects.
[0090] According to an eleventh aspect, the present invention relates to an encoding device having the function of performing the behavior according to either the first aspect or an embodiment of the method of the first aspect. The function may be implemented by hardware or by hardware that determines the corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned function. In a possible design, the encoding device includes an acquisition module configured to transform an original picture or residual picture into a feature space by using an encoder network and extract feature data for compression, wherein probabilistic estimation is performed on the feature data to obtain the probabilistic estimation results of the feature elements of the feature data; and an encoding module configured to use the probabilistic estimation results of the feature elements of the feature data to determine, based on certain conditions, whether entropy coding is performed on the feature elements of the feature data, and to complete the coding process for all feature elements of the feature data to obtain an encoded bitstream of the feature data. These modules may determine the corresponding function in the example of the method according to either the first aspect or an example of the method. For further details, see the detailed description in the example of the method. Further details are not described here again.
[0091] According to a twelfth aspect, the present invention relates to a decoding device having the function of performing the behavior according to either a second aspect or an embodiment of the method of the second aspect. The function may be implemented by hardware or by hardware that determines the corresponding software. The hardware or software includes one or more modules corresponding to the function described above. In a possible design, the decoding device includes an acquisition module configured to acquire a bitstream of feature data to be decoded, perform probabilistic estimations based on the bitstream of feature data to be decoded to obtain the probabilistic estimation results of the feature elements of the feature data, and a decoding module configured to use the probabilistic estimation results of the feature elements of the feature data to determine, based on certain conditions, whether entropy decoding is performed on the feature elements of the feature data, complete the decoding process for all feature elements of the feature data to obtain the feature data, and decode the feature data to obtain a reconstructed picture or machine-ready task data. These modules may determine the corresponding function in the example of the method according to either a second aspect or an example of the method of either aspect. For further details, see the detailed description in the example of the method. Further details are not described here again.
[0092] According to the 13th aspect, A step of obtaining feature data to be encoded, wherein the feature data includes multiple feature elements, and each of the multiple feature elements includes a first feature element. The steps include obtaining side information of the feature data, inputting the side information of the feature data into a joint network to obtain determination information for the first feature element, A step of determining whether to perform entropy coding on the first feature element based on the determination information of the first feature element, The step of performing entropy coding on the first feature element only when it is determined that entropy coding should be performed on the first feature element. A feature data encoding method is provided, which includes [the following].
[0093] Feature data is one-dimensional, two-dimensional, or multi-dimensional data output by an encoder network, and each piece of data is a feature element.
[0094] Possibly, side information of feature data can be encoded into a bitstream. Side information is further extracted feature information by inputting the feature data into a neural network, and the number of feature elements contained in the side information is less than the number of feature elements in the feature data.
[0095] The first feature element is any feature element of the feature data.
[0096] Possibly, the set of judgment information for feature elements of feature data can be represented in a manner such as a judgment map. A judgment map is a one-dimensional, two-dimensional, or multi-dimensional picture data, and the size of the judgment map matches the size of the feature data.
[0097] In terms of probability, the joint network further outputs the probability estimation result for the first feature element. The probability estimation result for the first feature element includes the probability value of the first feature element, and / or the first and second parameters of the probability distribution.
[0098] In the aforementioned technical solution, whether or not entropy coding needs to be performed is determined for each feature element to be coded, thereby allowing the entropy coding process to be skipped for some feature elements, significantly reducing the number of elements that need to be coded. In this way, the complexity of entropy coding can be reduced.
[0099] In the possibilities, if the value corresponding to the position of the first feature element in the decision map is a pre-set value, then entropy coding needs to be performed on the first feature element. If the value corresponding to the position of the first feature element in the decision map is not a pre-set value, then entropy coding does not need to be performed on the first feature element.
[0100] According to the 14th aspect, A step of obtaining a bitstream of feature data to be decoded and side information of feature data to be decoded, Steps include: the feature data to be decoded contains multiple feature elements, and each of the multiple feature elements contains a first feature element; The steps include inputting side information of the feature data to be decoded into a joint network to obtain determination information for the first feature element, A step of determining whether to perform entropy decoding on the first feature element based on the determination information of the first feature element, The step of performing entropy decoding on the first feature element only when it is determined that entropy decoding needs to be performed on the first feature element. A feature data decoding method is provided, which includes the following.
[0101] In a possible scenario, the bitstream of feature data that will be decoded is decoded to obtain side information. The number of feature elements contained in the side information is less than the number of feature elements in the feature data.
[0102] The first feature element is any feature element of the feature data.
[0103] In some cases, the judgment information for feature elements of feature data can be represented using methods such as judgment maps. A judgment map is a one-dimensional, two-dimensional, or multi-dimensional picture data, and the size of the judgment map matches the size of the feature data.
[0104] In terms of probability, the joint network further outputs the probability estimation result for the first feature element. The probability estimation result for the first feature element includes the probability value of the first feature element, and / or the first and second parameters of the probability distribution.
[0105] In the possibilities, if the value corresponding to the position of the first feature element in the decision map is a pre-set value, then entropy decoding must be performed on the first feature element. If the value corresponding to the position of the first feature element in the decision map is not a pre-set value, then entropy decoding does not need to be performed on the first feature element, and the feature value of the first feature element is set to k, where k is an integer.
[0106] In the aforementioned technical solution, whether or not entropy decoding needs to be performed is determined for each feature element to be encoded, thereby allowing the entropy decoding process to be skipped for some feature elements, significantly reducing the number of elements for which entropy decoding needs to be performed. In this way, the complexity of entropy decoding can be reduced.
[0107] In existing mainstream end-to-end feature data encoding and decoding solutions, the processes of entropy encoding and decoding or arithmetic encoding and decoding are excessively complex. In this application, information about the probability distribution of feature points in the feature data to be encoded is used to determine whether entropy encoding needs to be performed for each feature element of the feature data to be encoded, and whether entropy decoding needs to be performed for each feature element of the feature data to be decoded, thereby making it possible to skip the entropy encoding and decoding processes for some feature elements and significantly reduce the number of elements that need to be encoded and decoded. This reduces the complexity of encoding and decoding. In another embodiment, the threshold can be flexibly set based on the requirement of the actual bitrate value of the bitstream to control the bitrate value of the generated bitstream.
[0108] Details of one or more embodiments are described in detail in the accompanying drawings and the following description. Other features, purposes, and advantages are evident from the description, drawings, and claims.
[0109] The following describes the accompanying drawings used in embodiments of this application. [Brief explanation of the drawing]
[0110] [Figure 1A] This is an exemplary block diagram of a picture decoding system. [Figure 1B] This is an implementation of the processing circuit for the picture decoding system. [Figure 1C] This is a schematic block diagram of the picture decoding device. [Figure 1D] This is a diagram showing the implementation of the device according to the embodiment of this application. [Figure 2A] This is a system architecture diagram of possible scenarios according to this application. [Figure 2B] This is a system architecture diagram of possible scenarios according to this application. [Figure 3A] This is a schematic block diagram of the encoder. [Figure 3B] This is a schematic block diagram of the encoder. [Figure 3C] This is a schematic block diagram of the encoder. [Figure 3D] This is a schematic block diagram of the encoder. [Figure 4A] This is a schematic diagram of the encoder network unit. [Figure 4B] This is a schematic diagram of the encoder network structure. [Figure 5] This is a schematic diagram of the structure of the coding determination implementation unit. [Figure 6] This is an illustrative output diagram of a joint network. [Figure 7] This is an illustrative output diagram of a generative network. [Figure 8] This is a schematic implementation diagram of the decryption detection implementation. [Figure 9] This is an illustrative diagram of the network structure of a decoder network. [Figure 10A] This is an illustrative diagram of a coding method according to an embodiment of this application. [Figure 10B] This is a schematic block diagram of a picture feature map decoder according to an embodiment of this application. [Figure 11A] This is an illustrative diagram of a coding method according to an embodiment of this application. [Figure 12] This is an illustrative diagram of the network structure of the side information extraction module. [Figure 13A] This is an illustrative diagram of a coding method according to an embodiment of this application. [Figure 13B] This is a schematic block diagram of a picture feature map decoder according to an embodiment of this application. [Figure 14] This is an illustrative diagram of a coding method according to an embodiment of this application. [Figure 15] This is an illustrative diagram of the network structure of a joint network. [Figure 16]This is a schematic block diagram of a picture feature map decoder according to an embodiment of this application. [Figure 17] This is an illustrative diagram of a coding method according to an embodiment of this application. [Figure 18] This is a schematic diagram illustrating an exemplary structure of the encoding device according to this application. [Figure 19] This is a schematic diagram illustrating an exemplary structure of a decoding device according to this application. [Modes for carrying out the invention]
[0111] In embodiments of this application, terms such as “first” and “second” are used solely for distinction and descriptive purposes and should not be understood as indicating or implying relative importance or order. In addition, the terms “include,” “comprise,” and any variations thereof are intended to cover non-exclusive inclusion, such as the inclusion of a set of steps or units. A method, system, product, or device is not necessarily limited to the steps or units explicitly listed, but may include other steps or units that are not explicitly listed and are specific to the process, method, product, or device.
[0112] In this application, it should be understood that “at least one (item)” refers to one or more, and “multiple” refers to two or more. The term “and / or” describes the relational relationship of the related objects and indicates that three relationships may exist. For example, “A and / or B” may indicate the following three cases: only A exists, only B exists, and both A and B exist, where A and B may be singular or plural. The letter “ / ” generally indicates an “or” relationship between the related objects. “At least one of the following items (elements)” or a similar expression refers to any combination of these items, including a single item (element) or any combination of multiple items (elements). For example, at least one of a, b, or c may express a, b, c, “a and b”, “a and c”, “b and c”, or “a, b, and c”, where a, b, and c may be singular or plural.
[0113] Embodiments of this application provide AI-based feature data coding and decoding techniques, and in particular, neural network-based picture feature mapping and / or audio feature variable coding and decoding techniques, specifically, end-to-end based picture feature mapping and / or audio feature variable coding and decoding systems.
[0114] In the field of picture coding, the terms "picture" and "image" can be used as synonyms. Picture coding (or simply coding) consists of two parts: picture encoding and picture decoding. Video is a representation of a sequence of pictures, containing multiple pictures. Picture encoding is determined on the source side and typically involves processing the original video picture (e.g., compressing) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Picture decoding is determined on the destination side and typically involves the reverse process compared to the encoder process to reconstruct the picture. Embodiments referring to "coding" of a picture or audio are understood as "encoding" or "decoding" of a picture or audio. The combination of the encoding and decoding parts is also called encoding and decoding (CODEC).
[0115] In lossless picture coding, the original picture can be reconstructed. In other words, the reconstructed picture has the same quality as the original picture (assuming no transmission loss or other data loss occurs during storage or transmission). In conventional lossy picture coding, further compression is decided, for example through quantization, to reduce the amount of data required to represent the video picture, and the video picture cannot be completely reconstructed on the decoder side. In other words, the quality of the reconstructed video picture is lower or worse than the quality of the original video picture.
[0116] Since the embodiments of this application relate to large-scale applications of neural networks, for ease of understanding, the following will explain the terms and concepts relating to neural networks that may be used in the embodiments of this application.
[0117] (1) Neural Network
[0118] A neural network can include neurons. Neurons take x as input. s And it may be an operation unit that uses an intercept of 1. The output of the operation unit may be as follows:
number
[0119] s = 1, 2, ..., or n, where n is a natural number greater than 1, and W s is x s The weights are , where b is the neuron's bias. f is the neuron's activation function, used to introduce nonlinear features into the neural network in order to convert the input signal within the neuron into an output signal. The output signal of the activation function may be used as the input to the next convolutional layer, and the activation function may be a sigmoid function. A neural network is a network constructed by connecting multiple single neurons together. Specifically, the output of one neuron may be the input of another neuron. The input of each neuron may be connected to the local receptive field of the previous layer in order to extract features of the local receptive field. The local receptive field may be a region containing several neurons.
[0120] (2) Deep neural networks
[0121] A deep neural network (DNN), also known as a multilayer neural network, can be understood as a neural network with multiple hidden layers. A DNN is divided based on the location of different layers. The neural network within a DNN can be classified into three types: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the intermediate layers are hidden layers. The layers are fully connected. Specifically, every neuron in the i-th layer is always connected to every neuron in the (i+1)th layer.
[0122] While DNNs may appear complex, the operations within each layer are not. Simply put, a DNN is based on the following linear relationship:
number
number
number
number
number
number
number
number
[0123] In conclusion, the coefficient from the kth neuron in layer (L-1) to the jth neuron in layer L is:
number
[0124] It should be noted that the input layer has no parameters W. In deep neural networks, more hidden layers allow the network to better describe complex real-world cases. Theoretically, models with more parameters have higher complexity and greater "capacity." This indicates that the model can complete more complex learning tasks. The process of training a deep neural network is the process of learning the weight matrix, and the ultimate goal of training is to obtain the weight matrix (the weight matrix containing the vector W across multiple layers) for all layers in the deep neural network being trained.
[0125] (3) Convolutional Neural Networks
[0126] A convolutional neural network (CNN) is a deep neural network that has a convolutional structure. A convolutional neural network includes a feature extractor, which consists of convolutional layers and subsampling layers, and the feature extractor can be thought of as a filter. A convolutional layer is a layer of neurons within a convolutional neural network where the convolutional operation is performed on the input signal. In the convolutional layers of a convolutional neural network, a single neuron may be connected to only some of the neurons in adjacent layers. A convolutional layer typically contains several feature planes, each of which may contain several neurons arranged in a rectangle. Neurons in the same feature plane share weights, and these shared weights constitute the convolutional kernel. Weight sharing can be understood as a picture information extraction scheme that is location-independent. The convolutional kernel can be initialized in the form of a random-size matrix. In the process of training a convolutional neural network, the convolutional kernel can acquire appropriate weights through learning. In addition, a direct benefit of weight sharing is that it reduces the number of connections between layers in a convolutional neural network, thereby lowering the risk of overfitting.
[0127] (4) Entropy coding
[0128] Entropy coding is used to obtain encoded data that can be output through the output in the form of an encoded bitstream, by applying an entropy coding algorithm or scheme (e.g., variable length coding, VLC, context adaptive VLC (CAVLC), arithmetic coding, binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique) to quantized coefficients and other syntax elements, so that a decoder or similar device can receive and use the parameters for decoding. The encoded bitstream can be sent to the decoder or stored in memory for subsequent transmission or retrieval by the decoder.
[0129] In the following embodiments of the coding system 10, the encoder 20A and decoder 30A are described with reference to Figures 1A to 15.
[0130] Figure 1A is a schematic block diagram illustrating an exemplary coding system 10, for example, a picture (or audio) coding system 10 (or simply coding system 10) that may utilize the techniques of this application. The encoder 20A and decoder 30A in the picture coding system 10 represent devices and similar that may be configured to determine various techniques based on the various examples described in this application.
[0131] As shown in Figure 1A, the coding system 10 includes a source device 12 configured to provide an encoded bitstream 21, for example, an encoded picture (or audio), to a destination device 14 configured to decode the encoded bitstream 21.
[0132] The source device 12 includes an encoder 20A and optionally includes a picture source 16, a preprocessor (or preprocessing unit) 18, a communication interface (or communication unit) 26, and a probability estimation (or probability estimation unit) 40.
[0133] The picture (or audio) source 16 may include, or may be, any type of picture capture device configured to capture real-world pictures (or audio), and / or any type of picture generation device, such as a computer graphics processing unit configured to generate computer-animated pictures, or any type of device configured to acquire and / or provide real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). The audio or picture source may be any type of memory or storage device that stores any of the aforementioned audio or pictures.
[0134] To distinguish from the preprocessor (or preprocessing unit) 18 and the processing determined by the preprocessor (or preprocessing unit) 18, the picture or audio (picture or audio data) 17 may also be called the original picture or audio (original picture or audio data) 17.
[0135] The preprocessor 18 is configured to receive (original) picture (or audio) data 17, perform preprocessing on the picture (or audio) data 17 to obtain a preprocessed picture or audio (or preprocessed picture or audio data) 19. For example, the preprocessing determined by the preprocessor 18 may include cropping, color format conversion (e.g., RGB to YCbCr), color correction, or noise reduction. It can be understood that the preprocessing unit 18 may be an optional component.
[0136] The encoder 20A includes an encoder network 20, an entropy coding 24, and optionally a preprocessor 22.
[0137] The picture (or audio) encoder network (or encoder network) 20 is configured to receive pre-processed picture (or audio) data 19 and provide encoded picture (or audio) data 21.
[0138] The preprocessor 22 is configured to receive feature data 21 to be encoded, preprocess the feature data 21 to be encoded, and obtain preprocessed feature data 23 to be encoded. For example, the preprocessing determined by the preprocessor 22 may include trimming, color format conversion (e.g., RGB to YCbCr), color correction, or denoising. It can be understood that the preprocessing unit 22 may be an optional component.
[0139] Entropy coding 24 is used to receive (or preprocess) the feature data to be coded 23 and generate the coded bitstream 25 based on the probability estimation result 41 provided by probability estimation 40.
[0140] The communication interface 26 of the source device 12 may be configured to receive the encoded bitstream 25 and transmit the encoded bitstream 25 (or any further processed version thereof) over the communication channel 27 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.
[0141] The destination device 14 includes a decoder 30A and may optionally include a communication interface (or communication unit) 28, a post-processor (or post-processing unit) 36, and a display device 38.
[0142] The communication interface 28 of the destination device 14 is configured to receive the encoded bitstream 25 (or any further processed version thereof) directly from the source device 12 or from any other source device, such as a storage device for encoded bitstream data, and to provide the encoded bitstream 25 to the decoder 30A.
[0143] Communication interfaces 26 and 28 may be configured to transmit or receive encoded bitstreams (or encoded bitstream data) 25 via a direct communication link between the source device 12 and the destination device 14, for example, via a direct wired or wireless connection, or via any type of network, for example, a wired or wireless network or any combination thereof, or any type of private and public network or any combination thereof.
[0144] The communication interface 26 may be configured to process the encoded bitstream 25, for example, by packaging the encoded bitstream 25 into an appropriate format, such as a packet, and / or by using any type of transmit encoding or processing for transmission over a communication link or communication network.
[0145] Communication interface 28 corresponds to communication interface 26 and may be configured, for example, to receive the transmitted data and process it by using any type of corresponding transmit decoding or processing and / or decapsulation to obtain the encoded bitstream 25.
[0146] Both communication interfaces 26 and 28 may be configured as one-way or two-way communication interfaces, as indicated by the arrow for the communication channel 27 in Figure 1A, from the source device 12 to the destination device 14, and may be configured to send and receive messages, establish connections, and acknowledge and exchange any other information regarding the communication link and / or data transmission, such as encoded picture data transmission.
[0147] The decoder 30A includes a decoder network 34, an entropy decoder 30, and optionally a post-processor 32.
[0148] Entropy decoding 30 is used to receive the encoded bitstream 25 and provide the decoded feature data 31 based on the probability estimation result 42 provided by probability estimation 40.
[0149] The post-processor 32 is configured to perform post-processing on the decoded feature data 31 to obtain post-processed decoded feature data 33. The post-processing determined by the post-processing unit 32 may include, for example, color format conversion (e.g., YCbCr to RGB), color correction, cropping, or resampling. It can be understood that the post-processing unit 32 may be an optional component.
[0150] The decoder network 34 is used to receive the decoded feature data 31 or the post-processed decoded feature data 33 and to provide the reconstructed picture data 35.
[0151] The post-processor 36 is configured to perform post-processing on the reconstructed picture data 35 to obtain the post-processed reconstructed picture data 37. The post-processing determined by the post-processing unit 36 may include, for example, color format conversion (e.g., YCbCr to RGB), color correction, cropping, or resampling. It can be understood that the post-processing unit 36 may be an optional component.
[0152] The display device 38 is configured to receive reconstructed picture data 35 or post-processed picture data 37 for displaying the picture to a user, viewer, or the like. The display device 38 may be or include any type of player or display for representing the reconstructed audio or picture, such as an integrated or external display or monitor. For example, the display may include a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a microLED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.
[0153] Figure 1A shows the source device 12 and destination device 14 as separate devices; however, the device embodiments may, alternatively, include both the source device 12 and the destination device 14, or include the functions of both the source device 12 and the destination device 14, i.e., include both the source device 12 or its corresponding function and the destination device 14 or its corresponding function. In these embodiments, the source device 12 or its corresponding function and the destination device 14 or its corresponding function may be implemented by using the same hardware and / or software, or by using separate hardware and / or software or any combination thereof.
[0154] Based on the description, the presence and (exact) division of different units or functions of the source device 12 and / or destination device 14 shown in Figure 1A may vary with the actual device and application. This will be apparent to those skilled in the art.
[0155] A feature data encoder 20A (for example, a picture feature map encoder or an audio feature variable encoder), a feature data decoder 30A (for example, a picture feature map decoder or an audio feature variable decoder), or both a feature data encoder 20A and a feature data decoder 30A may be implemented using a processing circuit as shown in Figure 1B, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), separate logic, hardware, a dedicated processor for picture encoding, or any combination thereof. The feature data encoder 20A may also be implemented using a processing circuit 56, and the feature data decoder 30A may also be implemented using a processing circuit 56. The processing circuit 56 may be configured to determine a variety of operations. If the technique is partially implemented in software, the device may store instructions for the software in a suitable non-temporary computer-readable storage medium, or the instructions may be determined in hardware by using one or more processors to determine the technique of the present invention. As shown in Figure 1B, one of the feature data encoder 20A and feature data decoder 30A can be integrated into a single device as part of a combined encoder / decoder (CODEC).
[0156] The source device 12 and destination device 14 may include any one of a variety of devices, including any type of handheld or stationary device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or content distribution server), a broadcast receiving device, a broadcast transmitting device, and the like, and may or may not use any type of operating system. In some cases, the source device 12 and destination device 14 may be equipped with components for wireless communication. Thus, the source device 12 and destination device 14 may be wireless communication devices.
[0157] In some cases, the coding system 10 shown in Figure 1A is merely an example. The technology provided in this application may also be applicable to picture feature map or audio feature variable coding settings (e.g., picture feature map coding or picture feature map decoding), the settings do not necessarily involve any data communication between the coding device and the decoding device. In another example, the data is retrieved from local memory, transmitted over a network, and so on. The picture feature map or audio feature variable coding device may encode the data and store the data in memory, and / or the picture feature map or audio feature variable decoding device may retrieve the data from memory and decode the data. In some examples, coding and decoding are determined by devices that do not communicate with each other but encode the data in memory and / or retrieve the data from memory and decode the data.
[0158] Figure 1B is a diagram illustrating an example of a coding system 50, which includes the feature data encoder 20A and / or the feature data decoder 30A of Figure 1A, according to an exemplary embodiment. The coding system 50 may include an imaging (or audio generation) device 51, the encoder 20A and decoder 30A (and / or feature data encoder / decoder implemented using processing circuitry 56), an antenna 52, one or more processors 53, one or more memory storage 54, and / or a display (or audio playback) device 55.
[0159] As shown in Figure 1B, the imaging (or audio generation) device 51, antenna 52, processing circuit 56, encoder 20A, decoder 30A, processor 53, memory storage 54, and / or display (or audio playback) device 55 can communicate with each other. In a different example, the coding system 50 may include only the encoder 20A or only the decoder 30A.
[0160] In some examples, the antenna 52 may be configured to transmit or receive an encoded bitstream of feature data. In addition, in some examples, the display (or audio playback) device 55 may be configured to present picture (or audio) data. The processing circuit 56 may include application-specific integrated circuit (ASIC) logic, graphics processing units, general-purpose processors, and the like. The coding system 50 may also include an optional processor 53. Similarly, the optional processor 53 may include application-specific integrated circuit (ASIC) logic, graphics processing units, audio processors, general-purpose processors, and the like. In addition, the memory storage 54 may be any type of memory, for example, volatile memory (e.g., static random access memory, SRAM, or dynamic random access memory, DRAM) or non-volatile memory (e.g., flash memory). In examples that are not limited, the memory storage 54 may be implemented by using cache memory. In another example, the processing circuit 56 may include memory (e.g., a cache) configured to implement a picture buffer.
[0161] In some examples, an encoder 20A implemented using logic circuits may include a picture buffer (for example, implemented using processing circuits 56 or memory storage 54) and a graphics processing unit (for example, implemented using processing circuits 56). The graphics processing unit may be communicatively coupled to the picture buffer. The graphics processing unit may include an encoder 20A implemented using processing circuits 56. The logic circuits may be configured to determine various operations in the specification.
[0162] In some examples, the decoder 30A may be implemented by using the processing circuit 56 in a similar manner to implement various modules described with reference to the decoder 30 shown in Figure 1B and / or any other decoder system or subsystem described in the specification. In some examples, the decoder 30A implemented by using logic circuits may include a picture buffer (for example, implemented by using the processing circuit 56 or memory storage 54) and a graphics processing unit (for example, implemented by using the processing circuit 56). The graphics processing unit may be communicatively coupled to the picture buffer. The graphics processing unit may include a picture decoder 30A implemented by using the processing circuit 56.
[0163] In some examples, antenna 52 may be configured to receive an encoded bitstream of picture data. As described above, the encoded bitstream may include data, indicators, index values, mode selection data, and similar data relating to audio or video frame encoding, as described in the specification, such as data relating to encoding segments. The coding system 50 may also include a decoder 30A coupled to antenna 52 and configured to decode the encoded bitstream. A display (or audio playback) device 55 may be configured to present the picture (or audio).
[0164] In this embodiment of the application, it should be understood that, in the example described with reference to encoder 20A, decoder 30A may be configured to determine the reverse process. For signaling syntax elements, decoder 30A may be configured to receive and parse the syntax elements and, in correspondence, decode the associated picture data. In some examples, encoder 20A may perform entropy coding on the syntax elements to obtain an encoded bitstream. In the example, decoder 30A may parse the syntax elements and, in correspondence, decode the associated picture data.
[0165] Figure 1C is a schematic diagram of a coding device 400 according to an embodiment of the present invention. The coding device 400 is suitable for implementing the disclosed embodiments described in the specification. In the embodiment, the coding device 400 may be a decoder, for example, the picture feature map decoder 30A in Figure 1A, or an encoder, for example, the picture feature map encoder 20A in Figure 1A.
[0166] The picture coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 configured to receive data, a processor, logic unit, or central processing unit (CPU) 430 configured to process the data (for example, the processor 430 may be a neural network processing unit 430), a transmitter unit (Tx) 440 and an exit port 450 (or output port 450) configured to transmit data, and a memory 460 configured to store data. The picture (or audio) coding device 400 may further include optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to the inlet port 410, a receiver unit 420, a transmitter unit 440, and an exit port 450 for the exit or inlet of optical or electrical signals.
[0167] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more processor chips, cores (e.g., a multi-core processor), FPGAs, ASICs, and DSPs. The processor 430 communicates with an inlet port 410, a receiver unit 420, a transmitter unit 440, an exit port 450, and memory 460. The processor 430 includes a coding module 470 (e.g., a coding module 470 based on a neural network NN). The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 determines, processes, prepares, or provides various coding operations. Thus, the inclusion of the coding module 470 significantly improves the functionality of the coding device 400 and affects the switching of the coding device 400 to different states. Alternatively, the coding module 470 is implemented by using instructions stored in memory 460 and determined by the processor 430.
[0168] Memory 460 includes one or more disks, tape drives, and solid-state drives and may be used as an overflow data storage device to store a program when such a program is selected for decision and to store instructions and data read during program decision. Memory 460 may be volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0169] Figure 1D is a simplified block diagram of an apparatus 500 that can be used as either or both of the source device 12 and destination device 14 in Figure 1A, according to an embodiment.
[0170] The processor 502 in the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device, existing or to be developed in the future, that can manipulate or process information, or any number of such devices. The disclosed implementation may be implemented by a single processor, such as the processor 502 shown in the figure, but by using one or more processors, advantages in speed and efficiency can be achieved.
[0171] In the implementation, the memory 504 in the device 500 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as memory 504. Memory 504 may include code and data 506 accessed by the processor 502 via the bus 512. Memory 504 may further include an operating system 508 and an application program 510, the application program 510 including at least one program that enables the processor 502 to determine the method in the specification. For example, the application program 510 may include applications 1 through N and further include a picture coding application that determines the method in the specification.
[0172] The device 500 may further include one or more output devices, such as a display 518. In this example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensing element that can be configured to sense touch input. The display 518 may be coupled to the processor 502 via a bus 512.
[0173] Although bus 512 in the device 500 is described as a single bus in this specification, bus 512 may comprise multiple buses. Furthermore, the secondary memory may be directly coupled to another component of the device 500, or it may be accessed via a network, and may comprise a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, the device 500 may have various configurations.
[0174] Figure 2A shows possible system architectures 1800 in picture feature mapping or audio feature variable coding and decoding scenarios, and includes the following: Capture device 1801: The video capture device completes the capture of the original video (or audio). Preprocessing 1802: A series of preprocessing steps are performed on the original video (or audio) capture to obtain video (or audio) data. Encoding 1803: Video (audio) encoding is used to reduce encoding redundancy and the amount of data transmitted in the picture feature map or audio feature variable compression process. Transmission 1804: Compressed and encoded bitstream data obtained through encoding is transmitted using the transmit module. Received 1805: Compressed and encoded bitstream data is received by the receiving module via network transmission. Bitstream Decoding 1806: Bitstream decoding is performed on the bitstream data, and Rendering (or playback) 1807: Rendering (or playback) is performed on the decoded data.
[0175] Figure 2B shows possible system architectures 1900 in a machine-oriented task scenario for picture feature maps (or audio feature variables), including: Feature extraction 1901: Feature extraction is performed on the picture (or audio) source. Side information extraction 1902: Side information extraction is performed on data obtained through feature extraction. Probability Estimation 1903: Side information is used as input for probability estimation, and probability estimation is performed on the feature map (or feature variable) to obtain the probability estimation result. Encoding 1904: Entropy coding is performed on the data obtained through feature extraction, referencing the probability estimation results to obtain the bitstream. Optionally, before encoding is performed, quantization or rounding operations are performed on the data obtained through feature extraction, and the quantized or rounded data obtained through feature extraction is encoded. Optionally, entropy coding is performed on the side information so that the bitstream includes the side information data. Decoding 1905: Entropy decoding is performed on the bitstream, referencing the probability estimation results to obtain a picture feature map (or audio feature variable). Optionally, if the bitstream contains side information encoded data, entropy decoding is performed on the side information encoded data, and the decoded side information data is used as input to a probability estimation to obtain the probability estimation result. When only side information is used as input for probability estimation, the probability estimation results of feature elements may be output in parallel, and when the input for probability estimation includes context information, the probability estimation results of feature elements must be output in serial order, and it should be noted that the side information is feature information further extracted by inputting a picture feature map or audio feature variable into a neural network, and the number of feature elements contained in the side information is less than the number of feature elements in the picture feature map or audio feature variable, and optionally, the side information of the picture feature map or audio feature variable may be encoded into a bitstream, and Machine vision task 1906: A machine vision (or auditory) task is performed on the decoded feature map (or feature variable).
[0176] Specifically, the decoded feature data is input to a machine vision (or auditory) task network, which outputs one-dimensional, two-dimensional, or multi-dimensional data such as classification, target recognition, and semantic segmentation related to the vision (or auditory) task.
[0177] In possible implementations, the System Architecture 1900 implementation process involves feature extraction and encoding processes performed on the terminal, while decoding and machine vision tasks are performed in the cloud.
[0178] The encoder 20A may be configured to receive a picture (or picture data) or audio (or audio data) 17 through an input 202 or similar. The received picture, picture data, audio, and audio data may, alternatively, be a pre-processed picture (or pre-processed picture data) or audio (or pre-processed audio data) 19. For ease of brevity, the following description will use picture (or audio) 17. Picture (or audio) 17 may, alternatively, be called the current picture or the picture to be encoded (in particular when the current picture is distinguished from other pictures in video encoding, for example, when the other pictures are in the same video sequence, i.e., include previous encoded and / or decoded pictures in the video sequence of the current picture), or the current audio or the audio to be encoded.
[0179] A (digital) picture is, or can be considered as, a two-dimensional array or matrix of samples having intensity values. A sample in the array may also be called a pixel (or pel) (an abbreviation for picture element). The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, three color components are typically employed. Specifically, a picture may be represented as, or contain, three sample arrays. In the RGB format or color space, a picture contains corresponding red, green, and blue sample arrays. Similarly, each pixel may be represented in a luminance or chromaticity format or color space, for example, YCbCr, which contains a luminance component represented by Y (sometimes L is used instead) and two chromaticity components represented by Cb and Cr. The luminance (luma) component Y represents brightness or gray level intensity (for example, these two are the same in a grayscale picture), while the two chrominance (chroma) components Cb and Cr represent chromaticity or color information components. Correspondingly, a picture in the YCbCr format contains a luminance sample array of luminance sample values (Y) and two chromaticity sample arrays (Cb and Cr) of chromaticity values. A picture in the RGB format may be converted to or from the YCbCr format, and vice versa; this process is also known as color conversion. If the picture is monochrome, the picture may contain only a luminance sample array. Correspondingly, a picture may be, for example, an array of luminance samples in a monochrome format, or two corresponding arrays of luminance samples and chromaticity samples in 4:2:0, 4:2:2, and 4:4:4 color formats. The picture encoder 20A does not limit the color space of the picture.
[0180] Possibly, an embodiment of encoder 20A may include a picture (or audio) partitioning unit (not shown in Figures 1A or 1B) configured to partition a picture (or audio) 17 into multiple (usually non-overlapping) picture blocks 203 or audio segments. These picture blocks may also be called root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB) or coding tree units (CTU) in the H.265 / HEVC and VVC standards. The partitioning unit may be configured to use the same block size for all pictures in the video sequence and a corresponding grid that defines the block size, or to vary the block size between pictures, picture subsets, or groups of pictures, partitioning each picture into a corresponding block.
[0181] In another possibility, the encoder may be configured to directly receive block 203 of picture 17, for example, one, some, or all of the blocks that make up picture 17. Picture block 203 may also be called the current picture block or the picture block to be encoded.
[0182] Like picture 17, picture block 203 is a two-dimensional array or matrix of samples having intensity values (sample values), although it is smaller in dimensions than picture 17, or can be considered as such. In other words, block 203 may contain, for example, one sample array (e.g., a luminance array for monochrome picture 17, or a luminance or chromaticity array for color picture), three sample arrays (e.g., one luminance array and two chromaticity arrays for color picture 17), or any other quantity and / or type of array depending on the applied color format. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Correspondingly, the block may be, for example, an array of M × N (M columns × N rows) samples, or an array of M × N transformation coefficients.
[0183] In another possibility, the encoder 20A shown in Figures 1A and 1B or Figures 3A to 3D is configured to encode the picture 17 block by block.
[0184] In another possibility, the encoder 20A shown in Figures 1A and 1B or Figures 3A to 3D is configured to encode the picture 17.
[0185] In another possibility, the encoder 20A shown in Figures 1A and 1B or Figures 3A to 3D may be further configured to segment or encode a picture by using slices (also called video slices), and the picture may be segmented or encoded by using one or more slices (usually non-overlapping). Each slice may contain one or more blocks (e.g., coding tree units CTUs), or one or more groups of blocks (e.g., tiles in the H.265 / HEVC / VVC standard or subpictures in the VVC standard).
[0186] In another possibility, the encoder 20A shown in Figures 1A and 1B or Figures 3A to 3D may be further configured to segment and / or encode a picture by using slice / tile groups (also called video tile groups) and / or tiles (also called video tiles), the picture may be segmented or encoded by using one or more slice / tile groups (usually non-overlapping), each slice / tile group may contain one or more blocks (e.g., CTUs) or one or more tiles. Each tile may be rectangular in shape and may contain one or more complete or fragmented blocks (e.g., CTUs).
[0187] Encoder network 20
[0188] The encoder network 20 is configured to obtain picture feature maps or audio feature variables by using the encoder network based on the input data.
[0189] Possibly, the encoder network 20 shown in Figure 4A includes multiple network layers. Any network layer may be a convolutional layer, a normalization layer, a nonlinear activation layer, or similar.
[0190] Possibly, the input to the encoder network 20 is at least one picture to be encoded or at least one picture block to be encoded. The picture to be encoded may be the original picture, a lossy picture, or a residual picture.
[0191] In a possible case, an example of the network structure of the encoder network in encoder network 20 is shown in Figure 4B. In this example, it can be understood that the encoder network includes five network layers, specifically three convolutional layers and two nonlinear activation layers.
[0192] Round 24
[0193] Rounding is used to round picture feature maps or audio feature variables, for example by using scalar quantization or vector quantization, in order to obtain rounded picture feature maps or audio feature variables.
[0194] Possibly, the encoder 20A may output a quantization parameter (QP), for example, by directly outputting the quantization parameter, or by outputting the quantization parameter after it has been encoded or compressed by an encoding determination implementation unit, so that, for example, the decoder 30A may receive and apply the quantization parameter for decoding.
[0195] Possibly, the output feature map or feature audio feature variable is preprocessed before rounding, and preprocessing may include trimming, color format conversion (e.g., RGB to YCbCr), color correction, noise reduction, or similar.
[0196] Probability Estimation 40
[0197] The probability estimation results for picture feature maps or audio feature variables are obtained through probability estimation based on the input feature map or feature variable information.
[0198] Probability estimation is used to perform probability estimation on rounded picture feature maps or audio feature variables.
[0199] Probability estimation may also be a probability estimation network, which is a convolutional network, and a convolutional network includes convolutional layers and nonlinear activation layers. Figure 4B is used as an example. The probability estimation network includes five network layers, specifically three convolutional layers and two nonlinear activation layers. Probability estimation can be achieved by using a convolutional non-network probability estimation method. Probability estimation methods include, but are not limited to, equi-maximum likelihood estimation, maximum posterior probability estimation, maximum likelihood estimation, and other statistical methods.
[0200] Encoding determination implementation 26
[0201] As shown in Figure 5, the coding decision implementation includes coding element determination and entropy coding. The picture feature map or audio feature variable is one-dimensional, two-dimensional, or multi-dimensional data output by the encoder network, and each of the data is a feature element.
[0202] Encoding element determination 261
[0203] Encoding element determination involves determining each feature element of a picture feature map or audio feature variable based on the probability estimation results of probability estimation, and then determining which feature elements will be entropy coded based on the determination results.
[0204] After the element determination process for the Pth feature element of the picture feature map or audio feature variable is completed, the element determination process for the (P+1)th feature element of the picture feature map begins, where P is a positive integer and P is less than M.
[0205] Entropy coding 262
[0206] Entropy coding can be performed using various disclosed entropy coding algorithms, such as variable length coding (VLC), context adaptive VLC (CAVLC), entropy coding schemes, binarization algorithms, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques. Encoded picture data 25, which may be output in the form of an encoded bitstream 25 or similar, is obtained through output 212, thereby allowing a decoder 30A or similar to receive and use parameters for decoding. The encoded bitstream 25 may be transmitted to the decoder 30A or stored in memory for subsequent transmission or retrieval by the decoder 30A.
[0207] In another possibility, entropy coding can be performed by using an entropy coding network, which is implemented, for example, by using a convolutional network.
[0208] In a possible scenario, since entropy coding does not know the actual character probabilities of the rounded feature map, the actual character probabilities of the rounded feature map or related information may be collected and added to the entropy coding, and this information is sent to the decoder.
[0209] Joint Network 44
[0210] The joint network obtains probability estimation results and decision information for picture feature maps or audio feature variables based on input side information. The joint network is a multilayer network, and may be a convolutional network, which includes convolutional layers and nonlinear activation layers. Any network layer of the joint network may be a convolutional layer, a normalization layer, a nonlinear activation layer, or similar.
[0211] The judgment information may be one-dimensional, two-dimensional, or multi-dimensional data, and the size of the judgment information may match the size of the picture feature map.
[0212] The decision information may be output after any of the network layers in the joint network.
[0213] The probability estimation results can be output after any of the network layers of the joint network.
[0214] Figure 6 shows an example of the output of the joint network's network structure. The network structure includes four network layers. Decision information is output after the fourth network layer, and probability estimation results are output after the second network layer.
[0215] Generative network 46
[0216] The generative network obtains judgment information for feature elements of a picture feature map based on the input probability estimation results. The generative network is a multilayer network, and may be a convolutional network, which includes convolutional layers and nonlinear activation layers. Any network layer of the generative network may be a convolutional layer, a normalization layer, a nonlinear activation layer, or similar.
[0217] Decision information can be output after any network layer of the generating network. Decision information can be one-dimensional, two-dimensional, or multi-dimensional data.
[0218] Figure 7 shows an example of outputting decision information based on the network structure of the generative network. The network structure includes four network layers.
[0219] Decryption detection implementation 30
[0220] As shown in Figure 8, the decoding decision implementation includes element determination and entropy decoding. The picture feature map or audio feature variable is one-dimensional, two-dimensional, or multi-dimensional data output by the decoding decision implementation, where each data element is a feature element.
[0221] Decoding element determination 301
[0222] Decoding element determination is the process of determining each feature element of a picture feature map or audio feature variable based on the probability estimation results of probability estimation, and then determining the specific feature elements on which entropy decoding will be performed based on the determination results. Decoding element determination is the process of determining each feature element of a picture feature map or audio feature variable, and then determining the specific feature elements on which entropy decoding will be performed based on the determination results. It can be seen as the reverse process of encoding element determination, which involves determining each feature element of a picture feature map and then determining the specific feature elements on which entropy coding will be performed based on the determination results.
[0223] Entropy Decoding 302
[0224] Entropy decoding can perform entropy decoding, e.g., variable length coding (VLC) scheme, context adaptive VLC (CAVLC) scheme, entropy decoding scheme, binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques using various disclosed entropy decoding algorithms. Encoded picture (or audio) data 25, which may be output in the form of an encoded bitstream 25 or similar, is obtained through output 212, thereby allowing a decoder 30A or similar to receive and use parameters for decoding. The encoded bitstream 25 may be transmitted to the decoder 30A or stored in memory for subsequent transmission or retrieval by the decoder 30A.
[0225] In another possibility, entropy decoding can be performed by using an entropy decoding network, which may be implemented, for example, by using a convolutional network.
[0226] Decoder Network 34
[0227] The decoder network is used to pass the decoded picture feature map or audio feature variable 31 or the post-processed decoded picture feature map or audio feature variable 33 through the decoder network 34 to obtain reconstructed picture (or audio) data 35 or machine-ready task data in the pixel region.
[0228] The decoder network includes multiple network layers. Any network layer may be a convolutional layer, a normalization layer, a nonlinear activation layer, or similar. Operations such as concatenation, addition, and subtraction may be present in the decoder network unit 306.
[0229] In possibility, the network layer structures of the decoder networks may be the same or different from each other.
[0230] An example of a decoder network structure is shown in Figure 9. In this example, it can be understood that the decoder network includes five network layers, specifically one normalization layer, two convolutional layers, and two nonlinear activation layers.
[0231] The decoder network outputs a reconstructed picture (or audio) or acquired machine-ready task data. Specifically, the decoder network may include a target recognition network, a classification network, or a semantic segmentation network.
[0232] It should be understood that in encoder 20A and decoder 30A, the processing result of the current step may be further processed and then output to the next step. For example, after the encoder unit or decoder unit, further calculations or processing, such as clipping or shifting operations or filtering operations, may be performed on the processing result of the encoder unit or decoder unit.
[0233] Based on the foregoing description, the following provides several picture feature maps or audio feature variable encoding and decoding methods according to the embodiments of this application. For the ease of explanation, the embodiments of the methods described below are expressed as a combination of a series of operation steps. However, those skilled in the art of this technology should understand that the specific implementation of the technical solution of this application is not limited to the order of the series of operation steps described.
[0234] The following will describe the procedure of this application in detail with reference to the accompanying drawings. It should be noted that the process on the encoder side in the flowchart may specifically be executed by the encoder 20A, and the process on the decoder side in the flowchart may specifically be executed by the decoder 30A.
[0235] In Embodiments 1 to 5, the first feature element or the second feature element is the feature element to be currently encoded or the feature element to be currently decoded, for example,
Number
[0236] In Embodiment 1 of this application, FIG. 10A represents a specific implementation process 1400. The execution steps are as follows.
[0237] Encoder side: Step 1401: Obtain a picture feature map.
[0238] This step is specifically carried out by the encoder network 204 in Figure 3A. For details, see the preceding description of encoder network 20. To output a picture feature map y, the picture is input to the feature extraction module, and the feature map y may be 3D data with dimensions w, x, h, x, and c. Specifically, the feature extraction module may be implemented by using an existing neural network, but is not limited to this. This step is an existing technique.
[0239] The feature quantization module quantizes each feature value of feature map y, rounds the floating-point feature values, and through rounding obtains integer feature values, resulting in the quantized feature map.
number
[0240] Step 1402: Feature Map
number
number
number
[0241] The parameters x, y, and i are positive integers, and the coordinates (x, y, i) indicate the position of the feature element to be currently encoded. Specifically, the coordinates (x, y, i) indicate the position of the feature element to be currently encoded in the current 3D feature map relative to the feature element at the top-left vertex. This step is specifically performed by the probability estimation 210 in FIG. 3A. For details, refer to the aforementioned description of the probability estimation 40. Specifically, a probability distribution model can be used to obtain the probability distribution. For example, a Gaussian single model (GSM) or a Gaussian mix model (GMM) is used for modeling. First, the side information
Number
Number
Number
Number
[0242] Step 1403: To obtain the compressed bitstream, the feature map
Number
[0243] This step is specifically implemented by the encoding determination implementation 208 in FIG. 3A. For details, refer to the foregoing description of the encoding determination implementation 26. The feature element to be currently encoded
Number
Number
Number
Number
[0244] Step 1404: The encoder transmits or stores the compressed bitstream.
[0245] Decoder side: Step st1411: Obtain the bitstream of the decoded picture feature map.
[0246] Step 1412: Perform probability estimation based on the bitstream to obtain the probability estimation result of the feature element.
[0247] This step is specifically carried out by probability estimation 302 in Figure 10B. For details, please refer to the above explanation of probability estimation 40. The probability estimation is performed on the feature map that will be decoded.
number
number
number
number
[0248] The diagram of the structure of the probability estimation network used by the decoder is the same as that of the probability estimation network on the encoder side in this embodiment.
[0249] Step 1413: Feature map to be decrypted
number
[0250] This step is specifically performed by the decoding decision implementation 304 in Figure 10B. For details, please refer to the above description of the decoding decision implementation 30. The probability P that the value of the feature element to be decoded is k, i.e., the probability estimation result P of the feature element to be decoded, is obtained based on the probability distribution of the feature element to be decoded. If the probability estimation result P does not satisfy the pre-set conditions, i.e., P is greater than the first threshold T0, then entropy decoding does not need to be performed on the feature element to be decoded, and the value of the feature element to be decoded is set to k. Otherwise, if the feature element to be decoded satisfies the pre-set conditions, i.e., P is less than or equal to the first threshold T0, then entropy decoding is performed on the bitstream and the value of the feature element to be decoded is obtained.
[0251] Index numbers can be obtained from the bitstream by parsing it and based on a first threshold T0. The decoder constructs a list of threshold candidates in the same manner as the encoder, and then obtains the corresponding threshold according to the correspondence between the thresholds and index numbers in the pre-configured list of threshold candidates. The index numbers are obtained from the bitstream, or in other words, from the sequence header, picture header, slice header, or SEI.
[0252] Alternatively, the bitstream may be parsed directly, and the threshold can be obtained from the bitstream. Specifically, the threshold can be obtained from the sequence header, picture header, slice header, or SEI.
[0253] Alternatively, a fixed threshold is set directly according to a threshold policy that matches the decoding.
[0254] Step 1414: Decoded Feature Map
number
[0255] Example 1: Feature map obtained through entropy decoding
number
[0256] Case Study 2: Feature Maps Obtained Through Entropy Decoding
number
[0257] The aforementioned decoder value k is set to the corresponding encoder value k.
[0258] Figure 11A shows a specific implementation process 1500 according to Embodiment 2 of this application. The execution steps are as follows:
[0259] It should be noted that in methods 1 to 6 of this embodiment, the probability estimation result includes a first parameter and a second parameter. When the probability distribution is a Gaussian distribution, the first parameter is the mean μ and the second parameter is the variance σ. When the probability distribution is a Laplace distribution, the first parameter is the position parameter μ and the second parameter is the scale parameter b.
[0260] Encoder side:
[0261] Step 1501: Obtain the picture feature map.
[0262] This step is specifically carried out by the encoder network 204 in Figure 3B. For details, see the preceding description of encoder network 20. To output a picture feature map y, the picture is input to the feature extraction module, and the feature map y may be 3D data with dimensions w, x, h, x, and c. Specifically, the feature extraction module may be implemented by using an existing neural network, but is not limited to this. This step is an existing technique.
[0263] The feature quantization module quantizes each feature value of feature map y, rounds the floating-point feature values to obtain integer feature values, and then generates the quantized feature map.
number
[0264] Step 1502: Picture Feature Map
number
number
[0265] This step is specifically carried out by the side information extraction unit 214 in Figure 3B. The side information extraction module can be implemented by using the network shown in Figure 12.
number
number
number
number
number
[0266] Entropy coding is used for side information
number
number
number
number
[0267] Step 1503: Feature Map
number
[0268] This step is specifically carried out by probability estimation 210 in Figure 3B. For details, please refer to the above explanation of probability estimation 40. A probability distribution model may be used to obtain the probability estimation results and probability distributions. The probability distribution model may be a single Gaussian model (GSM), an asymmetric Gaussian model, a Gaussian mix model (GMM), or a Laplace distribution model.
[0269] When the probability distribution model is a Gaussian model (single Gaussian model, asymmetric Gaussian model, or mixture Gaussian model), first, side information
number
number
number
[0270] When the probability distribution model is a Laplace distribution model, first, side information
number
number
number
[0271] As an alternative, side information
number
number
number
number
number
number
[0272] Probabilistic estimation networks may use deep learning-based networks, such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs), but are not limited to these.
[0273] Step 1504: Based on the probability estimation results, entropy coding will now encode the feature elements.
number
number
[0274] This step is specifically carried out by the coding decision implementation 208 in Figure 3B. For details, please refer to the above description of coding decision implementation 26. One or more of the following methods determine the feature element that will be currently coded by entropy coding based on the probability estimation result.
number
[0275] Method 1: When the probability distribution model is a Gaussian distribution, whether to perform entropy encoding for the feature element to be currently encoded is determined based on the probability estimation result of the first feature element. When the values of the mean parameter μ and variance σ of the Gaussian distribution of the feature element to be currently encoded do not satisfy the preset conditions, that is, the absolute value of the difference between the mean μ and k is smaller than the second threshold T1, and the variance σ is smaller than the third threshold T2, the entropy encoding process is for the feature element to be currently encoded
Number
Number
Number
[0276] In particular, when the value of k is 0, it is the optimal value. When the absolute value of the mean parameter μ of the Gaussian distribution is less than T1 and the variance σ of the Gaussian distribution is less than T2, the feature element to be currently encoded
Number
Number
Number
[0277] Method 2: When the probability distribution is a Gaussian distribution, the feature element to be currently encoded
Number
Number
Number
Number
[0278] Method 3: When the probability distribution is a Laplace distribution, the feature element to be currently encoded
Number
Number
Number
number
[0279] Method 4: When the probability distribution is a Laplace distribution, the feature elements that will be encoded
number
number
number
number
[0280] In particular, when the value of k is 0, it is the optimal value. The feature element to be encoded is when the absolute value of the position parameter μ is less than T5 and the scale parameter b is less than T6.
number
number
number
[0281] Method 5: When the probability distribution is a Gaussian mixture, the feature elements that will be encoded
number
number
number
number
[0282] Method 6: Feature elements to be encoded
number
[0283] In actual applications, it should be noted that in order to ensure platform consistency, the thresholds T1, T2, T3, T4, T5, and T6 may be rounded, that is, shifted to an integer and scaled.
[0284] It should be noted that one of the following methods may be used as an alternative to obtain the threshold. This is not limited here.
[0285] Method 1: Threshold T1 is used as an example, any value within the range of the value of T1 is used as the threshold T1, and the threshold T1 is written into the bitstream. Specifically, the threshold is written into the bitstream, stored in the sequence header, picture header, slice / slice header, or SEI, and can be transmitted to the decoder side. Alternatively, another method may be used. This is not limited here. A similar method can also be used for the remaining thresholds T0, T2, T3, T4, T5, and T6.
[0286] Method 2: The encoder uses a fixed threshold that matches the decoder. The fixed threshold does not need to be written to the bitstream and does not need to be sent to the decoder. For example, threshold T1 is used as an example, and any value within the range of T1 is used directly as the value of T1. A similar method can be used for the remaining thresholds T0, T2, T3, T4, T5, and T6.
[0287] Method 3: A threshold candidate list is constructed, and the most likely values within the range of T1 values are added to the threshold candidate list. Each threshold corresponds to a threshold index number, the optimal threshold is determined, and the optimal threshold is used as the value of T1. The index number of the optimal threshold is used as the threshold index number of T1, and the threshold index numbers of T1 are written to the bitstream. Specifically, the thresholds are written to the bitstream and can be stored in the sequence header, picture header, slice header, or SEI and sent to the decoder side. Alternatively, other methods may be used, and are not limited here. Similar methods may also be used for the remaining thresholds T0, T2, T3, T4, T5, and T6.
[0288] Step 1505: The encoder transmits or stores the compressed bitstream.
[0289] Decoder side:
[0290] Step 1511: Obtain the bitstream of the picture feature map that will be decoded.
[0291] Step 1512: Obtain the probability estimation results for the feature elements.
[0292] This step is specifically carried out by the probability estimation unit 302 in Figure 11A. For details, please refer to the above explanation of probability estimation 40. Entropy decoding is side information
number
number
number
number
number
number
[0293] Accordingly, it should be noted that the probability estimation method used by the decoder is the same as that used by the encoder in this embodiment, and the diagram of the structure of the probability estimation network used by the decoder is the same as that of the probability estimation network on the encoder side in this embodiment. Further details will not be explained here again.
[0294] Step 1513: This step is specifically performed by the decoding decision implementation 304 in Figure 11A. For details, please refer to the above description of the decoding decision implementation 30. Entropy decoding will now decode the feature element to be decoded.
number
number
[0295] One or more of the following methods, based on the probability estimation results, determine the feature element that will be currently decoded by entropy decoding.
number
[0296] Method 1: When the probability distribution model is a Gaussian distribution, the feature elements that will be decoded at present
number
number
number
number
number
[0297] Specifically, when the value of k is 0, it is the optimal value. When the absolute value of the mean parameter μ of the Gaussian distribution is less than T1 and the variance σ of the Gaussian distribution is less than T2, the feature element that will be currently decoded
Number
Number
Number
Number
[0298] are obtained based on the probability estimation results. When the relationship between the mean μ, variance σ, and k satisfies abs(μ - k)+σ < T3 (the pre-set condition is not met), where T3 is the fourth threshold value, and the current feature element to be decoded
Number
Number
Number
Number
Number
[0299] Method 3: When the probability distribution is a Laplace distribution, the values of the location parameter μ and the scale parameter b are obtained based on the probability estimation result. When the relationship between the location parameter μ, the scale parameter b, and k satisfies abs(μ - k)+σ < T4 (a preset condition is not satisfied), where T4 is the fourth threshold, the value of the feature element that will be currently decoded
Number
Number
Number
number
number
number
number
[0301] In particular, when the value of k is 0, it is the optimal value. The feature element that will be decoded is when the absolute value of the position parameter μ is less than T5 and the scale parameter b is less than T6.
number
number
number
number
[0302] Method 5: When the probability distribution is a Gaussian mixture, the feature elements that will be decoded are currently
number
number
number
number
number
[0303] Method 6: The probability P that the value of the feature element to be decoded is k, i.e., the probability estimation result P of the feature element to be decoded, is obtained based on the probability distribution of the feature element to be decoded. If the probability estimation result P does not satisfy a predetermined condition, i.e., P is greater than the first threshold T0, then entropy decoding does not need to be performed on the feature element to be decoded, and the value of the feature element to be decoded is set to k. Otherwise, if the feature element to be decoded satisfies a predetermined condition, i.e., P is less than or equal to the first threshold T0, then entropy decoding is performed on the bitstream and the value of the feature element to be decoded is obtained.
[0304] The aforementioned decoder value k is set to the corresponding encoder value k.
[0305] The method for obtaining thresholds T0, T1, T2, T3, T4, T5, T6, and T7 corresponds to the method on the encoder side, and one of the following methods may be used.
[0306] Method 1: The threshold is obtained from the bitstream. Specifically, the threshold is obtained from the sequence header, picture header, slice header, or SEI.
[0307] Method 2: The decoder uses a fixed threshold that matches that of the encoder.
[0308] Method 3: The threshold index number is obtained from the bitstream. Specifically, the threshold index number is obtained from the sequence header, picture header, slice header, or SEI. The decoder then constructs a list of threshold candidates in the same manner as the encoder and obtains the corresponding threshold from the list of threshold candidates based on the threshold index number.
[0309] In practical applications, it should be noted that, in order to ensure platform consistency, thresholds T1, T2, T3, T4, T5, and T6 may be rounded, i.e., shifted to integers and scaled.
[0310] Step 1514 is the same as step 1414.
[0311] Figure 13A shows a specific implementation process 1600 according to Embodiment 3 of this application. The execution steps are as follows:
[0312] Encoder side:
[0313] Step 1601 is the same as step 1501. This step is specifically performed by the encoder network 204 in Figure 3C. For details, please refer to the above description of encoder network 20.
[0314] Step 1602 is the same as step 1502. Specifically, this step is performed by the side information extraction 214 in Figure 3C.
[0315] Step 1603: Feature Map
number
[0316] This step can be specifically carried out by probability estimation 210 in Figure 3C. For details, please refer to the above explanation of probability estimation 40. A probability distribution model may be used to obtain the probability estimation results. The probability distribution model may be a single Gaussian model, an asymmetric Gaussian model, a mixture of Gaussian models, or a Laplace distribution model.
[0317] When the probability distribution model is a Gaussian model (single Gaussian model, asymmetric Gaussian model, or mixture Gaussian model), first, side information
number
number
number
[0318] When the probability distribution model is a Laplace distribution model, first, side information
number
number
number
[0319] Furthermore, the probability estimation results are input into the probability distribution model used to obtain the probability distribution.
[0320] As an alternative, side information
number
number
number
number
number
[0321] Probabilistic estimation networks may use deep learning-based networks, such as recurrent neural networks and convolutional neural networks, but are not limited to these.
[0322] Step 1604: Based on the probability estimation result, decide whether to perform entropy coding on the feature element to be currently coded. Based on the decision, entropy coding is performed on the feature element to be currently coded, and the feature element to be currently coded is written to the coded bitstream, or entropy coding is not performed. Entropy coding is performed on the feature element to be currently coded only when it is determined that entropy coding should be performed on the feature element to be coded.
[0323] This step is specifically carried out by the generative network 216 and coding decision implementation 208 in Figure 3C. For details, please refer to the above descriptions of the generative network 46 and coding decision implementation 26. The probability estimation result 211 is input to the decision module, and its dimension is the feature map.
number
number
number
number
number
number
number
number
[0324] In possible implementations, the probability estimation result or probability distribution of the feature element to be encoded is input to a decision module, which directly outputs decision information indicating whether entropy coding should be performed on the feature element to be encoded. For example, if the decision information output by the decision module is a pre-set value, it indicates that entropy coding should be performed on the feature element to be encoded. If the decision information output by the decision module is not a pre-set value, it indicates that entropy coding does not need to be performed on the feature element to be encoded. The decision module can be implemented using a network method. Specifically, the probability estimation result or probability distribution is input to a generative network shown in Figure 7, which outputs decision information, i.e., a pre-set value.
[0325] Method 1: The judgment information is a feature map whose dimension is
number
number
number
number
number
number
number
[0326] Method 2: The judgment information is a feature map whose dimension is
number
number
number
number
[0327] Method 3: The decision information can, alternatively, be an identifier or identifier value output directly by the joint network. When the decision information is a pre-set value, it indicates that entropy coding needs to be performed on the feature element currently being coded. When the decision information output by the decision module is not a pre-set value, it indicates that entropy coding does not need to be performed on the feature element currently being coded. For example, when the numerical choices for an identifier or identifier value are 0 and 1, the corresponding pre-set value is 0 or 1. When an identifier or identifier value can have multiple alternative values, the pre-set value is a specific set of values. For example, when the values of the identifier or identifier value range from 0 to 255, the pre-set value is a proper subset from 0 to 255.
[0328] A high probability indicates that the feature element currently being encoded is
number
[0329] Step 1605: The encoder transmits or stores the compressed bitstream.
[0330] Steps 1601 to 1604 are feature maps.
number
[0331] Decoder side:
[0332] Step 1611: Obtain the compressed bitstream that will be decrypted.
[0333] Step 1612: Feature map to be decrypted
number
[0334] This step can be specifically carried out by probability estimation 302 in Figure 13B. For details, please refer to the above explanation of probability estimation 40. (Side Information)
number
[0335] Step 1613: Obtain the decision information and, based on the decision information, decide whether to perform entropy decoding.
[0336] This step can be specifically carried out by the generation network 310 and the decoding decision implementation 304 in Figure 13B. For details, please refer to the above description of the generation network 46 and the decoding decision implementation 30. The decision information 311 is obtained in this embodiment by using the same method as the encoder side. When the decision map map[x][y][i] is a preset value, it is the feature element that will be decoded at the corresponding position by entropy decoding.
number
number
number
[0337] In possible implementations, the probability estimation result or probability distribution of the feature element to be decoded is input to a decision module, which directly outputs decision information indicating whether entropy decoding should be performed on the feature element to be decoded. For example, if the decision information output by the decision module is a pre-set value, it indicates that entropy decoding should be performed on the feature element to be decoded. If the decision information output by the decision module is not a pre-set value, it indicates that entropy decoding does not need to be performed on the feature element to be decoded, and the value of the feature element to be decoded is set to k. The decision module can be implemented using a network method. Specifically, the probability estimation result or probability distribution is input to a generative network shown in Figure 8, which outputs decision information, i.e., a pre-set value. The decision information indicates whether to perform entropy decoding on the feature element to be decoded, and the decision information may include a decision map.
[0338] Step 1614 is the same as step 1414.
[0339] The aforementioned decoder value k is set to the corresponding encoder value k.
[0340] Figure 14 shows a specific implementation process 1700 according to Embodiment 4 of this application. The execution steps are as follows:
[0341] Encoder side:
[0342] Step 1701 is the same as step 1501. This step can be specifically performed by the encoder network 204 in Figure 3D. For details, please refer to the above description of encoder network 20.
[0343] Step 1702 is the same as step 1502. Specifically, this step is performed by side information extraction 214 in Figure 3D.
[0344] Step 1703: Feature Map
number
[0345] This step can be specifically carried out by the joint network 218 in Figure 3D. For details, please refer to the above description of the joint network 34. Specifically, see the side information.
number
number
number
number
number
[0346] It should be noted that the specific structure of the joint network is not limited to this embodiment.
[0347] It should be noted that decision information, probability distributions, and / or probability estimation results can all be output from different layers of the joint network. For example, in example (1), the middle layer of the network outputs decision information, and the last layer outputs the probability distribution and / or probability estimation results. In example (2), the middle layer of the network outputs the probability distribution and / or probability estimation results, and the last layer outputs decision information. In example (3), the last layer of the network outputs decision information, probability distributions, and / or probability estimation results together.
[0348] When the probability distribution model is a Gaussian model (single Gaussian model, asymmetric Gaussian model, or mixture Gaussian model), first, side information
number
[0349] When the probability distribution model is a Laplace distribution model, first, side information
number
[0350] As an alternative, side information
number
number
number
[0351] Step 1704: Based on the decision information, a decision is made whether to perform entropy coding. Based on the decision result, entropy coding is performed and the feature element to be coded is written to the compressed bitstream (coded bitstream), or the entropy coding is skipped. Entropy coding is performed on the feature element to be coded only when it is determined that entropy coding needs to be performed on the feature element to be coded. This step can be specifically performed by the coding decision implementation 208 in Figure 3D. For details, please refer to the above description of the coding decision implementation 26.
[0352] Method 1: The judgment information is a feature map whose dimension is
number
number
number
number
number
number
number
[0353] Method 2: The judgment information is a feature map whose dimension is
number
number
number
number
[0354] Method 3: The decision information can, alternatively, be an identifier or the value of an identifier output directly by the joint network. When the decision information is a pre-set value, it indicates that entropy coding needs to be performed on the feature element that will be currently coded. When the decision information output by the decision module is not a pre-set value, it indicates that entropy coding does not need to be performed on the feature element that will be currently coded. When there are only two possible values for the feature element that will be currently coded in the decision map, output by the joint network, the pre-set value is a specific value. For example, when the possible values for the feature element that will be currently coded are 0 and 1, the pre-set value is 0 or 1. When there are multiple possible values for the feature element that will be currently coded in the decision map, output by the joint network, the pre-set value is a specific set of values. For example, when the possible values for the feature element that will be currently coded are from 0 to 255, the pre-set value is a proper subset from 0 to 255.
[0355] A high probability indicates that the feature element currently being encoded is
number
[0356] Step 1705: The encoder transmits or stores the compressed bitstream.
[0357] Decoder side:
[0358] Step 1711: Obtain the bitstream of the picture feature map to be decoded, and extract side information from the bitstream.
number
[0359] Step 1712: Feature Map
number
[0360] This step can be specifically carried out by the joint network 312 in Figure 16. For details, please refer to the above description of the joint network 34. Feature Map
number
[0361] Step 1713: Based on the decision information, a decision is made whether to perform entropy decoding, and based on the decision result, entropy decoding is performed or skipped. This step can be specifically performed by the decoding decision implementation 304 in Figure 16. For details, please refer to the above description of the decoding decision implementation 30.
[0362] Method 1: When the decision information is a decision map and the decision map map[x][y][i] is a pre-set value, then the entropy decoding will determine the feature element that will be decoded at the corresponding location.
number
number
number
[0363] Method 2: The judgment information is a feature map whose dimension is
number
number
number
number
number
[0364] Method 3: The decision information can, alternatively, be an identifier or the value of an identifier output directly by the joint network. When the decision information is a pre-set value, it indicates that entropy decoding needs to be performed on the feature element currently to be decoded. When the decision information output by the decision module is not a pre-set value, it indicates that entropy decoding does not need to be performed on the feature element currently to be decoded, and the value of the feature element currently to be decoded is set to k. When there are only two possible values for the feature element currently to be decoded in the decision map output by the joint network, the pre-set value is a specific value. For example, when the possible values for the feature element currently to be decoded are 0 and 1, the pre-set value is 0 or 1. When there are multiple possible values for the feature element currently to be decoded in the decision map output by the joint network, the pre-set value is a specific set of values. For example, when the possible values for the feature element currently to be decoded range from 0 to 255, the pre-set value is a proper subset from 0 to 255.
[0365] Step 1714 is the same as step 1414. Specifically, this step may be carried out by the decoder network unit 306 of the decoder 9C in the above-described embodiment. For details, please refer to the description of the decoder network unit 306 in the above-described embodiment.
[0366] The aforementioned decoder value k is set to the corresponding encoder value k.
[0367] Figure 17 shows a specific implementation process 1800 according to Embodiment 5 of this application. The execution steps are as follows:
[0368] Step 1801: Obtain the feature variables of the audio data to be encoded.
[0369] The audio signal to be encoded may be a time-domain audio signal. The audio signal to be encoded may be a frequency-domain signal obtained after a time-frequency transform has been performed on the time-domain signal. For example, the frequency-domain signal may be a frequency-domain signal obtained after an MDCT transform has been performed on the time-domain audio signal, and the time-domain audio signal is a frequency-domain signal obtained through an FFT transform. Alternatively, the signal to be encoded may be a signal obtained through QMF filtering. Alternatively, the signal to be encoded may be a residual signal, for example, another encoded residual signal or a residual signal obtained through LPC filtering.
[0370] Obtaining the feature variables of the audio data to be encoded may involve extracting feature vectors based on the audio signal to be encoded, for example, extracting Mel-cepstrum coefficients based on the audio signal to be encoded, quantizing the extracted feature vectors, and using the quantized feature vectors as the feature variables of the audio data to be encoded.
[0371] Alternatively, obtaining the feature variables of the audio data to be encoded may be done by using an existing neural network. For example, the audio signal to be encoded is processed by an encoding neural network to obtain latent variables, the latent variables output by the neural network are quantized, and the quantized latent variables are used as feature variables of the audio data to be encoded. The encoding neural network processing is pre-trained, and the specific network structure and training method of the encoding neural network are not limited in this invention. For example, a fully connected network or a CNN network may be selected for the encoding neural network. The number of layers included in the encoding neural network and the number of nodes in each layer are not limited in this invention.
[0372] The format of latent variables output by coding neural networks with different structures can vary. For example, if the coding neural network is a fully connected network, the output latent variables are vectors, where the dimension M of the vector is the latent size, e.g., y = [y(0), y(1), ..., y(M-1)]. If the coding neural network is a CNN network, the output latent variables are N*M dimensional matrices, where N is the number of channels in the CNN network, and M is the latent size of each channel in the CNN network, e.g.,
number
[0373] Both quantized feature vectors and quantized latent variables are
number
[0374] Step 1802: Feature variables of the audio data to be encoded
number
number
[0375] The side information extraction module can be implemented by using the network shown in Figure 12.
number
number
number
number
number
[0376] Entropy coding is used for side information
number
number
number
number
[0377] Step 1803: Feature Variables
number
[0378] Probability distribution models can be used to obtain probability estimation results and probability distributions. Probability distribution models can be single Gaussian models (GSM), asymmetric Gaussian models, Gaussian mix models (GMM), or Laplace distribution models.
[0379] The following are feature variables for explanatory purposes.
number
number
number
[0380] When the probability distribution model is a Gaussian model (single Gaussian model, asymmetric Gaussian model, or mixture Gaussian model), first, side information
number
number
number
[0381] Alternatively, the variance may be estimated. For example, when the probability distribution model is a Gaussian model (single Gaussian model, asymmetric Gaussian model, or mixture Gaussian model), first, side information
number
number
number
[0382] When the probability distribution model is a Laplace distribution model, first, side information
number
number
number
[0383] As an alternative, side information
number
number
number
number
number
number
[0384] Probabilistic estimation networks may use deep learning-based networks, such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs), but are not limited to these.
[0385] Step 1804: Based on the probability estimation result, determine whether entropy coding should be performed on the feature element to be coded. Based on the determination, perform entropy coding and write the feature element to be coded to the compressed bitstream (coded bitstream), or skip performing entropy coding.
[0386] One or more of the following methods, based on the probability estimation results, will entropy-code the feature elements that will be currently coded.
number
number
[0387] Entropy coding will now encode the feature elements
number
number
[0388] Method 1: When the probability distribution model is a Gaussian distribution, whether to perform entropy encoding on the feature element to be currently encoded is determined based on the probability estimation result of the first feature element. When the values of the mean parameter μ and variance σ of the Gaussian distribution of the feature element to be currently encoded satisfy the second condition, that is, the absolute value of the difference between the mean μ and k is less than the second threshold T1, and the variance σ is less than the third threshold T2, the entropy encoding process does not need to be performed on the feature element to be currently encoded. Otherwise, when the first condition is satisfied, that is, the absolute value of the difference between the mean μ and k is greater than or equal to the second threshold T1, or the variance σ is less than the third threshold T2, entropy encoding is performed on the feature element to be currently encoded, and the feature element to be currently encoded
Number
Number
Number
[0389]
Number
Number
Number
[0390] is written into the bitstream. The value of T2 is any number satisfying 0 < T2 < 1, for example, values such as 0.2, 0.3, 0.4, or the like. T1 is a number greater than or equal to 0 and less than 1, such as 0.01, 0.02, 0.001, and ......
Number
Number
Number
Number
[0391] When the probability distribution is a Gaussian distribution, the probability estimation is for each feature element of the feature variable [Number] If it is executed for each [Number] of the feature element that will be currently encoded [Number] only the value of the variance σ of the Gaussian distribution of is obtained. When the variance σ satisfies σ < T3 (the second condition), the entropy encoding process for the feature element that will be currently encoded [Number] is skipped. Otherwise, when the probability estimation result of the feature element that will be currently encoded satisfies σ ≥ T3 (the first condition), entropy encoding is performed for the feature element that will be currently encoded [Number] and the feature element that will be currently encoded [Number] is written into the bitstream. The fourth threshold T3 is a number greater than or equal to 0 and less than 1. For example, the value is 0.2, 0.3, 0.4, or the like.
[0392] Method 3: When the probability distribution is a Laplace distribution, the feature element that will be currently encoded
Number
Number
Number
Number
[0393] Method 4: When the probability distribution is a Laplace distribution, the feature element to be currently encoded <??>
Number
Number
number
number
[0394] In particular, when the value of k is 0, it is the optimal value. The feature element to be encoded is when the absolute value of the position parameter μ is less than T5 and the scale parameter b is less than T6.
number
number
number
[0395] Method 5: When the probability distribution is a mixture Gaussian distribution, the feature element to be currently encoded
Number
Number
Number
Number
[0396] Method 6: The feature element to be currently encoded
Number
[0397] In actual applications, it should be noted that in order to ensure platform consistency, the thresholds T1, T2, T3, T4, T5, and T6 may be rounded, that is, shifted to an integer and scaled.
[0398] It should be noted that the method for obtaining the threshold may alternatively use one of the following methods. This is not limited here.
[0399] Method 1: Threshold T1 is used as an example, any value within the range of the value of T1 is used as the threshold T1, and the threshold T1 is written into the bitstream. Specifically, the threshold is written into the bitstream, stored in the sequence header, picture header, slice / slice header, or SEI, and can be transmitted to the decoder side. Alternatively, another method can be used. This is not limited here. A similar method can also be used for the remaining thresholds T0, T2, T3, T4, T5, and T6.
[0400] Method 2: The encoder uses a fixed threshold that matches the decoder, and this fixed threshold does not need to be written to the bitstream or sent to the decoder. For example, threshold T1 is used as an example, and any value within the range of T1 is used directly as the value of T1. A similar method can be used for the remaining thresholds T0, T2, T3, T4, T5, and T6.
[0401] Method 3: A threshold candidate list is constructed, and the most likely values within the range of T1 values are added to the threshold candidate list. Each threshold corresponds to a threshold index number, the optimal threshold is determined, and the optimal threshold is used as the value of T1. The index number of the optimal threshold is used as the threshold index number of T1, and the threshold index numbers of T1 are written to the bitstream. Specifically, the thresholds are written to the bitstream and can be stored in the sequence header, picture header, slice header, or SEI and sent to the decoder side. Alternatively, other methods may be used, and are not limited here. Similar methods may also be used for the remaining thresholds T0, T2, T3, T4, T5, and T6.
[0402] Step 1805: The encoder transmits or stores the compressed bitstream.
[0403] Decoder side:
[0404] Step 1811: Obtain the bitstream of the audio feature variables that will be decoded.
[0405] Step 1812: Obtain the probability estimation results for the feature elements.
[0406] Entropy decoding is side information
number
number
number
number
number
number
number
number
number
number
number
[0407] Accordingly, it should be noted that the probability estimation method used by the decoder is the same as that used by the encoder in this embodiment, and the diagram of the structure of the probability estimation network used by the decoder is the same as that of the probability estimation network on the encoder side in this embodiment. Further details will not be explained here again.
[0408] 1813: Whether entropy decoding needs to be performed on the feature element to be decoded is determined based on the probability estimation result, and whether or not entropy decoding is performed on the decoded feature variable is determined based on the decision result.
number
[0409] One or more of the following methods, based on the probability estimation results, determine the feature element that will be currently decoded by entropy decoding:
number
number
[0410] Entropy decoding will now decode the feature elements
number
number
[0411] Method 1: When the probability distribution model is a Gaussian distribution, the feature elements that will be decoded at present
number
number
number
number
number
[0412] In particular, when the value of k is 0, it is the optimal value. The feature element that will be decoded is when the absolute value of the mean parameter μ of the Gaussian distribution is less than T1 and the variance σ of the Gaussian distribution is less than T2.
number
Number
Number
Number
[0413] Method 2: When the probability distribution is a Gaussian distribution, the mean value parameter μ and variance σ values of the feature element that will be currently decoded
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0414] Method 3: When the probability distribution is a Laplace distribution, the values of the location parameter μ and the scale parameter b are obtained based on the probability estimation result. When the relationship between the location parameter μ, the scale parameter b, and k satisfies abs(μ - k)+σ < T4 (the second condition), T4 is the fourth threshold, and the feature element that will be currently decoded
Number
Number
number
number
[0415] Method 4: When the probability distribution is a Laplace distribution, the values of the position parameter μ and the scale parameter b are obtained based on the probability estimation results. The current feature element to be decoded is obtained when the absolute value of the difference between the position parameter μ and k is less than the second threshold T5, and the scale parameter b is less than the third threshold T6 (second condition).
number
number
number
number
[0416] In particular, when the value of k is 0, it is the optimal value. The feature element that will be decoded is when the absolute value of the position parameter μ is less than T5 and the scale parameter b is less than T6.
number
number
number
number
[0417] Method 5: When the probability distribution is a Gaussian mixture, the feature elements that will be decoded are currently
number
number
number
number
number
[0418] Method 6: The probability P that the value of the feature element to be decoded is k, i.e., the probability estimation result P of the feature element to be decoded, is obtained based on the probability distribution of the feature element to be decoded. If the probability estimation result P satisfies the second condition, i.e., P is greater than the first threshold T0, then entropy decoding does not need to be performed on the first feature element, and the value of the feature element to be decoded is set to k. Otherwise, if the feature element to be decoded satisfies the first condition, i.e., P is less than or equal to the first threshold T0, then entropy decoding is performed on the bitstream and the value of the first feature element is obtained.
[0419] The aforementioned decoder value k is set to the corresponding encoder value k.
[0420] The method for obtaining thresholds T0, T1, T2, T3, T4, T5, T6, and T7 corresponds to the method on the encoder side, and one of the following methods may be used.
[0421] Method 1: The threshold is obtained from the bitstream. Specifically, the threshold is obtained from the sequence header, picture header, slice header, or SEI.
[0422] Method 2: The decoder uses a fixed threshold that matches that of the encoder.
[0423] Method 3: The threshold index number is obtained from the bitstream. Specifically, the threshold index number is obtained from the sequence header, picture header, slice header, or SEI. The decoder then constructs a list of threshold candidates in the same manner as the encoder and obtains the corresponding threshold from the list of threshold candidates based on the threshold index number.
[0424] In practical applications, it should be noted that, in order to ensure platform consistency, thresholds T1, T2, T3, T4, T5, and T6 may be rounded, i.e., shifted to integers and scaled.
[0425] Step 1814: Decoded feature variables
number
[0426] Example 1: Feature variables obtained through entropy decoding
number
[0427] Case 2: Feature variables obtained through entropy decoding
number
[0428] The aforementioned decoder value k is set to the corresponding encoder value k.
[0429] Figure 18 is a schematic diagram of an exemplary structure of the encoding device according to this application. As shown in Figure 18, the device in this example may correspond to an encoder 20A. The device may include an acquisition module 2001 and an encoding module 2002. The acquisition module 2001 may include the encoder network 204, rounding 206 (optional), probability estimation 210, side information extraction 214, generation network 216 (optional), and joint network 218 (optional) in the embodiments described above. The encoding module 2002 includes the encoding determination implementation 208 in the embodiments described above.
[0430] The acquisition module 2001 is configured to acquire feature data to be encoded, the feature data to be encoded includes multiple feature elements, the multiple feature elements include a first feature element, and the acquisition module is configured to acquire the probability estimation result of the first feature element. The encoding module 2002 is configured to determine whether to perform entropy coding on the first feature element based on the probability estimation result of the first feature element, and to perform entropy coding on the first feature element only when it is determined that entropy coding should be performed on the first feature element.
[0431] In possible implementations, deciding whether to perform entropy coding on the first feature element of the feature data includes the following: Entropy coding should be performed on the first feature element of the feature data if the probability estimation result of the first feature element of the feature data satisfies a predetermined condition. Entropy coding does not need to be performed on the first feature element of the feature data if the probability estimation result of the first feature element of the feature data does not satisfy the predetermined condition.
[0432] In possible implementations, the encoding module is further configured to determine, based on the probability estimation results of the feature data, that the probability estimation results of the feature data should be input to the generative network, and the network outputs decision information. When the decision information value for the first feature element is 1, the first feature element of the feature data should be encoded. When the decision information value for the first feature element is not 1, the first feature element of the feature data does not need to be encoded.
[0433] In a possible implementation, the predefined condition is that the probability value of the first feature element being k is less than or equal to a first threshold, where k is an integer.
[0434] In possible implementations, the predefined conditions are that the absolute difference between the mean of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to a second threshold, or the variance of the first feature element is greater than or equal to a third threshold, where k is an integer.
[0435] In another possible implementation, the predefined condition is that the sum of the variance of the probability distribution of the first feature element and the absolute difference between the mean of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to a fourth threshold, where k is an integer.
[0436] In a possible implementation, the probability value of the first feature element being k is the maximum probability value among all possible values of the first feature element.
[0437] In possible implementations, probability estimation is performed on the feature data to obtain probability estimation results for the feature elements of the feature data. The probability estimation result for the first feature element includes the probability value of the first feature element, and / or the first and second parameters of the probability distribution.
[0438] In possible implementations, the probability estimation results of the feature data are input to the generative network to obtain decision information for the first feature element. Whether or not to perform entropy coding on the first feature element is determined based on the decision information for the first feature element.
[0439] In possible implementations, if the decision information of the feature data is a decision map, and the value corresponding to the position of the first feature element in the decision map is a pre-set value, then it is determined that entropy coding should be performed on the first feature element. If the value corresponding to the position of the first feature element in the decision map is not a pre-set value, then it is determined that entropy coding should not be performed on the first feature element.
[0440] In a possible implementation, it is determined that entropy coding should be performed on the first feature element if the determination information for the feature data is a pre-set value. If the determination information is not a pre-set value, it is determined that entropy coding does not need to be performed on the first feature element. In a possible implementation, the coding module is further configured to construct a threshold candidate list for the first threshold, to put the first threshold into the threshold candidate list for the first threshold, to have an index number corresponding to the first threshold, and to write the index number of the first threshold into the coded bitstream, the length of the threshold candidate list for the first threshold may be set to T, where T is an integer greater than or equal to 1.
[0441] The apparatus in this embodiment may be used in a technical solution implemented by an encoder in the embodiment of the method shown in Figures 3A to 3D. The implementation principle and its technical effects are similar. Further details are not described here.
[0442] Figure 19 is a schematic diagram of an exemplary structure of a decoding device according to this application. As shown in Figure 19, the device in this example may correspond to a decoder 30. The device may include an acquisition module 2101 and a decoding module 2102. The acquisition module 2101 may include a probability estimation 302, a generation network 310 (optional), and a joint network 312 in the embodiments described above. The decoding module 2102 includes a decoding decision implementation 304 and a decoder network 306 in the embodiments described above.
[0443] The acquisition module 2101 is configured to acquire a bitstream of feature data to be decoded, and is configured to acquire the probability estimation result of the first feature element, where the feature data to be decoded includes multiple feature elements, and where the multiple feature elements include a first feature element. The decoding module 2102 is configured to decide whether to perform entropy decoding on the first feature element based on the probability estimation result of the first feature element, and to perform entropy decoding on the first feature element only when it is determined that entropy decoding should be performed on the first feature element.
[0444] In possible implementations, deciding whether to perform entropy decoding on the first feature element of the feature data includes the following: The first feature element of the feature data should be decoded if its probability estimate satisfies a predetermined condition. Alternatively, if the probability estimate of the first feature element of the feature data does not satisfy the predetermined condition, the first feature element of the feature data does not need to be decoded, and its feature value is set to k, where k is an integer.
[0445] In possible implementations, the decoding module is further configured to determine, based on the probability estimation results of the feature data, that the probability estimation results of the feature data are input to the decision network module, and the network outputs decision information. The first feature element of the feature data is decoded when the position value in the decision information corresponding to the first feature element of the feature data is 1. The first feature element of the feature data is not decoded when the position value in the decision information corresponding to the first feature element of the feature data is not 1, and the feature value of the first feature element is set to k, where k is an integer.
[0446] In a possible implementation, the predefined condition is that the probability value of the first feature element being k is less than or equal to a first threshold, where k is an integer.
[0447] In another possible implementation, the predefined conditions are that the absolute difference between the mean of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to a second threshold, or that the variance of the probability distribution of the first feature element is greater than or equal to a third threshold.
[0448] In another possible implementation, the pre-set condition is that the sum of the variance of the probability distribution of the first feature element and the absolute difference between the mean of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to a fourth threshold.
[0449] In possible implementations, probability estimation is performed on the feature data to obtain probability estimation results for the feature elements of the feature data. The probability estimation result for the first feature element includes the probability value of the first feature element, and / or the first and second parameters of the probability distribution.
[0450] In a possible implementation, the probability value of the first feature element being k is the maximum probability value among all possible values of the first feature element.
[0451] In a possible implementation, the probability estimation result for the Nth feature element includes, namely, the probability value of the Nth feature element, the first and second parameters of the probability distribution, and at least one of the decision information. The first feature element of the feature data is decoded when the position value in the decision information corresponding to the first feature element of the feature data is 1. The first feature element of the feature data is not decoded when the position value in the decision information corresponding to the first feature element of the feature data is not 1, and the feature value of the first feature element is set to k, where k is an integer.
[0452] In possible implementations, the probability estimation results of the feature data are input to the generative network to obtain decision information for the first feature element. If the value of the decision information for the first feature element is a preset value, it is determined that entropy decoding should be performed on the first feature element. If the value of the decision information for the first feature element is not a preset value, it is determined that entropy decoding does not need to be performed on the first feature element, and the feature value of the first feature element is set to k, where k is an integer and k is one of several candidate values for the first feature element.
[0453] In a possible implementation, the acquisition module is further configured to construct a list of threshold candidates for the first threshold, obtain the index number of the threshold candidate list for the first threshold by decoding the bitstream, and use the value of the position in the threshold candidate list for the first threshold, corresponding to the index number of the first threshold, as the value of the first threshold. The length of the threshold candidate list for the first threshold may be set to T, where T is an integer greater than or equal to 1.
[0454] The apparatus in this embodiment may be used in technical solutions implemented by a decoder in the embodiments of the method shown in Figures 10B, 13B, and 16. The implementation principle and its technical effects are similar. Further details are not described here.
[0455] Those skilled in the art will understand that the functions described with reference to the various illustrative logical blocks, modules, and algorithmic steps disclosed and described herein can be implemented by hardware, software, firmware, or any combination thereof. If implemented by software, the functions described with reference to the illustrative logical blocks, modules, and steps can be stored as one or more instructions or codes in a computer-readable medium, or transmitted over a computer-readable medium, and determined by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or it may include any communication medium that facilitates the transmission of a computer program from one place to another (for example, according to a communication protocol). Thus, the computer-readable medium can generally correspond to (1) a non-temporary tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the technology described in this application. A computer program product may include a computer-readable medium.
[0456] Such computer-readable storage media may include, but are not limited to, computer-readable storage media, RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other media that can be accessed by a computer and can store program code required in the form of instructions or data structures. In addition, any connection is appropriately called a computer-readable medium. For example, if instructions are transmitted from a website, server or another remote source via coaxial cable, optical fiber, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, radio, or microwave, then coaxial cable, optical fiber, twisted pair, DSL, or wireless technology such as infrared, radio, or microwave are included in the definition of a medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other temporary media, but in practice mean non-temporary tangible storage media. As used in this specification, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), and Blu-ray discs. A disk typically reproduces data magnetically, while a disc reproduces data optically using a laser. The combination of these should also be included within the scope of computer-readable media.
[0457] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or equivalent integrated circuits or individual logic circuits. Therefore, the term “processor” as used in this specification may refer to the aforementioned structures or any other structures applicable to implementations of the techniques described herein. In addition, in some embodiments, the functions described with reference to the illustrative logic blocks, modules, and steps described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Furthermore, the techniques may be fully implemented in one or more circuits or logic elements.
[0458] The technology described in this application may be implemented in a variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this application to highlight the functional aspects of devices configured to determine the disclosed technology, but implementation by different hardware units is not necessarily required. In practice, as described above, various units may be combined with a codec hardware unit in combination with appropriate software and / or firmware, or may be provided by interoperable hardware units (including one or more processors as described above).
[0459] The foregoing description is merely an exemplary specific implementation of this application and is not intended to limit the scope of protection of this application. Any modification or substitution readily understood by those skilled in the art within the technical scope disclosed in this application falls within the scope of protection of this application. Accordingly, the scope of protection of this application is subject to the scope of protection of the claims. [Explanation of Symbols]
[0460] 12 Source Devices 14 Destination Devices 17 Source Data 18 preprocessors 19 Preprocessed source data 20 Encoder Networks 21 Feature data to be encoded 22 preprocessors 23 Preprocessed feature data to be encoded 24 Entropy Decoding 25 Encoded bitstream 26 Communication Interfaces 27 Communication Channels 28 Communication Interfaces 29 Decrypted bitstream 30 Entropy Decoding 31 Decoded feature data 32 Post-Processors 33 Post-processed decoded feature data 34 Decoder Network 35 Reconstructed data 36 Post-processors 37 Post-processed and reconstructed data 38 Display Devices 40. Probability Estimation 41 Probability Estimation Results 42 Probability Estimation Results 44 Memory storage 50 Coding Systems 51 Imaging devices 52 Antennas 53 processors 55 Display Devices 56 Processing Units 202 inputs 204 Encoder Network 206 Round 208 Encoding Determination Implementation 210 Probability Estimation 212 Output 214 Side Information Extraction 216 Generating Network 218 Joint Network 261 Determination of Encoding Elements 262 Entropy coding 301 Determination of decryption element 302 Entropy Decoding 306 Decoder Network 310 Generating Network 400 coding devices 410 Entrance Port 420 Receiver Unit 430 processors 440 Transmitter Unit 450 Exit Port 460 memory 470 Coding Modules 502 Processors 504 memory 506 data 508 Operating Systems 510 Application Programs 512 Bus 518 displays 1801 Capture device 1802 Preprocessing before import 1803 Video Encoding 1804 Sent 1805 Received 1806 bitstream decryption 1807 Rendering and Display 1901 Feature Extraction 1902 Side Information Extraction 1903 Probability Estimation 1904 encoding 1905 Decryption 1906 Machine Vision (Auditory) Task 2001 Acquisition Module 2002 Encoding Module 2101 Acquisition Module 2102 Decoding Module
Claims
1. A feature data encoding method, A step of obtaining feature data to be encoded, wherein the feature data to be encoded includes a plurality of feature elements, and the plurality of feature elements include a first feature element. The steps include obtaining the probability estimation result of the first feature element, A step of determining whether to perform entropy coding on the first feature element based on the probability estimation result of the first feature element, The steps include: performing entropy coding on the first feature element only when it is determined that entropy coding should be performed on the first feature element; Methods that include...
2. The step of determining whether to perform entropy coding on the first feature element based on the probability estimation result of the first feature element is: A step in which it is determined that entropy coding should be performed on the first feature element when the probability estimation result of the first feature element satisfies a predetermined condition, or The method according to claim 1, further comprising the step of determining that entropy coding is not required for the first feature element when the probability estimation result of the first feature element does not satisfy a predetermined condition.
3. The method according to claim 2, wherein when the probability estimation result of the first feature element is a probability value that the value of the first feature element is k, the pre-set condition is that the probability value that the value of the first feature element is k is less than or equal to a first threshold, where k is an integer, and k is one of a plurality of candidate values of the first feature element.
4. When the probability estimation result of the first feature element includes the first parameter and the second parameter of the probability distribution of the first feature element, the pre-set condition is The absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to the second threshold. The second parameter of the probability distribution of the first feature element is greater than or equal to the third threshold, or The method according to claim 2, wherein the sum of the second parameter of the probability distribution of the first feature element and the absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to a fourth threshold, where k is an integer and k is one of a plurality of candidate values of the first feature element.
5. When the probability distribution is a Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean value of the Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the variance of the Gaussian distribution of the first feature element, or The method according to claim 4, wherein, when the probability distribution is a Laplace distribution, the first parameter of the probability distribution of the first feature element is the position parameter of the Laplace distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the scale parameter of the Laplace distribution of the first feature element.
6. The method described above is The method according to claim 3, further comprising the steps of constructing a threshold candidate list, adding a first threshold to the threshold candidate list, and writing an index number corresponding to the first threshold to an encoded bitstream, wherein the length of the threshold candidate list is T, and T is an integer of 1 or more.
7. When the probability estimation result of the first feature element is obtained through a Gaussian mixture distribution, the pre-set conditions are: The sum of the arbitrary variance of the Gaussian mixture distribution of the first feature element and the sum of the absolute values of the differences between all the mean values of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than or equal to a fifth threshold. The difference between any mean value of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than or equal to a sixth threshold, or The condition is that any variance of the mixture Gaussian distribution of the first feature element is greater than or equal to the seventh threshold, The method according to claim 2, wherein k is an integer and k is one of a plurality of candidate values of the first feature element.
8. When the probability estimation result of the first feature element is obtained through an asymmetric Gaussian distribution, the pre-set condition is, The absolute value of the difference between the mean value of the asymmetric Gaussian distribution of the first feature element and the value k of the first feature element is greater than or equal to an eighth threshold. The first variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to the ninth threshold, or The second variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to a tenth threshold. The method according to claim 2, wherein k is an integer and k is one of a plurality of candidate values of the first feature element.
9. The method according to any one of claims 3 to 8, wherein the probability value of the first feature element being k is the maximum probability value among the probability values of all candidate values of the first feature element.
10. The step of determining whether to perform entropy coding on the first feature element based on the probability estimation result of the first feature element is: The steps include: inputting the probability estimation results of the feature data into the generation network to obtain determination information for the first feature element; A step of determining whether to perform entropy coding on the first feature element based on the determination information of the first feature element. The method according to claim 1, including the method described in claim 1.
11. When the determination information for the feature data is a determination map, and the value corresponding to the position where the first feature element is located in the determination map is a preset value, it is determined that entropy coding needs to be performed on the first feature element. The method according to claim 10, wherein it is determined that entropy coding does not need to be performed on the first feature element when the value corresponding to the position in the determination map where the first feature element is located is not the preset value.
12. When the determination information for the feature data is a predetermined value, it is determined that entropy coding needs to be performed on the first feature element. The method according to claim 10, wherein it is determined that entropy coding is not required for the first feature element when the determination information is not the preset value.
13. The method according to any one of claims 1 to 12, wherein when the plurality of feature elements further include a second feature element and it is determined that entropy coding does not need to be performed on the second feature element, the entropy coding on the second feature element is skipped.
14. The method described above is The method according to any one of claims 1 to 13, further comprising the step of writing the entropy coding results of the plurality of feature elements, including the first feature element, to the coded bitstream.
15. A feature data decoding method, A step of obtaining a bitstream of feature data to be decoded, The feature data to be decoded includes a plurality of feature elements, and the plurality of feature elements include a first feature element, The steps include obtaining the probability estimation result of the first feature element, A step of determining whether to perform entropy decoding on the first feature element based on the probability estimation result of the first feature element, The steps include: performing entropy decoding on the first feature element only when it is determined that entropy decoding needs to be performed on the first feature element; Methods that include...
16. The step of determining whether to perform entropy decoding on the first feature element based on the probability estimation result of the first feature element is: A step in which it is determined that entropy decoding should be performed on the first feature element of the feature data when the probability estimation result of the first feature element satisfies a predetermined condition, or The method according to claim 15, comprising the step of determining that, when the probability estimation result of the first feature element does not satisfy a predetermined condition, it is not necessary to perform entropy decoding on the first feature element of the feature data, and setting the feature value of the first feature element to k, wherein k is an integer and k is one of a plurality of candidate values of the first feature element.
17. The method according to claim 16, wherein when the probability estimation result of the first feature element is a probability value that the value of the first feature element is k, the pre-set condition is that the probability value that the value of the first feature element is k is less than or equal to a first threshold, where k is an integer, and k is one of the plurality of candidate values of the first feature element.
18. When the probability estimation result of the first feature element includes the first parameter and the second parameter of the probability distribution of the first feature element, the pre-set condition is The absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to the second threshold. The second parameter of the probability distribution of the first feature element is greater than or equal to the third threshold, or The sum of the second parameter of the probability distribution of the first feature element and the absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to a fourth threshold. The method according to claim 16, wherein k is an integer and k is one of the plurality of candidate values of the first feature element.
19. When the probability distribution is a Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean value of the Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the variance of the Gaussian distribution of the first feature element, or The method according to claim 18, wherein, when the probability distribution is a Laplace distribution, the first parameter of the probability distribution of the first feature element is the position parameter of the Laplace distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the scale parameter of the Laplace distribution of the first feature element.
20. When the probability estimation result of the first feature element is obtained through a Gaussian mixture distribution, the pre-set conditions are: The sum of an arbitrary variance of the Gaussian mixture distribution of the first feature element and the sum of the absolute values of the differences between all the mean values of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than or equal to a fifth threshold. The difference between any mean value of the mixture Gaussian distribution of the first feature element and the value k of the first feature element is greater than or equal to a sixth threshold, or The condition is that any variance of the mixture Gaussian distribution of the first feature element is greater than or equal to the seventh threshold, The method according to claim 16, wherein k is an integer and k is one of the plurality of candidate values of the first feature element.
21. When the probability estimation result of the first feature element is obtained through an asymmetric Gaussian distribution, the pre-set condition is, The absolute value of the difference between the mean value of the asymmetric Gaussian distribution of the first feature element and the value k of the first feature element is greater than or equal to an eighth threshold. The first variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to the ninth threshold, or The second variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to a tenth threshold. The method according to claim 16, wherein k is an integer and k is one of the plurality of candidate values of the first feature element.
22. The apparatus according to any one of claims 16 to 21, wherein the probability value of the first feature element being k is the maximum probability value among the probability values of all candidate values of the first feature element.
23. The step of determining whether to perform entropy decoding on the first feature element based on the probability estimation result of the first feature element is: The steps include: inputting the probability estimation results of the feature data into the generation network to obtain determination information for the first feature element; A step of determining whether to perform entropy decoding on the first feature element based on the determination information of the first feature element. The method according to claim 15, including the method described in claim 15.
24. When the determination information for the feature data is a determination map, and the value corresponding to the position where the first feature element is located in the determination map is a preset value, it is determined that entropy decoding needs to be performed on the first feature element. The method according to claim 23, wherein when the value corresponding to the position in the determination map where the first feature element is located is not the preset value, it is determined that entropy decoding does not need to be performed on the first feature element.
25. When the determination information for the feature data is a predetermined value, it is determined that entropy decoding needs to be performed on the first feature element. The method according to claim 23, wherein when the determination information is not the preset value, it is determined that entropy decoding does not need to be performed on the first feature element.
26. The method described above is The method according to any one of claims 15 to 25, further comprising the step of obtaining the reconstructed data or machine-oriented task data, which is obtained after the feature data has passed through a decoder network.
27. A feature data encoding device, An acquisition module configured to acquire feature data to be encoded, wherein the feature data to be encoded includes a plurality of feature elements, the plurality of feature elements include a first feature element, and the acquisition module is configured to acquire the probability estimation result of the first feature element, An encoding module configured to determine whether to perform entropy coding on the first feature element based on the probability estimation result of the first feature element, and to perform entropy coding on the first feature element only when it is determined that entropy coding should be performed on the first feature element. A device including a device.
28. The determination of whether to perform entropy coding on the first feature element based on the probability estimation result of the first feature element is as follows: When the probability estimation result of the first feature element satisfies a predetermined condition, it is determined that entropy coding should be performed on the first feature element of the feature data, or The apparatus according to claim 27, further comprising determining that when the probability estimation result of the first feature element does not satisfy a predetermined condition, it is not necessary to perform entropy coding on the first feature element of the feature data.
29. The apparatus according to claim 28, wherein when the probability estimation result of the first feature element is a probability value that the value of the first feature element is k, the pre-set condition is that the probability value that the value of the first feature element is k is less than or equal to a first threshold, where k is an integer and k is one of a plurality of candidate values for the first feature element.
30. When the probability estimation result of the first feature element includes the first parameter and the second parameter of the probability distribution of the first feature element, the pre-set condition is The absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to the second threshold. The second parameter of the probability distribution of the first feature element is greater than or equal to the third threshold, or The sum of the second parameter of the probability distribution of the first feature element and the absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to a fourth threshold. The apparatus according to claim 28, wherein k is an integer and k is one of a plurality of candidate values of the first feature element.
31. When the probability distribution is a Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean value of the Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the variance of the Gaussian distribution of the first feature element, or The apparatus according to claim 30, wherein, when the probability distribution is a Laplace distribution, the first parameter of the probability distribution of the first feature element is the position parameter of the Laplace distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the scale parameter of the Laplace distribution of the first feature element.
32. The apparatus according to claim 29, wherein the encoding module is further configured to construct a threshold candidate list, to put a first threshold into the threshold candidate list, and to write an index number corresponding to the first threshold into an encoded bitstream, the length of the threshold candidate list being T, where T is an integer of 1 or more.
33. When the probability estimation result of the first feature element is obtained through a Gaussian mixture distribution, the pre-set conditions are: The sum of the arbitrary variance of the Gaussian mixture distribution of the first feature element and the sum of the absolute values of the differences between all the mean values of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than or equal to a fifth threshold. The difference between any mean value of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than or equal to a sixth threshold, or The condition is that any variance of the mixture Gaussian distribution of the first feature element is greater than or equal to the seventh threshold, The apparatus according to claim 28, wherein k is an integer and k is one of a plurality of candidate values of the first feature element.
34. When the probability estimation result of the first feature element is obtained through an asymmetric Gaussian distribution, the pre-set condition is, The absolute value of the difference between the mean value of the asymmetric Gaussian distribution of the first feature element and the value k of the first feature element is greater than or equal to an eighth threshold. The first variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to the ninth threshold, or The second variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to a tenth threshold. The apparatus according to claim 28, wherein k is an integer and k is one of a plurality of candidate values of the first feature element.
35. The method according to any one of claims 29 to 34, wherein the probability value of the first feature element being k is the maximum probability value among the probability values of all candidate values of the first feature element.
36. The determination of whether to perform entropy coding on the first feature element based on the probability estimation result of the first feature element is as follows: The probability estimation results of the feature data are input into the generative network to obtain the determination information for the first feature element, Based on the determination information of the first feature element, it is determined whether or not to perform entropy coding on the first feature element. The apparatus according to claim 27, including the apparatus described in claim 27.
37. When the determination information for the feature data is a determination map, and the value corresponding to the position where the first feature element is located in the determination map is a preset value, it is determined that entropy coding needs to be performed on the first feature element. The apparatus according to claim 36, wherein it is determined that entropy coding is not required for the first feature element when the value corresponding to the position in the determination map where the first feature element is located is not the preset value.
38. When the determination information for the feature data is a predetermined value, it is determined that entropy coding needs to be performed on the first feature element. The apparatus according to claim 36, wherein it is determined that entropy coding is not required for the first feature element when the determination information is not the preset value.
39. The apparatus according to any one of claims 27 to 38, wherein when the plurality of feature elements further include a second feature element and it is determined that entropy coding does not need to be performed on the second feature element, the performance of entropy coding on the second feature element is skipped.
40. The aforementioned encoding module, The apparatus according to any one of claims 27 to 39, further comprising writing the entropy coding results of the plurality of feature elements, including the first feature element, to the coded bitstream.
41. A feature data decoding device, An acquisition module configured to acquire a bitstream of feature data to be decoded, wherein the feature data to be decoded includes a plurality of feature elements, the plurality of feature elements include a first feature element, and the acquisition module is configured to acquire the probability estimation result of the first feature element, A decoding module configured to determine whether to perform entropy decoding on the first feature element based on the probability estimation result of the first feature element, and to perform entropy decoding on the first feature element only when it is determined that entropy decoding should be performed on the first feature element. A device including a device.
42. The determination of whether to perform entropy decoding on the first feature element based on the probability estimation result of the first feature element is as follows: When the probability estimation result of the first feature element satisfies a predetermined condition, it is determined that entropy decoding should be performed on the first feature element of the feature data, or The apparatus according to claim 41, comprising determining that entropy decoding does not need to be performed on the first feature element of the feature data when the probability estimation result of the first feature element does not satisfy a predetermined condition, and setting the feature value of the first feature element to k, wherein k is an integer and k is one of a plurality of candidate values.
43. The apparatus according to claim 42, wherein when the probability estimation result of the first feature element is a probability value that the value of the first feature element is k, the pre-set condition is that the probability value that the value of the first feature element is k is less than or equal to a first threshold, where k is an integer and k is one of the plurality of candidate values of the first feature element.
44. When the probability estimation result of the first feature element includes the first parameter and the second parameter of the probability distribution of the first feature element, the pre-set condition is The absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to the second threshold. The second parameter of the probability distribution of the first feature element is greater than or equal to the third threshold, or The sum of the second parameter of the probability distribution of the first feature element and the absolute value of the difference between the first parameter of the probability distribution of the first feature element and the value k of the first feature element is greater than or equal to a fourth threshold. The apparatus according to claim 42, wherein k is an integer and k is one of the plurality of candidate values of the first feature element.
45. When the probability distribution is a Gaussian distribution, the first parameter of the probability distribution of the first feature element is the mean value of the Gaussian distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the variance of the Gaussian distribution of the first feature element, or The apparatus according to claim 44, wherein, when the probability distribution is a Laplace distribution, the first parameter of the probability distribution of the first feature element is the position parameter of the Laplace distribution of the first feature element, and the second parameter of the probability distribution of the first feature element is the scale parameter of the Laplace distribution of the first feature element.
46. When the probability estimation result of the first feature element is obtained through a Gaussian mixture distribution, the pre-set conditions are: The sum of an arbitrary variance of the Gaussian mixture distribution of the first feature element and the sum of the absolute values of the differences between all the mean values of the Gaussian mixture distribution of the first feature element and the value k of the first feature element is greater than or equal to a fifth threshold. The difference between any mean value of the mixture Gaussian distribution of the first feature element and the value k of the first feature element is greater than or equal to a sixth threshold, or The condition is that any variance of the mixture Gaussian distribution of the first feature element is greater than or equal to the seventh threshold, The apparatus according to claim 42, wherein k is an integer and k is one of the plurality of candidate values of the first feature element.
47. When the probability estimation result of the first feature element is obtained through an asymmetric Gaussian distribution, the pre-set condition is, The absolute value of the difference between the mean value of the asymmetric Gaussian distribution of the first feature element and the value k of the first feature element is greater than or equal to an eighth threshold. The first variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to the ninth threshold, or The second variance of the asymmetric Gaussian distribution of the first feature element is greater than or equal to a tenth threshold. The apparatus according to claim 42, wherein k is an integer and k is one of the plurality of candidate values of the first feature element.
48. The apparatus according to any one of claims 42 to 47, wherein the probability value of the first feature element being k is the maximum probability value among the probability values of all candidate values of the first feature element.
49. The determination of whether to perform entropy decoding on the first feature element based on the probability estimation result of the first feature element is as follows: The probability estimation results of the feature data are input into the generative network to obtain the determination information for the first feature element, Based on the determination information of the first feature element, it is determined whether or not to perform entropy decoding on the first feature element. The apparatus according to claim 41, including the apparatus described in claim 41.
50. When the determination information for the feature data is a determination map, and the value corresponding to the position where the first feature element is located in the determination map is a preset value, then entropy decoding must be performed on the first feature element. The apparatus according to claim 49, wherein when the value corresponding to the position in the determination map where the first feature element is located is not the preset value, it is not necessary to perform entropy decoding on the first feature element.
51. When the determination information for the feature data is a predetermined value, it is determined that entropy decoding needs to be performed on the first feature element. The apparatus according to claim 49, wherein it is determined that entropy decoding does not need to be performed on the first feature element when the determination information is not the preset value.
52. The apparatus according to any one of claims 41 to 51, wherein the decoding module is further configured to obtain the reconstructed data or machine-oriented task data by passing the feature data through a decoder network.
53. A feature data encoding method, A step of obtaining feature data to be encoded, wherein the feature data includes a plurality of feature elements, and the plurality of feature elements include a first feature element. The steps include obtaining side information of the feature data, inputting the side information of the feature data into a joint network to obtain determination information for the first feature element, A step of determining whether to perform entropy coding on the first feature element based on the determination information of the first feature element, The steps include: performing entropy coding on the first feature element only when it is determined that entropy coding should be performed on the first feature element; Methods that include...
54. When the determination information of the feature data is a determination map, and the value corresponding to the position of the first feature element in the determination map is a preset value, it is determined that entropy coding needs to be performed on the first feature element. The method according to claim 53, wherein it is determined that entropy coding does not need to be performed on the first feature element when the value corresponding to the position where the first feature element is located in the determination map is not the preset value.
55. When the determination information of the feature data is a predetermined value, it is determined that entropy coding needs to be performed on the first feature element. The method according to claim 53, wherein it is determined that entropy coding is not required for the first feature element when the determination information is not the preset value.
56. The method according to any one of claims 53 to 55, wherein when the plurality of feature elements further include a second feature element and it is determined that entropy coding does not need to be performed on the second feature element, the entropy coding on the second feature element is skipped.
57. The method described above is The method according to any one of claims 53 to 56, further comprising the step of writing the entropy coding results of the plurality of feature elements, including the first feature element, to an encoded bitstream.
58. A feature data decoding method, A step of obtaining a bitstream of feature data to be decoded and side information of the feature data to be decoded, The feature data to be decoded includes a plurality of feature elements, and the plurality of feature elements include a first feature element, The steps include inputting the side information of the feature data to be decoded into a joint network to obtain determination information for the first feature element, A step of determining whether to perform entropy decoding on the first feature element based on the determination information of the first feature element, The steps include: performing entropy decoding on the first feature element only when it is determined that entropy decoding needs to be performed on the first feature element; Methods that include...
59. When the determination information of the feature data is a determination map, and the value corresponding to the position of the first feature element in the determination map is a preset value, it is determined that entropy decoding needs to be performed on the first feature element. The method according to claim 58, wherein when the value corresponding to the position where the first feature element is located in the determination map is not the preset value, it is determined that entropy decoding does not need to be performed on the first feature element, and the feature value of the first feature element is set to k, where k is an integer.
60. When the determination information of the feature data is a predetermined value, it is determined that entropy coding needs to be performed on the first feature element. The method according to claim 58, wherein when the determination information is not the preset value, it is determined that entropy coding does not need to be performed on the first feature element, the feature value of the first feature element is set to k, and k is an integer.
61. The method described above is The method according to any one of claims 58 to 60, further comprising the step of obtaining the reconstructed data or machine-oriented task data by passing the feature data through a decoder network.
62. An encoder comprising a processing circuit configured to perform the method according to any one of claims 1 to 14 and the method according to any one of claims 53 to 57.
63. A decoder comprising a processing circuit configured to perform the method according to any one of claims 15 to 26 and the method according to any one of claims 58 to 61.
64. A computer program product comprising program code, wherein, when the program code is determined in a computer or processor, the program code is used to determine the method according to any one of claims 1 to 26 and the method according to any one of claims 53 to 61.
65. A non-temporary computer-readable storage medium comprising a bitstream obtained by using the encoding method described in claim 14 or 57.
66. It is an encoder, One or more processors, An encoder comprising a non-temporary computer-readable storage medium coupled to the processor and storing a program determined by the processor, wherein when the program is determined by the processor, the encoder is enabled to perform the method according to any one of claims 1 to 14 and any one of claims 53 to 57.
67. It is a decoder, One or more processors, A decoder comprising a non-temporary computer-readable storage medium coupled to the processor and storing a program determined by the processor, wherein when the program is determined by the processor, the decoder is enabled to perform the method according to any one of claims 15 to 26 and any one of claims 58 to 61.
68. It is an encoder, One or more processors, An encoder comprising a non-temporary computer-readable storage medium coupled to the processor and storing a program determined by the processor, wherein when the program is determined by the processor, the encoder is enabled to perform the method according to any one of claims 1 to 14 and any one of claims 53 to 57.
69. A picture or audio processor comprising a processing circuit configured to perform the method according to any one of claims 1 to 26 and the method according to any one of claims 53 to 61.
70. A non-temporary computer-readable storage medium containing program code, wherein, when the program code is determined by a computer device, the program code is used to perform the method according to any one of claims 1 to 26 and the method according to any one of claims 53 to 61.