Data Encoding and Decoding Methods and Related Devices
The data encoding method addresses inefficiencies in existing compression algorithms by utilizing side information and optimized scaling processes to enhance compression performance and efficiency for diverse data types.
Patent Information
- Application Number
- JP2024575845
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-08
- Filing Date
- 2023-06-12
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-06-12
AI Technical Summary
Existing data compression algorithms face challenges in achieving high compression performance and efficiency, particularly in encoding and decoding processes for various data types such as image, video, audio, and point cloud data.
A data encoding method involving side information feature extraction, quantization, entropy encoding, and scaling processing of feature maps, combined with optimized network parameters, to improve data compression performance by reducing redundancy and information loss.
The method enhances data compression performance by reducing the volume of bitstreams and improving encoding and decoding efficiency, while maintaining or enhancing the quality of reconstructed data.
Smart Images

Figure 2025521640000001_ABST
Abstract
Description
Technical Field
[0001] This application claims priority to Chinese Patent Application No. 202210801030.7, titled "Data Encoding and Decoding Methods and Related Devices", filed with the State Intellectual Property Office of China on July 8, 2022, the entire disclosure of which is incorporated herein by reference in its entirety.
[0002] Embodiments of the present application relate to the field of data processing, and in particular, to data encoding and decoding methods and related devices.
Background Art
[0003] The purpose of data compression technology is to reduce redundant information in data so that data can be stored and transmitted in a more efficient format. In other words, data compression is a lossy or lossless representation of the original data in fewer bits. Since there is redundancy in the data, it can be compressed. The purpose of data compression is to reduce the number of bits required to represent data by eliminating data redundancy.
[0004] Methods for improving the compression performance of data compression algorithms are a hot topic being studied by those skilled in the art.
Summary of the Invention
[0005] The present application provides data encoding and decoding methods and related devices to improve data compression performance.
[0006] According to a first aspect, a data encoding method is provided, which is executed by a data encoding device. The method comprises the following steps.
[0007] Execute side information feature extraction on the first feature map of the current data to obtain a side information feature map. Next, perform quantization processing on the side information feature map to obtain a first quantized feature map. Perform entropy encoding on the first quantized feature map to obtain a first bitstream of the current data. Perform scaling processing on the second feature map based on the scaling factor to obtain a scaled feature map. Perform quantization processing on the scaled feature map to obtain a second quantized feature map. Obtain a scaling factor based on the first quantized feature map. Perform scaling processing on the first probability distribution parameter based on the scaling factor to obtain a second probability distribution parameter. Perform entropy encoding on the second quantized feature map based on the second probability distribution parameter to obtain a second bitstream of the current data.
[0008] The data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data. In this embodiment, the first feature map is a feature map obtained by performing feature extraction on the complete current data, and the second feature map is a feature map obtained based on the current data.
[0009] Side Information means that existing information Y is used to assist in the encoding of information X so that the encoding length of information X can be made shorter. In other words, the redundancy within information X is reduced. Information Y is the side information. In this embodiment of the present application, the side information is information extracted from the first feature map and used to assist in the encoding and decoding of the first feature map. Further, the bitstream is a bitstream generated after the encoding process.
[0010] In this solution, entropy encoding is performed based on the first quantized feature map to obtain the first bitstream of the current data. The first bitstream is the bitstream obtained after entropy encoding is performed on the first quantized feature map. Further, the scaling coefficient may be obtained through an estimation based on the first quantized feature map. In this way, scaling processing may be performed on the second feature map based on the scaling coefficient to obtain a scaled feature map. Next, quantization processing is performed on the scaled feature map to obtain a second quantized feature map. Scaling processing is performed on the first probability distribution parameter based on the scaling coefficient to obtain a second probability distribution parameter. Finally, entropy encoding is performed on the second quantized feature map based on the second probability distribution parameter to obtain the second bitstream of the current data. The second bitstream is the bitstream obtained after entropy encoding is performed on the second quantized feature map. The first bitstream and the second bitstream are used together as the total bitstream of the current data. The data encoding method in the prior art only includes the stage of performing scaling processing on the feature map. In this solution, the second feature map and the first probability distribution parameter are scaled by using the same scaling coefficient so that the degree of agreement between the second probability distribution parameter and the second quantized feature map is higher, thereby improving the encoding accuracy of the second quantized feature map, that is, improving the data compression performance.
[0011] In some possible embodiments of the first aspect, the second feature map is a residual feature map of the current data obtained based on the third feature map and the first probability distribution parameter of the current data, and the third feature map is obtained by performing feature extraction on the current data.
[0012] In this embodiment, the third feature map is a feature map obtained by performing feature extraction on the complete current data, and the third feature map is different from or the same as the first feature map. The residual of the third feature map, that is, the residual feature map, may be obtained based on the third feature map and the first probability distribution parameter, and the residual feature map is used as the second feature map of the current data.
[0013] In this solution, during encoding, after performing scaling processing on the residual feature map of the current data to obtain a scaled feature map, quantization processing is performed on the scaled feature map so that the quantization loss of the scaled feature map becomes smaller. That is, the information loss of the second bitstream becomes smaller. This helps to improve the data quality of the reconstructed data obtained through decoding based on the second bitstream. Furthermore, compared with the encoding network structure in the prior art, in this embodiment of the present application, after the training of the entire encoding network is completed, the network parameters of the entire encoding network including the network for generating the first bitstream and the network for generating the second bitstream can be optimized, so an encoding network structure including scaling for the residual feature map is used. Therefore, by using the encoding network structure in this embodiment of the present application, the data volume of the total bitstream of the current data can be reduced, and the encoding efficiency can be improved. That is, the scaling for the residual feature map and the scaling for the first probability distribution parameter are combined, thereby further improving the data compression performance.
[0014] In some possible embodiments of the first aspect, the second feature map is a feature map obtained by performing feature extraction on the current data. The first feature map is the same as or different from the second feature map.
[0015] In some possible embodiments of the first aspect, the first probability distribution parameter includes an average and / or a variance.
[0016] In some possible embodiments of the first aspect, the first probability distribution parameter is obtained based on a first quantization feature map.
[0017] In some possible embodiments of the first aspect, the first probability distribution parameter is a pre-set probability distribution parameter.
[0018] In some possible embodiments of the first aspect, the data encoding method further comprises the step of transmitting a first bit stream and a second bit stream.
[0019] After the first bit stream and the second bit stream of the current data are obtained by using the data encoding method in this solution means, the first bit stream and the second bit stream may be transmitted to another device according to requirements, whereby the other device can process the first bit stream and the second bit stream.
[0020] In some possible embodiments of the first aspect, the first bit stream and the second bit stream are stored in the form of a bit stream file.
[0021] According to a second aspect, the present application further provides a data encoding method, which is executed by a data encoding device. The method comprises the following steps.
[0022] Execute side information feature extraction on the first feature map of the current data to obtain a side information feature map; execute quantization processing on the side information feature map to obtain a first quantized feature map. Execute entropy encoding on the first quantized feature map to obtain a first bitstream of the current data. Execute scaling processing on the residual feature map based on the scaling coefficient to obtain a scaled feature map, and execute quantization processing on the scaled feature map to obtain a second quantized feature map. Obtain a residual feature map based on the second feature map of the current data and the first probability distribution parameter, and obtain a scaling coefficient based on the first quantized feature map. Execute entropy encoding on the second quantized feature map based on the first probability distribution parameter to obtain a second bitstream of the current data.
[0023] In this solution, during encoding, after performing scaling processing on the residual feature map of the current data to obtain a scaled feature map, quantization processing is performed on the scaled feature map so that the quantization loss of the scaled feature map becomes smaller. That is, the information loss of the second bitstream becomes smaller. This helps to improve the data quality of the reconstructed data obtained through decoding based on the second bitstream. Further, compared with the encoding network structure in the prior art, in this embodiment of the present application, after the training of the entire encoding network is completed, the network parameters of the entire encoding network including the network for generating the first bitstream and the network for generating the second bitstream can be optimized, and an encoding network structure including scaling for the residual feature map is used. Therefore, by using the encoding network structure in this embodiment of the present application, the data volume of the total bitstream of the current data can be reduced, and the encoding efficiency can be improved. Generally, the encoding method in this embodiment can further improve the data compression performance.
[0024] In some possible embodiments of the second aspect, the first feature map may be the same as or different from the second feature map.
[0025] In some possible embodiments of the second aspect, the data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data.
[0026] In some possible embodiments of the second aspect, the first probability distribution parameter includes an average and / or a variance.
[0027] In some possible embodiments of the second aspect, the first probability distribution parameter is obtained based on the first quantization feature map.
[0028] In some possible embodiments of the second aspect, the first probability distribution parameter is a pre-set probability distribution parameter.
[0029] According to a third aspect, the present application further provides a data decoding method, which is executed by a data decoding device. The method comprises the following steps.
[0030] Perform entropy decoding based on the first bit stream of the current data to obtain a third feature map. Obtain a scaling coefficient based on the third feature map. Perform scaling processing on the first probability distribution parameter based on the scaling coefficient to obtain a second probability distribution parameter. Perform entropy decoding on the second bit stream of the current data based on the second probability distribution parameter to obtain a fourth feature map. Perform scaling processing on the fourth feature map based on the scaling coefficient to obtain a fifth feature map. Obtain the reconstructed data of the current data based on the fifth feature map.
[0031] When performing scaling processing on the second feature map and scaling processing on the first probability distribution parameter during data encoding, correspondingly, during data decoding, by using the same scaling coefficient to process the first probability distribution parameter and the fourth feature map, the decoding accuracy of the fourth feature map is ensured. Further, descaling processing may be performed on the fourth feature map based on the scaling coefficient to obtain a fifth feature map, and reconstructed data may be obtained based on the fifth feature map, thereby making the accuracy and quality of the reconstructed data higher. That is, by combining the scaling of the first probability distribution parameter and the descaling processing of the fourth feature map, the accuracy and quality of data decoding can be improved.
[0032] In some possible embodiments of the third aspect, the fourth feature map is a residual feature map, the fifth feature map is a scaled residual feature map, and the fact that the reconstructed data of the current data is obtained based on the fifth feature map includes: Adding the first probability distribution parameter to the fifth feature map to obtain a sixth feature map; and obtaining the reconstructed data of the current data based on the sixth feature map including.
[0033] When performing scaling and quantization operations on the residual feature map during data encoding, the data amounts of the first bitstream and the second bitstream in the data decoding method in this solution means are smaller than those in the prior art. Therefore, the corresponding decoding processing amount is smaller, and this solution means can effectively improve the decoding efficiency. Further, the information loss of the second bitstream is smaller, whereby the data quality of the reconstructed data obtained by using this solution means is higher.
[0034] In some possible embodiments of the third aspect, the first probability distribution parameter includes an average and / or a variance.
[0035] In some possible embodiments of the third aspect, the data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data.
[0036] In some possible embodiments of the third aspect, the data decoding method further comprises receiving a first bitstream and a second bitstream of current data.
[0037] In some possible embodiments of the third aspect, the data decoding method further comprises obtaining the first probability distribution parameter based on a third feature map.
[0038] In some possible embodiments of the third aspect, the first probability distribution parameter is a pre-set probability distribution parameter.
[0039] According to a fourth aspect, the present application further provides a data decoding method, which is executed by a data decoding device. The method comprises the following steps.
[0040] Perform entropy decoding based on a first bitstream of current data to obtain a third feature map. Obtain a scaling coefficient based on the third feature map. Perform entropy decoding on a second bitstream of current data based on the first probability distribution parameter to obtain a fourth feature map. Perform scaling processing on the fourth feature map based on the scaling coefficient to obtain a fifth feature map. Obtain a sixth feature map based on the first probability distribution parameter and the fifth feature map. Obtain reconstructed data of current data based on the sixth feature map.
[0041] When performing scaling and quantization operations on the residual feature map during data encoding, the data amounts of the first bitstream and the second bitstream in the data decoding method of the present solution means are smaller than those in the prior art. Therefore, the corresponding decoding processing amount is smaller, and the present solution means can effectively improve the decoding efficiency. Furthermore, the information loss of the second bitstream is smaller, and thereby, the data quality of the reconstructed data obtained by using the present solution means is higher.
[0042] In some possible embodiments of the fourth aspect, the first feature map is the same as or different from the second feature map.
[0043] In some possible embodiments of the fourth aspect, the data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data.
[0044] In some possible embodiments of the fourth aspect, the first probability distribution parameter includes an average and / or a variance.
[0045] In some possible embodiments of the fourth aspect, the data decoding method further includes a step of obtaining a first probability distribution parameter based on a third feature map.
[0046] In some possible embodiments of the fourth aspect, the first probability distribution parameter is a pre-set probability distribution parameter.
[0047] According to a fifth aspect, the present application further provides a data encoder including a processing circuit configured to execute a data encoding method according to any one of the embodiments of the first aspect or the second aspect.
[0048] According to a sixth aspect, the present application further provides a computer-readable storage medium. The storage medium stores a bitstream, and the bitstream is generated according to a data encoding method according to any one of the embodiments of the first aspect or the second aspect.
[0049] According to a seventh aspect, the present application further provides a data decoder including a processing circuit configured to execute a data decoding method according to any one of the embodiments of the third aspect or the fourth aspect.
[0050] According to an eighth aspect, the present application further provides a computer program product including program code. When the program code is executed on a computer or a processor, the computer program product is configured to execute a method according to any one of the embodiments of the first aspect, the second aspect, the third aspect, or the fourth aspect.
[0051] According to a ninth aspect, the present application provides a data encoder comprising: one or more processors; and a computer-readable storage medium coupled to the one or more processors, wherein the computer-readable storage medium stores a program, and when the program is executed by the one or more processors, the data encoder is capable of executing a data encoding method according to any one of the embodiments of the first aspect or the second aspect and further provides a data encoder.
[0052] According to a tenth aspect, the present application provides a data decoder comprising: one or more processors; and a computer-readable storage medium coupled to the one or more processors, wherein the computer-readable storage medium stores a program, and when the program is executed by the one or more processors, the data decoder is capable of executing a data decoding method according to any one of the embodiments of the third aspect or the fourth aspect Further provided is a data decoder including the same.
[0053] According to an eleventh aspect, the present application further provides a computer-readable storage medium including program code. When the program code is executed by a computer device, the computer-readable storage medium is configured to execute a method according to any one of the embodiments of the first aspect, the second aspect, the third aspect, or the fourth aspect.
[0054] According to a twelfth aspect, the present application further provides a computer-readable storage medium storing a bitstream including program code. When the program code is executed by one or more processors, the decoder is capable of executing a data decoding method according to any one of the embodiments of the third aspect or the fourth aspect.
[0055] According to a thirteenth aspect, the present application further provides a chip. The chip includes a processor and a data interface. The processor reads instructions stored in a memory through the data interface and executes a method according to any one of the embodiments of the first aspect, the second aspect, the third aspect, or the fourth aspect.
[0056] Optionally, in one implementation, the chip may further include a memory. The memory stores instructions, and the processor is configured to execute the instructions stored in the memory. When the instructions are executed, the processor is configured to execute a method according to any one of the embodiments of the first aspect, the second aspect, the third aspect, or the fourth aspect.
Brief Description of the Drawings
[0057] Hereinafter, the accompanying drawings used in the embodiments of the present application will be described.
[0058]
Figure 1A
[0059]
Figure 1B
[0060]
Figure 1C
[0061]
Figure 2
[0062]
Figure 3
[0063]
Figure 4A
[0064]
Figure 4B
[0065]
Figure 4C
[0066]
Figure 4D
[0067]
Figure 4E
[0068]
Figure 4F
[0069]
Figure 4G
[0070]
Figure 4H
[0071]
Figure 4I
[0072]
Figure 5A
[0073]
Figure 5B
[0074]
Figure 5C
[0075]
Figure 5D
[0076]
Figure 5E
[0077]
Figure 5F
[0078]
Figure 6A
[0079]
Figure 6B
[0080]
Figure 7
[0081]
Figure 8
[0082]
Figure 9
[0083]
Figure 10
Mode for Carrying Out the Invention
[0084] Hereinafter, the technical solution of the present application will be described with reference to the accompanying drawings.
[0085] Embodiments of the present application relate to an application. Therefore, for ease of understanding, related concepts such as related terms in the embodiments of the present application will be first described below.
[0086] In the embodiments of the present application, expressions such as "example" or "for example" are used to give an example, illustration, or explanation. Any embodiment or design scheme described as an "example" or "for example" in the present application is not described as being more preferable or having more advantages than another embodiment or design scheme. Exactly, the use of expressions such as "example" or "for example" is intended to present relative concepts in a specific manner.
[0087] In the embodiments of the present application, "at least one" means one or more, and "a plurality of" means two or more. "At least one of the following items (elements)" or a similar expression indicates any combination of these items, including a single item (element) or any combination of a plurality of items (elements). For example, at least one of a, b, or c may indicate: a, b, c, (a and b), (a and c), (b and c), or (a, b, and c), where a, b, and c may be singular or plural. The term "and / or" represents the relationship between related objects and indicates that three relationships can exist. For example, A and / or B may indicate the following three cases: namely, only A exists, both A and B exist, and only B exists, and A and B may be singular or plural. The character " / " generally indicates an "or" relationship between related objects. The sequence numbers of the steps in the embodiments of the present application (for example, step S1 and step S21) are only used to distinguish different steps and do not limit the execution order of the steps.
[0088] Furthermore, unless otherwise specified, ordinal numbers such as "first" and "second" in the embodiments of the present application are used to distinguish a plurality of objects and are not intended to limit the order, time series, priority, or importance of the plurality of objects. For example, the first device and the second device are only for ease of explanation and do not indicate a difference in the structure and importance of the first device and the second device. In some embodiments, the first device and the second device may alternatively be the same device.
[0089] Depending on the context, the term "when" used in the foregoing embodiments may be interpreted as meaning "case", "after", "depending on the determination", or "depending on the detection". The foregoing description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the concept and principles of the present application shall be included within the protection scope of the present application.
[0090] For ease of understanding, terms and concepts related to the embodiments of the present application will first be described.
[0091] (1) Quantization
[0092] Quantization is used to convert a continuous signal into a discrete signal. In the compression process, quantization means converting continuous features into discrete features. In entropy encoding, the probability values of a probability distribution are typically changed from continuous values to discrete values.
[0093] (2) Entropy Encoding
[0094] Entropy encoding is an encoding process in which information is not lost according to the entropy principle. Information entropy is the average amount of information (a measure of uncertainty) in a source. Common entropy encodings include: Shannon coding, Huffman coding, run-length coding, LZW encoding, and arithmetic coding. The LZW encoding algorithm, also referred to as a string table compression algorithm, performs compression by establishing a string table and representing long strings using short codes.
[0095] (3) Neural Network
[0096] A neural network may include neurons. A neuron may be an arithmetic unit that uses xs and a bias of 1 as inputs. The output of the arithmetic unit may be as follows:
Equation
[0097] Here, s = 1, 2,..., or n, where n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neuron, f is the activation function of the neuron, which is used to introduce non-linear features into the neural network and convert the input signal within the neuron into an output signal. The output signal of the activation function may be used as the input to the next convolutional layer. The activation function may be a sigmoid function. A neural network is a network formed by linking a plurality of single neurons together. That is, the output of one neuron may be the input of another neuron. The input of each neuron may be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field may be a region containing a plurality of neurons.
[0098] (4) Deep Neural Network
[0099] A deep neural network (DNN), also referred to as a multi-layer neural network, can be understood as a neural network having a plurality of hidden layers. The DNN is divided based on the position of different layers such that the neural networks within the DNN can be classified into three types: an input layer, a hidden layer, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the intermediate layer is the hidden layer. The layers are fully connected to each other. Specifically, any neuron in the i-th layer is necessarily connected to any neuron in the (i + 1)-th layer.
[0100] Although the DNN appears complex, it is not complex with respect to the functions of each layer. Briefly speaking, the DNN is the following linear relationship expression:
Equation
Equation
[0101] In conclusion, the coefficient from the k-th neuron in the (L - 1)-th layer to the j-th neuron in the L-th layer is [Number] defined as.
[0102] Note that there is no parameter W in the input layer. In a deep neural network, more hidden layers increase the network's ability to describe complex cases in the real world. Theoretically, a model with more parameters has higher complexity and greater "capacity", indicating that the model can complete more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and the ultimate goal of training is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by the vector W in multiple layers).
[0103] (5) Convolutional Neural Network
[0104] A Convolutional Neural Network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network includes a feature extraction unit having convolutional layers and subsampling layers, and the feature extraction unit may be regarded as a filter. A convolutional layer is a neuron layer within the convolutional neural network that performs a convolution process on an input signal. In the convolutional layer of a convolutional neural network, one neuron may be connected to only some adjacent layer neurons. One convolutional layer usually has multiple feature planes, and each feature plane may include several neurons in a rectangular configuration. Neurons in the same feature plane share weights. The shared weights herein are convolutional kernels. Weight sharing may be understood as the image information extraction method being independent of position. The convolutional kernel may be initialized in the form of a matrix of random size. In the process of training a convolutional neural network, the convolutional kernel may obtain appropriate weights through learning. Furthermore, the benefit directly brought by weight sharing is that the connections between the layers of the convolutional neural network are reduced, and the risk of overfitting is reduced.
[0105] (6) Loss Function
[0106] In the process of training a deep neural network, since the output of the deep neural network is expected to be as close as possible to the actually expected predicted value, it is possible to compare the current predicted value of the network with the actually expected target value. Next, the weight vector of each layer of the neural network is updated based on the difference between the predicted value and the target value. (Of course, usually there is an initialization process before the first update. Specifically, parameters are pre-configured for all layers of the deep neural network.) For example, when the predicted value of the network is large, the weight vector is adjusted to decrease the predicted value, and the adjustment is continuously executed until the deep neural network can predict a value that is actually expected target value or a value very close to the actually expected target value. Therefore, it is necessary to pre-define a method for "obtaining the difference between the predicted value and the target value through comparison". This is the loss function (Loss Function) or objective function (Objective Function). The loss function and the objective function are important equations used to measure the difference between the predicted value and the target value. The loss function is used as an example. A larger output value (Loss) of the loss function indicates a larger difference. Therefore, the training of a deep neural network is a process of minimizing Loss as much as possible.
[0107] (7) Backpropagation Algorithm
[0108] In the training process, the neural network may correct the values of the parameters of the initial neural network model by using the error backpropagation (BP) algorithm so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, the input signal is fed forward until the error loss is generated at the output, and the parameters of the initial neural network model are updated through the backpropagation of the information regarding the error loss to converge the error loss. The backpropagation algorithm is a backpropagation motion centered on the error loss intended to obtain the parameters such as the weight matrix of the optimal neural network model.
[0109] In the prior art, it is necessary to improve the compression performance of the data compression algorithm. For example, the data occupies a large number of bits after the compression process. In view of this, the embodiments of the present application provide a data encoding and decoding method so as to effectively improve the data compression performance.
[0110] For example, the data in the data encoding method and / or data decoding method in the embodiments of the present application includes at least one of the following: image data, video data, motion vector (MV) data of video data, audio data, point cloud data, or text data. The image data may be one image or at least two images. The video data is a continuous image sequence and is essentially formed by a group of continuous images. For example, with respect to the image data or video data, the data encoding method and / or data decoding method in the embodiments of the present application may be executed on at least one image in the image data or video data to perform encoding and decoding processes on at least one image. As another example, with respect to the encoding and decoding of video data, the data encoding method and data decoding method in the embodiments of the present application may be used to process the frame of Image A to obtain a reconstructed frame corresponding to Image A. Next, the reconstructed frame may be used to predict the next frame of the image of Image A to obtain a predicted image of the next frame of the image. Next, the difference between the next frame of the image and the predicted image is compressed. The reconstruction result obtained during decoding is the sum of the predicted image of the next frame of the image and the reconstruction residual. With respect to video data, the pixel data of each video frame is usually encoded as a block of pixels (also referred to herein as "pixel block", "encoding unit", and "macroblock"). The motion vector is used to describe the offset vector of the position of the macroblock in the video frame with respect to the position of the macroblock in the reference frame. The motion vector data is at least one motion vector obtained based on the video data.
[0111] Point cloud data is a set of groups of vectors in a three-dimensional coordinate system. Scan data is recorded in point form, each point having three-dimensional coordinates, and some points may have color information or intensity information. In addition to geometric positions, some point cloud data also has color information. Color information is typically obtained by using a camera to acquire a color image, and then the color information of the pixels at the corresponding positions is assigned to the corresponding points in the point cloud. Intensity information is the echo intensity collected by a receiving device for a laser scanner. Intensity information is related to the surface material, roughness, incident angle direction of the target, radiation energy of the device, and laser wavelength. Further, text data is data that contains text, and the data format of the text data may be a TXT file, a PDF file, a Word file, an Excel file, etc. Further, the application scenarios of the data encoding method and / or the data decoding method in the embodiments of the present application include cloud storage services, cloud monitoring, live streaming, etc.
[0112] FIG. 1A is a diagram of the architecture of a data coding system according to an embodiment of the present application. For example, the data coding system in FIG. 1A includes a data encoder and a data decoder. The data encoder has an encoding unit, an entropy encoding unit, and a storage unit. The data decoder has a decoding unit, an entropy decoding unit, and a load unit.
[0113] The encoding unit is configured to convert the data to be processed into feature data having lower redundancy and obtain the probability distribution corresponding to the feature data. The data to be processed includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data.
[0114] The entropy encoding unit is configured to perform lossless encoding on the feature data based on the probability corresponding to the feature data, so as to further reduce the data transmission volume in the compression process.
[0115] The storage unit is configured to store the data file generated by the entropy encoding unit in the corresponding storage location.
[0116] The load unit is configured to load the data file from the corresponding storage location and input the data file into the entropy decoding unit.
[0117] The entropy decoding unit is configured to perform entropy decoding on the data file to obtain the processed data.
[0118] The decoding unit is configured to perform an inverse transformation on the processed data output by the entropy decoding unit to parse the processed data into reconstructed data.
[0119] For example, after a data collection device collects data to be processed, compression processing is performed on the data to be processed. The specific process is as follows: The encoding unit processes the data to be processed to obtain encoding target features and corresponding probability distributions, and inputs the encoding target features and probability distributions to the entropy encoding unit for processing to obtain a bitstream file, and the storage unit stores the bitstream file. When the bitstream file is decompressed, the specific process is as follows: The load unit loads the bitstream file from the storage location and inputs the bitstream file to the entropy decoding unit, and the entropy decoding unit and the decoding unit may cooperate with each other to obtain reconstructed data corresponding to the bitstream file. Further, for example, the reconstructed data may be output, for example, output for display.
[0120] In the following embodiments of the coding system 10, the encoder 20 and the decoder 30 are described based on FIGS. 1B and 1C.
[0121] FIG. 1B is a block diagram of an example of a coding system 10 that may use the technology of the present application, for example, a data coding system 10 (or simply referred to as the coding system 10). The data encoder 20 (or simply referred to as the encoder 20) and the data decoder 30 (or simply referred to as the decoder 30) in the coding system 10 represent devices and the like that may be configured to execute techniques based on various examples described in the present application.
[0122] As shown in FIG. 1B, the coding system 10 includes a source device 12. The source device 12 is configured to provide encoded data 21, such as an encoded image, encoded video, or encoded audio, to a destination device 14 that is configured to decode the encoded data 21.
[0123] The source device 12 has an encoder 20 and, in addition, i.e., optionally, may have a data source 16, a preprocessor (or preprocessing unit) 18, and a communication interface (or communication unit) 22.
[0124] The data source 16 may include and / or be any type of data acquisition device configured to acquire data and / or any type of data generation device. In this embodiment, the data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data.
[0125] For example, the data is image data. In this case, the data source 16 may include and / or be any type of image capture device configured to capture real-world images, and / or any type of image generation device, such as a computer graphics processing unit configured to generate computer animation images or any type of device configured to acquire and / or provide real-world images, computer-generated images (e.g., screen content or virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The data source may be any type of memory or storage that stores any of the aforementioned images.
[0126] For example, the data is video data. In this case, the data source 16 may include or be any type of video recording device configured to capture real-world images or the like to generate video, and / or any type of video generation device, such as a computer graphics processor configured to generate computer animation or any type of device configured to acquire and / or provide real-world video or computer-generated video (e.g., video obtained through screen recording). The data source may be or be any type of memory or storage that stores any of the aforementioned videos.
[0127] For example, the data is audio data. In this case, the data source 16 may include or be any type of audio capture device configured to capture real-world sounds or the like to generate audio, and / or any type of audio generation device, such as an audio processor configured to generate virtual audio (e.g., virtual human voice) or any type of device configured to acquire and / or provide real-world audio or computer-generated audio (e.g., audio obtained through screen recording). The data source may be or be any type of memory or storage that stores any of the aforementioned audio.
[0128] For example, the data is point cloud data. In this case, the data source 16 may include or be any type of device configured to acquire point cloud data, such as a three-dimensional laser scanner or an imaging scanner.
[0129] To distinguish the processing performed by the preprocessor (or preprocessing unit) 18, the data 17 may also be referred to as the original data 17.
[0130] The preprocessor 18 is configured to receive the (original) data 17, preprocess the data 17, and obtain the preprocessed data 19. For example, the data is image data. In this case, the preprocessing performed by the preprocessor 18 may include trimming, color format conversion (e.g., conversion from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optional component.
[0131] The encoder 20 is configured to receive the preprocessed data 19 and provide the encoded data 21 (further explanation will be given below based on FIGS. 4A, 4C, 4D, 4F, 4G, 4I, etc.).
[0132] The communication interface 22 of the source device 12 is configured to receive the encoded data 21 and transmit the encoded data 21 (or any further processed version thereof) through the communication channel 13 to another device such as the destination device 14 or any other device for storage or direct reconstruction.
[0133] The destination device 14 includes a decoder 30 and, in addition, i.e., optionally, a communication interface (or communication unit) 28, a postprocessor (or post - processing unit) 32, and a display device 34.
[0134] The communication interface 28 of the destination device 14 is configured to receive the encoded data 21 (or any further processed version thereof) directly from the source device 12 or from any other source device such as a storage device for the encoded data and provide the encoded data 21 to the decoder 30.
[0135] Communication interfaces 22 and 28 may be configured to transmit or receive encoded data 21 between source device 12 and destination device 14 through a direct communication link, e.g., a direct wired or wireless connection, or through any type of network, e.g., a wired or wireless network or any combination thereof, or any type of private and public network, or any combination of any type thereof.
[0136] For example, communication interface 22 may be configured to process the encoded data 21 by encapsulating the encoded data 21 into a suitable format, such as a packet, and / or through processing for transmission via any type of transmission encoding or communication link or communication network.
[0137] Communication interface 28 corresponds to communication interface 22 and may be configured to process the transmitted data, e.g., by receiving the transmitted data and through any type of corresponding transmission decoding or processing and / or decapsulation, to obtain the encoded data 21.
[0138] Communication interfaces 22 and 28 may both be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by the arrow directed from source device 12 to destination device 14 corresponding to communication channel 13 in FIG. 1B, and may be configured to transmit and receive messages, etc., establish connections, and confirm and exchange any other information, such as encoded data.
[0139] Decoder 30 is configured to receive the encoded data 21 and provide decoded data (or reconstructed data) 31 (further explanation will be given below based on FIGS. 4B, 4E, 4H, etc.).
[0140] The post-processor 32 is configured to post-process the decoded data 31 to obtain post-processed data 33. For example, the data is image data, and the post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., conversion from YCbCr to RGB), color correction, trimming, resampling, or any other processing such as generating the post-processed data 33 for display by a display device 34, etc.
[0141] The display device 34 is configured to receive the post-processed data 33 for displaying data to a user, viewer, etc. The display device 34 may be or include any type of display configured to represent the reconstructed data, for example, an integrated or external display screen or display. For example, the display screen may include a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.
[0142] The coding system 10 further includes a training engine 25. The training engine 25 is configured to train the neural network in the encoder 20 or the decoder 30 such that when the data 17 or the pre-processed data 19 is input, the encoder 20 can obtain the encoded data 21, or when the encoded data 21 is input, the decoder 30 can obtain the decoded data 31. Optionally, the input data further includes hyper prior information.
[0143] The training data may be stored in a database (not shown in FIG. 1B), and the training engine 25 obtains a neural network through training based on the training data. The neural network is a neural network in the encoder 20 or the decoder 30. It should be noted that the source of the training data is not limited in the embodiments of the present application. For example, the training data may be obtained from the cloud or another location to perform neural network training.
[0144] The neural network obtained by the training engine 25 through training may be applied to the coding system 10 or the coding system 40. For example, it may be applied to the source device 12 (e.g., the encoder 20) or the destination device 14 (e.g., the decoder 30) shown in FIG. 1B. For example, the training engine 25 may obtain a neural network through training on the cloud, and then the coding system 10 downloads the neural network from the cloud and uses the neural network.
[0145] FIG. 1B shows that the source device 12 and the destination device 14 are separate devices, but the device embodiment may include both the source device 12 and the destination device 14, or may include the functions of both the source device 12 and the destination device 14, that is, it may include both the source device 12 or the corresponding function and the destination device 14 or the corresponding function. In such an embodiment, the source device 12 or the corresponding function and the destination device 14 or the corresponding function may be implemented by using the same hardware and / or software or by using separate hardware and / or software or any combination thereof.
[0146] Based on the above description, it is apparent to those skilled in the art that the presence and (precise) partitioning of different units or functions of the source device 12 and / or the destination device 14 shown in FIG. 1B may vary depending on the actual device and application.
[0147] The encoder 20 or the decoder 30 or both may be implemented by using a processing circuit as shown in FIG. 1C, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding processors, or any combination thereof. The encoder 20 may be implemented by using the processing circuit 43 to embody the various modules described with reference to the encoder 20 in FIG. 1C and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented by using the processing circuit 43 to embody the various modules described with reference to the decoder 30 in FIG. 1C and / or any other decoder system or subsystem described herein. The processing circuit 43 may be configured to perform the various operations described below. As shown in FIG. 3, when some techniques are implemented in software, the device may store software instructions in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware by using one or more processors to execute the techniques in the present application. Either the encoder 20 or the decoder 30 may be integrated into a single device as part of a codec (Encoder / Decoder).
[0148] Source device 12 and destination device 14 may include any of a variety of devices, such as any type of handheld or stationary device, for example, a notebook or laptop computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, server, set-top box, television, display device, digital media player, video gaming console, video streaming device (such as a content service server or content delivery server), broadcast receiver device, broadcast transmitter device, etc., and may or may not use any type of operating system. In some cases, source device 12 and destination device 14 may include components for wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices.
[0149] In some cases, the coding system 10 shown in FIG. 1B is merely an example, and the techniques provided in the present application are applicable to coding settings, and these settings do not necessarily include all data communications between the encoding device and the decoding device. In another example, data is obtained from local storage and transmitted through a network, etc. The encoding device may encode data and store the data in storage, and / or the decoding device may obtain data from storage and decode the data. In some examples, encoding and decoding do not communicate with each other, but are simply performed by devices that encode data to storage and / or obtain data from storage and decode the data.
[0150] FIG. 1C is an explanatory diagram of an example of a coding system 40 including an encoder 20 and / or a decoder 30 according to an exemplary embodiment. The coding system 40 may include an encoder 20 and a decoder 30 (and / or an encoder / decoder implemented by using a processing circuit 43), an antenna 42, one or more memories 44, and / or a display device 45. For example, when the data is image data, the coding system may further include an imaging device 41.
[0151] As shown in FIG. 1C, the imaging device 41, the antenna 42, the processing circuit 43, the encoder 20, the decoder 30, the memory 44, and / or the display device 45 can communicate with each other. In different examples, the coding system 40 may include only the encoder 20 or only the decoder 30. Of course, the coding system 40 is not limited to the configuration shown in FIG. 1C and may include more or fewer components than those shown in FIG. 1C.
[0152] In some examples, antenna 42 may be configured to transmit or receive an encoded bitstream of data. Further, in some examples, display device 45 may be configured to present reconstructed data. Processing circuitry 43 may include, for example, Application-Specific Integrated Circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, and the like. Further, memory 44 may be any type of memory, such as volatile memory (e.g., Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM)), non-volatile memory (e.g., flash memory), and the like. In a non-limiting example, memory 44 may be implemented by a cache memory. In another example, processing circuitry 43 may include a memory (e.g., a cache) for implementing an image buffer.
[0153] In some examples, coding system 40 may further include a decoder 30 coupled to antenna 42 and configured to decode the encoded bitstream. Display device 45 is configured to present the reconstructed data.
[0154] In this embodiment of the present application, with respect to the examples described with reference to encoder 20, it should be understood that decoder 30 may be configured to perform the reverse process.
[0155] FIG. 2 is a diagram of a data coding device 200 according to an embodiment of the present application. The data coding device 200 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the data coding device 200 may be a decoder, such as decoder 30 in FIG. 1C, or an encoder, such as encoder 20 in FIG. 1C.
[0156] The data coding device 200 includes: an inlet port 210 (or input port 210) and a receiver unit 220 configured to receive data; a processor, logic unit, or central processing unit (CPU) 230 configured to process data. For example, the processor 230 in this specification may be a neural network processor 230; a transmitter unit (Tx) 240 and an outlet port 250 (or output port 250) configured to transmit data; and a memory 260 configured to store data. For example, the data coding device 200 may further include optoelectrical (OE) components and electro-optical (EO) components coupled to the inlet port 210, the receiver unit 220, the transmitter unit 240, and the outlet port 250 for the exit or entry of optical or electrical signals.
[0157] Processor 230 is implemented by using hardware and software. Processor 230 may be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 230 communicates with an input port 210, a receiver unit 220, a transmitter unit 240, an output port 250, and a memory 260. Processor 230 includes a coding module 270 (e.g., a coding module 270 based on a neural network NN). The coding module 270 implements the disclosed embodiments described above. For example, the coding module 270 executes, processes, prepares, or provides various coding operations. Thus, the coding module 270 substantially improves the function of the data coding device 200 and affects the switching of the data coding device 200 between different states. Alternatively, the coding module 270 is implemented by using instructions stored in the memory 260 and executed by the processor 230.
[0158] Memory 260 includes one or more magnetic disks, tape drives, and solid state drives and may be used as an overflow data storage device and is configured to store such programs when a program is selected for execution and to store instructions and data read during the execution of the program. Memory 260 may be volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0159] Figure 3 is a simplified block diagram of a data coding device 300 according to an exemplary embodiment. The device 300 may be used as either or both of the source device 12 and the destination device 14 in FIG. 1B.
[0160] The processor 302 in the device 300 can be a central processing unit. Alternatively, the processor 302 can be any other type of device or devices capable of manipulating or processing information that exists or will be developed in the future. The disclosed implementation can be carried out by using a single processor, e.g., the processor 302 shown in FIG. 3, but high speed and high efficiency can be achieved by using more than one processor.
[0161] In one implementation, the memory 304 in the device 300 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 304. The memory 304 may include code and data 306 that are accessed by the processor 302 through the bus 312. The memory 304 may further include an operating system 308 and an application 310. The application 310 includes at least one program that enables the processor 302 to execute the methods described herein. For example, the application 310 may include applications 1 through N and may further include a data coding application that executes the methods described herein, i.e., data encoding and / or data decoding.
[0162] The device 300 may further include one or more output devices, e.g., a display 314. In one example, the display 314 can be a touch-sensitive display that combines a display with a touch-sensitive element configured to sense touch inputs. The display 314 can be coupled to the processor 302 through the bus 312.
[0163] Although bus 312 in device 300 is shown as a single bus herein, bus 312 may include multiple buses. Further, secondary storage may be directly coupled to another component of device 300 or may be accessed through a network and may include a single integrated unit such as a storage card or multiple units such as multiple storage cards. Thus, device 300 may have a wide variety of configurations.
[0164] FIG. 4A is a diagram of the structure of a data encoder according to an embodiment of the present application. In an example of FIG. 4A, data encoder 20 includes an input terminal (or input interface) 401, an encoding network (Encoder) 402, a hyper-encoding network (HyperEncoder) 403, quantization unit 404, quantization unit 405, entropy encoding unit 406, entropy encoding unit 407, entropy estimation unit (Entropy) 408, hyper-entropy estimation unit (HyperEntropy) 409, and hyper-decoding network (HyperDecoder) 410. Quantization unit 404 and quantization unit 405 may be the same quantization unit or two independent quantization units. Similarly, entropy encoding unit 406 and entropy encoding unit 407 may be the same entropy encoding unit or two independent entropy encoding units. Entropy estimation unit 408 is also referred to as an entropy parameter model unit, and hyper-entropy estimation unit 409 is an entropy parameter model unit that uses a preset distribution.
[0165] Encoding target data may be received through data encoder 20, input terminal 401, etc. The encoding target data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data.
[0166] The encoding network 402 is configured to extract features from the data to be encoded and obtain a first feature map y1 and a second feature map y2. The first feature map y1 is different from the second feature map y2. The method for obtaining the first feature map y1 and the second feature map y2 is not particularly limited.
[0167] For example, the data is image data. For example, compared with the original image, the first feature map and the second feature map output by the encoding network 402 may have a changed size and redundant information is removed, whereby entropy encoding is more easily performed.
[0168] The super-prior encoding network 403 is configured to further extract concise information from the first feature map y1 and obtain a side information feature map z. For example, the size of the side information feature map z is smaller than that of the first feature map y1.
[0169] The quantization unit 404 performs quantization processing on the side information feature map z to obtain integer feature data, that is, the first quantized feature map
Number
[0170] The super-prior decoding network 410
Number
[0171] Furthermore, for example, the super prior decoding network 410 may be configured to only generate a scaling coefficient Δ based on a first quantized feature map
Number
[0172] The residual of the second feature map y2, that is, the residual feature map C, may be obtained based on the second feature map and the first probability distribution parameter X. Specifically, the residual feature map C may be obtained by subtracting X from y2. For example, the residual feature map C is obtained by subtracting the average μ from the second feature map y2.
[0173] A scaling process is performed on the residual feature map C based on the scaling coefficient Δ to obtain a scaled residual feature map C. Specifically, the scaled residual feature map C may be obtained by dividing the residual feature map C by the scaling coefficient Δ.
[0174] The quantization unit 405 is configured to perform quantization processing on the scaled residual feature map C to obtain integer feature data, that is, the second quantized feature map.
[0175] The super prior entropy estimation unit 409 is based on a preset distribution for the first quantized feature map
Number
[0176] The entropy encoding unit 406 is based on the probability distribution of the first quantized feature map estimated by the super prior entropy estimation unit 409
Number
Number
[0177] The entropy estimation unit 408 is configured to obtain the probability distribution of the second quantized feature map based on the first probability distribution parameter X.
[0178] The entropy encoding unit 407 is configured to perform entropy encoding on the second quantized feature map based on the probability estimated by the entropy estimation unit 408 to obtain a second bit stream B. The first bit stream A and the second bit stream B are used as the total bit stream of the data to be encoded. The data encoder may output the total bit stream of the data through an output terminal (or output interface) (not shown in FIG. 4A).
[0179] FIG. 4B is a diagram of the structure of a data decoder according to an embodiment of the present application. The data decoder shown in FIG. 4B is configured to decode the total bit stream of the data obtained through the processing by the data encoder shown in FIG. 4A. The data decoder 30 includes an entropy estimation unit 408, an entropy decoding unit 414, a super prior entropy estimation unit 409, an entropy decoding unit 413, a super prior decoding network 410, a decoding network 412, and an output terminal (or output interface) 411.
[0180] The data decoder 30 may obtain the total bit stream to be decoded, that is, the first bit stream A and the second bit stream B, through an input terminal (or input interface) (not shown in FIG. 4B).
[0181] The super prior entropy estimation unit 409 is configured to estimate the probability distribution of the first quantized feature map
Number
[0182] The entropy decoding unit 413 is configured to perform entropy decoding on the first bit stream A based on the probability estimated by the super prior entropy estimation unit 409 to obtain a third feature map. The entropy decoding unit 413 performs entropy decoding by using a distribution that matches the distribution used by the entropy encoding unit 406.
[0183] The super prior decoding network 410 is configured to generate a scaling coefficient Δ and a first probability distribution parameter X based on the third feature map. The first probability distribution parameter X represents the probability distribution of the fourth feature map, and X includes, but is not limited to, a mean μ and / or a variance σ. The first probability distribution parameter X is a matrix of the same scale as the scale of the fourth feature map. Each element in the matrix represents the probability distribution parameter of one of the elements in the fourth feature map. Each probability distribution parameter of the elements includes, but is not limited to, a mean and / or a variance. That is, one fourth feature map corresponds to one mean matrix and / or one variance matrix.
[0184] Furthermore, for example, the super prior decoding network 410 may be configured to generate only the scaling coefficient Δ based on the third feature map, and the first probability distribution parameter X is a probability distribution parameter set in advance according to the actual situation. This is not particularly limited.
[0185] The entropy estimation unit 408 is configured to obtain the probability distribution of the fourth feature map based on the first probability distribution parameter X.
[0186] The entropy decoding unit 414 is configured to perform entropy decoding on the second bit stream B based on the probability estimated by the entropy estimation unit 408 to obtain a fourth feature map.
[0187] The fifth feature map may be obtained based on the scaling factor △ and the fourth feature map. That is, the fifth feature map may be obtained by multiplying the second quantized feature map by the scaling factor △.
[0188] The sixth feature map may be obtained based on the fifth feature map and the first probability distribution parameter X. That is, the sixth feature map may be obtained by adding the first probability distribution parameter X to the fifth feature map. For example, corresponding to FIG. 4A, in this case, the sixth feature map may be obtained by adding the mean μ to the fifth feature map.
[0189] The decoding network 412 is configured to inverse-map the sixth feature map into reconstructed data.
[0190] The output terminal 411 is configured to output the reconstructed data.
[0191] FIG. 4C is a diagram of the structure of another data encoder according to an embodiment of the present application. The data encoder shown in FIG. 4C has the same configuration as the data encoder shown in FIG. 4A. The difference lies in that the encoding network 402 is configured to perform feature extraction on the data to be processed to obtain a feature map y (that is, in this case, the first feature map y1 is the same as the second feature map y2). In this case, the data decoder configured to decode the total bitstream of the data obtained through the process in FIG. 4C has the same structure as that in FIG. 4B.
[0192] FIG. 4D is a diagram of the structure of another data encoder according to an embodiment of the present application. The difference between the data encoder shown in FIG. 4D and the data encoder shown in FIG. 4A is that in FIG. 4D, there are scaling and quantization operations for the residual feature map and a scaling operation for the probability distribution parameters. Specifically, after the super-prior decoding network 410 obtains the first probability distribution parameter X and the scaling coefficient Δ, a scaling process is performed on the first probability distribution parameter X based on the scaling coefficient. That is, the first probability distribution parameter X is divided by the scaling coefficient Δ to obtain the second probability distribution parameter (X / Δ). The entropy encoding unit 407 is configured to perform entropy encoding on the second quantization feature map based on the second probability distribution parameter (X / Δ) in cooperation with the entropy estimation unit 408 to obtain the second bitstream B.
[0193] FIG. 4E is a diagram of the structure of another data decoder according to an embodiment of the present application. The data decoder shown in FIG. 4E is configured to decode the total bitstream of the data obtained through the processing by the data encoder shown in FIG. 4D. The data decoder shown in FIG. 4E has the same configuration as the data decoder shown in FIG. 4B. The difference lies in that after the super-prior decoding network 410 obtains the first probability distribution parameter X and the scaling coefficient Δ, a scaling process is performed on the first probability distribution parameter X based on the scaling coefficient. That is, the first probability distribution parameter X is divided by the scaling coefficient Δ to obtain the second probability distribution parameter (X / Δ). The entropy decoding unit 414 is configured to perform entropy decoding on the second bitstream B based on the second probability distribution parameter (X / Δ) in cooperation with the entropy estimation unit 408 to obtain the fourth feature map.
[0194] FIG. 4F is a diagram of the structure of another data encoder according to an embodiment of the present application. The difference between the data encoder shown in FIG. 4F and the data encoder shown in FIG. 4D is that, in this case, the first feature map y1 is the same as the second feature map y2. That is, the encoding network 402 performs feature extraction on the data to be processed to obtain the feature map y. In this case, the data decoder configured to decode the total bitstream of the data obtained through the process in FIG. 4F has the same structure as that in FIG. 4E.
[0195] FIG. 4G is a diagram of the structure of another data encoder according to an embodiment of the present application. The difference from the data encoder shown in FIG. 4D is that the data encoder shown in FIG. 4G does not generate a residual feature map and does not perform an operation of scaling the residual feature map. In FIG. 4G, after the encoding network 402 obtains the second feature map y2 of the data to be processed, a scaling process is performed on the second feature map y2 based on the scaling coefficient △. That is, the second feature map y2 is divided by the scaling coefficient △ to obtain a scaled feature map. The quantization unit 405 then performs quantization processing on the scaled feature map to obtain a second quantized feature map. FIG. 4H is a diagram of the structure of another data decoder according to an embodiment of the present application. The data decoder shown in FIG. 4H is configured to decode the total bitstream of the data obtained through the process by the data encoder shown in FIG. 4G. The data decoder 30 includes an entropy estimation unit 408, an entropy decoding unit 414, a hyper prior entropy estimation unit 409, an entropy decoding unit 413, a hyper prior decoding network 410, a decoding network 412, and an output terminal (or output interface) 411.
[0196] The data decoder 30 may obtain the total bit stream to be decoded, that is, the first bit stream A and the second bit stream B, through an input terminal (or input interface) (not shown in FIG. 4H).
[0197] The entropy decoding unit 413 is configured to perform entropy decoding on the first bit stream A in cooperation with the super-prior entropy estimation unit 409 to obtain a third feature map.
[0198] The super-prior decoding network 410 is configured to generate a scaling coefficient Δ and a first probability distribution parameter X based on the third feature map. The first probability distribution parameter X represents the probability distribution of the fourth feature map, and X includes, but is not limited to, an average μ and / or a variance σ.
[0199] Furthermore, for example, the super-prior decoding network 410 may be configured to only generate a scaling coefficient Δ based on the third feature map, and the first probability distribution parameter X is a probability distribution parameter set in advance according to the actual situation. This is not particularly limited.
[0200] Perform a scaling process on the first probability distribution parameter X based on the scaling coefficient Δ. That is, divide the first probability distribution parameter X by the scaling coefficient Δ to obtain a second probability distribution parameter (X / Δ).
[0201] The entropy decoding unit 414 is configured to perform entropy decoding on the second bit stream B based on the second probability distribution parameter (X / Δ) in cooperation with the entropy estimation unit 408 to obtain a fourth feature map.
[0202] The fifth feature map may be obtained based on the scaling coefficient △ and the fourth feature map, that is, it may be obtained by multiplying the fourth feature map by the scaling coefficient △.
[0203] The decoding network 412 is configured to inverse map the fifth feature map into reconstructed data.
[0204] The output terminal 411 is configured to output the reconstructed data.
[0205] FIG. 4I is a diagram of the structure of another data encoder according to an embodiment of the present application. The data encoder shown in FIG. 4I is the same as that shown in FIG. 4G. In FIG. 4I, there is a difference in that the first feature map y1 is the same as the second feature map y2. That is, the encoding network 402 performs feature extraction on the data to be processed to obtain the feature map y. The data decoder configured to decode the total bit stream of the data obtained through the processing by the data encoder shown in FIG. 4I has the same structure as that in FIG. 4H. Details will not be described again.
[0206] It should be noted herein that at least one of the above-described encoding network, super pre-encoding network, quantization unit, entropy encoding unit, entropy decoding unit, super pre-entropy estimation unit, entropy estimation unit, super pre-decoding network, and decoding network may be implemented by using a neural network, for example, a convolutional neural network.
[0207] For example, regarding the specific structure of the encoding network 402 in the data encoder shown in FIG. 4C, FIG. 4F, or FIG. 4I, the structure shown in FIG. 5A can be referred to. The encoding network 402 has a first convolutional (Conv) layer, a first non-linear unit (ResAU) layer, a second convolutional layer, a second non-linear unit (ResAU) layer, a third convolutional layer, a third non-linear unit (ResAU) layer, and a fourth convolutional layer. For example, the specific parameters of the first convolutional layer, the second convolutional layer, the third convolutional layer, and / or the fourth convolutional layer are 192×5×5 / 2↓, where the number of channels is 192, the size of the convolutional kernel is 5×5, and the stride is 2. The feature map y of the data to be processed may be obtained by using the encoding network shown in FIG. 5A.
[0208] For example, regarding the specific structure of the encoding network 402 in the data encoder shown in FIG. 4A, FIG. 4D, or FIG. 4G, the structure shown in FIG. 5B can be referred to. The encoding network 402 has a first encoder, a second encoder, and a third encoder. The first encoder is configured to first perform feature extraction on the data to be processed. The second encoder and the third encoder are configured to perform feature extraction again on the feature map obtained through the extraction by the first encoder to respectively obtain a first feature map y1 and a second feature map y2.
[0209] As another example, regarding the specific structure of the encoding network 402 in the data encoder shown in FIG. 4A, FIG. 4D, or FIG. 4G, feature extraction may be performed by using the encoding network shown in FIG. 5A to obtain a feature map y. Next, channel separation may be performed on the feature map y to obtain a first feature map y1 and a second feature map y2. The specific method of channel separation is not limited. For example, the feature map y obtained in FIG. 5A has 384 channels. The a (a is smaller than 384) channels of the feature map y may be used as the first feature map y1, and the remaining (384 - a) channels of the feature map y may be used as the second feature map y2. For example, the first 192 channels of the feature map y may be used as the first feature map y1, and the last 192 channels of the feature map y may be used as the second feature map y2.
[0210] For example, regarding the specific structure of the hyper pre-encoding network in the data encoder shown in FIG. 4A, FIG. 4C, FIG. 4D, FIG. 4F, FIG. 4G, or FIG. 4I, the structure shown in FIG. 5C may be referred to. The hyper pre-encoding network includes a first leaky ReLU layer, a first convolutional layer, a second leaky ReLU layer, a second convolutional layer, a third leaky ReLU layer, and a third convolutional layer. For example, the specific parameters of the first convolutional layer are 192×3×3, where the number of channels is 192 and the size of the convolutional kernel is 3×3. The specific parameters of the second convolutional layer and / or the third convolutional layer are 192×5×5 / 2↓, where the number of channels is 192, the size of the convolutional kernel is 5×5, and the stride is 2.
[0211] For example, regarding the specific structure of the super-prior decoding network in any one of FIGS. 4A to 4I, the structure shown in FIG. 5D may be referred to. The super-prior decoding network includes a first convolutional layer, a first leaky ReLU layer, a second convolutional layer, a second leaky ReLU layer, and a third convolutional layer. For example, the specific parameters of the first convolutional layer are 384×3×3, where the number of channels is 384 and the size of the convolutional kernel is 3×3. The specific parameters of the second convolutional layer are 288×5×5 / 2↑, where the number of channels is 288, the size of the convolutional kernel is 5×5, and the stride is 2. The specific parameters of the third convolutional layer are 192×5×5 / 2↑, where the number of channels is 192, the size of the convolutional kernel is 5×5, and the stride is 2.
[0212] For example, regarding the specific structure of the decoding network in the data decoder shown in FIG. 4B, FIG. 4E, or FIG. 4H, the structure shown in FIG. 5E may be referred to. The decoding network includes a first convolutional layer, a first ResAU layer, a second convolutional layer, a second ResAU layer, a third convolutional layer, a third ResAU layer, and a fourth convolutional layer. For example, the specific parameters of the first convolutional layer, the second convolutional layer, and / or the third convolutional layer are 192×5×5 / 2↑, where the number of channels is 192, the size of the convolutional kernel is 5×5, and the stride is 2. The specific parameters of the fourth convolutional layer are 384×5×5 / 2↑, where the number of channels is 384, the size of the convolutional kernel is 5×5, and the stride is 2.
[0213] For example, regarding the specific structure of the ResAU layer in the structure shown in FIG. 5A and / or FIG. 5E, FIG. 5F may be referred to. The ResAU layer includes a leaky ReLU layer, a convolutional layer, and a Tanh layer.
[0214] Hereinafter, as an example, a data coding system including FIGS. 4A and 4B (hereinafter referred to as the first data coding system) is used to compare the data coding system with the data coding system shown in FIG. 6A (hereinafter referred to as the second data coding system) for the purpose of explanation. In the first data coding system, there is a scaling operation on the residual feature map, while in the second data coding system, there is a scaling operation on the feature map.
[0215] For example, the data is image data. The Kodak test set is used as the data to be processed. The test set includes 24 PNG images with a resolution of 768×512 or 512×768. Each test image is processed by separately using the first data coding system and the second data coding system, and the BPP (Bits per pixel), Peak Signal-to-Noise Ratio (PSNR), and BD rate (Bjoentegaard-Delta bit Rate) are obtained. For details, refer to Table 1. The BPP and PSNR in Table 1 are the average values of the 24 images. BPP represents the average number of bits used by a pixel, and a smaller BPP indicates a smaller compression bit rate. The peak signal-to-noise ratio is an objective criterion for evaluating image quality, and a higher peak signal-to-noise ratio indicates better image quality. The BD rate indicates the bit rate reduction (increase in compression rate) in two comparison methods under the same image quality.
[0216] It can be seen from Table 1 that the results of the first data coding system are better than those of the second data coding system. The BD rate of the first data coding system is better, and the performance is improved by 3.87%. It can be seen that scaling and quantization on the residual feature map can improve the compression performance. Performance Comparison of the First Data Coding System and the Second Data Coding System in Table 1 [Table 1]
[0217] In the following, as an example, using the data coding system shown in Figure 6B (hereinafter referred to as the third data coding system), the data coding system is compared with the second data coding system shown in Figure 6A for the purpose of explanation. In the second data coding system, there is a scaling operation on the feature map, while in the third data coding system, there are a scaling operation on the feature map and a scaling operation on the probability distribution parameters. For example, in the third data coding system, both the mean μ and the variance σ are scaled.
[0218] For example, the data is image data. The Kodak test set is used as the data to be processed. The test set includes 24 PNG images with a resolution of 768×512 or 512×768. Each test image is processed by separately using the second data coding system and the third data coding system, and the BPP, PSNR, and BD rate are obtained. For details, please refer to Table 2. Similarly, the BPP and PSNR in Table 2 are the average values of 24 images.
[0219] It can be seen from Table 2 that the results of the third data coding system are better than those of the second data coding system. The BD rate of the third data coding system is better, and the performance has improved by 2.71%. It can be seen that scaling the probability distribution parameters can improve the compression performance. Performance Comparison of the Second Data Coding System and the Third Data Coding System in Table 2 [Table 2]
[0220] Below, for the sake of explanation, the first data coding system is compared with a data coding system including FIGS. 4D and 4E (hereinafter referred to as the fourth data coding system). In the first data coding system, there is a scaling operation on the residual feature map, while in the fourth data coding system, there are a scaling operation on the residual feature map and a scaling operation on the probability distribution parameters. For example, in the fourth data coding system, the variance σ is scaled.
[0221] For example, the data is image data. The Kodak test set is used as the data to be processed. The test set includes 24 PNG images having a resolution of 768×512 or 512×768. Each test image is processed by separately using the first data coding system and the fourth data coding system, and the BPP, PSNR, and BD rate are obtained. For details, refer to Table 3. Similarly, the BPP and PSNR in Table 3 are the average values of 24 images.
[0222] It can be seen from Table 3 that the results of the fourth data coding system are better than those of the first data coding system. The BD rate of the fourth data coding system is better, and the performance is improved by 0.65%. It can be seen that scaling the probability distribution parameters can improve the compression performance. Table 3 Performance comparison between the first data coding system and the fourth data coding system
Table 3
[0223] FIG. 7 is a schematic flowchart of a data encoding method according to an embodiment of the present application. The data encoding method 700 is executed by a data encoder 20. The method shown in FIG. 7 is described as a series of steps or operations. It should be understood that the steps or operations of the method may be executed in various orders and / or simultaneously, and are not limited to the execution order shown in FIG. 7.
[0224] As shown in FIG. 7, the data encoding method 700 includes the following steps.
[0225] 701: Execute side information feature extraction on the first feature map of the current data to obtain a side information feature map.
[0226] The data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data. Side information means that existing information Y is used to assist in encoding information X so that the encoding length of information X can be made shorter. In other words, the redundancy in information X is reduced. Information Y is side information. In this embodiment of the present application, the side information is information extracted from the first feature map and used to assist in encoding and decoding the first feature map.
[0227] 702: Execute quantization processing on the side information feature map to obtain a first quantized feature map.
[0228] 703: Execute entropy encoding on the first quantized feature map to obtain a first bit stream of the current data.
[0229] The bitstream is a bitstream generated after the encoding process. In this case, the first bitstream is the bitstream obtained after entropy encoding is performed on the first quantized feature map.
[0230] 704: Perform scaling processing on the residual feature map based on the scaling coefficient to obtain a scaled feature map, perform quantization processing on the scaled feature map to obtain a second quantized feature map. Obtain a residual feature map based on the second feature map of the current data and the first probability distribution parameter, and obtain a scaling coefficient based on the first quantized feature map. In this case, the scaling process is a downscaling process.
[0231] In this embodiment, the first feature map and the second feature map are different feature maps obtained by performing feature extraction on the complete current data.
[0232] 705: Perform entropy encoding on the second quantized feature map based on the first probability distribution parameter to obtain a second bitstream of the current data.
[0233] The first probability distribution parameter is a matrix of the same scale as the scale of the second quantized feature map. Each element within the matrix represents a probability distribution parameter of one of the elements within the second quantized feature map. Each probability distribution parameter of the elements includes, but is not limited to, the mean and / or variance. That is, one second quantized feature map corresponds to one mean matrix and / or one variance matrix. The second bitstream is the bitstream obtained after entropy encoding is performed on the second quantized feature map.
[0234] In this solution, during encoding, after performing scaling processing on the residual feature map of the current data to obtain a scaled feature map, quantization processing is performed on the scaled feature map so that the quantization loss of the scaled feature map becomes smaller. That is, the information loss of the second bitstream becomes smaller. This helps to improve the data quality of the reconstructed data obtained through decoding based on the second bitstream. Furthermore, compared with the encoding network structure in the prior art, in this embodiment of the present application, after the training of the entire encoding network is completed, the network parameters of the entire encoding network including the network for generating the first bitstream and the network for generating the second bitstream can be optimized, so that a network structure including scaling for the residual feature map is used. Therefore, by using the encoding network structure in this embodiment of the present application, the data volume of the total bitstream of the current data can be reduced, and the encoding efficiency can be improved. Generally, the encoding method in this embodiment can further improve the data compression performance.
[0235] For example, the first quantized feature map may be input into the super-prior decoding network, and the first probability distribution parameter and the scaling coefficient may be obtained through prediction. The first probability distribution parameter represents the probability distribution of the second quantized feature map.
[0236] As another example, the first quantized feature map may be input into the super-prior decoding network, and the scaling coefficient may be obtained through prediction. The first probability distribution parameter is a pre-set probability distribution parameter.
[0237] In some possible embodiments, the first probability distribution parameter includes an average and / or a variance.
[0238] In some possible embodiments, the first feature map is the same as the second feature map.
[0239] FIG. 8 is a schematic flowchart of another data encoding method according to an embodiment of the present application. The data encoding method 800 is executed by the data encoder 20. The method shown in FIG. 8 is described as a series of steps or operations. It should be understood that the steps or operations of the method may be executed in various orders and / or simultaneously, and are not limited to the execution order shown in FIG. 8.
[0240] As shown in FIG. 8, the data encoding method 800 includes the following steps.
[0241] 801: Perform side information feature extraction on the first feature map of the current data to obtain a side information feature map.
[0242] The data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data. Side information means that existing information Y is used to assist in encoding information X so that the encoding length of information X can be made shorter. In other words, the redundancy in information X is reduced. Information Y is side information. In this embodiment of the present application, the side information is information extracted from the first feature map and used to assist in encoding and decoding the first feature map.
[0243] 802: Perform quantization processing on the side information feature map to obtain a first quantized feature map.
[0244] 803: Perform entropy encoding on the first quantized feature map to obtain a first bitstream of the current data.
[0245] The bitstream is the bitstream generated after the encoding process. The first bitstream is the bitstream obtained after entropy encoding is performed on the first quantized feature map.
[0246] 804: Perform a scaling process on the second feature map based on the scaling factor to obtain a scaled feature map.
[0247] Obtain a scaling factor based on the first quantized feature map. In this embodiment, the first feature map and the second feature map are different feature maps obtained by performing feature extraction on the complete current data. For example, when the first feature map is different from the second feature map, for the specific method of obtaining the first feature map (y1 in FIG. 5B) and the second feature map (y2 in FIG. 5B), the relevant description in FIG. 5B may be referred to. In this case, the scaling process is a downscaling process.
[0248] 805: Perform a quantization process on the scaled feature map to obtain a second quantized feature map.
[0249] 806: Perform a scaling process on the first probability distribution parameter based on the scaling factor to obtain a second probability distribution parameter.
[0250] The first probability distribution parameter is obtained based on the first quantized feature map. Alternatively, the first probability distribution parameter is a pre-set probability distribution parameter. The first probability distribution parameter is a matrix with the same scale as the scale of the second quantized feature map. Each element in the matrix represents a probability distribution parameter of one of the elements in the second quantized feature map. Each probability distribution parameter of the element includes, but is not limited to, the mean and / or variance. That is, one second quantized feature map corresponds to one mean matrix and / or one variance matrix.
[0251] 807: Execute entropy encoding on the second quantization feature map based on the second probability distribution parameter to obtain the second bitstream of the current data.
[0252] The second bitstream is the bitstream obtained after entropy encoding is executed on the second quantization feature map.
[0253] In this solution, entropy encoding is performed based on the first quantization feature map to obtain the first bitstream of the current data. Further, the scaling coefficient and the first probability distribution parameter may be obtained through an estimation based on the first quantization feature map. In this way, scaling processing may be performed on the second feature map based on the scaling coefficient to obtain a scaled feature map. Next, quantization processing is performed on the scaled feature map to obtain a second quantization feature map. Scaling processing is performed on the first probability distribution parameter based on the scaling coefficient to obtain a second probability distribution parameter. Finally, entropy encoding is performed on the second quantization feature map based on the second probability distribution parameter to obtain the second bitstream of the current data. The first bitstream and the second bitstream are used together as the total bitstream of the current data. The data encoding method in the prior art only includes a stage of performing scaling processing on the feature map. In this solution, the second feature map and the first probability distribution parameter are scaled by using the same scaling coefficient such that the degree of match between the second probability distribution parameter and the second quantization feature map is higher, thereby improving the encoding accuracy of the second quantization feature map, that is, improving the data compression performance. In some possible embodiments, the second feature map is a residual feature map obtained based on the third feature map of the current data and the first probability distribution parameter of the current data, and the third feature map is obtained by performing feature extraction on the current data.
[0254] In this case, in this embodiment, the second feature map is a feature map obtained based on current data, that is, a residual feature map. In this embodiment, the third feature map is a feature map different from the first feature map obtained by performing feature extraction on the complete current data. The residual of the third feature map may be obtained based on the third feature map and the first probability distribution parameter. That is, the residual feature map may be obtained by subtracting the first probability distribution parameter from the third feature map, and the residual feature map is used as the second feature map of the current data. For example, when the first feature map is different from the third feature map, for a specific method for obtaining the first feature map (y1 in FIG. 5B) and the third feature map (y2 in FIG. 5B), reference may be made to the relevant description in FIG. 5B.
[0255] In this solution, during encoding, after performing scaling processing on the residual feature map of the current data to obtain a scaled feature map, quantization processing is performed on the scaled feature map so that the quantization loss of the scaled feature map becomes smaller. That is, the information loss of the second bitstream becomes smaller. This helps to improve the data quality of the reconstructed data obtained through decoding based on the second bitstream. Further, compared with the encoding network structure in the prior art, in this embodiment of the present application, after the training of the entire encoding network is completed, the network parameters of the entire encoding network including the network for generating the first bitstream and the network for generating the second bitstream can be optimized, and an encoding network structure including scaling for the residual feature map is used. Therefore, by using the encoding network structure in this embodiment of the present application, the data volume of the total bitstream of the current data can be reduced, and the encoding efficiency can be improved. That is, the scaling for the residual feature map and the scaling for the first probability distribution parameter are combined, thereby further improving the data compression performance.
[0256] Further, in some possible embodiments, the first feature map is the same as the third feature map. For example, for the specific method for obtaining the first feature map and the third feature map, reference may be made to the relevant description in FIG. 5A.
[0257] In some possible embodiments, the second feature map is a feature map obtained by performing feature extraction on the current data. The first feature map may be the same as or different from the second feature map. For example, for the specific method for obtaining the first feature map and the second feature map, reference may be made to the relevant description in FIG. 5A.
[0258] In some possible embodiments, the first probability distribution parameter includes an average and / or a variance.
[0259] In some possible embodiments, the data encoding method is: transmitting a first bitstream and a second bitstream and further comprises.
[0260] After obtaining the first bitstream and the second bitstream of the current data by using the data encoding method in this solution means, the first bitstream and the second bitstream may be transmitted to another device according to requirements, whereby the another device can process the first bitstream and the second bitstream.
[0261] In some possible embodiments, the first bitstream and the second bitstream are stored in the form of a bitstream file.
[0262] FIG. 9 is a schematic flowchart of a data decoding method according to an embodiment of the present application. The data decoding method 900 is executed by a data decoder 30. The method shown in FIG. 9 is described as a series of steps or operations. It should be understood that the steps or operations of the method may be executed in various orders and / or simultaneously, and are not limited to the execution order shown in FIG. 9.
[0263] As shown in FIG. 9, the data decoding method 900 comprises the following steps.
[0264] 901: Perform entropy decoding based on the first bitstream of the current data to obtain a third feature map.
[0265] 902: Obtain a scaling coefficient based on the third feature map.
[0266] 903: Perform entropy decoding on the second bitstream of the current data based on the first probability distribution parameter to obtain a fourth feature map.
[0267] The first probability distribution parameter is a matrix of the same scale as the scale of the fourth feature map. Each element within the matrix represents a probability distribution parameter of one of the elements in the fourth feature map. The probability distribution parameter of each element includes, but is not limited to, the mean and / or variance. That is, one fourth feature map corresponds to one mean matrix and / or one variance matrix.
[0268] 904: Perform a scaling process on the fourth feature map based on a scaling coefficient to obtain a fifth feature map. In this case, the scaling process is an upscaling process.
[0269] 905: Obtain a sixth feature map based on the first probability distribution parameter and the fifth feature map.
[0270] 906: Obtain reconstructed data of the current data based on the sixth feature map.
[0271] When performing scaling and quantization operations on the residual feature map during data encoding, the data volume of the first bitstream and the second bitstream in the data decoding method of this solution means is smaller than that in the prior art. Therefore, the corresponding decoding processing volume is smaller, and this solution means can effectively improve the decoding efficiency. Furthermore, the information loss of the second bitstream is smaller, and thereby, the data quality of the reconstructed data obtained by using this solution means is higher.
[0272] In some possible embodiments, the first feature map is the same as or different from the second feature map.
[0273] In some possible embodiments, the data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data.
[0274] In some possible embodiments, the first probability distribution parameter includes an average and / or a variance.
[0275] In some possible embodiments, the data decoding method is as follows: Obtaining a first probability distribution parameter based on a third feature map and further includes.
[0276] In some possible embodiments, the first probability distribution parameter is a pre-set probability distribution parameter.
[0277] FIG. 10 is a schematic flowchart of another data decoding method according to an embodiment of the present application. The data decoding method 1000 is executed by a data decoder 30. The method shown in FIG. 10 is described as a series of steps or operations. It should be understood that the steps or operations of the method may be executed in various orders and / or simultaneously, and are not limited to the execution order shown in FIG. 10.
[0278] As shown in FIG. 10, the data decoding method 1000 includes the following steps.
[0279] 1001: Perform entropy decoding based on the first bitstream of the current data to obtain a third feature map.
[0280] 1002: Obtain a scaling coefficient based on the third feature map.
[0281] 1003: Perform scaling processing on the first probability distribution parameter based on the scaling coefficient to obtain a second probability distribution parameter.
[0282] The first probability distribution parameter is a matrix of the same scale as the scale of the fourth feature map. Each of the elements in the matrix represents a probability distribution parameter of one of the elements in the fourth feature map. Each probability distribution parameter of the elements includes, but is not limited to, an average and / or a variance. That is, one fourth feature map corresponds to one average matrix and / or one variance matrix.
[0283] 1004: Perform entropy decoding on the second bitstream of the current data based on the second probability distribution parameter to obtain a fourth feature map.
[0284] 1005: Perform a scaling process on the fourth feature map based on a scaling coefficient to obtain a fifth feature map. In this case, the scaling process is an upscaling process.
[0285] 1006: Obtain reconstructed data of the current data based on the fifth feature map.
[0286] During data encoding, when performing a scaling process on the second feature map and a scaling process on the first probability distribution parameter, correspondingly, during data decoding, by using the same scaling coefficient to process the first probability distribution parameter and the fourth feature map, the decoding accuracy of the fourth feature map is ensured. Further, a descaling process may be performed on the fourth feature map based on the scaling coefficient to obtain a fifth feature map, and reconstructed data may be obtained based on the fifth feature map, thereby making the accuracy and quality of the reconstructed data higher. That is, by combining the scaling for the first probability distribution parameter and the descaling process for the fourth feature map, the accuracy and quality of data decoding can be improved.
[0287] In some possible embodiments, the fourth feature map is a residual feature map, and the step of obtaining the reconstructed data of the current data based on the fifth feature map is: The step of adding the first probability distribution parameter to the fifth feature map to obtain a sixth feature map; and The step of obtaining the reconstructed data of the current data based on the sixth feature map has.
[0288] When performing scaling and quantization operations on the residual feature map during data encoding, the data amounts of the first bitstream and the second bitstream in the data decoding method of this solution means are smaller than those in the prior art. Therefore, the corresponding decoding processing amount is smaller, and this solution means can effectively improve the decoding efficiency. Furthermore, the information loss of the second bitstream is smaller, so that the data quality of the reconstructed data obtained by using this solution means is higher.
[0289] In some possible embodiments, the first probability distribution parameter includes an average and / or a variance.
[0290] In some possible embodiments, the data includes at least one of the following: image data, video data, motion vector data of video data, audio data, point cloud data, or text data.
[0291] In some possible embodiments, the data decoding method is: The step of receiving the first bitstream and the second bitstream of the current data further includes.
[0292] In some possible embodiments, the data decoding method further includes the step of obtaining the first probability distribution parameter based on the third feature map.
[0293] In some possible embodiments, the first probability distribution parameter is a pre-set probability distribution parameter.
[0294] When the function of the method according to any one of the embodiments of the present application is implemented in the form of a software functional unit and sold or used as an independent product, the above function may be stored in a computer-readable storage medium. Based on such understanding, the data encoding method and / or data decoding method in the present application, which are essentially or contribute a part of the prior art or a part of the technical solution, may be implemented in the form of a computer program product. The computer program product includes a plurality of instructions stored in a storage medium for instructing an electronic device to execute all or some of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes any medium that can store program code, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0295] The present application further provides a computer-readable storage medium. The storage medium stores a bitstream, and the bitstream is generated according to the data encoding method according to any one of the above embodiments.
[0296] The present application further provides a computer-readable storage medium that stores a bitstream including program code. When the program code is executed by one or more processors, a decoder can execute the data decoding method according to any one of the above embodiments.
[0297] One embodiment of the present application further provides a chip. The chip is used in an electronic device. The chip includes one or more processors, and the processors are configured to call computer instructions so that the electronic device can execute the data encoding method and / or the data decoding method according to any one of the foregoing embodiments.
[0298] One embodiment of the present application further provides a computer program product including instructions. When the computer program product is executed on an electronic device, the electronic device can execute the data encoding method and / or the data decoding method according to any one of the foregoing embodiments.
[0299] It can be understood that the computer storage medium, the chip, and the computer program product provided above are all configured to execute the data encoding method and / or the data decoding method according to any one of the foregoing embodiments. Therefore, for the beneficial effects that can be achieved, reference may be made to the beneficial effects of the data encoding method and / or the data decoding method according to any one of the foregoing embodiments. Details will not be described again here.
[0300] Those skilled in the art can recognize that the units and algorithm steps in the examples described with reference to the embodiments disclosed in this specification can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software depends on the specific application and the design constraints of the technical solution. Those skilled in the art may use different methods to realize the functions described for each specific application, but the implementation form should not be regarded as exceeding the scope of the present application.
[0301] In some embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the described embodiments of the apparatus are merely examples. For example, the division into multiple units is only a logical function division, and during actual implementation, other divisions may be possible. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the shown or described interconnections or direct connections or communication connections may be implemented through some interfaces. The indirect coupling or communication connection between devices or units can be implemented in electronic, mechanical or other forms.
[0302] Each unit described as a separate part may or may not be physically separated, and the part shown as a unit may or may not be a physical unit. In other words, it may be located in one place or distributed over multiple network units. Some or all of the units may be selected according to the actual requirements for achieving the objectives of the solution of the embodiment.
[0303] Furthermore, the functional units in the embodiments of this patent application may be integrated into one processing unit, or each of these units may physically exist alone, or two or more units may be integrated into one unit.
[0304] The foregoing description is only a specific implementation form of the present application and is not intended to limit the protection scope of the present application. Any deformation or substitution that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall follow the protection scope of the claims.
Claims
1. A data encoding method, the method comprising the following steps: Performing side information feature extraction on a first feature map of current data to obtain a side information feature map; Performing quantization processing on the side information feature map to obtain a first quantized feature map; Performing entropy encoding on the first quantized feature map to obtain a first bitstream of the current data; Performing scaling processing on a second feature map based on a scaling coefficient to obtain a scaled feature map, wherein the scaling coefficient is obtained based on the first quantized feature map; Performing quantization processing on the scaled feature map to obtain the second quantized feature map; Performing scaling processing on a first probability distribution parameter based on the scaling coefficient to obtain a second probability distribution parameter; and Performing entropy encoding on the second quantized feature map based on the second probability distribution parameter to obtain a second bitstream of the current data A method comprising the above steps.
2. The second feature map is a residual feature map of the current data obtained based on a third feature map of the current data and the first probability distribution parameter, and the third feature map is obtained by performing feature extraction on the current data. The method according to claim 1.
3. The second feature map is a feature map obtained by performing feature extraction on the current data. The method according to claim 1.
4. The first feature map is the same as the third feature map. The method according to claim 2.
5. The first feature map is the same as the second feature map. The method according to claim 3.
6. The first probability distribution parameter includes an average and / or a variance. The method according to any one of claims 1 to 5.
7. The method further comprises: Transmitting the first bitstream and the second bitstream The method according to any one of claims 1 to 6.
8. The method according to any one of claims 1 to 7, wherein the data includes at least one of the following: image data, video data, motion vector data of the video data, audio data, point cloud data, or text data.
9. A data decoding method, the method comprising the following steps: Performing entropy decoding based on a first bitstream of current data to obtain a third feature map; Obtaining a scaling coefficient based on the third feature map; Performing scaling processing on a first probability distribution parameter based on the scaling coefficient to obtain a second probability distribution parameter; Performing entropy decoding on a second bitstream of the current data based on the second probability distribution parameter to obtain a fourth feature map; Performing scaling processing on the fourth feature map based on the scaling coefficient to obtain a fifth feature map; and Obtaining reconstructed data of the current data based on the fifth feature map A method comprising.
10. The fourth feature map is a residual feature map, the fifth feature map is a scaled residual feature map, and the step of obtaining the reconstructed data of the current data based on the fifth feature map is: Adding the first probability distribution parameter to the fifth feature map to obtain a sixth feature map; and Obtaining the reconstructed data of the current data based on the sixth feature map The method according to claim 9, comprising.
11. The method according to claim 9 or 10, wherein the first probability distribution parameter includes an average and / or a variance.
12. The method according to any one of claims 9 to 11, wherein the data includes at least one of the following: image data, video data, motion vector data of the video data, audio data, point cloud data, or text data.
13. The method comprises: Receiving the first bitstream and the second bitstream of the current data The method according to any one of claims 9 to 12, further comprising.
14. A data encoder comprising a processing circuit configured to execute the data encoding method according to any one of claims 1 to 8.
15. A computer-readable storage medium, wherein the storage medium stores a bit stream, and the bit stream is generated according to the data encoding method according to any one of claims 1 to 8.
16. A data decoder comprising a processing circuit configured to execute the data decoding method according to any one of claims 9 to 13.
17. A computer program product comprising program code, wherein when the program code is executed on a computer or a processor, the computer program product is configured to execute the method according to any one of claims 1 to 13.
18. A data encoder, one or more processors; and a computer-readable storage medium coupled to the one or more processors, wherein the computer-readable storage medium stores a program, and when the program is executed by the one or more processors, the data encoder is capable of executing the data encoding method according to any one of claims 1 to 8 A data encoder comprising the above.
19. A data decoder, one or more processors; and a computer-readable storage medium coupled to the one or more processors, wherein the computer-readable storage medium stores a program, and when the program is executed by the one or more processors, the data decoder is capable of executing the data decoding method according to any one of claims 9 to 13 A data decoder comprising the above.
20. A computer-readable storage medium comprising program code, wherein when the program code is executed by a computer device, the computer-readable storage medium is configured to execute the method according to any one of claims 1 to 13.
21. A computer-readable storage medium storing a bitstream including program code, wherein when the program code is executed by one or more processors, a decoder is capable of executing the data decoding method according to any one of claims 9 to 13.
Citation Information
Patent Citations
Image coding method and apparatus and image decoding method and apparatus
JP2020191077A
Image processing method and related device
JP2023512570A
Video decoding method, apparatus, and computer program, and video encoding method
JP2023527664A
Image coding method and apparatus and image decoding method and apparatus
US20200374522A1
Content-adaptive online training with scaling factors and / or offsets in neural image compression
US20220353522A1
Cited By
Decoding, encoding methods, apparatus, devices and media
JP2026512152A
Decoding, encoding methods, apparatus, devices and media
JP7866154B2