Image encoding method, image decoding method, device, and storage medium

Through the combination of feature transformation, quantization and entropy coding, the problem of feature information loss in the existing image encoding methods is solved, efficient feature coding and decoding is achieved, and the performance of machine vision tasks is improved.

CN115361559BActive Publication Date: 2025-07-11ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210772560.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-07-11
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

The existing image encoding methods cannot effectively take into account the performance requirements of machine vision tasks. The traditional encoding methods have lossy encoding resulting in lossy feature information, and the lossy encoding of hybrid encoders is inconsistent, which affects the visual analysis effect.

Method used

The combination of feature transformation, quantization and entropy coding is adopted to achieve efficient coding of features through dimensionality reduction network and configuration parameter optimization, including dimensionality reduction network, quantization processing and entropy coding. The feature transformation is performed using convolutional layer, full connection layer, attention mechanism and self-attention mechanism, and a probability model is constructed for entropy coding in combination with linear and nonlinear quantization methods.

Benefits of technology

The feature encoding rate is improved, ensuring that the performance of machine vision tasks is not lost, and the efficient transmission and reconstruction of feature information is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115361559B_ABST
    Figure CN115361559B_ABST
Patent Text Reader

Abstract

The present application discloses an image encoding method, an image decoding method, an apparatus, and a computer storage medium. The image encoding method includes: obtaining an encoding target feature of an image to be processed; performing feature transformation on the encoding target feature to obtain a transformed feature, where the feature dimension of the transformed feature is lower than that of the encoding target feature; performing quantization processing on the transformed feature based on configuration parameters to obtain a quantized feature; and performing feature encoding on the quantized feature to obtain a feature bitstream. The image encoding method of the present application can further improve the encoding rate of features through a simple and effective quantization method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of feature encoding, and particularly to an image encoding method, an image decoding method, an apparatus, and a computer storage medium. Background Art

[0002] Traditional image encoding technologies are designed for human visual characteristics. With the excellent performance of deep neural networks demonstrated in various machine vision tasks, such as image classification, object detection, semantic segmentation, etc., a large number of artificial intelligence applications based on machine vision have emerged. To ensure that the performance of machine vision tasks is not impaired by the image encoding process, a mode of analyzing first and then encoding is adopted to meet the machine vision requirements, that is, at the image acquisition end, a lossless image is directly subjected to feature extraction through a neural network, and then the extracted features are encoded and transmitted. The decoding end directly uses the decoded features to input into subsequent network structures to complete different machine vision tasks. Therefore, in order to save transmission bandwidth resources, it is necessary to study an image encoding method for machine vision.

[0003] Currently, there are mainly two categories of feature encoding algorithms: traditional encoding methods and learning-based schemes. Among them, the traditional encoding methods mainly include the following. One is to use low-precision data types to replace high-precision data types, thereby reducing the space occupied by the original feature data. However, in essence, it is not a real encoding of the feature data, but an implementation from the perspective of computer storage. The second is to extract the main data component information of the original feature data through dimensionality reduction methods, such as PCA (Principal Component Analysis), so that low-dimensional data can generally represent the information of the original data, which belongs to lossy encoding. The third is a hybrid encoder scheme, that is, first quantize the deep features, and then use encoders such as High Efficiency Video Coding (HEVC), H.266 / VVC, etc. to perform lossy encoding on the quantized features. The disadvantage of this scheme is that the lossy encoding degradation of the hybrid encoder is inconsistent with the performance degradation of the features during visual analysis tasks, resulting in the features being unable to provide important information required for visual analysis. Summary of the Invention

[0004] This application provides an image encoding method, an image decoding method, an image encoding apparatus, and a computer storage medium.

[0005] One technical solution adopted by this application is to provide an image encoding method, and the image encoding method includes:

[0006] Obtain the feature to be encoded of the image to be processed;

[0007] Perform a feature transformation on the feature to be encoded to obtain a transformed feature, where the feature dimension of the transformed feature is lower than the feature dimension of the feature to be encoded;

[0008] Quantize the transformation feature based on configuration parameters to obtain a quantized feature;

[0009] Perform feature encoding on the quantized feature to obtain a feature bitstream.

[0010] Among them, the obtaining of the transformation feature by performing feature transformation on the to-be-encoded feature includes:

[0011] Input the to-be-encoded feature into a dimensionality reduction network, and perform downsampling on the to-be-encoded feature through the convolutional layer and / or fully connected layer of the dimensionality reduction network to obtain the transformation feature.

[0012] Among them, the convolutional layer of the dimensionality reduction network is a one-dimensional convolutional layer or a two-dimensional convolutional layer.

[0013] Among them, the dimensionality reduction network further includes one or more of a spatial feature transformation sub-network, a channel attention mechanism sub-network, and a self-attention mechanism sub-network.

[0014] Among them, the inputting of the to-be-encoded feature into the dimensionality reduction network includes:

[0015] Input the to-be-encoded feature into several dimensionality reduction sub-networks of the dimensionality reduction network in sequence, and each dimensionality reduction layer sub-network includes a fully connected layer, a normalization layer, and an activation layer connected in series in sequence.

[0016] Among them, the obtaining of the transformation feature by performing feature transformation on the to-be-encoded feature includes:

[0017] Perform feature sparsification processing on the to-be-encoded feature based on at least one of an unsupervised dimensionality reduction algorithm and a supervised dimensionality reduction algorithm to obtain the transformation feature.

[0018] Among them, the quantizing of the transformation feature based on configuration parameters to obtain a quantized feature includes:

[0019] Obtain a preset linear transformation function, and assign values to the non-learning parameters in the preset linear transformation function based on the configuration parameters;

[0020] Use the assigned preset linear transformation function and a preset bit depth to map the transformation feature to obtain the quantized feature.

[0021] Among them, before the using of the assigned preset linear transformation function and a preset bit depth to map the transformation feature to obtain the quantized feature, the image encoding method further includes:

[0022] Perform a non-linear transformation on the transformation feature using a preset non-linear function to obtain a non-linearly transformed transformation feature.

[0023] Among them, after quantizing the transform features based on the configuration parameters to obtain quantized features,

[0024] The image encoding method further includes:

[0025] Performing inverse quantization on the quantized features to obtain inverse quantized features;

[0026] Obtaining a quantization loss value based on the difference information between the transform features and the inverse quantized features;

[0027] Training the learning parameters in the preset linear transformation function by using the quantization loss value.

[0028] Among them, after quantizing the transform features based on the configuration parameters to obtain quantized features, the image encoding method further includes:

[0029] Using an entropy coding model to extract context feature information of the quantized features;

[0030] Predicting the quantized features based on the context feature information of the quantized features to obtain entropy-coded features of the quantized features;

[0031] Performing feature encoding based on the entropy-coded features to obtain the feature bitstream.

[0032] Among them, the entropy coding model includes a probability model constructed by using a hyperprior network, where the probability model is a single Gaussian model, a mixture Gaussian model, a Laplace model, a logistic regression model, or a combined model of one or more of them.

[0033] Another technical solution adopted by this application is to provide an image decoding method, and the image decoding method includes:

[0034] Performing feature decoding on the feature bitstream to obtain decoded features;

[0035] Based on configuration parameters, performing inverse quantization on the decoded features to obtain inverse quantized features;

[0036] Performing feature inverse transformation on the inverse quantized features to obtain inverse transformed features, where the feature dimension of the inverse transformed features is higher than the feature dimension of the inverse quantized features;

[0037] Performing feature reconstruction on the inverse transformed features to obtain a reconstructed image.

[0038] Another technical solution adopted by this application is to provide an image encoding device, and the image encoding device includes a memory and a processor coupled to the memory;

[0039] Among them, the memory is used to store program data, and the processor is used to execute the program data to implement the image encoding method as described above.

[0040] Another technical solution adopted by this application is to provide an image decoding device, which includes a memory and a processor coupled to the memory;

[0041] Among them, the memory is used to store program data, and the processor is used to execute the program data to implement the image decoding method as described above.

[0042] Another technical solution adopted by this application is to provide a computer storage medium, which is used to store program data. When the program data is executed by a computer, it is used to implement the image encoding method and / or the image encoding method as described above.

[0043] The beneficial effect of this application is as follows: The image encoding device obtains the to-be-encoded features of the to-be-processed image; performs feature transformation on the to-be-encoded features to obtain transformed features, where the feature dimension of the transformed features is lower than that of the to-be-encoded features; based on configuration parameters, performs quantization processing on the transformed features to obtain quantized features; performs feature encoding on the quantized features to obtain a feature bitstream. The image encoding method of this application can further improve the encoding rate of features through a simple and effective quantization method. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0045] Figure 1 It is a schematic flowchart of an embodiment of the image encoding method provided by this application;

[0046] Figure 2 It is a schematic diagram of the overall framework structure of the feature encoding provided by this application;

[0047] Figure 3 It is a schematic diagram of the structure of the dimensionality reduction network based on two-dimensional convolution provided by this application;

[0048] Figure 4 It is a schematic diagram of an embodiment of the SFT structure provided by this application;

[0049] Figure 5 It is a schematic diagram of an embodiment of the channel attention structure provided by this application;

[0050] Figure 6 It is a schematic structural diagram of an embodiment of the fully connected dimensionality reduction network provided by this application;

[0051] Figure 7 It is a schematic diagram of the feature encoding framework including the entropy model provided by this application;

[0052] Figure 8 It is a schematic flowchart of an embodiment of the image decoding method provided by this application;

[0053] Figure 9 It is a schematic structural diagram of an embodiment of the image encoding device provided by this application;

[0054] Figure 10 It is a schematic structural diagram of an embodiment of the image decoding device provided by this application;

[0055] Figure 11 It is a schematic structural diagram of an embodiment of the computer storage medium provided by this application. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0057] Specifically, please refer to Figure 1 and Figure 2 , Figure 1 is a schematic flowchart of an embodiment of the image encoding method provided by this application, Figure 2 is a schematic diagram of the overall framework structure of the feature encoding provided by this application.

[0058] As Figure 2 shown, Figure 2 represents the overall framework structure of the image encoding method and the image decoding method provided by this application, and the image encoding method and the image decoding method are essentially inverse processes. Specifically, the overall framework structure sequentially includes a feature transformation module, a quantization module, an entropy encoding module, an entropy decoding module, an inverse quantization module, and a feature reconstruction module.

[0059] Among them, the feature transformation module performs a compact space transformation on the input original features to obtain a compact representation of the dimensionality-reduced features. The dimensionality reduction method of the feature transformation module may include, but is not limited to: traditional feature dimensionality reduction methods, deep learning-based dimensionality reduction methods. And the processing process of the feature reconstruction module is the inverse process of the processing process of the feature transformation module.

[0060] The quantization module assigns quantization parameters to the transformed features and performs quantization, which can further compress the size of the feature data. The processing process of the dequantization module is the inverse process of the processing process of the quantization module.

[0061] The entropy encoding module is an optional module in the overall framework structure. The entropy encoding module can construct a probability model based on the context information of the features, enabling the probability model to accurately predict the probability of each character appearing in the feature data, thereby reducing the redundancy of the feature data. The processing process of the entropy decoding module is the inverse process of the processing process of the entropy encoding module.

[0062] Next, in conjunction with Figure 2 the overall framework structure shown below, the image encoding method and the image decoding method provided by this application will be further introduced:

[0063] As Figure 1 shown, the image encoding method of the embodiment of this application includes the following steps:

[0064] Step S11: Obtain the features to be encoded of the image to be processed.

[0065] Step S12: Perform feature transformation on the features to be encoded to obtain transformed features, where the feature dimension of the transformed features is lower than the feature dimension of the features to be encoded.

[0066] In the embodiment of this application, the image encoding device performs feature transformation on the input features to be encoded, that is, a compact space transformation, to obtain transformed features, that is, a compact representation of the features after dimensionality reduction. The specific process can be: The image encoding device performs a compact space transformation on the input original features through an indirect rate constraint, that is, directly performs a rate constraint on the quantized multi-channel feature map, and indirectly realizes a compact space transformation on the input original features to obtain a compact feature representation of the original features.

[0067] Considering the differences between features and images / videos, audio, text, etc., this application can also design some more effective structures to capture the semantic information of the feature data. The main ideas are divided into: traditional feature dimensionality reduction methods and deep learning-based dimensionality reduction methods.

[0068] Next, the deep learning-based dimensionality reduction method will be introduced first:

[0069] Deep learning-based methods usually rely on VAE (Variational AutoEncoder) or GAN (Generative Adversarial Networks) to construct. Considering the characteristics of the feature data (the feature data contains more abstract semantic information, is sparser, and may no longer have the spatial correlation in the original image), the following several solutions are proposed:

[0070] (1) Use two-dimensional convolution based on the attention mechanism or spatial correlation to achieve dimensionality reduction and improve the encoding efficiency.

[0071] (2) Use one-dimensional convolution / full connection for dimensionality reduction, which is more friendly to feature information and can better capture the semantic information of feature data.

[0072] In addition, adopting some strategies for network design can improve the performance of the network.

[0073] For example, please refer to Figure 3 , Figure 3 , which is a schematic structural diagram of the dimensionality reduction network based on two-dimensional convolution provided by this application. The dimensionality reduction network based on two-dimensional convolution uses the convolutional layer to perform convolutional downsampling on the input features, thereby reducing the feature dimension of the input features.

[0074] In one of the embodiments such as Figure 3 , the input features are sequentially passed through convolutional downsampling, activation, convolutional downsampling, activation, and several residual blocks to obtain the transformed features after dimensionality reduction.

[0075] Specifically, in the dimensionality reduction network based on two-dimensional convolution, the convolutional kernel can be 3x3, 5x5, 7x7, etc., and the size and number of convolutional kernels are not limited here. In addition, in order to further improve the accuracy of feature dimensionality reduction, a spatial feature transformation layer (SFT layer), a channel attention mechanism, a self-attention mechanism (the attention mechanism in transformer), or one or more of them can be added to the dimensionality reduction network structure. In other embodiments, other network layers can also be added, which will not be listed one by one here.

[0076] For example, an SFT structure and a channel attention structure can be inserted into the dimensionality reduction network structure. For details, please refer to Figure 4 and Figure 5 , Figure 4 , which is a schematic structural diagram of an embodiment of the SFT structure provided by this application, Figure 5 , which is a schematic structural diagram of an embodiment of the channel attention structure provided by this application.

[0077] Among them, Figure 4 the D in represents dot product. The SFT structure multiplies or adds the original features with the environmental features after different convolutional processes, thereby performing spatial transformation on the original features. In addition, each channel of the feature represents a dedicated detector. Therefore, channel attention focuses on what kind of features are meaningful. The channel attention structure can set different weights for different features in different channels of the original features, and then fuse the different channel features of the original features according to different weights, thereby outputting more accurate features.

[0078] The usage modes of the SFT structure and the channel attention structure in the dimensionality reduction network structure include, but are not limited to:

[0079] (1) Flexibly insert it into any position of the dimensionality reduction network structure, such as in the residual block.

[0080] (2) Adopt the SFT structure and the channel attention structure to form a deeper or wider network structure. For example:

[0081] i) Add the SFT structure and the channel attention structure, etc. in the residual block, and vertically stack multiple residual blocks to form a deeper network structure.

[0082] ii) Add the SFT structure and the channel attention structure, etc. in the residual block, and draw on the idea of inception to horizontally combine residual blocks with different convolution kernel sizes at one or several layers, and finally concatenate all the results as the input of the next layer.

[0083] In addition, the image encoding device can also adopt a one-dimensional convolution or fully connected manner that is more friendly to feature information to construct the network. The convolution kernel of the one-dimensional convolution can be set larger, such as 25x1, etc., so as to have a larger receptive field to perceive the relationship between different positions of the features. Among them, the one-dimensional convolution dimensionality reduction network example can adopt a network structure Figure 3 similar to, just need to change the two-dimensional convolution kernel to one-dimensional, which will not be elaborated here.

[0084] When adopting the fully connected manner to construct the network, it is possible to combine a normalization layer, such as at least one of the BatchNorm layer, LayerNorm layer, InstanceNorm layer, and GroupNorm layer, to normalize the features to ensure the stability of the data feature distribution. For specific network structure examples, please refer to Figure 6 , Figure 6 which is the structural schematic diagram of an embodiment of the fully connected dimensionality reduction network provided by this application.

[0085] In Figure 6 the described fully connected dimensionality reduction network, the fully connected dimensionality reduction network includes at least several groups of fully connected dimensionality reduction layer groups, and each group of fully connected dimensionality reduction layer groups includes a fully connected layer, a normalization layer, an activation layer, etc. connected in sequence.

[0086] In other embodiments, it is also possible to adopt the manner of a fully connected layer plus a convolutional layer to construct the dimensionality reduction network, which will not be listed one by one here.

[0087] In addition, the image encoding device can also use the following method to design the dimensionality reduction network to further form a diverse network structure:

[0088] (1) Change the number of channels in the feature transformation process (for example, double it in the middle and then reduce it back to the original number of channels).

[0089] (2) Change the activation function (such as relu / leakyrelu / gdn / gelu).

[0090] (3) Change the upsampling method (in the feature reconstruction stage, such as transposed convolution, pixelshuffle, interpolate, etc.).

[0091] (4) Use batchnorm layer, layernorm layer, dropout layer, etc. to increase the convergence speed of the network, control gradient explosion and prevent overfitting.

[0092] The following continues to introduce traditional feature dimensionality reduction methods:

[0093] Traditional feature transformation methods usually focus on how to make features sparse so as to achieve the purpose of dimensionality reduction. There are many methods available, such as: unsupervised methods: PCA (Principal Component Analysis) dimensionality reduction, SVD (Singular Value Decomposition), Laplacian graph method, LASSO (Least absolute shrinkage and selection operator), manifold learning, etc.; supervised methods: LDA (Linear Discriminant Analysis); and some frequency domain transformation methods: wavelet analysis, Fourier transform, DCT (Discrete Cosine Transform) transformation, etc.

[0094] Step S13: Quantize the transformed features based on the configuration parameters to obtain quantized features.

[0095] In the embodiments of the present application, if the feature transformation in step S12 is implemented using a deep learning-based solution, then in the prior art, generally there is no quantization module or after adding a quantization module, the network design becomes fragmented and it is difficult to perform end-to-end joint optimization. The embodiments of the present application can consider a simpler and more effective quantization method and integrate it into the entire network structure to facilitate joint optimization of network performance. Specifically, linear or non-linear quantization methods can be used to implement:

[0096] (1) In general, the linear quantization method linearly maps the transformed feature data to a certain preset bit depth. The entire quantization process can be integrated into the entire neural network by combining a parameter - learnable strategy, thereby achieving joint optimization and reducing the loss caused by quantization. Suppose the transformed feature data needs to be quantized into n bits (n is an integer), and its data range is from 0 to 2 n - 1, then the following preset linear transformation function needs to be considered:

[0097]

[0098] Among them,

[0099]

[0100] where n is the preset bit depth, x i is the transformed feature, and x′ i is the quantized feature.

[0101] Since in the actual training process, max{x i} and min{x i} cannot be obtained, therefore, at least one of the parameters α and β in the above formula can be set as a learnable parameter, and the remaining parameters are set as configuration parameters with fixed values, so as to train the learnable parameters.

[0102] Similarly, in other embodiments, at least one of min{x i} and max{x i} can also be set as a learnable parameter. In this way, the quantization process of the features can be embedded into the entire neural network and participate in backpropagation.

[0103] (2) The non - linear quantization method first performs a non - linear transformation on the transformed feature data using a certain non - linear function, and then linearly maps it to a certain bit range. Among them, the linear mapping can be implemented in the above - mentioned linear quantization manner, which will not be elaborated here.

[0104] Similarly, assuming that the transformed feature data is quantized into n bits, first use a certain non - linear function f(x) to perform a non - linear transformation on the transformed feature data, and then linearly map it to the corresponding n - bit data range.

[0105] At the same time, in order to make the quantization process more controllable, the quantization loss can also be added to the overall loss value of the neural network for optimization, so as to improve the performance of the quantized features.

[0106] (3) If the feature transformation in step S12 is implemented using a traditional feature transformation method, then in addition to the uniform quantization method similar to the above, non-uniform quantization can also be designed according to some statistical information of the original feature data.

[0107] For example, typically, assume that the transformed feature data needs to be quantized to 8 bits, then the range of its quantized data is 0 - 255. At the same time, set the parameters α and max{x i} as learnable parameters. For example, in pytorch, the parameters to be learned can be set as learnable parameters in the way of nn.Parameter(). For the other two parameters β and min{x i}, for example, both are set to 0, and then calculate the loss value between the transformed feature before quantization and the quantized feature after dequantization, that is, calculate the difference information between the transformed feature and the quantized feature, such as the second norm. The loss value between the transformed feature and the quantized feature is added to the overall loss value as part of the overall loss of the neural network. The specific formula is as follows:

[0108] (aD1 + bD2) + λR

[0109] Where D1 and D2 represent the overall distortion and the distortion of the quantization part respectively, and a and b represent their proportions in the total distortion respectively. In another embodiment, a and b can both be taken as 0.5.

[0110] In other embodiments, the non-linear quantization specifically has the following several examples:

[0111] (1) Use the sigmoid function to normalize the transformed feature to between [0, 1], and then map the normalized feature to a new data range and round it, that is, use the clip function to limit the data range, and the final quantization value can be obtained:

[0112] x′ i =round(sigmoid(x i )×(2 n - 1))

[0113] x′ i =clip(x′ i , 0, 2 n - 1)

[0114] (2) Use the tanh function to map the transformed feature to between [-1, 1], and then linearly map the result to a new data range and round it, that is, use the clip function to limit the data range to obtain the quantization result:

[0115]

[0116] x' i = clip(x' i , 0, 2 n - 1)

[0117] (3) Use the relu function to map the transformed features above 0, set the upper limit of the numerical range to a learnable parameter, then amplify the result to a new data range through linear mapping and round it, that is, use the clip function to limit the data range to obtain the quantization result:

[0118]

[0119] x' i = clip(x' i , 0, 2 n - 1)

[0120] (4) Use the softplus function to map the transformed features above 0, set the upper limit of the numerical range to a learnable parameter, then amplify the result to a new data range through linear mapping and round it and clip to obtain the quantization result:

[0121]

[0122] x' i = clip(x' i , 0, 2 n - 1)

[0123] Among them, in (3) the relu function and (4) the softplus function, maxx represents the maximum value of the learned x i The loss function part can also be designed for non - linear quantization, which will not be elaborated here.

[0124] Furthermore, in the traditional quantization scheme, it is also possible to consider analyzing the input feature data and using the analysis results to allocate different numbers of bits for quantization of different data. For example:

[0125] Method 1: Analyze the distribution law or statistical law of the data, divide the data interval according to the density of the feature data distribution or sort and divide the data according to information such as the mean and variance of the feature data, and then quantize the data in different intervals to different bit ranges. For example, divide the feature distribution interval into 4 parts, and use 2, 4, 6, and 8 bits respectively to quantize each part.

[0126] Method 2: Similarly, first count the data pattern, and then adopt fixed-bit quantization. Only different numbers of values are assigned to the data in different intervals for quantization. For example, for 8-bit quantization, the feature distribution interval A is mapped to 0 - 140, the feature distribution interval B is mapped to 141 - 210, and the feature distribution interval C is mapped to 211 - 255.

[0127] After step S13, in order to further reduce the data volume of the features, entropy coding processing can also be performed on the quantized features, that is, as Figure 2 shown in the overall framework structure, an entropy coding module is added after the quantization module, and an entropy decoding module is added before the inverse quantization module.

[0128] Specifically, the entropy coding module and the entropy decoding module are optional modules of the overall framework structure, which can further compress the size of the bitstream without loss, thus ensuring the coding performance of the feature coding.

[0129] For example, as Figure 7 shown, in the framework based on deep learning, the entropy coding module and the entropy decoding module can adopt a hyperprior network to construct a probability model for the entropy coding process. For specific network structure examples, please refer to Figure 7 , Figure 7 which is a schematic diagram of the feature coding framework including the entropy model provided by this application.

[0130] Among them, the probability model in the embodiments of this application can be constructed by models such as single Gaussian model, mixture Gaussian model, Laplace model, and logistic regression, which will not be listed one by one here.

[0131] For example, in the traditional feature coding framework, the entropy coding module and the entropy decoding module can also adopt relatively mature entropy coding schemes such as CAVLC, CABAC, and Huffman coding.

[0132] Step S14: Perform feature coding on the quantized features to obtain a feature bitstream.

[0133] In the embodiments of this application, the image coding device can directly perform feature coding on the quantized features obtained in step S13, or perform feature coding on the entropy-coded features output by the entropy coding model, which will not be elaborated here.

[0134] In the embodiments of this application, the image coding device acquires the features to be coded of the image to be processed; performs feature transformation on the features to be coded to obtain transformed features, where the feature dimension of the transformed features is lower than the feature dimension of the features to be coded; based on configuration parameters, performs quantization processing on the transformed features to obtain quantized features; performs feature coding on the quantized features to obtain a feature bitstream. The image coding method of this application can further improve the coding rate of the features through a simple and effective quantization method.

[0135] Please continue to refer to Figure 8 , Figure 8 which is a schematic flowchart of an embodiment of the image decoding method provided in this application.

[0136] As Figure 8 shown, the image decoding method of the embodiment of this application includes the following steps:

[0137] Step S21: Perform feature decoding on the feature bitstream to obtain decoded features.

[0138] Step S22: Based on the configuration parameters, perform inverse quantization processing on the decoded features to obtain inverse quantization features.

[0139] Step S23: Perform feature inverse transformation on the inverse quantization features to obtain inverse transformation features, where the feature dimension of the inverse transformation features is higher than that of the inverse quantization features.

[0140] Step S24: Perform feature reconstruction on the inverse transformation features to obtain a reconstructed image.

[0141] In the embodiment of this application, as Figure 2 , Figure 3 , Figure 6 and Figure 7 , it can be understood that the image encoding method and the image decoding method of the embodiment of this application are inverse processes of each other. Therefore, the technical solutions of the image encoding method can be adaptively applied to the image decoding method of the embodiment of this application, and the specific technical solutions will not be elaborated here.

[0142] This application proposes an image encoding method and an image decoding method. The image encoding method and the image decoding method can be implemented using traditional or deep learning frameworks, and have universality, rather than being targeted at a specific visual task; at the same time, the feature encoding scheme based on deep learning can achieve end-to-end joint optimization; this application also proposes two specific feature encoding frameworks according to whether an entropy model is included, that is Figure 2 and Figure 7 the feature encoding frameworks shown.

[0143] This application further proposes to use the following methods for dimensionality reduction during the feature transformation process: (1) Use two-dimensional convolution based on the attention mechanism or spatial feature transformation to improve the network performance; (2) Use structures such as one-dimensional convolution and fully connected layers to capture the semantic information of the feature data (3) Consider deeper or wider network structures to improve the network performance.

[0144] This application proposes new simple and effective quantization methods, including linear quantization, non-linear quantization, non-uniform quantization, etc., and designs a new loss function to improve the coding rate of features.

[0145] The above embodiments are merely one common case of the present application and do not impose any restrictions on the technical scope of the present application. Therefore, any minor modifications, equivalent changes, or decorations made to the above content based on the essence of the present application's solution still fall within the scope of the technical solution of the present application.

[0146] Please continue to refer to Figure 9 , Figure 9 which is a schematic structural diagram of an embodiment of an image encoding device provided by the present application. The image encoding device 500 in the embodiment of the present application includes a processor 51, a memory 52, an input / output device 53, and a bus 54.

[0147] The processor 51, the memory 52, and the input / output device 53 are respectively connected to the bus 54. Program data is stored in the memory 52, and the processor 51 is configured to execute the program data to implement the image encoding method described in the above embodiments.

[0148] In the embodiment of the present application, the processor 51 may also be referred to as a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field-programmable gate array (FPGA, FieldProgrammable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 51 may also be any conventional processor, etc.

[0149] Please continue to refer to Figure 10 , Figure 10 which is a schematic structural diagram of an embodiment of an image decoding device provided by the present application. The image decoding device 600 in the embodiment of the present application includes a processor 61, a memory 62, an input / output device 63, and a bus 64.

[0150] The processor 61, the memory 62, and the input / output device 63 are respectively connected to the bus 64. Program data is stored in the memory 62, and the processor 61 is configured to execute the program data to implement the image decoding method described in the above embodiments.

[0151] The present application also provides a computer storage medium. Please continue to refer to Figure 11 , Figure 11It is a schematic structural diagram of an embodiment of a computer storage medium provided by this application. Program data 71 is stored in the computer storage medium 700. When the program data 71 is executed by a processor, it is used to implement the image encoding method and / or the image decoding method of the above embodiments.

[0152] When the embodiments of this application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0153] The above are only the embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, are similarly included in the patent protection scope of this application.

Claims

1. An image encoding method, characterized in that, The described image encoding method includes: Obtaining the feature to be encoded of the image to be processed; Performing feature transformation on the feature to be encoded to obtain transformed features, wherein the feature dimension of the transformed features is lower than that of the feature to be encoded; Based on configuration parameters, performing quantization processing on the transformed features to obtain quantized features; Performing feature encoding on the quantized features to obtain a feature bitstream; The step of based on configuration parameters, performing quantization processing on the transformed features to obtain quantized features includes: Obtaining a preset linear transformation function, and based on the configuration parameters, assigning values to the non-learning parameters in the preset linear transformation function; Using the preset linear transformation function after value assignment and a preset bit depth to map the transformed features to obtain the quantized features.

2. The image encoding method according to claim 1, wherein The step of performing feature transformation on the feature to be encoded to obtain transformed features includes: Inputting the feature to be encoded into a dimensionality reduction network, and performing downsampling on the feature to be encoded through the convolutional layer and / or fully connected layer of the dimensionality reduction network to obtain the transformed features.

3. The image encoding method according to claim 2, wherein The convolutional layer of the dimensionality reduction network is a one-dimensional convolutional layer or a two-dimensional convolutional layer.

4. The image encoding method according to claim 2 or 3, wherein The dimensionality reduction network further includes one or more of a spatial feature transformation sub-network, a channel attention mechanism sub-network, and a self-attention mechanism sub-network.

5. The image encoding method according to claim 2, wherein The step of inputting the feature to be encoded into the dimensionality reduction network includes: Sequentially inputting the feature to be encoded into several dimensionality reduction sub-networks of the dimensionality reduction network, and each dimensionality reduction sub-network includes a fully connected layer, a normalization layer, and an activation layer connected in series in sequence.

6. The image encoding method according to claim 1, wherein The step of performing feature transformation on the feature to be encoded to obtain transformed features includes: Based on at least one of an unsupervised dimensionality reduction algorithm and a supervised dimensionality reduction algorithm, performing feature sparsification processing on the feature to be encoded to obtain the transformed features.

7. The image encoding method according to claim 1, wherein Before using the preset linear transformation function after value assignment and a preset bit depth to map the transformed features to obtain the quantized features, the image encoding method further includes: Performing non-linear transformation on the transformed features using a preset non-linear function to obtain the transformed features after non-linear transformation.

8. The image encoding method according to claim 1 or 7, characterized in that, After the step of based on configuration parameters, performing quantization processing on the transformed features to obtain quantized features, The image encoding method further includes: Performing inverse quantization processing on the quantized features to obtain inverse quantized features; Based on the difference information between the transformed features and the inverse quantized features, obtaining a quantization loss value; Using the quantization loss value to train the learning parameters in the preset linear transformation function.

9. The image encoding method according to claim 1, wherein After quantizing the transformed feature based on the configuration parameter to obtain a quantized feature, the image encoding method further includes: extracting context feature information of the quantized feature by using an entropy coding model; predicting the quantized feature based on the context feature information of the quantized feature to obtain an entropy coding feature of the quantized feature; performing feature coding based on the entropy coding feature to obtain the feature bitstream.

10. The image encoding method according to claim 9, wherein the entropy coding model includes a probability model constructed by using a hyperprior network, wherein the probability model is a single Gaussian model, a mixture Gaussian model, a Laplace model, a logistic regression model, or a combined model of one or more of them.

11. An image decoding method, characterized in that, The image decoding method includes: performing feature decoding on the feature bitstream to obtain a decoded feature; performing inverse quantization processing on the decoded feature based on the configuration parameter to obtain an inverse quantized feature; performing feature inverse transformation on the inverse quantized feature to obtain an inverse transformed feature, wherein the feature dimension of the inverse transformed feature is higher than the feature dimension of the inverse quantized feature; performing feature reconstruction on the inverse transformed feature to obtain a reconstructed image; The step of quantizing the transformed feature based on the configuration parameter to obtain a quantized feature includes: obtaining a preset linear transformation function, and assigning values to non-learning parameters in the preset linear transformation function based on the configuration parameter; mapping the transformed feature by using the preset linear transformation function with assigned values and a preset bit depth to obtain the quantized feature.

12. An image encoding device, characterized in that, The image encoding device includes a memory and a processor coupled to the memory; wherein, the memory is used for storing program data, and the processor is used for executing the program data to implement the image encoding method according to any one of claims 1 to 10.

13. An image decoding device, characterized in that, The image decoding device includes a memory and a processor coupled to the memory; wherein, the memory is used for storing program data, and the processor is used for executing the program data to implement the image decoding method according to claim 11.

14. A computer storage medium, characterized in that, The computer storage medium is used for storing program data, and when the program data is executed by a computer, it is used to implement the image encoding method according to any one of claims 1 to 10 and / or the image decoding method according to claim 11.

Citation Information

Patent Citations

  • Method and apparatus for processing image for machine vision

    WO2022075754A1