Image coding device, probability model generating device, and image decoding device

By employing a pyramidal resize module and inception coder network with multi-scale dilated convolution, the method enhances feature extraction and latent representation, addressing the bitrate-distortion trade-off in image compression.

JP7679903B2Active Publication Date: 2025-05-20FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024060706
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-05-22
Filing Date
2024-04-04
Publication Date
2025-05-20
Estimated Expiration
2040-05-11

AI Technical Summary

Technical Problem

Existing image compression methods using deep neural networks face challenges in achieving a balance between bitrate and distortion, as they struggle to accurately extract image features and obtain competitive latent representations.

Method used

The use of a pyramidal resize module and an inception coder network for feature extraction, combined with a multi-scale dilated convolution unit and probabilistic model generation to enhance feature extraction and latent representation accuracy.

Benefits of technology

This approach allows for more accurate extraction of image features and reconstruction of images with reduced distortion, improving the balance between bitrate and quality in image compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679903000001
    Figure 0007679903000001
  • Figure 0007679903000002
    Figure 0007679903000002
  • Figure 0007679903000003
    Figure 0007679903000003
Patent Text Reader

Abstract

To provide an image coding apparatus, a probability model generating apparatus, and an image decoding apparatus according to the present invention.SOLUTION: An image coding apparatus includes: a first feature extraction unit configured to perform feature extraction on an input image and to obtain feature maps of N channels; a second feature extraction unit configured to perform feature extraction on input image whose size is adjusted at K times obtain feature maps of N channels for each of the adjusted image; and a first concatenating unit configured to merge the feature map of N channels from the first feature extraction unit and the feature maps of K×N channels from the second feature extraction unit and to output a merged map. Hence, features on image may be accurately extracted and more competitive latent representations may be obtained.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the technical fields of image compression and deep learning. [Background technology]

[0002] In recent years, deep learning has taken a leading position in the field of computer vision. In image recognition and super-resolution restoration, deep learning has become a key technology in image research, but its capabilities are not limited to these tasks. Deep learning technology has also been applied to the field of image compression, which has become a hot research topic.

[0003] Currently, image compression based on deep neural networks aims to generate high-quality images using as few code streams as possible, which leads to a rate-distortion trade-off problem. To achieve a good balance between bitrate and distortion, two research approaches have been developed: (1) finding the most approximate entropy model for the latent representation to optimize the length of the bitstream (to reduce the bitrate); and (2) obtaining a more efficient latent representation to accurately reconstruct the image (to reduce distortion). Summary of the Invention [Problem to be solved by the invention]

[0004] The embodiments of the present invention provide an image coding method and apparatus, a probability model generating method and apparatus, an image decoding method and apparatus, and an image compression system, which uses a pyramidal resize module and an inception coder network to accurately extract image features and obtain more competitive latent representations. [Means for solving the problem]

[0005] According to a first aspect of an embodiment of the present invention there is provided an image coding apparatus comprising: a first feature extraction unit that performs feature extraction on the input image to obtain an N-channel feature map; a second feature extraction unit that performs feature extraction on the input image whose size has been adjusted K times to obtain feature maps of N channels, respectively; and The first feature extraction unit includes a first combination unit for combining and outputting the N-channel feature map from the first feature extraction unit and the K×N-channel feature map from the second feature extraction unit.

[0006] According to a second aspect of an embodiment of the present invention there is provided a probabilistic model generation apparatus, the apparatus comprising: a multi-scale dilated convolution unit that performs feature extraction on the output of the hyperdecoder to obtain multi-scale auxiliary information; a context model processing unit that takes as input the latent representation of the input image from the quantizer and obtains content-based predictions; and and an entropy model processing unit that processes an output of the context model processing unit and an output of the multi-scale dilated convolution unit to obtain a probability model of prediction.

[0007] According to a third aspect of an embodiment of the present invention, there is provided an image decoding apparatus, the image decoding apparatus comprising: a multi-scale dilated convolution unit that performs feature extraction on the output of the hyperdecoder to obtain multi-scale auxiliary information; a combiner for combining the latent representation of the input image from an arithmetic decoder and the multi-scale auxiliary information from the multi-scale dilated convolution unit; and A decoder is included for decoding an output from the combiner to obtain a reconstructed image of the input image.

[0008] According to a fourth aspect of an embodiment of the present invention, there is provided an image coding method, the method comprising: Use multiple Inception units to perform feature extraction on the input image to obtain N-channel feature maps; Using multiple convolutional layers, perform feature extraction on each resized input image to obtain N channel feature maps; and The method includes combining and outputting the N-channel feature maps from the Inception unit and the corresponding N-channel feature maps from the multiple convolutional layers.

[0009] According to a fifth aspect of an embodiment of the present invention, there is provided a method for generating a probabilistic model, the method comprising: Using a multi-scale dilated convolution unit, perform feature extraction on the output of the hyper-decoder to obtain multi-scale auxiliary information; Using a context model, taking as input the latent representation of the input image from the quantizer, to obtain a content-based prediction; and The method includes processing an output of the context model and an output of the multi-scale dilated convolution unit using an entropy model to obtain a probability model of prediction.

[0010] According to a sixth aspect of an embodiment of the present invention, there is provided an image decoding method, the method comprising the steps of: A multi-scale dilated convolution unit is used to perform feature extraction on the output of the hyperdecoder to obtain multi-scale auxiliary information; using a combiner to combine the latent representation of the input image from the arithmetic decoder and the multi-scale auxiliary information from the multi-scale dilated convolution unit; and Using a decoder, decoding the output from the combiner to obtain a reconstructed image of the input image.

[0011] According to another aspect of an embodiment of the present invention, there is provided a computer readable program which, when executed in an image processing device, causes the image processing device to perform a method according to any one of the fourth, fifth and sixth aspects above.

[0012] According to another aspect of an embodiment of the present invention, a storage medium is provided having a computer readable program stored thereon, the computer readable program causing an image processing device to execute a method according to any one of the fourth, fifth and sixth aspects described above.

[0013] The beneficial effects of the embodiments of the present invention are as follows: the image coding method and apparatus in the embodiments of the present invention can accurately extract image features and obtain a more competitive latent representation; the image decoding method and apparatus in the embodiments of the present invention can fuse multi-scale auxiliary information to reconstruct images more accurately. [Brief description of the drawings]

[0014] [Figure 1] FIG. 1 illustrates an image compression system according to a first embodiment. [Diagram 2] FIG. 11 is a diagram illustrating an image coding device according to a second embodiment. [Diagram 3] FIG. 3 is a diagram showing a network configuration in an embodiment of an Inception unit of a first feature extraction unit of the image coding device shown in FIG. 2. [Figure 4] FIG. 3 is a diagram showing a network configuration of an embodiment of the second feature extraction unit of the image coding device shown in FIG. 2. [Diagram 5] FIG. 3 is a diagram showing a network configuration in an embodiment of the image coding device shown in FIG. 2. [Figure 6] FIG. 11 is a diagram showing an image decoding device in a third embodiment. [Figure 7] FIG. 13 is a diagram showing a network configuration in one embodiment of a multi-scale dilated convolution unit. [Figure 8] FIG. 13 is a diagram illustrating a probabilistic model generating device according to a fourth embodiment. [Figure 9] FIG. 13 is a diagram showing an image coding method in a fifth embodiment. [Figure 10] FIG. 13 is a diagram showing an image decoding method in Example 6. [Figure 11] FIG. 23 is a diagram illustrating a method for generating a probabilistic model in a seventh embodiment. [Figure 12] FIG. 13 is a diagram illustrating an image processing device according to an eighth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. EXAMPLES

[0016] An embodiment of the present invention provides an image compression system, and FIG. 1 is a diagram illustrating an image compression system in an embodiment of the present invention. As shown in FIG. 1, the image compression system 100 in an embodiment of the present invention includes an image coding device 101, a probability model generating device 102, and an image decoding device 103. The image coding device 101 can perform downsampling on an input image and convert the input image into a latent representation. The probability model generating device 102 can make a prediction on the probability distribution of the latent representation and obtain a probability model of the latent representation. The image decoding device 103 can perform upsampling on the latent representation obtained by decoding based on the probability model, and map the latent representation to an input image.

[0017] In the embodiment of the present invention, as shown in Fig. 1, an image coding device 101, which may be referred to as a coder 101, can perform compression coding on an input image, i.e., map the input image into a latent code space. The network configuration and implementation manner of the coder 101 will be described later.

[0018] In an embodiment of the present invention, as shown in FIG. 1, the image compression system 100 may further include a quantizer (Q) 104, an arithmetic coder (AE) 105, and an arithmetic decoder (AD) 106. The quantizer 104 quantifies the output from the coder 101, so that the latent representation from the coder 101 can be quantized to generate a discrete value vector. The arithmetic coder 105 can code the output from the quantizer 104 based on the probability model (i.e., the prediction probability distribution) generated by the above-mentioned probability model generating device 102, i.e., compress the above-mentioned discrete value vector into a bit stream. The arithmetic decoder 106 is the inverse of the arithmetic coder 105, which can decode the received bit stream based on the probability model generated by the above-mentioned probability model generating device 102, i.e., decompress the above-mentioned bit stream into a quantized latent representation, and provide it to the image decoding device 103.

[0019] In an embodiment of the present invention, as shown in Fig. 1, the image compression system 100 may further include a hypercoder 107, a quantizer (Q) 108, an arithmetic coder (AE) 109, an arithmetic decoder (AD) 110, and a hyperdecoder 111. The hypercoder 107 can further code the output from the coder 101. The processing of the quantizer 108, the arithmetic coder 109, and the arithmetic decoder 110 is similar to that of the quantizer 104, the arithmetic coder 105, and the arithmetic decoder 106, with the difference being that the arithmetic coder 109 and the arithmetic decoder 110 do not use the above-mentioned probability model when performing compression and decompression, and other specific processing processes are omitted here. The hyperdecoder 111 can further decode the output from the arithmetic decoder 109. Regarding the configuration and implementation method of the network of the hypercoder 107, the quantizer (Q) 108, the arithmetic coder (AE) 109, the arithmetic decoder (AD) 110, and the hyperdecoder 111, reference may be made to the prior art, and detailed description thereof will be omitted here.

[0020] In an embodiment of the present invention, as shown in FIG. 1, the image decoding device 103 includes a multi-scale dilated convolution unit (pyramid atrous) 1031, a combiner 1032, and a decoder 1033. The multi-scale dilated convolution unit 1031 can generate multi-scale auxiliary information. The combiner 1032 can combine the above-mentioned multi-scale auxiliary information and the output from the arithmetic decoder 106. The decoder 1033 can restore the input image by decoding the output from the combiner 1032, that is, the discrete elements of the latent representation can be transformed back into the data space to obtain a reconstructed image. Note that the network configuration and implementation manner of the multi-scale dilated convolution unit 1031 will be described later.

[0021] In an embodiment of the present invention, as shown in FIG. 1, the probability model generating device 102 includes a context model and an entropy model, and the context model can obtain a content-based prediction based on the output (latent representation) of the quantizer 104. The entropy model can be responsible for learning a probability model of the latent representation. In an embodiment of the present invention, the entropy model can generate the probability model based on multi-scale auxiliary information from the multi-scale dilated convolution unit 1031 and the output from the context model. The multi-scale auxiliary information can modify the context-based prediction. In one embodiment, the entropy model can generate the mu part (mean parameter 'mean') of the probability model based on the mu part of the context model and the above-mentioned multi-scale auxiliary information, and can also generate the sigma part (ratio parameter 'scale') of the probability model based on the sigma part of the context model and the above-mentioned multi-scale auxiliary information, but the embodiment of the present invention is not limited thereto. The entropy model can also directly generate the mean parameter and ratio parameter of the above-mentioned probability model based on the output of the context model and the multi-scale auxiliary information without distinguishing between the mu part and the sigma part.

[0022] The configurations of the image coding device 101, the image decoding device 103, and the probability model generating device 102 shown in Fig. 1 are merely examples, and the embodiments of the present invention are not limited thereto. For example, the hypercoder 107 and the hyperdecoder 111 may be part of the probability model generating device 102 or part of the image decoding device 103. Also, for example, the multi-scale dilated convolution unit 1032 may be part of the image decoding device 103 or part of the probability model generating device 102.

[0023] In the embodiment of the present invention, the distortion between the original image and the reconstructed image is directly related to the quality of the extracted features, and generally speaking, the more features are extracted, the smaller the distortion. In order to obtain as many latent representations containing features as possible, the embodiment of the present invention uses the above-mentioned coder 101 to configure a multi-scale network, which can effectively extract the features of the input image.

[0024] FIG. 2 is a diagram illustrating an image coding device 101 according to an embodiment of the present invention. As shown in FIG. 2, the image coding device 101 according to an embodiment of the present invention includes a first feature extraction unit 201, a second feature extraction unit 202, and a first combination unit 203. The first feature extraction unit 201, the second feature extraction unit 202, and the first combination unit 203 constitute the coder 101 shown in FIG. 1. In the embodiment of the present invention, the first feature extraction unit 201 can perform feature extraction on an input image to obtain feature maps of N channels. The second feature extraction unit 202 can perform feature extraction on an input image whose size has been adjusted K times to obtain feature maps of N channels, respectively. The first combination unit 203 can combine and output the feature maps of N channels from the first feature extraction unit 201 and the feature maps of K×N channels from the second feature extraction unit 202.

[0025] Usually, when a feature map is extracted from an image using a convolutional neural network, a relatively deep layer represents global information and high-level information, while a relatively shallow layer represents local information and detailed information, such as edges. Therefore, in the embodiment of the present invention, the above-mentioned first feature extraction unit 201 is used to obtain global information and high-level information from the original input image, and the above-mentioned second feature extraction unit 202 is used to obtain detailed features from the resized input image. The first feature extraction unit 201 can be a multi-layer network, for example a four-layer network, and the second feature extraction unit 202 can be a convolutional layer network, which will be described below respectively.

[0026] In an embodiment of the present invention, the first feature extraction unit 201 may include multiple inception units, each of which is sequentially connected to perform feature extraction on the above-mentioned input image or the feature map from the previous inception unit, so as to obtain the above-mentioned global information and high-level information of the input image. The working principle of the inception unit can be referred to the prior art, for example, "Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1-9, 2015", and the description thereof is omitted here.

[0027] FIG. 3 is a diagram showing a network configuration according to an embodiment of an Inception unit in an embodiment of the present invention. As shown in FIG. 3, in this embodiment, the Inception unit includes three convolution layers (also called third feature extraction units) 301, one pooling layer (also called pooling units) 302, one combination layer (also called second combination unit) 303, and one convolution layer (also called fourth feature extraction unit) 304. The three convolution layers 301 can perform feature extraction on the above-mentioned input image or feature map from the previous Inception unit using different convolution kernels (3×3, 5×5, 7×7) and the same number of channels (N), respectively, to obtain feature maps of N channels. The pooling layer 302 can also perform a dimensionality reduction process on the above-mentioned input image or feature map from the previous Inception unit to obtain feature maps of N channels. The combination layer 303 can combine the feature maps of N channels from the three convolution layers 301 and the feature maps of N channels from the pooling layer 302 to obtain 4N channel feature maps. The convolution layer 304 can further perform dimension reduction processing on the feature maps from the combination layer 303 to obtain feature maps of N channels. In the embodiment of the present invention, the pooling layer 302 adopts a maximum pooling method as an example, but the embodiment of the present invention is not limited thereto. In addition, the working principle of the pooling layer can refer to the prior art, and the description thereof will be omitted here.

[0028] In the embodiment of the present invention, the Inception unit can use multi-scale features to help reconstruct the image. In addition, the Inception unit in the embodiment of the present invention can obtain more features from the original input image by using multi-scale features with different kernels. In the embodiment of the present invention, convolutional layers 301 with different kernels use the same number of channels to combine their results, and one convolutional layer 304 with a kernel of 1×1 is used to determine which one is more important, so that the output of the current layer can be obtained.

[0029] The network configuration of the Inception Unit shown in FIG. 3 is merely an example, and the embodiments of the present invention are not limited thereto.

[0030] In an embodiment of the present invention, the second feature extraction unit 202 may include a size adjustment unit and a feature extraction unit (also referred to as a fifth feature extraction unit), in which the size adjustment unit performs size adjustment on the input image, and the fifth feature extraction unit performs feature extraction on the size-adjusted input image to obtain an N-channel feature map.

[0031] In an embodiment of the present invention, the resizing unit and the fifth feature extraction unit may be one or more sets, i.e., one resizing unit and one fifth feature extraction unit are one set of feature extraction modules, and the second feature extraction unit 202 may include one or more sets of feature extraction modules, where the resizing units of different sets can use different ratios to resize the input image, and the fifth feature extraction units of different sets can use different convolution kernels to extract features on the resized input image. The second feature extraction unit 202 can form a convolution layer network.

[0032] 4 is a diagram showing a network configuration of an embodiment of the second feature extraction unit 202. As shown in FIG. 4, the second feature extraction unit 202 includes three size adjustment units 401 and three convolution layers 402, i.e., three sets of feature extraction modules, in which the three size adjustment units 401, 401', 401'' respectively perform size adjustment of 1 / 2, 1 / 4, and 1 / 8 on the input image, thereby adjusting the input image three times, i.e., K=3, in which H is the height of the input image and W is the width of the input image, and the three convolution layers 402, 402', 402'' are used as the fifth feature extraction unit to perform feature extraction on the size-adjusted input image using different kernels (9×9, 5×5, 3×3), and obtain N-channel feature maps to output to the first combination unit 203. In the embodiment of the present invention, the three resizing units 401, 401', 401'' perform different resizing ratios on the input image, so the number of dimensional reductions performed by the three convolutional layers 402, 402', 402'' are also different. For example, for 1 / 2 of the input image, the convolutional layer 402 performs 8-dimensional reduction processing, for 1 / 4 of the input image, the convolutional layer 402' performs 4-dimensional reduction processing, and for 1 / 8 of the input image, the convolutional layer 402'' performs 2-dimensional reduction processing, so that it is possible to ensure that the dimension of the feature map input from the second feature extraction unit 202 to the first combination unit 203 is the same as the dimension of the feature map input from the first feature extraction unit 201 to the first combination unit 203.

[0033] In an embodiment of the present invention, as shown in FIG. 2, the image coding device 101 may further include a weighting unit 204 and a sixth feature extraction unit 205. The weighting unit 204 can provide a weight to the feature map of each channel from the first combining unit 203. The sixth feature extraction unit 205 can perform a dimension reduction process on the feature map from the weighting unit 204 to obtain and output feature maps of M channels. In an embodiment of the present invention, the weighting unit 204 can provide a weight to the feature map of each channel to retain useful features and suppress unusable features, and the sixth feature extraction unit can perform a dimension reduction process on the input feature map to reduce the amount of calculation.

[0034] In the embodiment of the present invention, there is no limitation on the network configuration of the weighting unit 204, and the structure related to the weighting layer in the prior art can function as the weighting unit 204 in the embodiment of the present invention. In the embodiment of the present invention, the sixth feature extraction unit 205 can be realized by one convolution layer with a kernel of 1×1, but the embodiment of the present invention is not limited thereto.

[0035] FIG. 5 is a diagram showing a network configuration according to an embodiment of the image coding device 101 in the embodiment of the present invention. As shown in FIG. 5, the first feature extraction unit 201 of the image coding device 101 is realized by four inception units, forming a four-layer network architecture, and can extract global information and high-level information from the original input image. The second feature extraction unit 202 of the image coding device 101 has three sets of feature extraction modules, each of which can further perform feature extraction after performing size adjustment on the original input image, and the specific network configuration is described in FIG. 4, so the description thereof is omitted here. The first combination unit 203 of the image coding device 101 may be realized by a concat function. The weight unit 204 of the image coding device 101 can be realized by a weight layer. The sixth feature extraction unit 205 of the image coding device 101 is realized by a 1×1 convolution layer, in this example, N=192, M=128.

[0036] FIG. 6 is a diagram illustrating an image decoding device 103 according to an embodiment of the present invention. As shown in FIG. 6, the image decoding device 103 according to the embodiment of the present invention includes a multi-scale dilated convolution unit 601, a combiner 602, and a decoder 603. The multi-scale dilated convolution unit 601 can perform feature extraction on the output of the hyper-decoder 111 to obtain multi-scale auxiliary information. The combiner 602 can perform combination on the latent representation of the input image from the arithmetic decoder 106 and the multi-scale auxiliary information from the multi-scale dilated convolution unit 601. The decoder 603 can decode the output from the combiner 602 to obtain a reconstructed image of the input image. Note that the network configuration and the implementation manner of the hyper-decoder 111 and the arithmetic decoder 106 are the same as those of the hyper-decoder 111 and the arithmetic decoder 106 shown in FIG. 1, and can refer to the prior art, and the detailed description thereof will be omitted here.

[0037] In an embodiment of the present invention, the multi-scale dilated convolution unit 602 may include multiple feature extraction units, and the feature extraction units may be realized by dilated convolution layers, for example, by three dilated convolution layers, which use different dilation ratios (i.e., dilated convolution kernels with different dilation ratios) and the same number of channels to perform feature extraction on the output of the hyper-decoder to obtain the above-mentioned multi-scale auxiliary information.

[0038] Figure 7 is a diagram showing a network configuration of an embodiment of the multi-scale dilated convolution unit 601. As shown in Figure 7, the multi-scale dilated convolution unit 601 is realized by three 3x3 dilated convolution layers with different dilation ratios, the dilation ratios are 1, 2, and 3, respectively, and the channel numbers of the three convolution layers are all N, so that multi-scale auxiliary information can be obtained. Note that the implementation manner of the dilated convolution layer can be referred to the prior art, and the description thereof is omitted here.

[0039] In an embodiment of the present invention, a multi-scale dilated convolution unit 601 is added after the hyperdecoder 111 to obtain multi-scale auxiliary information from the hypernetwork (hypercoder and hyperdecoder), and the combiner 602 combines this information with the quantized latent representation (the output of the arithmetic decoder 106) to obtain more features and feed them back to the decoder network (decoder 603).

[0040] FIG. 8 is a diagram illustrating a probability model generating device 102 in an embodiment of the present invention. As shown in FIG. 8, the probability model generating device 102 in an embodiment of the present invention includes a multi-scale dilated convolution unit 801, a context model processing unit 802, and an entropy model processing unit 803. The multi-scale dilated convolution unit 801 can perform feature extraction on the output of the hyper-decoder 111 to obtain multi-scale auxiliary information. The context model processing unit 802 can receive the latent representation of the input image from the quantizer 104 as input and obtain a content-based prediction. The entropy model processing unit 803 can process the output of the context model processing unit 802 and the output of the multi-scale dilated convolution unit 801 to obtain a probability model of prediction, and provide it to the arithmetic coder 105 and the arithmetic decoder 106. Note that the network configuration and implementation manner of the arithmetic coder 105 and the arithmetic decoder 106 can refer to the prior art, and the description thereof will be omitted here.

[0041] In the embodiment of the present invention, the network configuration of the multi-scale dilated convolution unit 801 is not limited. Although FIG. 7 shows an example, the embodiment of the present invention is not limited thereto.

[0042] The image compression system according to the embodiment of the present invention can accurately extract image features and obtain more competitive latent representations. EXAMPLES

[0043] An embodiment of the present invention provides an image coding device, and Fig. 2 is a diagram showing an image coding device 101 in an embodiment of the present invention. Fig. 3 is a diagram showing a network configuration according to an embodiment of an inception unit of a first feature extraction unit 201 of an image coding device in an embodiment of the present invention, and Fig. 4 is a diagram showing a network configuration according to an embodiment of a second feature extraction unit 202 of an image coding device in an embodiment of the present invention. Fig. 5 is a diagram showing a network configuration according to an embodiment of an image coding device in an embodiment of the present invention. Since the image coding device has been described in detail in the first embodiment, the contents thereof are incorporated herein, and the description thereof is omitted here.

[0044] The image coding apparatus according to the embodiment of the present invention can accurately extract image features and obtain more competitive latent representations. EXAMPLES

[0045] An embodiment of the present invention provides an image decoding device, and Fig. 6 is a diagram showing an image decoding device 103 in an embodiment of the present invention. Fig. 7 is a diagram showing a network configuration of an embodiment of a multi-scale dilated convolution unit 601 of the image decoding device 103. In the first embodiment, the image decoding device is described in detail, and the contents thereof are incorporated herein, and the description thereof is omitted here.

[0046] The image decoding device according to the embodiment of the present invention can obtain more auxiliary information and realize more accurate image reconstruction. EXAMPLES

[0047] An embodiment of the present invention provides a probability model generating device, and Fig. 8 is a diagram showing a probability model generating device according to an embodiment of the present invention. Fig. 7 is a diagram showing a network configuration of an embodiment of a multi-scale dilated convolution unit 801 of the probability model generating device. In the first embodiment, the probability model generating device is described in detail, and the contents thereof are incorporated herein, and the description thereof is omitted here.

[0048] The probabilistic model generating apparatus in the embodiment of the present invention can better predict the probability distribution of latent representations after adding multi-scale auxiliary information. EXAMPLES

[0049] An embodiment of the present invention provides an image coding method, the principle of which the method solves the problem is similar to that of the apparatus of embodiment 2 and has been described in embodiment 1. Therefore, for specific implementation, reference may be made to the implementation of the apparatus of embodiment 1 and embodiment 2, and duplicated descriptions of the same contents will be omitted.

[0050] FIG. 9 is a diagram illustrating an image coding method according to an embodiment of the present invention. As shown in FIG. 9, the image coding method includes the following operations.

[0051] 901: Using multiple inception units, perform feature extraction on an input image to obtain feature maps of N channels; 902: Using multiple convolution layers, perform feature extraction on the resized input images, respectively, to obtain feature maps of N channels; and 903: Concatenate and output the N-channel feature maps from the Inception unit and the corresponding N-channel feature maps from the multiple convolutional layers.

[0052] In this embodiment of the present invention, the implementation of each operation in FIG. 9 can refer to the implementation of each unit in FIG. 2 in the first embodiment, and the description thereof will be omitted here.

[0053] In operation 901 in this embodiment of the present invention, the above-mentioned multiple Inception units can be sequentially combined to perform feature extraction on the input image or feature maps from previous Inception units to obtain global and high-level information of the input image.

[0054] In one embodiment, each Inception Unit includes three convolutional layers and one pooling layer, in which the three convolutional layers perform feature extraction on the input image or the feature map from the previous Inception Unit using different convolution kernels and the same number of channels to obtain a feature map with N channels, respectively, and the pooling layer performs dimensionality reduction on the input image or the feature map from the previous Inception Unit to obtain a feature map with N channels.

[0055] In some embodiments, each Inception unit may further include one combination layer and one convolution layer, in which the combination layer combines the corresponding N-channel feature maps from the above three convolution layers with the N-channel feature map from the pooling layer to obtain 4N-channel feature maps, and the convolution layer can perform a dimensionality reduction process on the feature maps from the combination layer to obtain N-channel feature maps.

[0056] In operation 902 in an embodiment of the present invention, the input image may first be resized in different proportions, and then feature extraction may be performed on each resized input image by multiple convolution layers, where each convolution layer corresponds to one resized input image, thereby obtaining feature maps of N channels respectively.

[0057] In some embodiments, the above-mentioned multiple convolutional layers may use different convolution kernels and the same number of channels, thus ensuring that each convolutional layer performs the same number of dimensionality reductions on the resized input image, making combination convenient.

[0058] In operation 903 in an embodiment of the present invention, a concatenation layer or a concatenation function (concat) may be used to combine the feature maps extracted by each of the above feature extraction units.

[0059] In the embodiment of the present invention, a weight is further applied to the combined feature map of each channel, and then a dimensionality reduction process is performed on the feature map after the weighting, so that the feature map outputs of M channels can be obtained, thereby reducing the number of pixels waiting to be processed and saving the amount of calculation.

[0060] The image coding method according to the embodiment of the present invention can accurately extract image features and obtain more competitive latent representations. EXAMPLES

[0061] An embodiment of the present invention provides an image decoding method, and the principle by which the method solves the problem is the same as that of the apparatus of embodiment 3, and has been described in embodiment 1. Therefore, for specific implementation, reference may be made to the implementation of the apparatuses of embodiments 1 and 3, and duplicated descriptions of the same contents will be omitted.

[0062] 10 is a diagram illustrating an image decoding method in an embodiment of the present invention. As shown in FIG. 10, the image decoding method includes the following operations:

[0063] 1001: Using a multi-scale dilated convolution unit, perform feature extraction on the output of the hyperdecoder to obtain multi-scale auxiliary information; 1002: Using a combiner, combine the latent representation of the input image from the arithmetic decoder and the multi-scale auxiliary information from the multi-scale dilated convolution unit; and 1003: Using a decoder, perform decoding on the output from the combiner to obtain a reconstructed image of the input image.

[0064] In an embodiment of the present invention, the above-mentioned multi-scale dilated convolution unit may include three dilated convolution layers, which use different dilation ratios and the same number of channels to perform feature extraction on the output of the hyper-decoder to obtain the multi-scale auxiliary information.

[0065] In an embodiment of the present invention, the above-mentioned combiner may be a combination layer in a convolutional neural network, and other implementation methods are omitted.

[0066] The image decoding method in the embodiment of the present invention can obtain more auxiliary information and achieve more accurate image reconstruction. EXAMPLES

[0067] An embodiment of the present invention provides a probabilistic model generation method, and the principle by which the method solves the problem is similar to that of the apparatus of embodiment 4 and has been described in embodiment 1. Therefore, for specific implementation, reference can be made to the implementation of the apparatus of embodiment 1 and embodiment 4, and duplicated explanations of the same content will be omitted.

[0068] FIG. 11 is a diagram illustrating a method for generating a probability model in an embodiment of the present invention. As shown in FIG. 11, the method for generating a probability model includes the following operations:

[0069] 1101: Using a multi-scale dilated convolution unit, perform feature extraction on the output of the hyperdecoder to obtain multi-scale auxiliary information; 1102: Using a context model, taking as input the latent representation of an input image from a coder, and obtaining content-based predictions; and 1103: Process the output of the context model and the output of the multi-scale dilated convolution unit using an entropy model to obtain a probability model of prediction.

[0070] In an embodiment of the present invention, the above-mentioned multi-scale dilated convolution unit may include three dilated convolution layers, which use different dilation ratios and the same number of channels to perform feature extraction on the output of the hyper-decoder to obtain the multi-scale auxiliary information.

[0071] In an embodiment of the present invention, the above-mentioned context model and the above-mentioned entropy model may be the context model and the entropy model in an image compression system using a convolutional neural network, and other implementation methods are omitted.

[0072] The method for generating a probabilistic model according to an embodiment of the present invention can better predict the probability distribution of latent representations after adding multi-scale auxiliary information. EXAMPLES

[0073] An embodiment of the present invention provides an image processing apparatus, which includes the image coding apparatus according to the first and second embodiments, or includes the image decoding apparatus according to the first and third embodiments, or includes the probability model generating apparatus according to the first and fourth embodiments, or includes the image coding apparatus, the image decoding apparatus and the probability model generating apparatus simultaneously. When including the image decoding apparatus and the probability model generating apparatus simultaneously, the aforementioned multi-scale dilated convolution unit can be shared.

[0074] In the first to fourth embodiments, the image coding device, the probability model generating device, and the image decoding device have been described in detail, so the contents thereof are incorporated herein and the description thereof will be omitted here.

[0075] Fig. 12 is a diagram showing an image processing device in an embodiment of the present invention. As shown in Fig. 12, the image processing device 1200 may include a central processing unit (CPU) 1201 and a memory 1202, and the memory 1202 is connected to the central processing unit 1201. The memory 1202 may store various data, and may further store a program for information processing, and can execute the program under the control of the central processing unit 1201.

[0076] In one embodiment, the functionality of the image coding device and / or the probability model generating device and / or the image decoding device may be integrated into the central processing unit 1201. The central processing unit 1201 may be configured to implement the methods described in the fifth and / or sixth and / or seventh embodiments.

[0077] In another embodiment, the image coding device and / or the probability model generating device and / or the image decoding device may be arranged separately from the central processing unit 1201, for example, the image coding device and / or the probability model generating device and / or the image decoding device may be configured as a chip connected to the central processing unit 1201, and the functions of the image coding device and / or the probability model generating device and / or the image decoding device may be realized under the control of the central processing unit 1201.

[0078] 12, the image processing device may further include an input / output (I / O) device 1203 and a display 1204, and the functions of these components are similar to those of the prior art, and therefore the description thereof will be omitted here. Note that the image processing device does not need to include all the components shown in Fig. 12. The image processing device may further include components not shown in Fig. 12, and the prior art can be referred to for this.

[0079] An embodiment of the present invention provides a computer readable program, which when executed in an image processing device, causes the image processing device to perform the methods described in embodiments 5 and / or 6 and / or 7.

[0080] An embodiment of the present invention provides a storage medium having a computer readable program stored thereon, the computer readable program causing an image processing device to execute the method according to the fifth and / or sixth and / or seventh embodiments.

[0081] In addition, the apparatus, method, etc. according to the embodiments of the present invention may be realized by software, hardware, or a combination of hardware and software. The present invention also relates to such a computer readable program, i.e., the program, when executed by a logic component, can cause the logic component to realize the above-mentioned apparatus or component, or cause the logic component to realize the above-mentioned method or steps thereof. Furthermore, the present invention also relates to a storage medium, such as a hard disk, a magnetic disk, an optical disk, a DVD, a flash memory, etc., storing the above-mentioned program.

[0082] Although the preferred embodiment of the present invention has been described above, the present invention is not limited to this embodiment, and any modification to the present invention falls within the technical scope of the present invention as long as it does not depart from the spirit of the present invention.

Claims

1. 1. An apparatus for decoding an image, comprising: a multi-scale dilated convolution unit that performs feature extraction on the output of the hyperdecoder to obtain multi-scale auxiliary information; a combiner for combining the latent representation of the input image from an arithmetic decoder and the multi-scale auxiliary information from the multi-scale dilated convolution unit; and a decoder for decoding an output from the combiner to obtain a reconstructed image of the input image.

2. 2. The apparatus of claim 1, The multi-scale dilated convolution unit includes three feature extraction units; The three feature extraction units perform feature extraction on the output of the hyper-decoder using dilated convolution kernels with different dilation ratios and the same number of channels to obtain the multi-scale auxiliary information.

3. An apparatus for generating a probabilistic model, comprising: a multi-scale dilated convolution unit that performs feature extraction on the output of the hyperdecoder to obtain multi-scale auxiliary information; a context model processing unit that receives the latent representation of the input image from the quantizer and obtains a content-based prediction; and and an entropy model processing unit that processes an output of the context model processing unit and an output of the multi-scale dilated convolution unit to obtain a probabilistic model of prediction.

4. 4. The apparatus of claim 3, The multi-scale dilated convolution unit includes three feature extraction units; The three feature extraction units perform feature extraction on the output of the hyper-decoder using dilated convolution kernels with different dilation ratios and the same number of channels to obtain the multi-scale auxiliary information.