Training method, image encoding method, image decoding method and device
The training method for image encoders and decoders using a cost function L allows for efficient bitrate adjustments by modifying the quantization step length Q, simplifying the process and enhancing the flexibility of bitrate management in image coding.
Patent Information
- Application Number
- JP2021024862
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-23
- Filing Date
- 2021-02-19
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2041-02-19
AI Technical Summary
Existing image coding methods based on deep neural networks require cumbersome adjustments to bitrate by repeatedly training the encoder network for different λ values to achieve desired bitrates, complicating the process.
A training method and apparatus that utilize a cost function L to train image encoders and decoders, involving latent variable acquisition, quantization, and entropy encoding/decoding, allowing for easy adjustment of bitrates through quantization step length Q without retraining the network.
Enables quick and easy adjustment of bitrates in image encoding and decoding processes, improving efficiency and reducing the complexity associated with traditional methods.
Smart Images

Figure 0007714885000056 
Figure 0007714885000057 
Figure 0007714885000058
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of image processing. [Background technology]
[0002] With the development of computer technology, the application of images is becoming more and more widespread. To efficiently store or transmit an image file, the image must be encoded, and the encoded result can be converted into a bitstream. The image can be reproduced by decoding the bitstream.
[0003] Deep neural networks have become a promising research topic in the field of image coding. Nonlinear transform coding methods designed based on deep neural networks have better performance than traditional image coding methods, such as the Better Portable Graphics (BPG) coding method.
[0004] In image coding methods based on deep neural networks, a key issue is how to achieve a compromise between bit rate and distortion. The bit rate can reflect the magnitude of the image bitstream relative to the image size. For example, the bit rate can be equal to the quotient obtained by dividing the length of the bitstream by the product of the length and width of the image. The distortion can reflect the difference between the image obtained after decoding and the original image.
[0005] Generally, the Lagrange multiplier can be introduced to achieve a trade - off between the bitrate and distortion. For example, when training an encoder based on a deep neural network, training can be performed based on the loss function (R + λ * D), where R represents the bitrate, D represents the degree of distortion, and λ is an adjustable parameter.
Summary of the Invention
Problems to be Solved by the Invention
[0006] After obtaining the encoder network by training based on the loss function (R + λ * D), the bitrate and the degree of distortion of the image are determined.
[0007] The inventor has discovered the following. That is, when it is necessary to adjust the bitrate, usually, the value of λ is corrected multiple times, and for each value of λ, the encoder network is trained again, and it is necessary to determine the encoder network whose bitrate is closest to the required bitrate. Such a method for adjusting the bitrate is relatively cumbersome.
[0008] Embodiments of the present invention provide a training method, an image encoding method, an image decoding method, and an apparatus. The image encoder obtained by the training method can easily achieve adjustment of different bitrates.
Means for Solving the Problems
[0009] According to a first aspect of an embodiment of the present invention, there is provided a training device for an image processing apparatus that trains an image encoder and an image decoder using training images. The training device includes: A first acquisition unit that acquires a latent variable z obtained by the image encoder encoding input training image data; A second acquisition unit that acquires first restored image data obtained by the image decoder performing decoding on the latent variable z, and second restored image data obtained by the image decoder performing decoding on the sum (z + ε) of the latent variable z and noise ε; and A training unit that performs training on the image encoder and the image decoder based on a cost function L, where the cost function L includes a deviation between the input training image data x and the first restored image data, and a deviation between the first restored image data and the second restored image data.
[0010] According to a second aspect of an embodiment of the present invention, an image encoding apparatus is provided, which An image encoder that performs encoding on input image data x to obtain a latent variable z, where the image encoder is obtained by training of the training apparatus described in the first aspect above; A quantizer that performs quantization processing on the latent variable z based on a quantization step length Q to generate a quantized latent variable; and An entropy encoder that performs entropy encoding on the quantized latent variable using an entropy model to form a bitstream.
[0011] According to a third aspect of an embodiment of the present invention, an image decoding apparatus is provided, which An entropy decoder that performs entropy decoding on the bitstream using an entropy model to form a quantized latent variable; A de - quantizer that performs de - quantization processing on the quantized latent variable based on the quantization step length Q to generate a reconstructed latent variable; and An image decoder that performs decoding processing on the reconstructed latent variable to obtain restored image data, where the image decoder is obtained by training of the training apparatus described in the first aspect above.
[0012] According to a fourth aspect of an embodiment of the present invention, there is provided a training method for an image processing method, for training an image encoder and an image decoder using training images, the training method comprising: Obtain a latent variable z obtained by the image encoder performing encoding on the input training image data; Obtaining first restored image data obtained by the image decoder performing decoding on the latent variable z, and second restored image data obtained by the image decoder performing decoding on the sum (z+ε) of the latent variable z and noise ε; and The image encoder and the image decoder are trained based on a cost function L, wherein the cost function L is related to the deviation between input training image data x and the first restored image data, and the deviation between the first restored image data and the second restored image data.
[0013] According to a fifth aspect of an embodiment of the present invention, there is provided an image coding method, comprising: an image encoder performs encoding on input image data x to obtain latent variables z, and the image encoder is trained by the training method described in the fourth aspect; A quantizer performs a quantization process on the latent variable z based on a quantization step length Q to generate a quantized latent variable; and The entropy encoder performs entropy encoding on the quantized latent variables using the entropy model to form a bitstream.
[0014] According to a sixth aspect of an embodiment of the present invention, there is provided an image decoding method, comprising: an entropy decoder performs entropy decoding on the bitstream using the entropy model to form quantized latent variables; An inverse quantizer performs an inverse quantization process on the quantized latent variables based on a quantization step length Q to generate reconstructed latent variables; and an image decoder performs a decoding process on the reconstructed latent variables to obtain restored image data; The image decoder is obtained by training according to the training method described in the fourth aspect above.
[0015] The advantageous effects of the embodiments of the present invention are as follows, that is, the image encoder obtained by the training method can easily realize the adjustment of different bitrates.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0017] Hereinafter, with reference to the accompanying drawings, preferred embodiments for carrying out the present invention will be described in detail.
[0018] <Embodiment of the First Aspect> In the embodiment of the first aspect of the present invention, an image encoding device and an image decoding device are provided. FIG. 1 is a diagram showing the image encoding device and the image decoding device.
[0019] As shown in FIG. 1, the image encoding device 1 can process the image data x to form a bit stream 100, and the bit stream 100 can be stored or transmitted to the image decoding device 2 by transmission means. The image decoding device 2 processes the received bit stream 100 and can generate restored image data (External 1) TIFF0007714885000001.tif7153. Thus, the image data x input to the image encoding device 1 can be reproduced as the image data (External 2) TIFF0007714885000002.tif7153 at the image decoding device 2.
[0020] As shown in FIG. 1, the image encoding device 1 may include an image encoder 11, a quantizer 12, and an entropy encoder 13.
[0021] The image encoder 11 can perform an encoding process on the input image data x to obtain a latent variable z. The image encoder 11 can perform the encoding process based on a deep neural network. For example, the image encoder 11 can be realized by using a basic convolution layer, and / or a deconvolution layer, and / or generalized divisive normalization (GDN) / inverse generalized divisive normalization (IGDN) as an activation function. For the concept and content of the deep neural network, reference can be made to related technologies.
[0022] The quantizer 12 performs a quantization process on the latent variable z output by the image encoder 11 based on the quantization step length Q, and the quantized latent variable (External 3) TIFF0007714885000003.tif8153 can be generated. The latent variable z is floating-point data, and quantization processing can convert floating-point data into finite-length data.
[0023] The entropy coder 13 uses an entropy model 14 to generate the quantized latent variables (outside 4) TIFF0007714885000004.tif9153 can be entropy coded to form a bitstream 100. The bitstream 100, also called a bit stream, is a data stream containing a plurality of bits. Entropy coding reduces the number of quantized latent variables, which are difficult to store and transmit. (outside 5) TIFF0007714885000005.tif8153 can be converted into a bitstream 100 that is easy to store and transmit. Entropy coding is a coding method based on the entropy principle that does not lose information. Therefore, the information contained in the bitstream 100 is expressed as a quantized latent variable (outside 6) It can completely reflect the information in TIFF0007714885000006.tif8153.
[0024] In at least one embodiment, the entropy model 14 can be used to estimate the entropy of the latent variable z. The entropy encoder 13 can perform entropy coding based on the entropy estimation result of the latent variable z by the entropy model 14. In particular, the entropy model 14 can be, for example, a factorized entropy model.
[0025] The bit rate R of the bit stream 100 generated by the entropy encoder 13 can be expressed as R=n / (W*H), where n indicates the length of the bit stream 100, and W and H indicate the width and length of the image corresponding to the image data x, respectively, and both the width and length can be expressed in terms of the number of pixels.
[0026] The bitstream 100 produced by the entropy coder 13 can be stored or transmitted to the image decoding device 2 .
[0027] As shown in FIG. 1, the image decoding device 2 may include an image decoder 21, an inverse quantizer 22, and an entropy decoder .
[0028] The entropy decoder 23 performs entropy decoding on the received bitstream 100 using the entropy model 14, and generates the quantized latent variables (outer 7) TIFF0007714885000007.tif7153 can be formed. The entropy decoding process may be the reverse process of the entropy encoding process of the entropy encoder 13.
[0029] The inverse quantizer 22 calculates the quantized latent variables (outside 8) Inverse quantization is performed on TIFF0007714885000008.tif7153, and the reconstructed latent variables (outer 9) TIFF0007714885000009.tif7153 can be generated. The inverse quantization process may be the inverse process of the quantization process.
[0030] The image decoder 21 reconstructs the latent variables (Outside 10) Decoding is performed on TIFF0007714885000010.tif8153, and the restored image data is (Outside 11) TIFF0007714885000011.tif7153 can be obtained. The image decoder 21 can perform decoding processing based on a deep neural network. For example, the image decoder 21 can be realized by using a basic convolution layer, and / or a deconvolution layer, and / or generalized divisive normalization (GDN) / inverse generalized divisive normalization (IGDN) as an activation function. For the concept and content of the deep neural network, reference can be made to the related technology.
[0031] In at least one embodiment, the image encoder 11 and the image decoder 21 may be an image encoder and an image decoder based on the RaDOGAGA (Rate-Distortion Optimization Guided Autoencoder for Generative Analysis) model. For the principle of the RaDOGAGA model, reference can be made to the related technology, for example, :https: / / arxiv.org / abs / 1910.04329.
[0032] In at least one embodiment, the image encoder 11 and the image decoder 21 can be trained using a training device based on the RaDOGAGA model.
[0033] FIG. 2 is a diagram showing a training device in an embodiment of the present invention. As shown in FIG. 2, the training device 3 may include a first acquisition unit 31, a second acquisition unit 32, and a training unit 33.
[0034] As shown in FIG. 2, the first acquisition unit 31 can obtain a latent variable z obtained by the image encoder 11 encoding the input training image data x. For example, z can be expressed as in Equation (1). z=f θ (x) (1)
[0035] In Equation (1), fθ indicates the encoding process of the image encoder 11, and the encoding process has θ as a parameter.
[0036] The second acquisition unit 32 acquires the first restored image data obtained by the image decoder 21 performing decoding on the latent variable z. (Outside 12) TIFF0007714885000012.tif6153 and the second restored image data obtained by the image decoder 21 performing decoding on the sum (z+ε) of the latent variable z and the noise ε. (Outside 13) You can get TIFF0007714885000013.tif7153. For example, (Outside 14) TIFF0007714885000014.tif6153 and (Outside 15) TIFF0007714885000015.tif9153 can be expressed as in equation (2).
number
[0037] In equation (2), g φ represents the decoding process of the image decoder 21, and the decoding process has φ as a parameter. Furthermore, the noise ε may be uniform noise.
[0038] The training unit 33 can train the image encoder 11 and the image decoder 21 based on a cost function L, where L is the input training image data x and the first reconstructed image data (Outside 16) Deviation of TIFF0007714885000017.tif10153( (Outside 17) TIFF0007714885000018.tif11153), and the first restored image data (outside 18) TIFF0007714885000019.tif9153 and the second restored image data (Outside 19) Deviation from TIFF0007714885000020.tif9153 ( (outside 20) TIFF0007714885000021.tif10153). Furthermore, the training unit 33 training the image encoder 11 and the image decoder 21 means that the training unit 33 trains the network in the image encoder 11 and the network in the image decoder 21.
[0039] In at least one embodiment, the cost function L can be expressed as follows:
number
[0040] The first term of equation (3) is log(P z,ψ In (z), P z,ψ (z) represents the probability of the latent variable z, and this probability has the latent variables z and ψ as parameters. The entropy model 14 shown in FIG. 1 can obtain the cumulative density function (CDF) of the latent variable z, and the cumulative density function CDF can be used to calculate the probability P z (z) can be estimated.
[0041] In the entropy model 14, the cumulative density function CDF can satisfy the relationships shown in equations (4a) and (4b).
number
[0042] In equations (4a) and (4b), α denotes the quantization step length of the bit rate of the latent variable z, and R zdenotes the bit rate of the latent variable z, where H and W denote the height and width of the input image, respectively.
[0043] In equation (3), the second term (outside 21) TIFF0007714885000024.tif10153 is used to calculate the reconstruction loss of the image encoder 11 and the image decoder 21, and the third term (outside 22) TIFF0007714885000025.tif10153 can reflect the scaling relationship between the image and the latent space. λ1 is used to control the degree of reconstruction, and λ2 is used to control the scaling ratio between the image and the latent space.
[0044] The second term of Equation (3) (outside 23) TIFF0007714885000026.tif11153 and the third paragraph (outside 24) In TIFF0007714885000027.tif11153, D(x1,x2) is the distortion function of the difference between x1 and x2. Distortion parameters used in the image coding field may be mean square error (MSE), peak signal-to-noise ratio (PSNR), multi-scale structural similarity (MS-SSIM) index, or structural similarity (SSIM) index. Corresponding to the above distortion parameters, the distortion function D(x1,x2) may be mean square error (MSE) distortion function, peak signal-to-noise ratio (PSNR) distortion function, multi-scale structural similarity (MS-SSIM) index distortion function, or structural similarity (SSIM) index distortion function.
[0045] In the second term of Equation (3), h(D) may be log(D). As a result, the curve of the loss function becomes steeper in the vicinity of log(D)=0. Therefore, the image encoder 11 and the image decoder 21 can obtain better reconstruction characteristics and orthogonality. Note that the present invention is not limited to this, and h(D) may be D.
[0046] In one specific example, the shape of the input training image x is H*W*3, where H is the height of the training image x, W is the width of the training image x, 3 indicates three channels, the noise ε takes values from -0.5 to 0.5, the value of α is 0.2, and the shape of each feature image generated by the image encoder 11 is H / 16*W / 16. In the first stage of training, the distortion function D(x1,x2) adopts the minimum mean square error (MSE) distortion function and h(D)=D. In the second stage of training, the distortion function D(x1,x2) adopts the multi-scale structural similarity (MS-SSIM) index distortion function MS SSIM (x1,x2) and h(D)=log(D). That is, in the second stage of training, the image encoder 11 and the image decoder 21 are trained using the loss function L of Equation (5).
Equation
[0047] In Equation (5), λ1 may be 1, and λ2 may be greater than 100.
[0048] Based on FIG. 2 above, the process by which the training device 3 trains the image encoder 11 and the image decoder 21 has been described. The model obtained by training using the above cost function L has a relationship where the feature layer space and the MS-SSIM space are equidistant. That is, the feature layer is optimized to be orthogonal to the inner product space of the distortion function, and its function is similar to the discrete cosine transform (DCT) adopted by JPEG (Joint Photographic Experts Group). In the training stage, MSE(x1,x2), SSIM(x1,x2), etc. may be used as D(x1,x2). For example, when using MSE(x1,x2) as the distortion function, an effect similar to MS-SSIM can also be obtained. That is, different quantization step lengths can be used to obtain a PSNR value equivalent to the independent training model (R + λ*D).
[0049] In the first aspect of the embodiment of the present invention, through the training of the training device 3, the image encoder 11 and the image decoder 21 can be obtained. The image encoding device 1 having the image encoder 11 can easily realize the adjustment of different bitrates, and the image decoding device 2 having the image decoder 21 can adapt to different bitrates.
[0050] Hereinafter, the operations related to the quantization process of the image encoding device 1 and the image decoding device 2 will be described.
[0051] In at least one embodiment, the quantization process of the quantizer 12 may be a non-uniform quantization process. Among them, the non-uniform quantization process uses the latent variable z corresponding to the peak value (or median) of the probability distribution of the latent variable z as the zero point, and the latent variables z in the first range including the zero point are the first quantized latent variables (Outer 25) TIFF0007714885000029.tif11153 are made to correspond, and the first quantized latent variable (Outer 26) TIFF0007714885000030.tif9153 and other quantized latent variables (Outer 27) For TIFF0007714885000031.tif10153, each quantized latent variable (outside 28) TIFF0007714885000032.tif11153 corresponds to the latent variable z in a second range, where the second range is not larger than the first range, and the probability distribution peak value of the latent variable z can be obtained based on the entropy model 14.
[0052] For example, the quantizer 12 can perform the quantization process using equation (6).
number
[0053] Wherein, sign(z) indicates the sign of the latent variable z. For example, if z is greater than 0, sign(z) is positive; if z is less than 0, sign(z) is negative; floor indicates truncation; abs(z) means taking the absolute value of z; offset is a preset deviation, where 0≦offset≦0.5.
[0054] In the present invention, offset is used to set the length of the first range, i.e., the length of the first range is 2*(1-offset)*Q. The length of the second range is equal to the quantization step length Q.
[0055] In at least one embodiment, the offset is not equal to 0.5, the length of the second range is less than the length of the first range, and the quantization performed by the quantizer 12 is a non-uniform quantization. (outside 29) The entropy of TIFF0007714885000034.tif11153 becomes smaller. Note that the present invention is not limited to this, and for example, when offset is equal to 0.5, the length of the second range is equal to the length of the first range, and the quantization process performed by the quantizer 12 is a uniform quantization process.
[0056] The quantized latent variable generated by the quantizer 12 (External 30) TIFF0007714885000035.tif10153 can form a bit stream 100 through the entropy encoding of the entropy encoder 13. When the bit stream 100 is entropy decoded by the entropy decoder 23, in the image decoding device 2, the quantized latent variable (External 31) TIFF0007714885000036.tif11153 can be obtained
[0057] In at least one embodiment, the inverse quantizer 22 can perform a dequantized process using the quantization step length Q. For example, the inverse quantizer 22 uses Equation (7) to process the quantized latent variable output by the entropy decoder 23 (External 32) TIFF0007714885000037.tif10153 is dequantized to obtain a reconstructed latent variable (External 33) TIFF0007714885000038.tif10153 as follows [Number]
[0058] Based on the entropy model 14, the reconstructed latent variable (External 34) The cumulative density function (CDF) of TIFF0007714885000040.tif9153 can be obtained, and when the quantizer 12 quantizes z, z is mapped to a corresponding representative value based on the quantization step length (External 35) TIFF0007714885000041.tif10153 can be quantized, where the representative value (External 36) The upper bound of the z interval corresponding to TIFF0007714885000042.tif8153 is z highand the low bound is z low That is, [z low ,z high ] All z's in the interval (Outside 37) TIFF0007714885000043.tif11153, of which
number
[0059] and 0<ω<1.
number
number
[0060] z high and zlow Based on the reconstructed latent variables (Outside 38) Bitrate of TIFF0007714885000047.tif9153 (Outside 39) TIFF0007714885000048.tif11153 can be obtained by equation (10).
number
[0061] 3 is a diagram showing the quantization process of the quantizer 12 and the inverse quantization process of the inverse quantizer 22. As shown in FIG. 3, an arrow 31 indicates the quantization process of the quantizer 12, and an arrow 32 indicates the inverse quantization process of the inverse quantizer 22.
[0062] As shown in FIG. 3, for example, by the quantization process of Equation (6), the latent variable z is converted into the quantized latent variable (outside 40) TIFF0007714885000050.tif12153. For example, the latent variables z in the first range (interval) indicated by 301 are all quantized latent variables with a value of 0. (outside 41) TIFF0007714885000051.tif12153, and outside the first range, the latent variable z is divided evenly into a plurality of second ranges (intervals) 302, and within the second range 302, the latent variable z is divided into a quantized latent variable (outside 42) Mapped to TIFF0007714885000052.tif11153.
[0063] As shown in FIG. 3, for example, by the inverse quantization process of Equation (7), each quantized latent variable (outside 43) TIFF0007714885000053.tif10153 is the corresponding reconstructed latent variable (outside 44) It can be mapped to TIFF0007714885000054.tif10153.
[0064] 1, the image encoding device 1 may further include a first quantization step length adjuster 15. The first quantization step length adjuster 15 can adjust the bit rate of the bitstream 100 by adjusting the quantization step length Q used by the quantizer 12.
[0065] 1, the image decoding device 2 may further include a second quantization step length adjuster 25. The second quantization step length adjuster 25 can adjust the quantization step length Q used by the inverse quantizer 22. For example, the second quantization step length adjuster 25 can adjust the quantization step length Q used by the inverse quantizer 22 based on the quantization step length Q adjusted by the first quantization step length adjuster 15. This allows the inverse quantizer 22 and the quantizer 12 to use the same step length Q.
[0066] In the image coding device 1 of the present invention, the image encoder 11 is an image encoder based on the RaDOGAGA model, and can adjust the bit rate by adjusting the quantization step length Q. This allows for quick and easy bit rate adjustment. In contrast, in conventional methods, the value of the loss function λ needs to be modified multiple times, and the encoder network needs to be retrained for each λ value to determine the encoder network whose bit rate is closest to the required bit rate. This makes the bit rate adjustment process relatively complicated and time-consuming.
[0067] In order to compare the performance of the image coding device 1 of the present invention with that of the conventional image coding device, experiments were conducted on the image coding device 1 of the present invention and the conventional image coding device based on the general-purpose test data collection Kodak, and bit rate-distortion degree curves (RD curves) of both were created. (outside 45) The same coding network structure as TIFF0007714885000055.tif10153 is adopted. To plot the RD curve of the conventional image coding device, we train the image codec network for different λ∈{4, 8, 16, 32, 64, 96}, respectively, and calculate the distortion parameter MS-SSIM. dB The distortion level of each image codec network is shown using MS_SSIM. dB =-10log2(1-MS_SSIM). The R and D corresponding to each of the six image codec networks are fitted as the first curve.
[0068] In the image coding device 1 of the present invention, the network structure of the image coder 11 does not need to be trained multiple times, and it is sufficient to adjust the quantization step length Q, where Q∈{0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 3.5, 4}, and R and D corresponding to each quantization step length are calculated. R and D corresponding to each of these multiple quantization step lengths Q are fitted as a second curve.
[0069] Fig. 4 is a diagram showing the first and second curves. In Fig. 4, the horizontal axis represents the bit rate R in units of bpp (bits per pixel), and the vertical axis represents the degree of distortion represented by MS-SSIM in units of dB (decibels).
[0070] In FIG. 4 , the point “λ=64” on the first curve 41 indicates that the loss function for training the model is (R+λ*D) as described in the Background Art, where λ=64, i.e., the loss function is (R+64*D), and the rest can be inferred based on this. The point “Q=1” on the second curve indicates that the quantization step length is Q=1, and the rest can be inferred based on this.
[0071] 4, the RD characteristics of the second curve 42 are close to those of the first curve 41. That is, the image coding device 1 of the present invention can adjust the bit rate without training the network structure of the image coder 11 multiple times, and the RD characteristics do not deteriorate, simply by adjusting the quantization step length Q. This allows the image coding device 1 of the present invention to adjust the bit rate in a simple and quick manner.
[0072] <Example of the second aspect> In the embodiments of the present invention, an image encoding method, an image decoding method, and a training method are provided.
[0073] 5 is a diagram showing an image coding method according to an embodiment of the second aspect of the present invention. As shown in FIG. 5, the image coding method includes the following operations (steps):
[0074] Operation 51: The image encoder performs encoding on the input image data x to obtain a latent variable z; Operation 52: A quantizer performs a quantization process on the latent variable z based on a quantization step length Q to generate a quantized latent variable; and Operation 53: The entropy encoder uses the entropy model to perform entropy coding on the quantized latent variables to form a bitstream.
[0075] As shown in FIG. 5, the image encoding method further includes the following operations.
[0076] Operation 54: A first quantization step length adjuster adjusts the quantization step length Q, thereby adjusting the bit rate of the bitstream.
[0077] In at least one embodiment, the quantization process of the quantizer is a non-uniform quantization process, which includes the following operations: The latent variable z corresponding to the peak value of the probability distribution of the latent variable z is set as a zero point, and the latent variables z in a first range including the zero point correspond to first quantized latent variables, and for quantized latent variables other than the first quantized latent variable, each quantized latent variable corresponds to a latent variable z in a second range, and the second range is not larger than the first range.
[0078] Wherein, the probability distribution peak value of the latent variable z can be obtained based on the entropy model.
[0079] For the explanation of each operation in FIG. 5, please refer to the explanation of the corresponding unit in FIG.
[0080] 6 is a diagram illustrating an image decoding method according to an embodiment of the second aspect of the present invention. As shown in FIG. 6, the image decoding method includes the following operations:
[0081] Operation 61: The entropy decoder performs entropy decoding on the bitstream using the entropy model to form quantization latent variables; Operation 62: The inverse quantizer performs inverse quantization processing on the quantization latent variables based on the quantization step length Q to generate reconstruction latent variables; and Operation 63: The image decoder performs decoding processing on the reconstruction latent variables to obtain restored image data.
[0082] The inverse quantizer in Operation 62 performs inverse quantization processing based on the quantization step length.
[0083] As shown in FIG. 6, the image decoding method may further include the following operations, that is, Operation 64: The second quantization step length adjuster adjusts the quantization step length Q.
[0084] For the description of each operation in FIG. 6, reference may be made to the description of the corresponding unit in FIG. 1.
[0085] FIG. 7 is a diagram showing a training method in an embodiment of the second aspect of the present invention. As shown in FIG. 7, the training method includes the following operations, that is, Operation 71: The image encoder obtains the latent variable z obtained by encoding the input training image data; Operation 72: The first restored image data obtained by the image decoder decoding the latent variable z, and the second restored image data obtained by the image decoder decoding the sum of the latent variable z and noise (z + ε) are obtained; and Operation 73: Training is performed on the image encoder and the image decoder based on the cost function L, and the cost function L is related to the deviation between the input training image data x and the first restored image data, and the deviation between the first restored image data and the second restored image data.
[0086] For the description of each operation in FIG. 7, reference may be made to the description of the corresponding unit in FIG. 2.
[0087] <Example of the third side> In an embodiment of the present invention, an electronic device is further provided, which includes the image encoding device 1 and / or the image decoding device 2 and / or the training device 3 described in the embodiment of the first side. In this embodiment, its content is incorporated herein. The electronic device may be, for example, a computer, a server, a workstation, a laptop computer, a smartphone, etc. Note that the embodiments of the present invention are not limited to this.
[0088] FIG. 8 is a configuration diagram of an electronic device in an embodiment of the present invention. As shown in FIG. 8, the electronic device 800 may include a processor (for example, a central processing unit CPU) 810 and a memory 820, and the memory 820 is connected to the central processing unit 810. Among them, the memory 820 can store various data, further store a program for information processing, and can also execute the program under the control of the processor 810.
[0089] In one embodiment, the functions of the image encoding device 1 and / or the image decoding device 2 and / or the training device 3 may be integrated into the processor 810. Among them, the processor 810 may be configured to implement the image encoding method and / or the image decoding method and / or the training method described in the embodiment of the second side.
[0090] In another embodiment, the image encoding device 1 and / or the image decoding device 2 and / or the training device 3 may be separately arranged from the processor 810. For example, the image encoding device 1 and / or the image decoding device 2 and / or the training device 3 may be configured as a chip connected to the processor 810, and the functions of the image encoding device 1 and / or the image decoding device 2 and / or the training device 3 may be realized under the control of the processor 810.
[0091] Note that for the specific implementation of the processor 810, reference may be made to the embodiments of the first side and the second side, and the detailed description thereof is omitted here.
[0092] Also, as shown in FIG. 8, the electronic device 800 may further include a transmission / reception unit 830 and the like. Among them, since the functions of these components are similar to those of the prior art, detailed descriptions thereof are omitted here. Note that the electronic device 800 does not necessarily include all the components shown in FIG. 8. Further, the electronic device 800 may further include components not shown in FIG. 8, and for this, reference may be made to the prior art.
[0093] Embodiments of the present invention further provide a computer-readable program, wherein when the program is executed in an image encoding device and / or an image decoding device and / or a training device, the program causes a computer to execute the image encoding method and / or the image decoding method and / or the training method in the embodiments of the second aspect described above in the image encoding device and / or the image decoding device and / or the training device.
[0094] Embodiments of the present invention further provide a storage medium storing a computer-readable program, wherein the computer-readable program causes a computer to execute the image encoding method and / or the image decoding method and / or the training method in the embodiments of the second aspect described above in the image encoding device and / or the image decoding device and / or the training device.
[0095] The methods, apparatuses, etc. described in the embodiments of the present invention can be implemented by hardware, software modules executed by a processor, or a combination of both. For example, one or more functions in the functional block diagrams shown in FIGS. 1 to 13 and / or a combination of one or more functions in the functional block diagrams may correspond to each software module in a computer program or to each hardware module. Further, these software modules can respectively correspond to each step shown in FIG. 14. These hardware modules can be realized, for example, by solidifying these software modules using an FPGA (field-programmable gate array).
[0096] Furthermore, the apparatus, method, etc. according to the embodiments of the present invention may be realized by software, hardware, or a combination of hardware and software. The present invention also relates to such a computer-readable program, i.e., the program, when executed by a logic component, can cause the logic component to realize the above-mentioned apparatus or component, or cause the logic component to implement the above-mentioned method or steps thereof. Furthermore, the present invention also relates to a storage medium, such as a hard disk, magnetic disk, optical disk, DVD, flash memory, etc., on which the above-mentioned program is stored.
[0097] Furthermore, the following supplementary notes are disclosed regarding the above-mentioned embodiments.
[0098] (Appendix 1) 1. A training device for an image processing device that trains an image encoder and an image decoder using training images, comprising: a first acquisition unit for acquiring a latent variable z obtained by the image encoder performing encoding on input training image data; a second acquisition unit for acquiring first restored image data obtained by the image decoder performing decoding on the latent variable z, and second restored image data obtained by the image decoder performing decoding on the sum (z+ε) of the latent variable z and noise ε; and A method comprising: a training unit that trains the image encoder and the image decoder based on a cost function L, the cost function L being related to the deviation between input training image data x and the first reconstructed image data, and the deviation between the first reconstructed image data and the second reconstructed image data.
[0099] (Appendix 2) An image encoding device, an image encoder that encodes input image data x to obtain a latent variable z, the image encoder being obtained by training using the training device of Supplementary Note 1; A quantizer that performs quantization processing on the latent variable z based on the quantization step length Q to generate a quantized latent variable; and An apparatus including an entropy coder that performs entropy coding on the quantized latent variable using an entropy model to form a bit stream.
[0100] (Appendix 3) An image coding apparatus according to Appendix 2, further including a first quantization step length adjuster that adjusts the bit rate of the bit stream by adjusting the quantization step length Q.
[0101] (Appendix 4) An image coding apparatus according to Appendix 2, wherein the quantization processing of the quantizer is non-uniform quantization processing.
[0102] (Appendix 5) An image coding apparatus according to Appendix 4, wherein the non-uniform quantization processing uses the latent variable z corresponding to the peak value of the probability distribution of the latent variable z as a zero point, and makes the latent variable z in a first range including the zero point correspond to a first quantized latent variable; and for other quantized latent variables other than the first quantized latent variable, makes each quantized latent variable correspond to a latent variable z in a second range, and includes that the second range is not larger than the first range.
[0103] (Appendix 6) An image coding apparatus according to Appendix 5, wherein the peak value of the probability distribution of the latent variable z is obtained based on the entropy model.
[0104] (Appendix 7) An image decoding apparatus, including an entropy decoder that performs entropy decoding on a bit stream using an entropy model to form a quantized latent variable; An inverse quantizer that performs inverse quantization processing on the quantization latent variable based on the quantization step length Q to generate a reconstructed latent variable; and An image decoder that performs decoding processing on the reconstructed latent variable to obtain restored image data, wherein the image decoder is obtained by training with the training device described in Appendix 1, and the apparatus includes the image decoder.
[0105] (Appendix 8) An image decoding apparatus according to Appendix 7, wherein the inverse quantizer performs the inverse quantization processing based on the quantization step length.
[0106] (Appendix 9) An image decoding apparatus according to Appendix 7, further including a second quantization step length adjuster that adjusts the quantization step length Q.
[0107] (Appendix 10) A training method for an image processing apparatus that performs training on an image encoder and an image decoder using training images, wherein the image encoder obtains a latent variable z obtained by encoding input training image data; the image decoder obtains first restored image data obtained by decoding the latent variable z, and second restored image data obtained by decoding the sum (z + ε) of the latent variable z and noise ε; and performs training on the image encoder and the image decoder based on a cost function L, and the cost function L includes being related to the deviation between the input training image data x and the first restored image data, and the deviation between the first restored image data and the second restored image data.
[0108] (Appendix 11) An image encoding method, wherein an image encoder encodes input image data x to obtain a latent variable z, and the image encoder is obtained by training with the training method described in Appendix 10; The quantizer performs quantization processing on the latent variable z based on the quantization step length Q to generate a quantized latent variable; and An entropy coder forms a bit stream by performing entropy coding on the quantized latent variable using an entropy model, the method comprising:
[0109] (Appendix 12) An image coding method according to Appendix 11, further comprising: A first quantization step length adjuster adjusts the quantization step length Q to adjust the bit rate of the bit stream, the method comprising:
[0110] (Appendix 13) An image coding method according to Appendix 11, wherein the quantization processing of the quantizer is non-uniform quantization processing, the method comprising:
[0111] (Appendix 14) An image coding method according to Appendix 13, wherein the non-uniform quantization processing comprises: taking the latent variable z corresponding to the peak value of the probability distribution of the latent variable z as a zero point, and making the latent variable z in a first range including the zero point correspond to a first quantized latent variable; and for other quantized latent variables other than the first quantized latent variable, making each quantized latent variable correspond to a latent variable z in a second range, and the second range not being larger than the first range, the method comprising:
[0112] (Appendix 15) An image coding method according to Appendix 14, wherein the peak value of the probability distribution of the latent variable z is obtained based on the entropy model, the method comprising:
[0113] (Appendix 16) An image decoding method, comprising: An entropy decoder performs entropy decoding on a bit stream using an entropy model to form a quantized latent variable; The inverse quantizer performs inverse quantization processing on the quantized latent variable based on the quantization step length Q to generate a reconstructed latent variable; and An image decoder performs decoding processing on the reconstructed latent variable to obtain restored image data, and the image decoder is obtained by being trained by the training method described in Appendix 10. A method comprising this.
[0114] (Appendix 17) An image decoding method according to Appendix 16, wherein The inverse quantizer performs the inverse quantization processing based on the quantization step length. A method.
[0115] (Appendix 18) An image decoding method according to Appendix 16, wherein The method further includes a second quantization step length adjuster adjusting the quantization step length Q.
[0116] As described above, the preferred embodiments of the present invention have been described. However, the present invention is not limited to this embodiment, and any changes to the present invention belong to the technical scope of the present invention as long as the gist of the present invention is not departed from.
Claims
1. An image encoding device, comprising: An image encoder that encodes input image data and obtains a latent variable; A quantizer that performs quantization processing on the latent variable based on a quantization step size to generate a quantized latent variable; An entropy encoder that performs entropy encoding on the quantized latent variable using an entropy model to form a bit stream; and A first quantization step size adjuster that adjusts the bit rate of the bit stream by adjusting the quantization step size, wherein the image encoder is A training device for an image processing device that trains an image encoder and an image decoder using training images, A first acquisition unit that obtains a latent variable obtained by the image encoder encoding input training image data; A second acquisition unit that obtains first restored image data obtained by the image decoder decoding the latent variable, and second restored image data obtained by the image decoder decoding the sum of the latent variable and noise; and A training unit that trains the image encoder and the image decoder based on a cost function, wherein the cost function includes a deviation between input training image data and the first restored image data, and a deviation between the first restored image data and the second restored image data, the image encoding device obtained by training by the training unit An image encoding device.
2. The image encoding device according to claim 1, wherein the quantization processing of the quantizer is non-uniform quantization processing.
3. The image encoding device according to claim 2, wherein the non-uniform quantization processing uses the latent variable corresponding to the peak value of the probability distribution of the latent variable as a zero point, and enables the latent variables in a first range including the zero point to correspond to first quantized latent variables; and for other quantized latent variables other than the first quantized latent variables, enables each quantized latent variable to correspond to latent variables in a second range, wherein the second range is not larger than the first range.
4. The image encoding device according to claim 3, wherein the peak value of the probability distribution of the latent variable is obtained based on the entropy model.
5. An image decoding device, comprising: An entropy decoder that performs entropy decoding on a bit stream using an entropy model to form a quantized latent variable; An inverse quantizer that performs inverse quantization processing on the quantized latent variable based on the quantization step length to generate a reconstructed latent variable; An image decoder that performs decoding processing on the reconstructed latent variable to obtain restored image data; and Including a second quantization step length adjuster that adjusts the quantization step length, The image decoder, A training device for an image processing apparatus that trains an image encoder and an image decoder using training images, the training device comprising: A first acquisition unit that acquires a latent variable obtained by the image encoder encoding input training image data; A second acquisition unit that acquires first restored image data obtained by the image decoder decoding the latent variable, and second restored image data obtained by the image decoder decoding the sum of the latent variable and noise; and A training unit that trains the image encoder and the image decoder based on a cost function, the cost function being related to the deviation between the input training image data and the first restored image data, and the deviation between the first restored image data and the second restored image data, the training device including the training unit An image decoding apparatus obtained by training by the above.
6. The image decoding apparatus according to claim 5, The inverse quantizer performs the inverse quantization processing based on the quantization step length, the image decoding apparatus.
Citation Information
Patent Citations
Method for improving image quality
JP2020010331A