Methods for training image processing networks, encoding methods, decoding methods, and electronic devices

By dividing images into blocks and training the image processing network based on loss, the checkerboard effect is minimized, enhancing image quality in deep learning-based image processing.

JP7853515B2Active Publication Date: 2026-04-28HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2023-05-18
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Deep learning-based image processing networks often produce a checkerboard effect in images, significantly degrading visual quality.

Method used

The proposed method involves training an image processing device to address the technical problem of the checkerboard effect by dividing the image into blocks and calculating the loss between the image before and after processing the image using the trained image processing device to remove the checkerboard effect by training the image processing network based on the loss.

Benefits of technology

The method effectively reduces the checkerboard effect in images, improving their visual quality by compensating for periodic noise and enhancing image processing networks in tasks like image super-resolution and image restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007853515000012
    Figure 0007853515000012
  • Figure 0007853515000013
    Figure 0007853515000013
  • Figure 0007853515000014
    Figure 0007853515000014
Patent Text Reader

Abstract

[0003] The present application provides a method, an encoding method, a decoding method, and an electronic device for training an image processing network. The training method includes the steps of: obtaining a first training image and a first predicted image, and obtaining a period of a checkerboard effect, where the first predicted image is generated by performing image processing on the first training image based on an image processing network; dividing the first training image into M first image blocks and the first predicted image into M second image blocks based on the period, where both the size of the first image block and the size of the second image block are related to the period; determining a first loss based on the M first image blocks and the M second image blocks; and then training the image processing network based on the first loss. After the image processing network is trained based on the training method of the present application, the checkerboard effect of an image obtained by processing performed using the trained image processing network can be removed to a certain extent, and the visual quality of the image obtained by processing performed using the trained image processing network can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to Chinese Patent Application No. 202210945704.0, titled "Method, encoding method, decoding method, and electronic device for training an image processing network," filed with the China National Intellectual Property Administration on 8 August 2022, which is incorporated herein by reference in its entirety.

[0002] Embodiments of this application relate to the field of image processing, and more particularly to methods, encoding methods, decoding methods, and electronic devices for training image processing networks. [Background technology]

[0003] Deep learning is being applied to various image processing tasks such as image compression, image restoration, and image super-resolution, based on its performance, which is far superior to that of conventional image algorithms in many fields, including image recognition and target detection.

[0004] Typically, in many image processing scenarios (e.g., image compression, image restoration, and image super-resolution), a checkerboard effect appears in images obtained by processing performed using deep learning networks. In other words, a grid very similar to a checkerboard appears in parts or all of the obtained image, resulting in a significant decrease in the visual quality of the image obtained by the image processing. [Overview of the Initiative] [Means for solving the problem]

[0005] This application provides a method, encoding method, decoding method, and electronic device for training an image processing network. After the image processing network is trained based on the training method, the checkerboard effect in images obtained by processing performed using the trained image processing network can be removed to some extent.

[0006] According to a first aspect, one embodiment of the present application provides a method for training an image processing network. The method includes the steps of: first acquiring a first training image and a first prediction image and acquiring a period of the checkerboard effect, the first prediction image being generated by performing image processing on the first training image based on an image processing network; dividing the first training image into M first image blocks and the first prediction image into M second image blocks based on the period, the size of both the first and second image blocks being related to the period, where M is an integer greater than 1; determining a first loss based on the M first image blocks and the M second image blocks; and then training an image processing network based on the first loss.

[0007] Because the checkerboard effect is periodic, in this application, images before and after processing performed using an image processing network (the first training image is the image before processing performed using the image processing network, and the first prediction image is the image after processing performed using the image processing network) are divided into image blocks based on the period of the checkerboard effect. Next, the loss is calculated by comparing the difference between the image blocks before and after processing performed using the image processing network (M first image blocks are the image blocks before processing performed using the image processing network, and M second image blocks are the image blocks after processing performed using the image processing network). The image processing network is trained based on the loss to effectively compensate each image block processed using the image processing network, and as a result, reduce the difference between each first image block and the corresponding second image block. Both the size of the first image blocks and the size of the second image blocks are related to the period of the checkerboard effect. As the difference between each first image block and the corresponding second image block decreases, the checkerboard effect in each period is also eliminated to some extent. In this way, after the image processing network is trained based on the training method of this application, the checkerboard effect of images acquired by processing performed using the trained image processing network can be removed to some extent, and as a result, the visual quality of images acquired by processing performed using the trained image processing network can be improved.

[0008] For example, the checkerboard effect is periodic noise on an image (periodic noise is noise associated with a spatial domain and a specific frequency), and the period of the checkerboard effect is the period of the noise.

[0009] For example, the period of the checkerboard effect is a two-dimensional period (the checkerboard effect is a phenomenon in which a grid very similar to a checkerboard appears in a partial or whole region of an image; that is, the checkerboard effect is two-dimensional, and a two-dimensional period means that the period of the checkerboard effect contains two dimensions, where the two dimensions correspond to the length and width of the image). The number of periods of the checkerboard effect is M.

[0010] For example, the shape of the period of the checkerboard effect may be rectangular, the size of the period of the checkerboard effect may be expressed as p*q, where p and q are positive integers, the units of p and q are px (pixels), and p and q may or may not be equal. This is not limited to the present application.

[0011] It should be understood that the periodicity of the checkerboard effect may be of a different shape (e.g., triangular, elliptical, or irregular). This is not limited to the present application.

[0012] For example, the first loss is used to compensate for the checkerboard effect in the image.

[0013] For example, the sizes of the M first image blocks may be the same or different. This is not limited to this application.

[0014] For example, the sizes of the M second image blocks may be the same or different. This is not limited to this application.

[0015] For example, the sizes of the first image block and the second image block may be the same or different. This is not limited to the present application.

[0016] For example, the size of both the first and second image blocks, which are related to the period, may indicate that the sizes of the first and second image blocks are determined based on the period of the checkerboard effect. For example, the sizes of both the first and second image blocks may be greater than, less than, or equal to the period of the checkerboard effect.

[0017] For example, image processing networks can be applied to image super-resolution. Image super-resolution is the process of restoring a low-resolution image or video to a high-resolution image or video.

[0018] For example, image processing networks can be applied to image restoration. Image restoration is the process of restoring an image or video that has blurred areas to an image or video that has sharp details in those areas.

[0019] For example, image processing networks can be applied to the encoding and decoding of images.

[0020] For example, when an image processing network is applied to the encoding and decoding of images, the bitrate points at which the checkerboard effect appears can be reduced. In other words, it can be seen that, compared to the prior art, there are fewer bitrate points at which the checkerboard effect appears in images obtained by encoding and decoding performed using an image processing network trained according to the training method of this application. In addition, at medium bitrates (for example, the bitrate may be 0.15 Bpp (bits per pixel) to 0.3 Bpp, and may be specifically set according to requirements), images obtained by encoding and decoding performed using an image processing network trained according to the training method of this application have higher quality.

[0021] For example, when an image processing network includes an upsampling layer, the period of the checkerboard effect can be determined based on the number of upsampling layers.

[0022] In relation to the first aspect, the step of determining a first loss based on M first image blocks and M second image blocks includes the steps of: obtaining a first feature block based on M first image blocks, wherein the characteristic value of the first feature block is obtained by calculation based on pixels at corresponding positions in the M first image blocks; obtaining a second feature block based on M second image blocks, wherein the characteristic value of the second feature block is obtained by calculation based on pixels at corresponding positions in the M second image blocks; and determining a first loss based on the first and second feature blocks. In this way, the information from the M first image blocks is aggregated into one or more first feature blocks, and the information from the M second image blocks is aggregated into one or more second feature blocks. Next, the first loss is calculated by comparing the first and second feature blocks in order to more deliberately compensate for the periodic checkerboard effect and, as a result, to implement a better effect in eliminating the checkerboard effect.

[0023] Please understand that the number of the first feature blocks and the number of the second feature blocks are not limited in this application.

[0024] In relation to the first aspect or any one embodiment of the first aspect, the step of obtaining a first feature block based on M first image blocks includes the step of obtaining a first feature block based on N first image blocks out of the M first image blocks, wherein the characteristic value of the first feature block is obtained by calculation based on pixels at corresponding positions in the N first image blocks, where N is a positive integer less than or equal to M. The step of obtaining a second feature block based on M second image blocks includes the step of obtaining a second feature block based on N second image blocks out of the M second image blocks, wherein the characteristic value of the second feature block is obtained by calculation based on pixels at corresponding positions in the N second image blocks. Thus, the calculation may be performed based on pixels at corresponding positions in some or all of the first image blocks in order to obtain a first feature block, and the calculation may be performed based on pixels at corresponding positions in some or all of the second image blocks in order to obtain a second feature block. When the first feature block is obtained by calculation based on pixels at corresponding positions in a portion of the first image block, and the second feature block is obtained by calculation based on pixels at corresponding positions in a portion of the second image block, less information is used to calculate the first loss, and as a result, the efficiency of calculating the first loss is improved. When the first feature block is obtained by calculation based on pixels at corresponding positions in all of the first image blocks, and the second feature block is obtained by calculation based on pixels at corresponding positions in all of the second image blocks, the information used to calculate the first loss is more comprehensive, and as a result, the accuracy of the first loss is improved.

[0025] In relation to the first aspect or any one embodiment of the first aspect, the step of obtaining a first feature block based on M first image blocks includes the step of performing a calculation based on first target pixels at corresponding first positions in all of the M first image blocks in order to obtain a characteristic value for a corresponding first position in the first feature block, wherein the number of first target pixels is less than or equal to the total number of pixels in the first image block, and one characteristic value is obtained correspondingly for the first target pixels at the same first position in the M first image blocks. The step of obtaining a second feature block based on M second image blocks includes the step of performing a calculation based on second target pixels at corresponding second positions in all of the M second image blocks in order to obtain a characteristic value for a corresponding second position in the second feature block, wherein the number of second target pixels is less than or equal to the total number of pixels in the second image block, and one characteristic value is obtained correspondingly for the second target pixels at the same second position in the M second image blocks. In this way, a feature block (first feature block / second feature block) can be calculated based on some or all pixels within an image block (first image block / second image block). When the calculation is performed based on some pixels within the image block, less information is used to calculate the first loss, resulting in improved efficiency. When the calculation is performed based on all pixels within the image block, the information used to calculate the first loss is more comprehensive, resulting in improved accuracy of the first loss.

[0026] For example, when N is less than M, the step of obtaining a first feature block based on N first image blocks out of M first image blocks may include the step of performing a calculation based on the first target pixels at the corresponding first positions in all of the N first image blocks in order to obtain a characteristic value for the corresponding first position in the first feature block, wherein one characteristic value is obtained correspondingly for the first target pixels at the same first position in the N first image blocks. The step of obtaining a second feature block based on N second image blocks out of M second image blocks may include the step of performing a calculation based on the second target pixels at the corresponding second positions in all of the N second image blocks in order to obtain a characteristic value for the corresponding second position in the second feature block, wherein one characteristic value is obtained correspondingly for the second target pixels at the same second position in the N second image blocks.

[0027] In relation to the first aspect or any one embodiment of the first aspect, the step of obtaining a first feature block based on M first image blocks includes the step of determining a characteristic value of a corresponding position in the first feature block based on the mean value of pixels at the corresponding position in all of the M first image blocks, wherein one characteristic value is obtained with respect to the mean value of pixels at the same position in the M first image blocks. The step of obtaining a second feature block based on M second image blocks includes the step of determining a characteristic value of a corresponding position in the second feature block based on the mean value of pixels at the corresponding position in all of the M second image blocks, wherein one characteristic value is obtained with respect to the mean value of pixels at the same position in the M second image blocks. In this way, the characteristic value of a feature block (first feature block / second feature block) is determined in a manner that calculates the mean value of pixels. The calculation is simple and as a result improves the efficiency of calculating the first loss.

[0028] In this application, it should be noted that in addition to the aforementioned method of calculating the pixel average value, another linear calculation method (for example, linear weighting) may be used to perform calculations on pixels to determine the characteristic value of the feature block. This is not limited in this application.

[0029] In this application, it should be noted that the calculation may alternatively be performed on pixels using a non-linear calculation method to determine the characteristic value of the feature block. For example, the non-linear calculation may be performed according to the following formula. F i =A1*e 1i 2 +A2*e 2i 2 +A3*e 3i 2 +…+AM*e Mi 2

[0030] Here, F i represents the characteristic value at the i-th position in the first feature map. e 1i represents the pixel at the i-th position in the first image block of the first one, and e 2i represents the pixel at the i-th position in the second image block of the first one, and e 3i represents the pixel at the i-th position in the third image block of the first one, …, and e Mi represents the pixel at the i-th position in the M-th image block of the first one. A1 represents the weight coefficient corresponding to the pixel at the i-th position in the first image block of the first one, A2 represents the weight coefficient corresponding to the pixel at the i-th position in the second image block of the first one, A3 represents the weight coefficient corresponding to the pixel at the i-th position in the third image block of the first one, …, AM represents the weight coefficient corresponding to the pixel at the i-th position in the M-th image block of the first one, and i is an integer in the range of 1 to M (i may be equal to 1 or M).

[0031] In this application, it should be noted that calculations may be performed on pixels based on convolutional layers to determine the characteristic values ​​of feature blocks. For example, M first image blocks may be input to convolutional layers (one or more layers) and fully connected layers to obtain first feature blocks output by fully connected layers, and M second image blocks may be input to convolutional layers (one or more layers) and fully connected layers to obtain second feature blocks output by fully connected layers.

[0032] It should be understood that alternative calculation methods may be used to perform calculations on pixels to determine the characteristic values ​​of feature blocks. This is not limited to this application.

[0033] For example, when N is less than M, the step of obtaining a first feature block based on N first image blocks out of M first image blocks may include the step of determining a characteristic value for a corresponding position in the first feature block based on the average value of pixels at corresponding positions in all of the N first image blocks, wherein one characteristic value is obtained with respect to the average value of pixels at the same position in the N first image blocks. The step of obtaining a second feature block based on N second image blocks out of M second image blocks may include the step of determining a characteristic value for a corresponding position in the second feature block based on the average value of pixels at corresponding positions in all of the N second image blocks, wherein one characteristic value is obtained with respect to the average value of pixels at the same position in the N second image blocks.

[0034] In relation to the first aspect or any one embodiment of the first aspect, the step of determining a first loss based on a first feature block and a second feature block includes the step of determining a first loss based on the point-to-point loss between the first feature block and the second feature block.

[0035] For example, point-based loss may include Ln distance (e.g., L1 distance (Manhattan distance), L2 distance (Euclidean distance), or L-Inf distance (Chebyshev distance)). This is not limited to the present application.

[0036] In relation to the first aspect or any one embodiment of the first aspect, the step of determining a first loss based on a first feature block and a second feature block includes the step of determining a first loss based on a feature-based loss between the first feature block and the second feature block.

[0037] For example, feature-based loss may include SSIM (Structural Similarity), MSSSIM (Multi-Scale Structural Similarity), and LPIPS (Learned Perceptual Image Patch Similarity), but is not limited to these.

[0038] For example, the first and second feature blocks may be further input into a neural network (e.g., a convolutional network or a VGG network (Visual Geometry Group Network)), and the neural network will output the first feature of the first feature block and the second feature of the second feature block. Next, the distance between the first and second features is calculated to obtain a feature-based loss.

[0039] In relation to the first aspect or any one embodiment of the first aspect, the image processing network includes an encoding network and a decoding network. Prior to the step of acquiring a first training image and a first prediction image, the method further includes the steps of acquiring a second training image and a second prediction image, the second prediction image being acquired by encoding the second training image based on an untrained encoding network and then decoding the encoded result of the second training image based on an untrained decoding network; determining a second loss based on the second prediction image and the second training image; and pre-training the untrained encoding network and the untrained decoding network based on the second loss. By pre-training the image processing network in this way, the image processing network can converge more quickly and better in the subsequent training process.

[0040] For example, to obtain a second loss, the bitrate loss and mean squared error loss may be determined based on the second predicted image and the second training image, and then the bitrate loss and mean squared error loss are weighted.

[0041] Bitrate loss indicates the size of the bitstream. Mean squared error loss can be the mean squared error between the second predicted image and the second training image, and can be used to improve objective metrics of the image (e.g., PSNR (Peak Signal to Noise Ratio)).

[0042] For example, the weighting coefficients corresponding to bitrate loss and mean squared error loss, respectively, may be the same or different. This is not limited to the present application.

[0043] According to the first embodiment or any one of the embodiments of the first embodiment, a first predicted image is obtained by encoding a first training image based on a pre-trained coding network, and then decoding the encoded result of the first training image based on a pre-trained decoding network. The step of training an image processing network based on a first loss is to determine a third loss based on the discrimination result obtained by the discrimination network with respect to the first training image and the first predicted image, wherein the third loss is a loss of a Generative Adversarial Network (GAN), and the GAN network includes a discrimination network and a decoding network, and the step of training a pre-trained coding network and a pre-trained decoding network based on the first loss and the third loss. In this way, the checkerboard effect can be better compensated for by referencing the first loss and the third loss in order to eliminate the checkerboard effect to a greater extent.

[0044] For example, the decoding network may include an upsampling layer, and the period of the checkerboard effect may be determined based on the upsampling layer within the decoding network.

[0045] In relation to the first aspect or any one embodiment of the first aspect, the step of training an image processing network based on a first loss further includes the step of determining a fourth loss, wherein the fourth loss includes at least one of the following: namely, L1 loss, bitrate loss, perceptual loss, and edge loss. or The step of training a pre-trained coding network and a pre-trained decoding network based on a third loss includes the step of training a pre-trained coding network and a pre-trained decoding network based on a first loss, a third loss, and a fourth loss.

[0046] L1 loss can be used to improve objective metrics of an image (e.g., PSNR (Peak Signal to Noise Ratio)). Bitrate loss indicates the size of the bitstream. Perceptual loss can be used to improve the visual effect of an image. Edge loss can be used to prevent edge distortion. In this way, by training an image processing network by referencing multiple losses, the quality of images processed by the trained image processing network (including objective quality and subjective quality, which may also be called visual quality) can be improved.

[0047] For example, the weighting coefficients corresponding to the first loss, third loss, L1 loss, bitrate loss, perceived loss, and edge loss may be the same or different. This is not limited to the present application.

[0048] In relation to the first aspect or any one embodiment of the first aspect, the image processing network further includes a hyperplier coding network and a hyperplier decoding network, the hyperplier decoding network and the decoding network each including an upsampling layer. The period includes a first period and a second period. The first period is determined based on the number of upsampling layers in the decoding network. The second period is determined based on the first period and the number of upsampling layers in the hyperplier decoding network. The second period is greater than the first period.

[0049] In relation to the first aspect or any one embodiment of the first aspect, the step of dividing a first training image into M first image blocks and a first prediction image into M second image blocks based on a period includes the step of dividing a first training image into M first image blocks based on a first period and a second period, wherein the M first image blocks include M1 third image blocks and M2 fourth image blocks, the size of the third image block is related to the first period and the size of the fourth image block is related to the second period, M1 and M2 are positive integers and M1 + M2 = M, and the step of dividing a first prediction image into M second image blocks based on a first period and a second period, wherein the M second image blocks include M1 fifth image blocks and M2 sixth image blocks, the size of the fifth image block is related to the first period and the size of the sixth image block is related to the second period. When the first loss includes the fifth loss and the sixth loss, the step of determining the first loss based on M first image blocks and M second image blocks includes the step of determining the fifth loss based on M1 third image blocks and M1 fifth image blocks, and the step of determining the sixth loss based on M2 fourth image blocks and M2 sixth image blocks.

[0050] In relation to the first aspect or any one embodiment of the first aspect, the step of training an image processing network based on a first loss includes the steps of performing weighting calculations on the fifth and sixth losses to obtain a seventh loss, and training a pre-trained coding network and a pre-trained decoding network based on the seventh loss.

[0051] When an image processing network includes an encoding network and a decoding network, and further includes a hyperplier encoding network and a hyperplier decoding network, the hyperplier decoding network causes a checkerboard effect with longer periods. Therefore, to compensate for this longer-period checkerboard effect, the loss is determined after the image is divided into blocks with longer periods. In this way, the longer-period checkerboard effect is removed to some extent, resulting in further improvement of image quality.

[0052] For example, the weighting coefficients corresponding to the fifth loss and the sixth loss, respectively, may be the same or different. This is not limited to the present application.

[0053] For example, the weighting coefficient corresponding to the fifth loss may be greater than the weighting coefficient corresponding to the sixth loss.

[0054] It should be understood that the period of the checkerboard effect can include many more periods distinct from the first and second periods. For example, the period of the checkerboard effect may include k periods (where k periods can be the first period, the second period, ..., and the kth period, and k is an integer greater than 2). In this way, the first training image can be divided into M first image blocks based on the first period, the second period, ..., and the kth period. The M first image blocks may include M1 image blocks 11, M2 image blocks 12, ..., and Mk image blocks 1k. The size of image block 11 relates to the first period, the size of image block 12 relates to the second period, ..., and the size of image block 1k relates to the kth period, with M1 + M2 + ... + Mk = M. The first predicted image can be divided into M second image blocks based on the first period, the second period, ..., and the kth period. The M second image blocks may include M1 image blocks 21, M2 image blocks 22, ..., and Mk image blocks 2k. The size of image block 21 relates to the first period, the size of image block 22 relates to the second period, ..., and the size of image block 2k relates to the kth period. Then, Loss 1 may be determined based on M1 image blocks 11 and M1 image blocks 21, Loss 2 may be determined based on M2 image blocks 12 and M2 image blocks 22, ..., and Lossk may be determined based on Mk image blocks 1k and Mk image blocks 2k. Subsequently, a pre-trained coding network and a pre-trained decoding network may be trained based on Loss 1, Loss 2, ..., and Lossk.

[0055] According to a second aspect, one embodiment of the present application provides an encoding method. The method includes the steps of: acquiring an image to be encoded; then inputting the image to be encoded into the encoding network to acquire a feature map output by the encoding network and processing the image to be encoded by the encoding network; and performing entropy coding on the feature map to acquire a first bitstream. The encoding network is acquired by training performed using either the first aspect or one embodiment of the first aspect. Correspondingly, a decoder decodes the first bitstream using a decoding network acquired by training performed using either the first aspect or one embodiment of the first aspect. Furthermore, the present application allows the same image to be encoded at a lower bitrate than the prior art, provided that it is guaranteed that the reconstructed image does not exhibit a checkerboard effect. In addition, when the same image is encoded using the same bitrate (e.g., a medium bitrate), the encoding quality in the present application is higher than that in the prior art.

[0056] In relation to the second aspect, the step of performing entropy coding on a feature map to obtain a first bitstream includes the steps of inputting the feature map into a hyperplia coding network to obtain hyperplia features and processing the feature map by the hyperplia coding network, inputting the hyperplia features into a hyperplia decoding network and processing the hyperplia features by the hyperplia decoding network and then outputting a probability distribution, and performing entropy coding on the feature map based on the probability distribution to obtain a first bitstream. Both the hyperplia coding network and the hyperplia decoding network are acquired by training performed using the first aspect and any one embodiment of the first aspect. Hyperplia coding networks and hyperplia decoding networks trained using prior art training methods result in a checkerboard effect with a larger period. In comparison, the present application can avoid a checkerboard effect with a larger period on the reconstructed image.

[0057] In relation to the second aspect or any one embodiment of the second aspect, entropy coding is performed on the hyperplier features to obtain a second bitstream. Thus, after subsequently receiving the first bitstream and the second bitstream, in order to assist the decoder in performing decoding, the decoder may determine a probability distribution based on the hyperplier features obtained by decoding the second bitstream, and then decode and reconstruct the first bitstream based on the probability distribution to obtain a reconstructed image.

[0058] According to a third aspect, one embodiment of the present application provides a decoding method. The method includes the steps of: acquiring a first bitstream, the first bitstream being a bitstream of feature maps; then performing entropy decoding on the first bitstream to acquire feature maps; and inputting the feature maps into the decoding network and processing the feature maps by the decoding network to acquire a reconstructed image output by the decoding network. The decoding network is acquired by training performed using either the first aspect or one embodiment of the first aspect. Correspondingly, the first bitstream is acquired by encoding performed by an encoder based on an encoding network acquired by training performed using either the first aspect or one embodiment of the first aspect. Furthermore, when the bitrate of the first bitstream is lower than the bitrate of the prior art, the checkerboard effect does not appear in the reconstructed image in the present application. In addition, when the bitrate of the first bitstream is the same as the bitrate of the prior art (e.g., a medium bitrate), the quality of the reconstructed image in the present application is higher.

[0059] According to a third aspect, the method further includes the step of acquiring a second bitstream, wherein the second bitstream is a bitstream of hyperpliamentary features. The step of performing entropy decoding on the bitstream to acquire a feature map includes the step of performing entropy decoding on the second bitstream to acquire hyperpliamentary features, the step of inputting the hyperpliamentary features into a hyperpliamentary decoding network and processing the hyperpliamentary features by the hyperpliamentary decoding network to acquire a probability distribution, and the step of performing entropy decoding on the first bitstream based on the probability distribution to acquire a feature map. In this way, the decoder can directly decode the bitstream to acquire a probability distribution without recalculation, thereby improving decoding efficiency. In addition, the hyperpliamentary decoding network is acquired by training performed using the first aspect and any one of the embodiments of the first aspect. A hyperpliamentary decoding network trained using a prior art training method results in a checkerboard effect with a larger period. In comparison, the present application can avoid a checkerboard effect with a larger period on the reconstructed image.

[0060] According to a fourth aspect, one embodiment of the present application provides an electronic device including a memory and a processor. The memory is coupled to the processor. The memory stores program instructions. When the program instructions are executed by the processor, the electronic device is enabled to perform a training method according to the first aspect or any one of the possible embodiments of the first aspect.

[0061] Each of the fourth aspect and any one of its embodiments corresponds to either the first aspect or any one of its embodiments. For technical effects corresponding to each of the fourth aspect and any one of its embodiments, please refer to the technical effects corresponding to either the first aspect or any one of its embodiments. Further details will not be provided here.

[0062] According to a fifth aspect, one embodiment of the present application provides a chip including one or more interface circuits and one or more processors. The interface circuit is configured to receive signals from the memory of an electronic device and transmit signals to a processor, the signals including computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device is enabled to perform a training method according to the first aspect or any one of the possible embodiments of the first aspect.

[0063] Each of the fifth aspect and any one of its embodiments corresponds to either the first aspect or any one of its embodiments. For technical effects corresponding to each of the fifth aspect and any one of its embodiments, please refer to the technical effects corresponding to either the first aspect or any one of its embodiments. Further details will not be provided here.

[0064] According to a sixth aspect, one embodiment of the present application provides a computer-readable storage medium for storing a computer program. When the computer program is run on a computer or processor, the computer or processor is enabled to perform a training method according to the first aspect or any one of possible embodiments thereof.

[0065] Each of the sixth aspect and any one of its embodiments corresponds to either the first aspect or any one of its embodiments. For technical effects corresponding to each of the sixth aspect and any one of its embodiments, please refer to the technical effects corresponding to either the first aspect or any one of its embodiments. Further details will not be provided here.

[0066] According to a seventh aspect, one embodiment of the present application provides a computer program product, the computer program product including a software program, when the software program is executed by a computer or processor, the computer or processor is enabled to perform a training method in any one of the first aspects or possible embodiments thereof.

[0067] Each of the seventh aspect and any one of its embodiments corresponds to either the first aspect or any one of its embodiments. For technical effects corresponding to each of the seventh aspect and any one of its embodiments, please refer to the technical effects corresponding to either the first aspect or any one of its embodiments. Further details will not be provided here.

[0068] According to the eighth aspect, one embodiment of the present application provides a bitstream storage device. The device includes a receiver and at least one storage medium. The receiver is configured to receive a bitstream. At least one storage medium is configured to store a bitstream. The bitstream is generated according to either the second aspect or one embodiment of the second aspect.

[0069] Each of the eighth aspect and any one of its embodiments corresponds to either the second aspect or any one of its embodiments. For technical effects corresponding to each of the eighth aspect and any one of its embodiments, please refer to the technical effects corresponding to either the second aspect or any one of its embodiments. Further details will not be provided here.

[0070] According to the ninth aspect, one embodiment of the present application provides a bitstream transmission device. The device includes a transmitter and at least one storage medium. The at least one storage medium is configured to store a bitstream. The bitstream is generated according to either the second aspect or one embodiment of the second aspect. The transmitter is configured to retrieve the bitstream from the storage medium and to transmit the bitstream to a terminal device using the transmitter.

[0071] Each of the ninth aspect and any one of its embodiments corresponds to either the second aspect or any one of its embodiments. For technical effects corresponding to each of the ninth aspect and any one of its embodiments, please refer to the technical effects corresponding to each of the second aspect and any one of its embodiments. Further details will not be provided here.

[0072] According to a tenth aspect, one embodiment of the present application provides a bitstream distribution system. The system includes at least one storage medium configured to store at least one bitstream, the at least one bitstream being generated according to any one of the second aspect and embodiments of the second aspect, and a streaming media device configured to retrieve a target bitstream from the at least one storage medium and transmit the target bitstream to a terminal-side device, the streaming media device including a content server or a content distribution server.

[0073] Each of the tenth aspect and any one of its embodiments corresponds to either the second aspect or any one of its embodiments. For technical effects corresponding to each of the tenth aspect and any one of its embodiments, please refer to the technical effects corresponding to each of the second aspect and any one of its embodiments. Further details will not be provided here. [Brief explanation of the drawing]

[0074] [Figure 1a] This is a diagram illustrating an example of a main framework for artificial intelligence. [Figure 1b] This is a diagram illustrating an example of an application scenario. [Figure 1c] This is a diagram illustrating an example of an application scenario. [Figure 1d] This is a diagram illustrating an example of an application scenario. [Figure 1e] This is an example of the checkerboard effect. [Figure 2a] This is a diagram illustrating an example of the training process. [Figure 2b] This is a diagram illustrating an example of the deconvolution process. [Figure 3a] This is a diagram illustrating an example of an end-to-end image compression framework. [Figure 3b] This is a diagram illustrating an example of an end-to-end image compression framework. [Figure 3c(1)] This is a diagram illustrating an example of a network structure. [Figure 3c(2)] This is a diagram illustrating an example of a network structure. [Figure 4a] This is a diagram illustrating an example of the training process. [Figure 4b] This is a diagram illustrating an example of the chunking process. [Figure 4c] This is a diagram illustrating an example of the process for generating the first feature block. [Figure 5] This is a diagram illustrating an example of the training process. [Figure 6] This is a diagram illustrating an example of the coding process. [Figure 7] This is a diagram illustrating an example of the decryption process. [Figure 8a]This is an example diagram showing the results of an image quality comparison. [Figure 8b] This is an example diagram showing the results of an image quality comparison. [Figure 8c] This is an example diagram showing the results of an image quality comparison. [Figure 9] This is a diagram illustrating an example of the device's structure. [Modes for carrying out the invention]

[0075] The following describes the technical solutions in the embodiments of this application with reference to the accompanying drawings. Certainly It is clear that the embodiments described are only a part of, and not all, of the embodiments of this application. All other embodiments that can be obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0076] In this specification, the term "and / or" describes only the relational relationship used to describe the related objects, indicating that three relationships may exist. For example, A and / or B may represent the following three cases: A exists alone, both A and B exist, and B exists alone.

[0077] In the specification and claims of the embodiments of this application, terms such as "first" and "second" are intended to distinguish different subjects, but not to indicate a specific order of subjects. For example, "first target subject" and "second target subject" are used to distinguish different target subjects, but not to describe a specific order of subjects.

[0078] In the embodiments of this application, words such as “example” and “for example” are used to indicate that an example, illustration, or explanation is being given. No embodiment or design described as “example” or “for example” in the embodiments of this application is described as being more preferable or having more advantages than another embodiment or design. More precisely, the use of words such as “example” and “for example” is intended to present relative concepts in a specific way.

[0079] In the description of embodiments of this application, unless otherwise stated, “multiple” means two or more. For example, “multiple processing units” means two or more processing units, and “multiple systems” means two or more systems.

[0080] Figure 1a is a diagram illustrating an example of an artificial intelligence main framework. The main framework describes the overall operating procedure of an artificial intelligence system and is applicable to the requirements of the general field of artificial intelligence.

[0081] The following explains the aforementioned main framework of artificial intelligence from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis).

[0082] The "intelligent information chain" reflects a series of processes from data acquisition to data processing. For example, the process may be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a "data-information-knowledge-intelligence" refinement process.

[0083] The "IT value chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure and artificial intelligence information (technology provision and processing implementation) to the industrial environment processes of the system.

[0084] (1) Infrastructure The infrastructure provides computing power support to the artificial intelligence system, enables communication with the outside world, and provides support using the underlying platform. The infrastructure communicates with the outside world using sensors. Computing power is provided by smart chips (hardware acceleration chips such as CPUs, NPUs, GPUs, ASICs, or FPGAs). The underlying platform includes related platform assurance and support such as distributed computing frameworks and networks, and may include cloud storage and computing, as well as interconnection and interworking networks. For example, sensors communicate with the outside world to acquire data, and the data is provided to intelligent chips in a distributed computing system, which are provided for computing by the underlying platform.

[0085] (2) Data The higher layers of infrastructure data represent data sources in the field of artificial intelligence. This data includes graphs, images, audio, and text, as well as data from the Internet of Things on traditional devices, service data from existing systems, and perceptual data such as force, displacement, liquid level, temperature, and humidity.

[0086] (3) Data processing Data processing typically includes data training, machine learning, deep learning, search, inference, and decision-making.

[0087] Machine learning and deep learning can be interpreted as performing symbolic and formal intelligent information modeling, extraction, preprocessing, and training on data.

[0088] Reasoning is the process by which human intellectual reasoning is simulated in a computer or intelligent system, and machine thinking and problem solving are performed using formal information according to reasoning control policies. Typical functions include retrieval and matching.

[0089] Decision-making is the process of making decisions after intelligent information has been inferred, and typically involves functions such as classification, ranking, and prediction.

[0090] (4) General abilities After the data processing described above has been performed on the data, several general capabilities, such as algorithms or general systems for translation, text analysis, computer vision processing, speech recognition, image recognition, and texture mapping generation, may be further formed based on the data processing results.

[0091] (5) Intelligent products and industrial applications Smart products and industrial applications are products and applications of artificial intelligence systems in various fields, representing a package of comprehensive AI solutions where intelligent information-based decision-making is commercialized and implemented. Application areas primarily include smart manufacturing, smart transportation, smart homes, smart healthcare, smart security, autonomous driving, and smart devices.

[0092] The image processing networks described in this application may be used to perform machine learning, deep learning, search, inference, and decision-making, among other things. The image processing networks described in this application may include, but are not limited to, several types of neural networks, such as deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), residual networks, neural networks using transformer models, or other neural networks.

[0093] The operation of each layer of a neural network is described by mathematical formulas.

number

number

[0094] The goal of training a neural network is to ultimately obtain a weight matrix (a weight matrix formed by vectors W across multiple layers) for all layers of the trained neural network. Therefore, the neural network training process is essentially a method of learning to control spatial transformations, and more specifically, a method of learning the weight matrix.

[0095] For example, according to the training method provided in this application, all image processing networks that can be used for image processing, and all image processing networks in which a checkerboard effect appears in the image obtained after image processing, can be trained.

[0096] Figure 1b is an example of an application scenario. For example, the image processing network in Figure 1b is an image super-resolution network, and the application scenario corresponding to Figure 1b is training the image super-resolution network. The image super-resolution network can be used for image super-resolution, that is, to restore a low-resolution image or video to a high-resolution image or video.

[0097] Referring to Figure 1b, for example, Image 1 is input to the image super-resolution network, which processes Image 1 to output Image 2, where the resolution of Image 2 is greater than that of Image 1. Next, a Loss can be determined based on Images 1 and 2, and then the image super-resolution network is trained based on the Loss. The specific training process will be described later.

[0098] Figure 1c is an example of an application scenario. For example, the image processing network in Figure 1c is an image restoration network, and the application scenario corresponding to Figure 1c is training the image restoration network. The image restoration network can be used for image restoration, that is, to restore an image or video with blurred subregions to an image or video with sharp details in those subregions.

[0099] Referring to Figure 1c, for example, Image 1 (with a blurred portion within Image 1) is input to the image restoration network, and Image 2 is output after the restoration performed by the image restoration network. Next, a Loss can be determined based on Images 1 and 2, and then the image restoration network is trained based on the Loss. The specific training process will be described later.

[0100] Figure 1d is a diagram illustrating an example application scenario. The image processing network in Figure 1d includes an encoding network and a decoding network, and the application scenario corresponding to Figure 1d is training the encoding network and the decoding network. Images / videos can be encoded based on the encoding network, and then the encoded images / videos are decoded based on the decoding network.

[0101] Referring to Figure 1d, for example, Image 1 (i.e., the original image) is input to the coding network, which transforms Image 1 and outputs Feature Map 1 to the quantization module. Next, the quantization module may quantize Feature Map 1 to obtain Feature Map 2. Subsequently, the quantization module may input Feature Map 2 to the entropy estimation network, which performs entropy estimation and outputs the entropy estimation information of the feature points included in Feature Map 2 to the entropy coding module and the entropy decoding module. The quantization module may input Feature Map 2 to the entropy coding module, which performs entropy coding on the feature points included in Feature Map 2 based on the entropy estimation information of the feature points included in Feature Map 2 to obtain a bitstream, and then inputs the bitstream to the entropy decoding module. Next, the entropy decoding module performs entropy decoding on the bitstream based on the entropy estimation information of all feature points included in Feature Map 2 and outputs Feature Map 2 to the decoding network. Next, the decoding network can transform feature map 2 to obtain image 2 (i.e., the reconstructed image). Subsequently, Loss1 may be determined based on images 1 and 2, and Loss2 may be determined based on the entropy estimation information determined by the entropy estimation network. The coding network and decoding network are trained based on Loss1 and Loss2. The specific training process will be described later.

[0102] The following uses Figure 1d as an example to illustrate the training process of an image processing network.

[0103] Figure 1e is an example of the checkerboard effect. For example, (1) in Figure 1e is the original image, and (2) in Figure 1e is the original image. dThis is a reconstructed image obtained by encoding and decoding the original image (1) in Figure 1e according to the encoding and decoding process. Referring to Figure 1e, a checkerboard effect appears in the reconstructed image (the white box in (2) in Figure 1e represents one period of the checkerboard effect), and it can be seen that the checkerboard effect is periodic.

[0104] Based on this, in the process of training the image processing network in this application, the images before and after processing performed using the image processing network may be divided into multiple image blocks based on the period of the checkerboard effect. Next, the loss is calculated by comparing the difference between the image blocks before and after processing performed using the image processing network. The image processing network is trained based on the loss to compensate the image blocks after processing performed using the image processing network to a different degree, and as a result, to remove the regularity of the checkerboard effect in the image processed using the image processing network. In this way, the visual quality of the image obtained by processing performed using the trained image processing network can be improved. A specific training process for the image processing network may be as follows: Figure 2a is a diagram illustrating an example of the training process.

[0105] S201: The first training image and the first prediction image are obtained, the period of the checkerboard effect is obtained, and the first prediction image is generated by performing image processing on the first training image based on the image processing network.

[0106] For example, multiple images may be acquired to be used to train an image processing network. For ease of explanation, the images used to train the image processing network may be referred to as the first training images. In this application, an example in which an image processing network is trained using one first training image is used for illustrative purposes.

[0107] For example, a first training image may be input to an image processing network, which then performs forward computation (i.e., image processing) to output a first predicted image.

[0108] For example, when the image processing network is an image super-resolution network, the image processing performed using the image processing network is image super-resolution, and the resolution of the first predicted image is higher than the resolution of the first training image. When the image processing network is an image restoration network, the image processing performed using the image processing network is image restoration, and the first training image is a partially blurred image. When the image processing network includes an encoding network and a decoding network, the image processing performed using the image processing network may include encoding and decoding. Specifically, the first training image may be encoded based on the encoding network, and then the encoded result of the first training image is decoded based on the decoding network. The first training image is the image to be encoded, and the first predicted image is the reconstructed image.

[0109] In possible ways, the first predicted image output by the image processing network has a checkerboard effect, and the period of the checkerboard effect can be determined by analyzing the first predicted image.

[0110] In possible ways, image processing networks include an upsampling layer (e.g., a decoding network includes an upsampling layer). The image processing network performs an upsampling operation (e.g., deconvolution) in the image processing process, and this type of operation results in different computational modes for adjacent pixels. Thus, ultimately, an intensity / color difference occurs between adjacent pixels, resulting in a periodic checkerboard effect. The following uses one-dimensional deconvolution as an example for explanation.

[0111] Figure 2b is an example of a deconvolution process.

[0112] Referring to Figure 2b, for example, suppose the inputs are a, b, c, d, and e (wherein "0" in Figure 2b is not an input but is interpolated by a zero-padding operation in the deconvolution process), and the convolution kernel used for deconvolution is assumed to be a 1*3 matrix (w1, w2, w3). Deconvolution is performed on the convolution kernel and inputs, and the output obtained is x1, y1, x2, y2, x3, and y3. Here, x1, y1, x2, y2, x3, and y3 are calculated as follows: x1=w2×b, y1=w1×b+w3×c, x2=w2×c, y2=w1×c+w3×d, x3=w2×d, and y3=w1×d+w3×e.

[0113] Referring to Figure 2b, since x and y are calculated in different ways in the calculation process, the modes (intensity / color, etc.) represented by x1, x2, and x3 are also different from the modes represented by y1, y2, and y3. As a result, a periodic checkerboard effect appears in the result. It should be understood that the same is true for 2D deconvolution. Details are not explained here.

[0114] In addition, by testing the period of the checkerboard effect and the number of upsampling layers, it can be seen that the period of the checkerboard effect is related to the number of upsampling layers. Furthermore, in possible ways, the period of the checkerboard effect can be determined based on the number of upsampling layers included in the image processing network.

[0115] For example, the period of the checkerboard effect is a two-dimensional period. Based on Figure 1e, it can be seen that the checkerboard effect is two-dimensional. Therefore, the two-dimensional period can be shown to indicate that the period of the checkerboard effect contains two-dimensional values, where the two dimensions correspond to the length and width of the image.

[0116] In possible ways, the shape of the period of the checkerboard effect may be rectangular, the size of the period of the checkerboard effect may be expressed as p*q, where p and q are positive integers, the units of p and q are px (pixels), and p and q may be equal or unequal. This is not limited to the present application. The relationship between the number of upsampling layers included in the image processing network and the period of the checkerboard effect may be as follows: T checkboard =p*q=2 C *2 C

[0117] Here, T checkboard is the period of the checkerboard effect, C (where C is a positive integer) is the number of upsampling layers, and p=q=2 C That is the case.

[0118] Please understand that p and q do not necessarily have to be equal. This is not limited to this application.

[0119] It should be noted that the period of the checkerboard effect may also take on a different shape (e.g., circular, triangular, elliptical, or irregular). Further details will not be provided here.

[0120] S202: Based on the period, the first training image is divided into M first image blocks, and the first prediction image is divided into M second image blocks, where both the size of the first and second image blocks are related to the period of the checkerboard effect, and M is an integer greater than 1.

[0121] For example, the number of periods in the checkerboard effect can be M.

[0122] For example, after the period of the checkerboard effect is determined, the first training image may be divided into M first image blocks based on the period of the checkerboard effect, and the first predicted image may be divided into M first images based on the period of the checkerboard effect. 2 It can be divided into image blocks.

[0123] It should be noted that the sizes of the M second image blocks may be the same or different. This is not limited to this application. The size of each second image block may be greater than, equal to, or less than the period of the checkerboard effect. For example, regardless of the relationship between the size of the second image block and the period of the checkerboard effect, it is only necessary that the second image block has only one complete or incomplete period of the checkerboard effect.

[0124] It should be noted that the sizes of the M first image blocks may be the same or different. This is not limited to this application. The size of each first image block may be greater than, equal to, or less than the period of the checkerboard effect. For example, regardless of the relationship between the size of the first image block and the period of the checkerboard effect, it is only necessary to ensure that the regions in the first predicted image that are at the same regional location of the first image block have only one complete or incomplete period of the checkerboard effect.

[0125] For example, the sizes of the first image block and the second image block may or may not be equal. This is not limited to this application.

[0126] S203: Determine the first loss based on M first image blocks and M second image blocks.

[0127] For example, the first loss may be determined by comparing the differences (or similarities) between M first image blocks and M second image blocks, and the first loss is used to compensate for the checkerboard effect of the images.

[0128] In possible ways, M first image blocks may be merged into one or more first feature blocks, and M second image blocks may be merged into one or more second feature blocks, such that both the number of first and second feature blocks are less than M. Next, the difference (or similarity) between the first and second feature blocks is determined by comparing the first and second feature blocks. Subsequently, the first loss may be determined based on the difference (or similarity) between the first and second feature blocks.

[0129] S204: Train the image processing network based on the first loss.

[0130] For example, an image processing network can be trained (i.e., backpropagated) based on a first loss to adjust the network parameters of the image processing network.

[0131] Furthermore, the image processing network may be trained based on S201-S204 using a plurality of first training images and a plurality of first prediction images until the image processing network satisfies a first pre-set condition. The first pre-set condition is a condition for stopping the training of the image processing network and may be set according to requirements. For example, reaching a pre-set number of training iterations or the loss becoming less than a pre-set loss. This is not limited to the present application.

[0132] Because the checkerboard effect is periodic, in this application, images before and after processing performed using an image processing network (the first training image is the image before processing, and the first predicted image is the image after processing) are divided into image blocks based on the period of the checkerboard effect. Next, the loss is calculated by comparing the difference between the image blocks before and after processing performed using the image processing network (i.e., by comparing M first image blocks with M second image blocks). The image processing network is trained based on the loss to effectively compensate each image block processed using the image processing network (i.e., the second image block), and as a result, reduce the difference between each first image block and the corresponding second image block. Both the size of the first image blocks and the size of the second image blocks are related to the period of the checkerboard effect. As the difference between each first image block and the corresponding second image block decreases, the checkerboard effect in each period is also eliminated to some extent. In this way, after the image processing network is trained based on the training method of this application, the checkerboard effect of images acquired by processing performed using the trained image processing network can be removed to some extent, and as a result, the visual quality of images acquired by processing performed using the trained image processing network can be improved.

[0133] Figure 3a is an example of an end-to-end image compression framework.

[0134] Referring to Figure 3a, for example, the end-to-end image compression framework includes an encoding network, a decoding network, a hyperplier encoding network, a hyperplier decoding network, an entropy estimation module, a quantization module (including quantization module A1 and quantization module A2), an entropy encoding module (including entropy encoding module A1 and entropy encoding module A2), and an entropy decoding module (including entropy decoding module A1 and entropy decoding module A2). The image processing network includes an encoding network, a decoding network, a hyperplier encoding network, and a hyperplier decoding network.

[0135] For example, the entropy estimation network in Figure 1d may include the entropy estimation module, hyperplier coding network, and hyperplier decoding network in Figure 3a.

[0136] For example, the hyperplier coding network and hyperplier decoding network are configured to generate probability distributions, and the entropy estimation module is configured to perform entropy estimation based on the probability distributions in order to generate entropy estimation information.

[0137] Figure 3b is a diagram of an example of an end-to-end image compression framework. In Figure 3b, a discriminator network (also called a discriminator) is added based on Figure 3a. In this case, the decoding network may also be called a generator network (also called a generator), and the discriminator network and decoding network may form a generative adversarial network (GAN network).

[0138] For example, the image processing network may first be pre-trained based on the end-to-end image compression framework shown in Figure 3a. Specifically, an untrained coding network, an untrained decoding network, an untrained hyperplier coding network, and an untrained hyperplier decoding network are pre-trained together. Next, after the pre-training of the image processing network shown in Figure 3a is completed, the image processing network is trained based on the end-to-end image compression framework of Figure 3b, which is based on Figure 3a. Specifically, a pre-trained coding network, a pre-trained decoding network, a pre-trained hyperplier coding network, and a pre-trained hyperplier decoding network are trained together. Note that the discriminator network in Figure 3b may be a trained discriminator network or an untrained discriminator network (in which case the discriminator network may be trained first, and the image processing network may be trained after the discriminator network has been trained). This is not limited to the present application.

[0139] By pre-training the image processing network, the subsequent training process can lead to faster convergence of the image processing network and higher encoding quality for the trained network.

[0140] It should be noted that pre-training of the image processing network may be skipped, and the image processing network is trained directly based on the end-to-end image compression framework shown in Figure 3b. This is not limited to this application. This application is illustrated using examples in which the image processing network is pre-trained and trained.

[0141] Figures 3c(1) and 3c(2) illustrate examples of network structures. The embodiments of Figures 3c(1) and 3c(2) show the network structures of an encoding network, a decoding network, a hyperplier encoding network, a hyperplier decoding network, and an identification network.

[0142] Referring to Figure 3c(1), for example, the coding network may include convolutional layers A1, A2, A3, and A4. Please understand that Figure 3c(1) is merely an example of a coding network in this application. The coding network in this application may include more or fewer convolutional layers than those in Figure 3c(1), or other network layers.

[0143] For example, the size of the convolution kernels for convolutional layers A1, A2, A3, and A4 may be 5*5, and the convolution step is 2. (The convolution kernel may also be called a convolution operator. A convolution operator may essentially be a weight matrix. Weight matrices are usually predefined. Image processing is used as an example. Different weight matrices are used to extract different features in an image. For example, one weight matrix is ​​used to extract image edge information, another to extract a particular color in an image, and yet another to blur unwanted noise in an image. The weight values ​​for these weight matrices need to be obtained by extensive training in a real application. Each weight matrix formed using the weight values ​​obtained by training can be used to extract information from the input data, thereby enabling the network to make correct predictions.) It should be understood that the size of the convolution kernels and the convolution step for convolutional layers A1, A2, A3, and A4 are not limited in this application. Note that the convolution steps for convolutional layers A1, A2, A3, and A4 are 2, which indicates that convolutional layers A1, A2, A3, and A4 perform a downsampling operation simultaneously with the convolution operation.

[0144] Referring to Figure 3c(1), for example, the decoding network may include upsampling layers D1, D2, D3, and D4. Please understand that Figure 3c(1) is merely an example of a decoding network in this application. The decoding network in this application may include more or fewer upsampling layers than those in Figure 3c(1), or other network layers.

[0145] For example, upsampling layers D1, D2, D3, and D4 may each include a deconvolution unit. The size of the convolution kernel of the deconvolution unit included in upsampling layers D1, D2, D3, and D4 may be 5*5, and the number of convolution steps is 2. It should be understood that the size of the convolution kernel and the number of convolution steps of the deconvolution unit included in upsampling layers D1, D2, D3, and D4 are not limited in this application.

[0146] Referring to Figure 3c(1), for example, the hyperpliamentary coding network may include convolutional layers B1, B2, and B3. Please understand that Figure 3c(1) is merely an example of a hyperpliamentary coding network in this application. The hyperpliamentary coding network in this application may include more or fewer convolutional layers than those in Figure 3c(1), or other network layers.

[0147] For example, convolutional layers B1 and B2 may each include a convolution unit and an activation unit (which may be a Leaky ReLU (Leaky Normalized Linear Unit)). The size of the convolution kernel of the convolution unit may be 5*5, and the number of convolution steps may be 2. The size of the convolution kernel of convolutional layer B3 may be 3*3, and the number of convolution steps may be 1. It should be understood that the size of the convolution kernel and the number of convolution steps of the convolution units of convolutional layers B1 and B2, as well as the size of the convolution kernel and the number of convolution steps of convolutional layer B3, are not limited in this application. Note that the number of convolution steps of convolutional layer B3 is 1, which indicates that convolutional layer B3 performs only a convolution operation and does not perform a downsampling operation.

[0148] Referring to Figure 3c(1), for example, the hyperplia decoding network may include an upsampling layer C1, an upsampling layer C2, and a convolutional layer C3. Please understand that Figure 3c(1) is merely an example of a hyperplia decoding network in this application. The hyperplia decoding network in this application may include more or fewer upsampling layers than those in Figure 3c(1), or other network layers.

[0149] For example, upsampling layer C1 and upsampling layer C2 may each include an inverse convolution unit and an activation unit (which may be a Leaky ReLU (Leaky Normalized Linear Unit)). The size of the convolution kernel of the inverse convolution unit included in upsampling layer C1 and upsampling layer C2 may be 5*5, and the number of convolution steps may be 2. The size of the convolution kernel of convolution layer C3 may be 3*3, and the number of convolution steps may be 1. It should be understood that the size and number of convolution steps of the convolution kernel of the inverse convolution unit included in upsampling layer C1 and upsampling layer C2, as well as the size and number of convolution steps of the convolution kernel of convolution layer C3, are not limited in this application.

[0150] Referring to Figure 3c(2), for example, the identification network may include convolutional layers E1, E2, E3, E4, E5, E6, E7, E8, and E9. Please understand that Figure 3c(2) is merely an example of an identification network in this application. The identification network in this application may include more or fewer convolutional layers than those in Figure 3c(2), or other network layers.

[0151] For example, convolutional layer E1 may include convolutional units and activation units (which may be Leaky ReLU (Leaky Normalized Linear Units)). The size of the convolution kernel of the convolutional unit in convolutional layer E1 may be 3*3, and the number of convolution steps may be 1. For example, convolutional layers E2, E3, E4, E5, E6, E7, and E8 may each include convolutional units, activation units (Leaky ReLU (Leaky Normalized Linear Units)), and BN (Batch Normalization) units. The size of the convolution kernel of the convolutional units included in convolutional layers E2, E3, E4, E5, E6, E7, and E8 may be 3*3. The convolution steps of the convolutional units included in convolutional layers E3, E5, and E7 may be 1. The convolution steps of the convolutional units included in convolutional layers E2, E4, E6, and E8 may be 2. The size of the convolution kernel of convolutional layer E9 may be 3*3, and the convolution steps may be 1. It should be understood that the size and convolution steps of the convolution kernels of the inverse convolutional units included in convolutional layers E1 to E8, as well as the size and convolution steps of the convolution kernel of convolutional layer E9, are not limited in this application.

[0152] Based on Figures 3a and 3b, the following describes the pre-training and training processes of the image processing network.

[0153] Figure 4a is a diagram illustrating an example of the training process.

[0154] The image processing network can be pre-trained based on the framework shown in Figure 3a. See S401-S403.

[0155] S401: A second training image and a second prediction image are obtained, the second prediction image being obtained by encoding the second training image based on an untrained coding network, and then decoding the encoded result of the second training image based on an untrained decoding network.

[0156] For example, multiple images may be acquired to be used to pre-train an image processing network. For ease of explanation, the images used to pre-train the image processing network may be referred to as second training images. In this application, an example in which an image processing network is pre-trained using one second training image is used for illustrative purposes.

[0157] Referring to Figure 3a, for example, a second training image may be input to the coding network, which transforms the second training image and outputs feature map 1 to quantization module A1. Next, quantization module A1 may quantize feature map 1 to obtain feature map 2. Subsequently, quantization module A1 may input feature map 2 to the hyperplier coding network, which performs hyperplier coding on feature map 2 and outputs hyperplier features. Next, the hyperplier features may be input to quantization module A2. Quantization module A2 quantizes the hyperplier features and inputs the quantized hyperplier features to entropy coding module A2, which performs entropy coding on the quantized hyperplier features to obtain bitstream 2. Subsequently, bitstream 2 may be input to entropy decoding module A2 for entropy decoding to obtain quantized hyperplier features, and the quantized hyperplier features are input to the hyperplier decoding network. Next, the hyperplier decoding network can perform hyperplier decoding on the quantized hyperplier features and output the probability distribution to the entropy estimation module. The entropy estimation module performs entropy estimation based on the probability distribution and outputs the entropy estimation information of the feature points included in feature map 2 to the entropy coding module A1 and the entropy decoding module A1. The quantization module A1 can input feature map 2 to the entropy coding module A1, which, based on the entropy estimation information of the feature points included in feature map 2, performs entropy coding on the feature points included in feature map 2 to obtain bitstream 1, and then inputs bitstream 1 to the entropy decoding module A1. Next, the entropy decoding module A1 can perform entropy decoding on the bitstream based on the entropy estimation information of all feature points included in feature map 2 and output feature map 2 to the decoding network.Next, the decoding network can transform feature map 2 to obtain a second predicted image (i.e., a reconstructed image).

[0158] Next, the image processing network can be pre-trained based on a second predicted image and a second training image. See S402 and S403.

[0159] S402: Determine the second loss based on the second predicted image and the second training image.

[0160] S403: Pre-train the untrained image processing network based on the second loss.

[0161] For example, after a second predicted image is determined, an untrained image processing network can be pre-trained based on the second predicted image and the second training image.

[0162] For example, the second loss may be determined based on the second predicted image and the second training image. The formula for calculating the second loss may be given by the following equation (1). L fg =L rate +α*L mse (1)

[0163] Here, L fg This represents the second loss, L rate represents bitrate loss, L mse represents the MSE (Mean Squared Error) loss, and α is L mse These are the corresponding weighting coefficients.

[0164] For example, L rate The formula for calculating this can be shown by the following formula (2).

number

[0165] Here, S i1H1 is the entropy estimate information for the i-th feature point in feature map 2, and H1 is the total number of feature points in feature map 2.

[0166] Furthermore, the entropy estimation information for each feature point in feature map 2, estimated by the entropy estimation module, can be obtained, and then the entropy estimation information for all feature points in feature map 2 is obtained, with a bitrate loss L rate To obtain this, it is added according to formula (2).

[0167] For example, L mse The formula for calculating this can be shown by the following formula (3).

number

[0168] Here, H2 is the number of pixels in the second predicted image or the second training image, and Y 2i is the pixel at the i2th position in the second predicted image (i.e., the pixel value of the pixel at the i2th position in the second predicted image),

number

[0169] Furthermore, L mse This can be obtained computationally according to equation (3), the pixels of the second training image, and the pixels of the second prediction image.

[0170] In this way, the bitrate loss L rate and MSE loss L mse After that is determined, in order to obtain the second loss, the bitrate loss L rate The corresponding weighting coefficient (i.e., "1") and MSE loss L mse Based on the corresponding weighting coefficient (i.e., "α"), the bitrate loss L follows equation (1). rate and MSE loss L mseWeight calculations may be performed on these values.

[0171] Next, an untrained image processing network may be pre-trained based on a second loss. In the method described above, the untrained image processing network may be pre-trained using a plurality of second training images and a plurality of corresponding second prediction images until the image processing network satisfies a second pre-set condition. The second pre-set condition is a condition for stopping the pre-training of the image processing network and may be set according to requirements. For example, reaching a pre-set number of training iterations or the loss becoming less than a pre-set loss. This is not limited to the present application.

[0172] MSE loss function L mse However, this is used as a loss function to pre-train the image processing network, thereby improving the objective quality (e.g., PSNR) of the images obtained by image processing performed using the pre-trained image processing network. In this way, the image processing network is pre-trained based on the second loss, thereby improving the objective quality of the images obtained by image processing performed using the pre-trained image processing network. In addition, by pre-training the image processing network, the image processing network can converge more quickly and better in the subsequent training process.

[0173] Next, the image processing network can be trained based on the framework shown in Figure 3b. See S404-S411.

[0174] S404: The first training image and the first prediction image are obtained, the period of the checkerboard effect is obtained, and the first prediction image is generated by performing image processing on the first training image based on the image processing network.

[0175] For example, regarding S404, please refer to the explanation for S201. Further details will not be provided here.

[0176] For example, there may be multiple first training images, and there may also be multiple second training images.

[0177] In a possible way, there is a common set of images between a set containing multiple first training images and a set containing multiple second training images.

[0178] In possible ways, there is no common set between a set containing multiple first training images and a set containing multiple second training images. This is not limited to the present application.

[0179] The following example is used for illustrative purposes, where the decoding network within the image processing network contains C upsampling layers, and correspondingly, the period of the checkerboard effect is T checkboard =p*q=2 C *2 C That is the case.

[0180] For example, if the decoding network within an image processing network includes two upsampling layers, the period of the checkerboard effect is 4*4.

[0181] S405: Based on the period, the first training image is divided into M first image blocks, and the first prediction image is divided into M second image blocks.

[0182] The following example is used for illustrative purposes, where both the size of the first image block and the size of the second image block are equal to the period of the checkerboard effect.

[0183] For example, the first training image is 2 C *2 C The first predicted image can be divided into M first image blocks, each having a size of 2. C *2 C It can be divided into M second image blocks.

[0184] Figure 4b is an example of a chunking process. In Figure 4b, when the decoding network in the image processing network includes two upsampling layers, the period of the checkerboard effect is 4*4.

[0185] Referring to (1) in Figure 4b, for example, (1) in Figure 4b is the first training image, and the size of the first training image is 60*40. The first training image can be divided into 150 (i.e., M=150) first image blocks, each with a size of 4*4.

[0186] Referring to (2) in Figure 4b, for example, (2) in Figure 4b is the first predicted image, and the size of the first predicted image is 60*40. prediction The image can be divided into 150 second image blocks, each measuring 4x4 (i.e., M=150).

[0187] S406: To obtain the characteristic value of a corresponding position within a first feature block, a calculation is performed based on the pixels at the corresponding positions within M first image blocks, and one characteristic value is obtained correspondingly for the pixels at the same positions within M first image blocks.

[0188] For example, a method for fusing M first image blocks into a first feature block might involve performing calculations based on the pixels at corresponding positions within the M first image blocks to obtain the characteristic values ​​at those corresponding positions within the first feature block. In this way, the information from the M first image blocks is aggregated into one or more first feature blocks, and the information from the M second image blocks is aggregated into one or more second feature blocks. Next, a first loss is calculated by comparing the first and second feature blocks to more deliberately compensate for the periodic checkerboard effect and, as a result, achieve a better effect in eliminating the checkerboard effect.

[0189] In possible ways, all or part of the first image blocks can be merged. In this way, N first image blocks can be selected from M first image blocks, and then the characteristic values ​​of the corresponding positions in the first feature blocks can be obtained by calculation based on the pixels of the corresponding positions in the N first image blocks, where N is a positive integer less than or equal to M. When the first feature blocks are obtained by calculation based on the pixels of the corresponding positions in part of the first image blocks, and the second feature blocks are obtained by calculation based on the pixels of the corresponding positions in part of the second image blocks, less information is used to calculate the first loss, and as a result, the efficiency of calculating the first loss is improved. When the first feature blocks are obtained by calculation based on the pixels of the corresponding positions in all the first image blocks, and the second feature blocks are obtained by calculation based on the pixels of the corresponding positions in all the second image blocks, the information used to calculate the first loss is more comprehensive, and as a result, the accuracy of the first loss is improved.

[0190] In possible ways, all or some pixels of each of the N first image blocks can be merged. Furthermore, in order to obtain the characteristic value of the corresponding first position in the first feature block, calculations can be performed based on the first target pixels of the corresponding first position in all of the N first image blocks. The number of first target pixels is less than or equal to the total number of pixels contained in the first image block. The first position and the number of first target pixels can be set according to requirements, but are not limited to this application. When calculations are performed based on some pixels in an image block, less information is used to calculate the first loss, resulting in improved efficiency in calculating the first loss. When calculations are performed based on all pixels in an image block, the information used to calculate the first loss is more comprehensive, resulting in improved accuracy of the first loss.

[0191] It should be noted that there may be one or more first feature blocks. This is not limited to this application. The following example is used for illustrative purposes, where N=M, the number of first feature blocks is 1, the number of first target pixels is equal to the total number of pixels in the first image block, and the first position is the position of all pixels in the first image block.

[0192] Figure 4c is a diagram illustrating an example of the process for generating a first feature block. The embodiment in Figure 4c describes the process of fusing M first image blocks (the method for dividing the first training image into M first image blocks is shown in Figure 4b) into a single first feature block.

[0193] Referring to Figure 4c, for example, the M first image blocks obtained by splitting the first training image may be called B1, B2, B3, ..., and BM, respectively. The 16 pixels contained in B1 are e 11 , e 12 , e 13 , ..., and e 116 It is called and the 16 pixels included in B2 are e 21 , e 22 , e 23 , ..., and e 216 It is called and the 16 pixels included in B3 are e 31 , e 32 , e 33 , ..., and e 316 It is called, and the 16 pixels included in BM are E M1 , e M2 , e M3 , ..., and e M16 It is called [name].

[0194] For example, to obtain the characteristic value of a corresponding position within a first feature block, a linear calculation may be performed based on the pixels at the corresponding positions within M first image blocks.

[0195] In a possible way, to obtain the characteristic value of the corresponding position in the first feature block, a linear calculation may be performed based on the pixels at the corresponding positions in the M first image blocks, with reference to equation (4) below.

number

[0196] Here, F i represents the characteristic value at the i-th position within the first characteristic block, and sum is the summation function. 1i represents the pixel at the i-th position in the first image block, and e 2i represents the pixel at the i-th position in the second first image block, and e 3i represents the pixel at the i-th position in the third first image block, ..., e Mi This represents the pixel at the i-th position within the M-th first image block.

[0197] An example is used to illustrate how to calculate the characteristic value F1 at a first position within a first feature block. For example, pixel e of B1. 11 , B2 pixel e 21 , B3 pixel e 31 ..., and the pixels of BM M1 The average value can be calculated, and the obtained average value is used as the characteristic value F1 of the first position in the first feature block. Similarly, M characteristic values ​​for M positions in the first feature block can be obtained.

[0198] In this way, the characteristic value of the first feature block is determined by calculating the pixel average value. This is a simple calculation and, as a result, improves the efficiency of calculating the first loss.

[0199] It should be understood that equation (4) is merely an example of linear calculation. In this application, linear calculation may be performed in another way (e.g., linear weighting) based on the pixels at corresponding positions in M ​​first image blocks in order to obtain the characteristic values ​​of the corresponding positions in the first feature block.

[0200] It should be noted that in this application, the nonlinear calculation may be performed based on the pixels of corresponding positions in M ​​first image blocks in order to obtain the characteristic values ​​of the corresponding positions in the first feature block. See, for example, equation (5). F i =A1*e 1i 2 +A2*e 2i 2 +A3*e 3i 2 +…+AM*e Mi 2 (5)

[0201] A1 represents the weight coefficient corresponding to the pixel at the i-th position in the first image block (B1), A2 represents the weight coefficient corresponding to the pixel at the i-th position in the second image block (B2), A3 represents the weight coefficient corresponding to the pixel at the i-th position in the third image block (B3), ..., AM represents the weight coefficient corresponding to the pixel at the i-th position in the Mth image block (BM).

[0202] In this application, it should be noted that calculations may be performed on pixels based on convolutional layers to determine the characteristic values ​​of feature blocks. For example, M first image blocks may be input to convolutional layers (one or more layers) and fully connected layers to obtain first feature blocks output by fully connected layers.

[0203] It should be understood that, in this application, or a combination of other methods may be used. This is not limited to this application.

[0204] S407: To obtain the characteristic value of the corresponding position in the second feature block, a calculation is performed based on the pixels at the corresponding positions in M ​​second image blocks, and one characteristic value is obtained correspondingly for the pixels at the same position in the M second image blocks.

[0205] For example, regarding S407, please refer to the explanation for S408. Further details will not be explained here.

[0206] S408: Determine the first loss based on the first and second feature blocks.

[0207] In possible ways, the first loss can be determined based on the point-to-point loss between the first and second feature blocks.

[0208] For example, the point-to-point loss between the first feature block and the second feature block can be obtained based on the pixel at the third position in the first feature block and the pixel at the third position in the second feature block.

[0209] In possible ways, the point-to-point loss may include the distance Ln. For example, the distance Ln is the distance L1, which can be calculated by referring to equation (6) below.

number

[0210] Here, L1 is the L1 distance, and F i This represents the characteristic value at the i-th position within the first characteristic block,

number

[0211] It should be understood that the Ln distance may further include the L2 distance and the L-Inf distance, etc. This is not limited to this application. The formulas for calculating the L2 distance and the L-Inf distance are the same as those for calculating the L1 distance. Further details are not provided here.

[0212] For example, the point-to-point loss can be determined as the first loss, which can be expressed by equation (7) below.

number

[0213] Here, L pc is the first loss.

[0214] In a possible way, the first loss can be determined based on a feature-based loss between the first feature block and the second feature block.

[0215] For example, the feature-based loss may include SSIM, MSSSIM, and LPIPS, etc. (For the calculation formulas of SSIM, MSSSIM, and LPIPS, reference can be made to the description of the prior art, and details are not described here). This is not limited in this application.

[0216] For example, the first feature block and the second feature block can be further input into a neural network (e.g., a convolutional network or a VGG network (Visual Geometry Group Network)), and the neural network outputs the first feature of the first feature block and the second feature of the second feature block. Next, in order to obtain the feature-based loss, the distance between the first feature and the second feature is calculated.

[0217] For example, the reciprocal of the feature-based loss can be used as the first loss. For example, when the feature-based loss is SSIM, the relationship between the first loss and SSIM can be shown by the following formula (8).

Number

[0218] Here, L pc is the first loss, F is the first feature block,

Number

[0219] For example, a method of training an image processing network based on a first loss can be shown in S409 and S410 below.

[0220] S409: Determine a third loss based on the discrimination result obtained by the discrimination network for the first training image and the first prediction image, where the third loss is the loss of the adversarial generation network GAN.

[0221] For example, the calculation formula for the GAN loss can be shown by the following formula (9). L GAN =ΣE(logD d )+ΣE[log(1-D(G(p)))] (9)

[0222] Here, E is the expected value, D d is the discrimination result of the discrimination network for the first training image, and D(G(p)) is the discrimination result of the discrimination network for the first prediction image. In the process of training the image processing network, E(logD d ) can be a constant term.

[0223] Furthermore, the third loss can be calculated according to formula (9) based on the discrimination result of the discrimination network for the first training image and the discrimination result of the discrimination network for the first prediction image.

[0224] S410: Determine a fourth loss, where the fourth loss includes at least one of the following, namely, L1 loss, bitrate loss, perceptual loss, or at least one of edge loss.

[0225] For example, the final loss L fg used to train the image processing network can be calculated using the following formula (10). L fg =L rate +α*L1+β*(L percep +γL GAN )+δ*L edge +η*L pc (10)

[0226] Here, L rate represents the bitrate loss, L1 represents the L1 distance, and α is the weighting coefficient of L1. L percep represents the perceptual loss, L GAN represents the GAN loss, and β is the weighting coefficient of L percep and L GAN and is the weighting coefficient of 、L edge represents the edge loss, and δ is the weighting coefficient of L edge and L pc represents the first loss, and η is the weighting coefficient of L pc and is the weighting coefficient of.

[0227] The final loss L fg used to train the image processing network is understood to be calculated based on L pc and L GAN as well as L1, L rate , L percep , and L edge of any one or more of. It is not limited in this application. The following examples are used for illustration, and the final loss L fg used to train the image processing network is calculated according to Equation (10). fg

[0228] To calculate the bitrate loss L rate , calculate the L1 distance, calculate the first loss L pc , and calculate the third loss L GAN , please refer to the above description. Details will not be explained again here.

[0229] For example, L percep can be calculated based on a method of calculating the LPIPS loss using the first training image and the first predicted image as calculation data. For details, please refer to the method of calculating the LPIPS loss in the prior art. Details will not be explained here.

[0230] For example, L edgeThis can be the L1 distance. The L1 distance can be calculated with respect to the edge region, and the edge loss can be obtained. For example, an edge detector (which may also be called an edge detection network) can be used to detect the first edge region of the first training image and the second edge region of the first predicted image. Then, in order to obtain the edge loss, the L1 distance between the first edge region and the second edge region can be calculated.

[0231] S411: Train the pre-trained image processing network based on the first loss, the third loss, and the fourth loss.

[0232] For example, after the first, third, and fourth losses have been obtained, weighting calculations may be performed on the first, third, and fourth losses based on the weight coefficients corresponding to the first loss (e.g., "η" in equation (10)), the weight coefficients corresponding to the third loss (e.g., "β*γ" in equation (10)), and the weight coefficients corresponding to the fourth loss (e.g., in equation (10), the weight coefficient corresponding to the bitrate loss is "1", the weight coefficient corresponding to the L1 loss is "α", the weight coefficient corresponding to the perception loss is "β", and the weight coefficient corresponding to the edge loss is "δ"). The image processing network is then trained based on the final loss.

[0233] It should be understood that the image processing network may be trained based on only the first and third losses. Compared to using only the first loss, the checkerboard effect can be better compensated for by referencing the first and third losses to eliminate the checkerboard effect to a greater extent.

[0234] When an image processing network is trained based on the first, third, and fourth losses, the quality of the images processed by the trained image processing network (including objective quality and subjective quality, which may also be called visual quality) can be improved.

[0235] For example, the upsampling layer of the decoding network causes a checkerboard effect, and the upsampling layer of the hyperprior decoding network also causes a checkerboard effect (which is weaker than the checkerboard effect caused by the decoding network and has a checkerboard effect with a longer period). Further, in order to further improve the visual quality of the reconstructed image, the period of the checkerboard effect caused by the upsampling layer of the decoding network (hereinafter referred to as the first period) and the period of the checkerboard effect caused by the upsampling layer of the hyperprior decoding network (hereinafter referred to as the second period) can be obtained. Next, based on the first period and the second period, the first training image can be divided into M first image blocks, and the first predicted image can be divided into M second image blocks. Thereafter, a pre-trained image processing network (including an encoding network, a decoding network, a hyperprior encoding network, and a hyperprior decoding network) is trained based on a first loss determined based on the M first image blocks and the M second image blocks. The specific process can be as follows.

[0236] FIG. 5 is a diagram of an example of the training process.

[0237] S501: Obtain a second training image and a second predicted image, where the second predicted image is obtained by encoding the second training image based on an untrained encoding network and then decoding the encoding result of the second training image based on an untrained decoding network.

[0238] S502: Determine a second loss based on the second predicted image and the second training image.

[0239] S503: Pre-train an untrained image processing network based on the second loss.

[0240] For example, for S501-S503, please refer to the explanation for S401-S403. Further details will not be explained here.

[0241] S504: A first training image and a first prediction image are obtained, and the period of the checkerboard effect is obtained. The first prediction image is generated by performing image processing on the first training image based on an image processing network, and the period of the checkerboard effect includes a first period and a second period.

[0242] For example, the first period is smaller than the second period.

[0243] For example, the first period may be determined based on the number of upsampling layers in the decoding network. The first period may be represented as p1*q1, where p1 and q1 are positive integers, the unit of p1 and q1 is px (pixels), and p1 and q1 may or may not be equal. This is not limited to the present application.

[0244] For example, if the number of upsampling layers in the decoding network is C, then the first period is T1 checkboard =p1*q1=2 C *2 C That is the case.

[0245] For example, the second period may be determined based on the first period and the number of upsampling layers in the hyperplier decoding network. The second period may be expressed as p2*q2, where p2 and q2 are positive integers, the unit of p2 and q2 is px (pixels), and p2 and q2 may or may not be equal. This is not limited to the present application.

[0246] In possible ways, the second period can be an integer multiple of the first period, i.e., T2 checkboard=p2*q2=(G*p1)*(G*q1), where G is an integer greater than 1. In possible ways, G is directly proportional to the number of upsampling layers in the hyperplier decoding network. Specifically, a larger number of upsampling layers in the hyperplier decoding network results in a larger value of G, and a smaller number of upsampling layers results in a smaller value of G.

[0247] For example, if G=2, and the first period of the checkerboard effect is 16*16, then the second period of the checkerboard effect is 32*32.

[0248] It should be understood that the first and second periods may also be determined based on the analysis of the first predicted image output by the image processing network. This is not limited to the present application.

[0249] S505: The first training image is divided into M first image blocks based on the first and second periods, and each of the M first image blocks contains M1 third image blocks and M2 fourth image blocks, the size of the third image block is related to the first period and the size of the fourth image block is related to the second period.

[0250] For example, the first training image may be divided into M1 third image blocks based on the first period, and the first training image may be divided into M2 fourth image blocks based on the second period. For specific division methods, please refer to the explanation above in S405. Details will not be explained again here. Here, M1 and M2 are positive integers, and M1 + M2 = M.

[0251] S506: The first predicted image is divided into M second image blocks based on the first and second periods, and the M second image blocks include M1 fifth image blocks and M2 sixth image blocks, the size of the fifth image block is related to the first period and the size of the sixth image block is related to the second period.

[0252] For example, the first predicted image may be divided into M1 fifth image blocks based on the first period, and the first predicted image may be divided into M2 sixth image blocks based on the second period. For specific division methods, please refer to the explanation above in S405. Details will not be explained again here.

[0253] S507: Determine the fifth loss based on M1 third image blocks and M1 fifth image blocks.

[0254] For example, regarding S507, please refer to the explanations in S406-S408. Further details will not be explained here.

[0255] S508: Determine the sixth loss based on M2 fourth image blocks and M2 sixth image blocks.

[0256] For example, regarding S508, please refer to the explanations for S406-S408. Further details will not be explained here.

[0257] S509: A third loss is determined based on the classification results obtained by the discriminative network for the first training image and the first prediction image, and the third loss is the loss of the generative adversarial network (GAN).

[0258] S510: Determine the fourth loss, which is at least one of the following: L1 loss, bitrate loss, perceived loss, or It includes at least one of the edge losses.

[0259] S511: Train a pre-trained image processing network based on the fifth loss, sixth loss, third loss, and fourth loss.

[0260] For example, the final loss L used to train an image processing network fg This can be calculated using the following formula (11). Lfg =L rate +α*L1+β*(L percep +γL GAN )+δ*L edge +η(L1 pc +εL2 pc ) (11)

[0261] Here, L1 pc represents the fifth loss, and η is L1 pc This is the weighting coefficient. L2 pc represents the sixth loss, and η*ε is L2 pc These are the weighting coefficients.

[0262] In possible ways, η is greater than η*ε, meaning that the weighting coefficient corresponding to the fifth loss is greater than the weighting coefficient corresponding to the sixth loss.

[0263] In possible cases, η is less than η*ε, meaning that the weighting coefficient corresponding to the fifth loss is smaller than the weighting coefficient corresponding to the sixth loss.

[0264] In any possible way, η is equal to η*ε, that is, the weighting coefficient corresponding to the fifth loss is equal to the weighting coefficient corresponding to the sixth loss.

[0265] For details on S509-S511, please refer to the explanations for S409-S411 mentioned above. Further details will not be provided here.

[0266] It should be understood that the period of the checkerboard effect can include many more periods distinct from the first and second periods. For example, the period of the checkerboard effect may include k periods (where k periods can be the first period, the second period, ..., and the kth period, and k is an integer greater than 2). In this way, the first training image can be divided into M first image blocks based on the first period, the second period, ..., and the kth period. The M first image blocks may include M1 image blocks 11, M2 image blocks 12, ..., and Mk image blocks 1k. The size of image block 11 relates to the first period, the size of image block 12 relates to the second period, ..., and the size of image block 1k relates to the kth period, with M1 + M2 + ... + Mk = M. The first predicted image can be divided into M second image blocks based on the first period, the second period, ..., and the kth period. The M second image blocks may include M1 image blocks 21, M2 image blocks 22, ..., and Mk image blocks 2k. The size of image block 21 relates to the first period, the size of image block 22 relates to the second period, ..., and the size of image block 2k relates to the kth period. Then, Loss 1 may be determined based on M1 image blocks 11 and M1 image blocks 21, Loss 2 may be determined based on M2 image blocks 12 and M2 image blocks 22, ..., and Lossk may be determined based on Mk image blocks 1k and Mk image blocks 2k. Subsequently, a pre-trained coding network and a pre-trained decoding network may be trained based on Loss 1, Loss 2, ..., and Lossk.

[0267] The following describes the coding and decoding processes based on image processing modules (including coding networks, decoding networks, hyperplier coding networks, and hyperplier decoding networks) acquired through training.

[0268] Figure 6 is a diagram illustrating an example of the coding process. In the embodiment of Figure 6, the coding network, the hyperplier coding network, and the hyperplier decoding network are acquired using the training method described above.

[0269] S601: Retrieves the image to be encoded.

[0270] S602: The image to be encoded is input to the encoding network. The encoding network processes the image to be encoded in order to obtain the feature map output by the encoding network.

[0271] For example, the fact that an encoding network processes an image to be encoded can be interpreted as the encoding network transforming the image to be encoded.

[0272] S603: Entropy coding is performed on the feature map to obtain the first bitstream.

[0273] Referring to Figure 3a, for example, the image to be encoded may be input to the encoding network, which transforms the image and outputs feature map 1 to quantization module A1. Next, quantization module A1 may quantize feature map 1 to obtain feature map 2. Subsequently, quantization module A1 may input feature map 2 to the hyperplier encoding network, which performs hyperplier encoding on feature map 2 and outputs hyperplier features. Next, the hyperplier features may be input to quantization module A2. Quantization module A2 quantizes the hyperplier features and inputs the quantized hyperplier features to entropy encoding module A2, which performs entropy encoding on the quantized hyperplier features to obtain a second bitstream (corresponding to bitstream 2 in Figure 3a). Subsequently, the second bitstream may be input to the entropy decoding module A2 for entropy decoding in order to obtain quantized hyperplier features, and the quantized hyperplier features are input to the hyperplier decoding network. Next, the hyperplier decoding network may perform hyperplier decoding on the quantized hyperplier features and output the probability distribution to the entropy estimation module. The entropy estimation module performs entropy estimation based on the probability distribution and outputs the entropy estimation information of the feature points included in feature map 2 to the entropy coding module. The quantization module A1 may input feature map 2 to the entropy coding module A1, and the entropy coding module A1 performs entropy coding on the feature points included in feature map 2 based on the entropy estimation information of the feature points included in feature map 2 to obtain the first bitstream (corresponding to bitstream 1 in Figure 3a).

[0274] For example, after the hyperplier features output by the hyperplier coding network are obtained, entropy coding may be performed on the hyperplier features to obtain a second bitstream.

[0275] Where possible, the first bitstream and the second bitstream can be stored. Where possible, the first bitstream and the second bitstream can be sent to the decoder.

[0276] It should be noted that the first bitstream and the second bitstream may be packaged into a single bitstream for storage / transmission, or, of course, the first bitstream and the second bitstream may be stored / transmitted as two separate bitstreams. This is not limited to the present application.

[0277] One embodiment of the present application further provides a bitstream distribution system. The bitstream distribution system includes at least one storage medium configured to store at least one bitstream, the at least one bitstream being generated based on the encoding method described above, and a streaming media device configured to retrieve a target bitstream from the at least one storage medium and transmit the target bitstream to a terminal-side device, wherein the streaming media device includes a content server or a content distribution server.

[0278] Figure 7 is a diagram illustrating an example of the decoding process. The decoding process in Figure 7 corresponds to the encoding process in Figure 6. In the embodiment of Figure 7, the decoding network and the hyperplier decoding network are acquired using the training method described above.

[0279] S701: The first bitstream is obtained, and the first bitstream is a bitstream of feature maps.

[0280] S702: Entropy decoding is performed on the first bitstream to obtain a feature map.

[0281] S703: The feature map is input to the decoding network. The decoding network processes the feature map to obtain the reconstructed image output by the decoding network.

[0282] Referring to Figure 3a, for example, the decoder may receive a first bitstream and a second bitstream. Entropy decoding module A2 may first perform entropy decoding on the second bitstream to obtain quantized hyperplier features. Next, the hyperplier features are input to the hyperplier decoding network, which processes the quantized hyperplier features to obtain a probability distribution. The probability distribution can then be input to the entropy estimation module. Based on the probability distribution, the entropy estimation module determines the entropy estimation information for the feature points to be decoded in the feature map (i.e., feature map 2) and inputs the entropy estimation information for the feature points to be decoded in the feature map to entropy decoding module A1. Subsequently, entropy decoding module A1 may perform entropy decoding on the first bitstream based on the entropy estimation information for the feature points to be decoded in the feature map to obtain the feature map. The feature map can then be input to the decoding network. The decoding network transforms the feature map to obtain a reconstructed image.

[0283] Figure 8a shows an example of the image quality comparison results.

[0284] Referring to (1) in Figure 8a, (1) in Figure 8a is a reconstructed image obtained by encoding image 1 using an encoding network trained based on a prior art training method to obtain bitstream 1 (see the description of the embodiment in Figure 6 for the specific encoding process), and decoding bitstream 1 using a decoding network trained based on a prior art training method (see the description of the embodiment in Figure 7 for the specific decoding process). The bitrate of bitstream 1 is 0.406 Bpp.

[0285] Referring to Figure 8a(2), Figure 8a(2) is a reconstructed image obtained by encoding image 1 using an encoding network trained according to the training method of this application to obtain bitstream 2 (see the description of the embodiment in Figure 6 for a specific encoding process), and decoding bitstream 2 using a decoding network trained according to the training method of this application (see the description of the embodiment in Figure 7 for a specific decoding process). The bitrate of bitstream 2 is 0.289 Bpp.

[0286] When Figure 8a(1) is compared with Figure 8a(2), the checkerboard effect is visible within the white circles in Figure 8a(1), while it is not visible within the white circles in Figure 8a(2). Compared to the prior art, it can be seen that the number of bitrate points in which the checkerboard effect appears is smaller in images obtained by processing performed using an image processing network trained according to the training method of this application.

[0287] Figure 8b is an example of the image quality comparison results.

[0288] Referring to (1) in Figure 8b, (1) in Figure 8b is a reconstructed image obtained by encoding image 1 using an encoding network trained based on a prior art training method to obtain bitstream 1 (see the description of the embodiment in Figure 6 for the specific encoding process), and decoding bitstream 1 using a decoding network trained based on a prior art training method (see the description of the embodiment in Figure 7 for the specific decoding process). The bitrate of bitstream 1 is 0.353 Bpp.

[0289] Referring to (2) in Figure 8b, (2) in Figure 8b is a reconstructed image obtained by encoding image 1 using an encoding network trained according to the training method of this application to obtain bitstream 2 (see the description of the embodiment in Figure 6 for the specific encoding process), and decoding bitstream 2 using a decoding network trained according to the training method of this application (see the description of the embodiment in Figure 7 for the specific decoding process). The bitrate of bitstream 2 is 0.351 Bpp.

[0290] The bitrates of bitstream 1 and bitstream 2 are approximately equal. However, when Figure 8a(1) is compared with Figure 8a(2), a checkerboard effect is visible within the white circles in Figure 8a(1), while no checkerboard effect is visible within the white circles in Figure 8a(2). It can be seen that, at relatively low bitrates, the checkerboard effect in images obtained by processing performed using an image processing network trained according to the training method of this application can be removed to some extent.

[0291] Figure 8c shows an example of the image quality comparison results.

[0292] Referring to Figure 8c(1), Figure 8c(1) is a reconstructed image obtained by encoding image 1 using an encoding network trained based on a prior art training method to obtain bitstream 1 (see the description of the embodiment in Figure 6 for a specific encoding process), and decoding bitstream 1 using a decoding network trained based on a prior art training method (see the description of the embodiment in Figure 7 for a specific decoding process). The bitrate of bitstream 1 is a medium bitrate (for example, it may be 0.15 Bpp to 0.3 Bpp, and may be specifically set according to requirements).

[0293] Referring to Figure 8c(2), Figure 8c(2) is a reconstructed image obtained by encoding image 1 using an encoding network trained according to the training method of this application to obtain bitstream 2 (see the description of the embodiment in Figure 6 for the specific encoding process), and decoding bitstream 2 using a decoding network trained according to the training method of this application (see the description of the embodiment in Figure 7 for the specific decoding process). The bitrate of bitstream 2 is medium bitrate.

[0294] The bitrates of both bitstream 1 and bitstream 2 are medium bitrate. In Figure 8c(1) and Figure 8c(2), the checkerboard effect is not visible. However, when the region enclosed by the ellipse in Figure 8c(1) is compared with the region enclosed by the ellipse in Figure 8c(2), Figure 8 c The region enclosed by the ellipse in (2) has more detail than the region enclosed by the ellipse in (1) of Figure 8a. At medium bitrates, it can be seen that images obtained by encoding and decoding performed using an image processing network trained according to the training method of this application have higher quality.

[0295] For example, Figure 9 is a block diagram of a device 900 according to one embodiment of the present application. The device 900 may include a processor 901 and a transceiver / transceiver pin 902, and optionally further include a memory 903.

[0296] The components of device 900 are coupled to each other via bus 904. In addition to the data bus, bus 904 further includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the various buses are referred to as bus 904 in the diagram.

[0297] Optionally, the memory 903 may be configured to store the instructions in the method embodiment described above. The processor 901 may be configured to execute the instructions in the memory 903, control the receive pin to receive signals, and control the transmit pin to transmit signals.

[0298] The apparatus 900 may be an electronic device or a chip of an electronic device in the method embodiment described above.

[0299] All relevant details of the steps in the aforementioned method embodiment may be referenced in the description of the function of the corresponding functional module. Further details will not be provided here.

[0300] One embodiment further provides a computer-readable storage medium that stores program instructions. When the program instructions are executed on an electronic device, the electronic device is enabled to perform the aforementioned related method steps in order to carry out the method of the aforementioned embodiment.

[0301] One embodiment further provides a computer program product. When the computer program product runs on a computer, the computer is enabled to perform the aforementioned related steps in order to carry out the method of the above embodiment.

[0302] In addition, one embodiment of the present application further provides an apparatus, which may specifically be a chip, a component, or a module. The apparatus may include a connected processor and memory. The memory is configured to store computer executable instructions. When the apparatus is operating, the processor may execute computer executable instructions stored in the memory to enable the chip to perform the method in the aforementioned method embodiment.

[0303] The electronic devices, computer-readable storage media, computer program products, or chips provided in the embodiments are configured to perform the corresponding methods provided above. Therefore, for the beneficial effects that can be achieved, please refer to the beneficial effects in the corresponding methods provided above. Further details are not provided here.

[0304] Based on the description of the embodiments described above, those skilled in the art will understand that, for the sake of simplicity, the division into functional modules described above is used as an example for illustrative purposes. In actual applications, the functions described above may be assigned to different functional modules and implemented on a case-by-case basis. In other words, the internal structure of the device is divided into different functional modules to implement all or some of the functions described above.

[0305] In some embodiments provided in this application, it should be understood that the disclosed apparatus and methods may be carried out in other ways. For example, the described apparatus embodiments are merely illustrative. For example, the division into modules or units is merely a division of logical functions, and other divisions may be used in actual embodiments. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not performed. In addition, the presented or described interconnections or direct connections or communication connections may be carried out using several interfaces. Indirect connections or communication connections between apparatus or units may be carried out electronically, mechanically, or in other forms.

[0306] Units described as separate parts may or may not be physically separate, and parts presented as units may be one or more physical units, may be located in one place, or may be distributed in different locations. Some or all of the units may be selected according to the actual requirements to achieve the objectives of the solution of the embodiment.

[0307] In addition, the functional units in the embodiments of this application may be integrated into a single processing unit, or each unit may exist physically independently, or two or more units may be integrated into a single unit. The integrated unit may be implemented in hardware form or in the form of a software functional unit.

[0308] Any content in the embodiments of this application and any content in the same embodiments can be freely combined. Any combination of the aforementioned content is within the scope of this application.

[0309] When an integrated unit is implemented in the form of a software function unit and sold or used as an independent product, the integrated unit may be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application may be implemented in the form of a software product, either essentially, or with respect to the prior art, or all or part of the technical solutions. The software product is stored in a storage medium and includes several instructions for instructing a device (which may be a single-chip microcomputer or chip, etc.) or processor to perform all or part of the steps of the method described in the embodiments of this application. The aforementioned storage medium includes any medium capable of storing program code, such as a USB flash drive, a removable hard disk, read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0310] The embodiments of this application have been described above with reference to the attached drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are examples, not limitations. Those skilled in the art may make further modifications to the inventions of this application without departing from the subject matter and scope of the claims, all modifications shall remain within the scope of the protection of this application.

[0311] The methods or algorithmic steps described in combination with those disclosed in this embodiment of the application may be implemented by hardware or by a processor by executing software instructions. Software instructions may include corresponding software modules. Software modules may be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable hard disks, compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art. For example, the storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.

[0312] Those skilled in the art will recognize that, in one or more of the above-mentioned examples, the functions described in the embodiments of this application may be implemented by hardware, software, firmware, or any combination thereof. When the functions are implemented by software, the above-mentioned functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes in a computer-readable medium. The computer-readable medium includes a computer-readable storage medium and a communication medium, the communication medium including any medium that enables a computer program to be transmitted from one place to another. The storage medium may be any available medium accessible to a general-purpose or dedicated computer.

[0313] The embodiments of this application have been described above with reference to the attached drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are examples, not limitations. Those skilled in the art may make further modifications to the inventions of this application without departing from the subject matter and scope of the claims, all modifications shall remain within the scope of the protection of this application. [Explanation of symbols]

[0314] 900 equipment 901 Processor 902 Transceiver / Transceiver Pin 903 memory 904 Bus

Claims

1. A method for training an image processing network using a computer, wherein the method is: A step of obtaining a first training image and a first prediction image, and obtaining the period of the checkerboard effect, wherein the first prediction image is generated by performing image processing on the first training image based on the image processing network, A step of dividing the first training image into M first image blocks and the first prediction image into M second image blocks based on the period, wherein both the size of the first image blocks and the size of the second image blocks are related to the period and M is an integer greater than 1, A step of determining a first loss based on the M first image blocks and the M second image blocks, The steps include training the image processing network based on the first loss and Methods that include...

2. The step of determining a first loss based on the M first image blocks and the M second image blocks is: A step of obtaining a first feature block based on the M first image blocks, wherein the characteristic value of the first feature block is obtained by calculation based on the pixels at corresponding positions within the M first image blocks. A step of obtaining a second feature block based on the M second image blocks, wherein the characteristic value of the second feature block is obtained by calculation based on the pixels at corresponding positions within the M second image blocks. A step of determining the first loss based on the first feature block and the second feature block. The method according to claim 1, including the method described in claim 1.

3. The step of obtaining a first feature block based on the M first image blocks is: A step of obtaining a first feature block based on N first image blocks out of the M first image blocks, wherein the characteristic value of the first feature block is obtained by calculation based on the pixels at corresponding positions in the N first image blocks. Includes, The step of obtaining a second feature block based on the M second image blocks is: A step of obtaining a second feature block based on N second image blocks out of the M second image blocks, wherein the characteristic value of the second feature block is obtained by calculation based on the pixels at corresponding positions in the N second image blocks. The method according to claim 2, including the method described in claim 2.

4. The step of obtaining a first feature block based on the M first image blocks is: A step of performing a calculation based on first target pixels at corresponding first positions in all of the M first image blocks in order to obtain a characteristic value for a corresponding first position in the first feature block, wherein the number of first target pixels is less than or equal to the total number of pixels in the first image block, and one characteristic value is obtained correspondingly for the first target pixels at the same first position in the M first image blocks. Includes, The step of obtaining a second feature block based on the M second image blocks is: A step of performing a calculation based on the second target pixels at the corresponding second positions in all of the M second image blocks in order to obtain a characteristic value for the corresponding second position in the second feature block, wherein the number of second target pixels is less than or equal to the total number of pixels in the second image block, and one characteristic value is obtained correspondingly for the second target pixels at the same second position in the M second image blocks. The method according to claim 2, including the method described in claim 2.

5. The step of obtaining a first feature block based on the M first image blocks is: A step of determining a characteristic value for a corresponding position in a first feature block based on the average value of pixels at corresponding positions in all of the M first image blocks, wherein one characteristic value is obtained with respect to the average value of pixels at the same position in the M first image blocks. Includes, The step of obtaining a second feature block based on the M second image blocks is: A step of determining a characteristic value for a corresponding position in a second feature block based on the average value of pixels at corresponding positions in all of the M second image blocks, wherein one characteristic value is obtained with respect to the average value of pixels at the same position in the M second image blocks. The method according to claim 2, including the method described in claim 2.

6. The step of determining the first loss based on the first feature block and the second feature block is: The step of determining the first loss based on the point-to-point loss between the first feature block and the second feature block. Includes, The method according to claim 2, wherein the point-to-point loss includes the L1 distance, the L2 distance, or the L-Inf distance.

7. The step of determining the first loss based on the first feature block and the second feature block is: The step of determining the first loss based on the feature-based loss between the first feature block and the second feature block. Includes, The method according to claim 2, wherein the feature-based loss includes structural similarity (SSIM), multiscale structural similarity (MSSSIM), or learned perceptual image patch similarity (LPIPS).

8. The image processing network comprises an encoding network and a decoding network, and prior to the step of acquiring a first training image and a first prediction image, the method A step of obtaining a second training image and a second prediction image, wherein the second prediction image is obtained by encoding the second training image based on an untrained coding network, and then decoding the encoded result of the second training image based on an untrained decoding network. A step of determining a second loss based on the second predicted image and the second training image, The steps include pre-training the untrained coding network and the untrained decoding network based on the second loss, and The method according to claim 1, further comprising:

9. The first predicted image is obtained by encoding the first training image based on a pre-trained coding network, and then decoding the encoded result of the first training image based on a pre-trained decoding network. The step of training the image processing network based on the first loss is: A step of determining a third loss based on the identification results obtained by the identification network with respect to the first training image and the first prediction image, wherein the third loss is the loss of a generative adversarial network (GAN), and the GAN comprises the identification network and the decoding network. A step of training the pre-trained coding network and the pre-trained decoding network based on the first loss and the third loss, The method according to claim 8, including the method described in claim 8.

10. The step of training the image processing network based on the first loss is: A step of determining a fourth loss, wherein the fourth loss includes at least one of the following: L1 loss, bitrate loss, perceived loss, or edge loss. It further includes, The step of training the pre-trained coding network and the pre-trained decoding network based on the first loss and the third loss is: The method according to claim 9, comprising the step of training the pre-trained coding network and the pre-trained decoding network based on the first loss, the third loss, and the fourth loss.

11. The image processing network further comprises a hyperplier coding network and a hyperplier decoding network, and the hyperplier decoding network and the decoding network each comprise an upsampling layer. The period includes a first period and a second period, the first period being determined based on the number of upsampling layers in the decoding network, and the second period being determined based on the first period and the number of upsampling layers in the hyperplier decoding network, wherein the second period is greater than the first period. The method according to claim 10.

12. The step of dividing the first training image into M first image blocks and the first prediction image into M second image blocks based on the aforementioned period is as follows: A step of dividing the first training image into M first image blocks based on the first period and the second period, wherein the M first image blocks include M1 third image blocks and M2 fourth image blocks, the size of the third image block is related to the first period, the size of the fourth image block is related to the second period, M1 and M2 are positive integers, and M1 + M2 = M, A step of dividing the first predicted image into M second image blocks based on the first period and the second period, wherein the M second image blocks include M1 fifth image blocks and M2 sixth image blocks, the size of the fifth image block is related to the first period and the size of the sixth image block is related to the second period, and Includes, When the first loss includes a fifth loss and a sixth loss, the step of determining the first loss based on the M first image blocks and the M second image blocks is: The steps include determining the fifth loss based on the M1 third image blocks and the M1 fifth image blocks, The steps include determining the sixth loss based on the M2 fourth image blocks and the M2 sixth image blocks, The method according to claim 11, including the method described in claim 11.

13. The step of training the image processing network based on the first loss is: A step of performing a weighting calculation on the fifth loss and the sixth loss in order to obtain a seventh loss, The steps include training the pre-trained coding network and the pre-trained decoding network based on the seventh loss, and The method according to claim 12, including the method described in claim 12.

14. An encoding method, wherein the method is Steps include obtaining the image to be encoded, A step of inputting the image to be encoded into the encoding network in order to obtain a feature map output by the encoding network, and processing the image to be encoded by the encoding network, wherein the encoding network is obtained by training using the method described in any one of claims 8 to 13, The steps include performing entropy coding on the feature map in order to obtain a first bitstream, and An encoding method that includes this.

15. The step of performing entropy coding on the feature map in order to obtain a first bitstream is: The steps include inputting the feature map into a hyperplier coding network to obtain hyperplier features, and processing the feature map by the hyperplier coding network. A step comprising inputting the hyperplier feature into a hyperplier decoding network, processing the hyperplier feature by the hyperplier decoding network, and then outputting a probability distribution, wherein both the hyperplier coding network and the hyperplier decoding network are acquired by training using the method described in any one of claims 9 to 13. To obtain the first bitstream, the steps include performing entropy coding on the feature map based on the probability distribution and The method according to claim 14, including the method described in claim 14.

16. The aforementioned method, The step of performing entropy coding on the hyperplier feature to obtain a second bitstream. The method according to claim 15, further comprising:

17. A decryption method, wherein the method is A step of obtaining a first bitstream, wherein the first bitstream is a bitstream of feature maps, The steps include performing entropy decoding on the first bitstream in order to obtain the feature map, A step of inputting the feature map into the decoding network and processing the feature map by the decoding network in order to obtain a reconstructed image output by the decoding network, wherein the decoding network is obtained by training using the method described in any one of claims 8 to 13. A decryption method that includes this.

18. The aforementioned method, A step of obtaining a second bitstream, wherein the second bitstream is a bitstream of hyperplier features. It further includes, The step of performing entropy decoding on the bitstream in order to obtain the feature map is, The steps include performing entropy decoding on the second bitstream in order to obtain the hyperplier features, A step of inputting the hyperplier features into a hyperplier decoding network to obtain a probability distribution, and processing the hyperplier features by the hyperplier decoding network, wherein the hyperplier decoding network is obtained by training using the method described in claim 9. To obtain the feature map, the steps include: performing entropy decoding on the first bitstream based on the probability distribution; The method according to claim 17, including the method described in claim 17.

19. It is an electronic device, The system comprises memory and a processor, wherein the memory is coupled to the processor. The memory stores program instructions, and when the program instructions are executed by the processor, the electronic device is enabled to perform the training method according to any one of claims 1 to 13.

20. A chip comprising one or more interface circuits and one or more processors, wherein the interface circuits are configured to receive signals from the memory of an electronic device and transmit the signals to the processor, the signals including computer instructions stored in the memory, and when the processor executes the computer instructions, the electronic device is enabled to perform the training method according to any one of claims 1 to 13.

21. A computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is run on a computer or processor, the computer or processor is enabled to perform the training method according to any one of claims 1 to 13.

22. A computer program that, when executed by a computer or processor, causes the computer or processor to perform the method according to any one of claims 1 to 13.

23. A bitstream storage device comprising a receiver and at least one storage medium, The aforementioned receiver is configured to receive a bitstream, The at least one storage medium is configured to store the bitstream, The bitstream is generated according to the encoding method described in claim 14, in a bitstream storage device.

24. A bitstream transmission device comprising a transmitter and at least one storage medium, The at least one storage medium is configured to store a bitstream, the bitstream is generated according to the encoding method described in claim 14, The transmission device is a bitstream transmission device configured to acquire the bitstream from the storage medium and transmit the bitstream to a terminal device using the transmission medium.

25. A bitstream distribution system, wherein the bitstream distribution system is A storage medium configured to store at least one bitstream, wherein the at least one bitstream is generated according to the encoding method described in claim 14, A streaming media device configured to acquire a target bitstream from at least one storage medium and transmit the target bitstream to a terminal-side device, wherein the streaming media device comprises a content server or a content distribution server. A bitstream distribution system equipped with the following features.

Citation Information

Patent Citations

  • Method and apparatus for block-by-block neural image compression with post-filtering

    CN114747207A

  • Encoding device, decoding device, encoding method, decoding method, encoding program, and decoding program

    JP2019205011A

  • Method and apparatus for block-wise neural image compression with post filtering

    US20220101492A1