Encoding method and apparatus, and decoding method and apparatus

Through the progressive encoding method, the feature map with the largest channel entropy or channel variance is selected as the first encoding feature map, and the number of feature maps is gradually increased, which solves the problem of high coding complexity of AI images and realizes efficient image encoding and decoding in different environments.

WO2025152469A1PCT designated stage expired Publication Date: 2025-07-24HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/117239
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2024-09-05
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The existing AI image encoding is complex, especially the entropy encoding and decoding process is time-consuming, which affects the implementation of the product.

Method used

The progressive encoding method is used to encode part of the feature map of the input image, and the feature map with the largest channel entropy or channel variance is selected as the first encoding feature map, and the number of feature maps is gradually increased as needed to reduce the encoding delay.

Benefits of technology

In the case of poor network environment or insufficient hardware resources, image reconstruction can be completed through only some feature maps, reducing encoding delay; when sufficient resources are sufficient, image quality is improved through multiple sets of feature maps and flexibly adapted to various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024117239_24072025_PF_FP_ABST
    Figure CN2024117239_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of media. Disclosed are an encoding method and apparatus, and a decoding method and apparatus, for use in reducing image decoding latency. The method comprises: performing feature extraction on an input image to obtain a first feature map of the input image; determining N groups of second feature maps of the input image on the basis of the first feature map; and performing entropy encoding on at least one group of second feature maps of the input image to obtain a code stream. N is a positive integer. It can be seen that in the method provided in the embodiments of the present application, a progressive encoding method is used, and some feature maps of the input image can be encoded. Compared with encoding all feature maps of the input image, encoding some feature maps of the input image can reduce image encoding latency. In addition, the first group of second feature maps of the input image are feature maps having the maximum channel entropy (variance) among the N groups of second feature maps of the input image, and a decoding end can obtain a complete reconstructed image on the basis of the feature maps.
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method and device

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 18, 2024, with application number 202410080945.2 and application name “Encoding Method”, and the Chinese patent application filed with the China Intellectual Property Office on February 8, 2024, with application number 202410176941.4 and application name “Encoding and Decoding Method and Device”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The embodiments of the present application relate to the field of media technology, and in particular to encoding and decoding methods and devices. Background Art

[0003] Breakthroughs have been made in artificial intelligence (AI) image coding, significantly improving compression efficiency compared to commonly used image coding standards while maintaining comparable subjective quality. AI image coding is widely used in various fields, including cloud storage, visual surveillance, autonomous vehicles and equipment, image acquisition, storage, and management, real-time monitoring of visual data, and media distribution.

[0004] The existing AI image coding is relatively complex, especially the entropy encoding and decoding process is time-consuming, which has a certain impact on product implementation.

[0005] Summary of the Invention

[0006] The embodiments of the present application provide a coding and decoding method and apparatus for reducing image coding and decoding delay. To achieve the above-mentioned purpose, the embodiments of the present application adopt the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides an encoding method, comprising: performing AI encoding on an input image to obtain a first feature map of the input image; determining N groups of second feature maps of the input image based on the first feature map; and entropy encoding a first group of second feature maps from the N groups of second feature maps into a bitstream. N is a positive integer.

[0008] It can be seen that in the method provided in the embodiment of the present application, a progressive encoding method is adopted to encode part of the feature map of the input image. Compared with encoding all the feature maps of the input image, encoding part of the feature map of the input image can reduce the image encoding delay. In addition, the first group of second feature maps of the input image is the feature map with the maximum channel entropy (variance) among the N groups of second feature maps of the input image, and the decoding end can obtain a complete reconstructed image based on the feature map.

[0009] For example, in a scenario with a poor network environment or insufficient decoding hardware resources, the embodiment of the present application adopts a progressive encoding method that does not require all feature maps. A complete reconstructed image can be obtained based only on decoding of the first group of second feature maps, thereby reducing the image encoding and decoding delay.

[0010] In a possible implementation, any one of the N groups of second feature maps can be decoded to obtain a reconstructed image.

[0011] In a possible implementation, the second group of second feature maps among the N groups of second feature maps may be entropy encoded into the bitstream.

[0012] It can be seen that the method provided in the embodiment of the present application adopts a progressive encoding method to encode two sets of feature maps of the input image. Compared with encoding all the feature maps of the input image, encoding part of the feature maps of the input image can reduce the image encoding delay.

[0013] In a possible implementation, the third group of second feature maps among the N groups of second feature maps may be entropy encoded into the bitstream.

[0014] As can be seen, the method provided in the embodiment of the present application adopts a progressive encoding method to encode the three sets of feature maps of the input image. Compared with encoding all the feature maps of the input image, encoding part of the feature maps of the input image can reduce the image encoding delay.

[0015] For example, in scenarios where the network environment is good or decoding hardware resources are sufficient, the embodiments of the present application adopt a progressive encoding method. On the one hand, a complete reconstructed image can be obtained based on decoding only the first set of second feature maps, and on the other hand, a reconstructed image with good image quality can be obtained based on decoding multiple sets of second feature maps. Therefore, it can be seen that the progressive encoding method provided by the embodiments of the present application can be flexibly adapted to various network environments and hardware.

[0016] In one possible implementation, the first feature map may be subjected to channel separation to obtain M second feature maps of the input image. The M second feature maps may be grouped according to channel information of the second feature map of the input image to obtain the N groups of second feature maps, where M is a positive integer.

[0017] It can be seen that in the method provided in the embodiment of the present application, after obtaining the feature map of the input image, the feature map of the input image can be channel-separated to divide the feature map of the input image into multiple feature maps in the channel direction dimension, and then the partial feature map in the channel dimension of the input image is encoded. Compared with encoding the entire feature map of the input image, using a progressive encoding method to encode a portion of the feature map of the input image can reduce the image encoding delay.

[0018] In a possible implementation, the channel information includes channel entropy and / or channel variance, where the channel entropy is the entropy of all elements in the channel, and the channel variance is the variance of all elements in the channel.

[0019] In one possible implementation, the channel information may further include at least one of a channel number or a channel group number. The channel number may be determined based on the channel entropy and / or the channel variance, and the channel group number may be determined based on the channel entropy and / or the channel variance.

[0020] It can be seen that in the method provided in the embodiment of the present application, after obtaining the feature map of the input image, the feature map of the input image can be channel-separated using information such as channel entropy or channel variance to divide the feature map of the input image into multiple feature maps in the channel direction dimension, and then the partial feature map of the input image in the channel dimension is encoded. Compared with encoding all the feature maps of the input image, using a progressive encoding method to encode part of the feature map of the input image can reduce the image encoding delay.

[0021] In a possible implementation, at least one group of the N groups of second feature maps may be entropy encoded into the bitstream in descending order of channel entropy or channel variance.

[0022] It is understood that the greater the channel entropy or channel variance of the feature map, the easier it is to restore and reconstruct the image. Therefore, at least one of the N sets of second feature maps can be entropy encoded into the bitstream in descending order of channel entropy or channel variance, so that a progressive encoding method can be used to encode some feature maps of the input image to reduce image encoding latency.

[0023] In one possible implementation, channel number information may be encoded into the bitstream. The channel number information indicates the number of channels of at least one set of second feature maps encoded into the bitstream. This allows a decoder to determine the number of channels of each set of second feature maps in the bitstream by decoding the channel number information, thereby facilitating recovery of the feature maps from the bitstream.

[0024] For example, the number of channels of the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps are 4, 3, and 3, respectively. Then, the channel number information of 4, 3, and 3 can be encoded into the above code stream.

[0025] In one possible implementation, first position information may be encoded into the bitstream, where the first position information indicates the position of at least one set of second feature maps encoded into the bitstream within the first feature map. This allows a decoder to determine the position of each set of second feature maps within the first feature map by decoding the first position information, thereby facilitating recovery of pre-group feature maps from the bitstream.

[0026] It should be noted that the decoding end needs to use the position of the second feature map in the above-mentioned first feature map to decode the second feature map, and the position of the second feature map in the code stream may be different from the position of the second feature map in the above-mentioned first feature map. For this reason, it is necessary to encode the first position information of at least one group of second feature maps encoded into the above-mentioned code stream in the above-mentioned first feature map into the code stream.

[0027] For example, the first feature map is separated into 10 second feature maps through channels. The positions of these 10 second feature maps in the first feature map are recorded as 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 respectively. These 10 second feature maps are grouped and sorted and encoded into the bitstream in the order of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 respectively. Therefore, the first position information of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 can be encoded into the bitstream. After obtaining this information, the decoder can know that the position of the first second feature map obtained by decoding in the first feature map is 2, the position of the second second feature map is 9, and so on. The position of the tenth second feature map is 8 in the first feature map.

[0028] In one possible implementation, second position information may be encoded into the bitstream. The second position information indicates the position of at least one set of second feature maps encoded into the bitstream. This allows a decoder to determine the position of each set of second feature maps in the bitstream using the second position information, thereby facilitating parallel decoding.

[0029] In one possible implementation, the code stream length information may be added to the code stream, and the code stream length information is used to indicate the length of the code stream. In this way, the decoding end can determine the length of the code stream through the code stream length information, thereby facilitating parallel decoding at the decoding end.

[0030] In a second aspect, an embodiment of the present application provides a decoding method, comprising: performing entropy decoding on a bitstream to obtain a first set of second feature maps of an input image. Performing AI decoding on the first set of second feature maps to obtain a first reconstructed image of the input image. The first set of second feature maps is a portion of a feature map corresponding to the input image.

[0031] It can be seen that in the method provided in the embodiment of the present application, a reconstructed image can be determined based on a partial feature map of the input image. Compared with determining the reconstructed image based on all the feature maps of the input image, the image decoding delay can be reduced by using a progressive decoding method to determine the reconstructed image based on a partial feature map of the input image. In addition, the first group of second feature maps of the input image is the feature map with the maximum channel entropy (variance) in the second feature map of the input image, and the decoding end can obtain a complete reconstructed image based on the feature map.

[0032] For example, in a scenario where the network environment is poor or decoding hardware resources are insufficient, the embodiment of the present application adopts a progressive decoding method to obtain a complete reconstructed image based only on the decoding of the first group of second feature maps.

[0033] It should be noted that the first reconstructed image is not a partial image of the input image, but a reconstructed image having the same size as the input image but different image quality.

[0034] In one possible implementation, the code stream can be entropy decoded to obtain a second group of second feature maps of the input image; and the first group of second feature maps and the second group of second feature maps can be AI decoded to obtain a second reconstructed image of the input image.

[0035] In one possible implementation, the code stream can be entropy decoded to obtain the third group of second feature maps of the input image; the above-mentioned first group of second feature maps, the above-mentioned second group of second feature maps, and the above-mentioned third group of second feature maps can be AI decoded to obtain a third reconstructed image of the above-mentioned input image.

[0036] For example, in scenarios where the network environment is good or decoding hardware resources are sufficient, the embodiments of the present application adopt a progressive decoding method. On the one hand, a complete reconstructed image can be obtained based on decoding of only the first set of second feature maps, and on the other hand, a reconstructed image with good image quality can be obtained based on decoding of multiple sets of second feature maps. Therefore, it can be seen that the progressive decoding method provided by the embodiments of the present application can be flexibly adapted to various network environments and hardware.

[0037] In a possible implementation, the image quality of the second reconstructed image is better than that of the first reconstructed image, and the image quality of the third reconstructed image is better than that of the second reconstructed image.

[0038] The image quality is characterized by any of the following variables: peak signal to noise ratio (PSNR), multi-scale structural similarity index (MS-SSIM), or learned perceptual image patch similarity (LPIPS).

[0039] In one possible implementation, the code stream may be decoded to obtain channel number information, where the channel number information is used to indicate the number of channels of at least one group of second feature maps obtained by decoding; and the first group of second feature maps may be AI decoded according to the channel number information to obtain a first reconstructed image.

[0040] For example, the number of channels of the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps are 4, 3, and 3, respectively. The channel number information of 4, 3, and 3 can be encoded into the above-mentioned code stream. Correspondingly, after the decoding end obtains the channel number information through decoding, it can know that the first decoding must obtain the second feature maps of 4 channels in order to decode the first group of second feature maps of the input image, the second decoding must obtain the second feature maps of 4 channels in order to decode the second group of second feature maps of the input image, and the third decoding must obtain the second feature maps of 3 channels in order to decode the third group of second feature maps of the input image.

[0041] In one possible implementation, the code stream may be decoded to obtain first position information, where the first position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the first feature map of the input image; and the first group of second feature maps may be AI decoded according to the first position information to obtain a first reconstructed image.

[0042] For example, the first feature map is separated into 10 second feature maps through channels. The positions of these 10 second feature maps in the first feature map are recorded as 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 respectively. These 10 second feature maps are grouped and sorted and encoded into the bitstream in the order of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 respectively. Therefore, the first position information of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 can be encoded into the bitstream. After obtaining this information, the decoder can know that the position of the first second feature map obtained by decoding in the first feature map is 2, the position of the second second feature map is 9, and so on. The position of the tenth second feature map is 8 in the first feature map.

[0043] The first second feature map obtained by decoding is placed at position 2 in the first feature map for image restoration, the second second feature map obtained by decoding is placed at position 9 in the first feature map for image restoration, ..., the tenth second feature map obtained by decoding is placed at position 8 in the first feature map for image restoration.

[0044] In one possible implementation, the code stream may be decoded to obtain second position information, where the second position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the code stream; the bit data corresponding to the first group of second feature maps in the code stream may be determined based on the second position information, and the corresponding bit data may be entropy decoded to obtain the first group of second feature maps.

[0045] For example, the corresponding positions of the bit data of the first, second, and third groups of second characteristic maps in the code stream are A1-A2, A3-A4, and A5-A6, respectively. The second position information of A1-A2, A3-A4, and A5-A6 can then be encoded into the code stream. Accordingly, after the decoder obtains the second position information through decoding, it can be determined that the bit data at positions A1-A2 in the decoded code stream can be used to obtain the first group of second characteristic maps, the bit data at positions A3-A4 in the decoded code stream can be used to obtain the second group of second characteristic maps, and the bit data at positions A5-A6 in the decoded code stream can be used to obtain the third group of second characteristic maps.

[0046] In one possible implementation, the above-mentioned code stream can be decoded to obtain code stream length information, where the above-mentioned code stream length information is used to indicate the length of the above-mentioned code stream; and the above-mentioned first group of second feature maps are AI decoded according to the above-mentioned code stream length information to obtain a first reconstructed image.

[0047] For example, the decoding time can be estimated based on the bitstream length. If the estimated decoding time exceeds a preset time, a progressive decoding method is used. Specifically, entropy decoding is performed on the bitstream to obtain a portion of the second feature map of the input image. AI decoding is then performed on this portion of the second feature map to obtain a reconstructed image of the input image.

[0048] In a third aspect, an embodiment of the present application provides an encoding device comprising: an encoding unit and a determination unit. The encoding unit is configured to perform AI encoding on an input image to obtain a first feature map of the input image. The determination unit is configured to determine N groups of second feature maps of the input image based on the first feature map, where N is a positive integer. The encoding unit is further configured to entropy encode a first group of second feature maps from the N groups of second feature maps into a bitstream.

[0049] In a possible implementation, the encoding unit is further configured to: entropy encode a second group of second feature maps among the N groups of second feature maps into the bitstream.

[0050] In a possible implementation, the encoding unit is further configured to: entropy encode the third group of second feature maps among the N groups of second feature maps into the bitstream.

[0051] In one possible implementation, the determination unit is specifically used to: perform channel separation on the first feature map to obtain M second feature maps of the input image, where M is a positive integer; and group the M second feature maps according to the channel information of the second feature map of the input image to obtain the N groups of second feature maps.

[0052] In a possible implementation, the encoding unit is further configured to entropy encode at least one of the N groups of second feature maps into the bitstream in descending order of channel entropy or channel variance.

[0053] In a possible implementation, the encoding unit is further configured to: encode channel number information into the bitstream, where the channel number information is used to indicate the number of channels of at least one set of second feature maps encoded into the bitstream.

[0054] In a possible implementation, the encoding unit is further configured to: encode first position information into the code stream, wherein the first position information is used to indicate a position of at least one set of second feature maps encoded into the code stream in the first feature map.

[0055] In a possible implementation, the encoding unit is further configured to: encode second position information into the code stream, where the second position information is used to indicate a position in the code stream of at least one set of second feature maps encoded into the code stream.

[0056] In a possible implementation, the encoding unit is further configured to: encode code stream length information into the code stream, where the code stream length information is used to indicate the length of the code stream.

[0057] In a fourth aspect, an embodiment of the present application provides a decoding device comprising: a decoding unit and a reconstruction unit. The decoding unit is configured to perform entropy decoding on a bitstream to obtain a first set of second feature maps of an input image, wherein the first set of second feature maps is a portion of a feature map corresponding to the input image. The reconstruction unit is configured to perform AI decoding on the first set of second feature maps to obtain a first reconstructed image of the input image.

[0058] In a possible implementation, the decoding unit is further configured to perform entropy decoding on the code stream to obtain a second group of second feature maps, where the second group of second feature maps is a part of the feature maps corresponding to the input image.

[0059] In a possible implementation, the reconstruction unit is further configured to perform AI decoding on the first group of second feature maps and the second group of second feature maps to obtain a second reconstructed image of the input image.

[0060] In a possible implementation, the decoding unit is further configured to perform entropy decoding on the code stream to obtain a third group of second feature maps, where the third group of second feature maps is a part of the feature map corresponding to the input image.

[0061] In a possible implementation, the reconstruction unit is further configured to perform AI decoding on the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps to obtain a third reconstructed image of the input image.

[0062] In one possible implementation, the image quality of the second reconstructed image is better than the image quality of the first reconstructed image, and the image quality of the third reconstructed image is better than the image quality of the second reconstructed image, and the image quality is characterized by any one of the following variables: PSNR, MS-SSIM or LPIPS.

[0063] In a possible implementation, the decoding unit is further configured to decode the code stream to obtain channel number information, where the channel number information is used to indicate the number of channels of at least one set of second feature maps obtained by decoding.

[0064] In a possible implementation, the reconstruction unit is further configured to perform AI decoding on the first group of second feature maps according to the channel number information to obtain a first reconstructed image.

[0065] In a possible implementation, the decoding unit is further configured to decode the code stream to obtain first position information, where the first position information is configured to indicate a position of at least one set of second feature maps obtained by decoding in the first feature map of the input image.

[0066] In a possible implementation, the reconstruction unit is further configured to perform AI decoding on the first group of second feature maps according to the first position information to obtain a first reconstructed image.

[0067] In a possible implementation, the decoding unit is further configured to decode the code stream to obtain second position information, where the second position information is used to indicate a position of at least one set of second feature maps obtained by decoding in the code stream.

[0068] In a possible implementation, the decoding unit is further used to determine the bit data corresponding to the first group of second feature maps in the code stream based on the second position information, and perform entropy decoding on the corresponding bit data to obtain the first group of second feature maps.

[0069] In a possible implementation, the decoding unit is further configured to decode the code stream to obtain code stream length information, where the code stream length information indicates the length of the code stream.

[0070] In a possible implementation, the reconstruction unit is further configured to perform AI decoding on the first group of second feature maps according to the bitstream length information to obtain a first reconstructed image.

[0071] In a fifth aspect, an embodiment of the present application further provides a coding device, which includes: at least one processor, which, when the at least one processor executes program code or instructions, implements the method described in the above first aspect or any possible implementation thereof.

[0072] Optionally, the apparatus may further include at least one memory configured to store the program code or instruction.

[0073] In a sixth aspect, an embodiment of the present application further provides a decoding device, which includes: at least one processor, which, when the at least one processor executes program code or instructions, implements the method described in the above second aspect or any possible implementation thereof.

[0074] In a seventh aspect, an embodiment of the present application further provides a code stream storage method, the method comprising: obtaining and storing a code stream obtained by the method described in the first aspect or any possible implementation thereof.

[0075] In an eighth aspect, an embodiment of the present application further provides a bitstream storage device, which is used to obtain and store the bitstream obtained by the method described in the first aspect or any possible implementation thereof.

[0076] In a ninth aspect, an embodiment of the present application further provides a code stream transmission method, the method comprising: acquiring and transmitting a code stream obtained by the method described in the first aspect or any possible implementation thereof.

[0077] In a tenth aspect, an embodiment of the present application further provides a code stream transmission device, which is used to obtain and transmit the code stream obtained by the method described in the first aspect or any possible implementation thereof.

[0078] In an eleventh aspect, an embodiment of the present application further provides a computer-readable storage medium, on which is stored a code stream obtained by the method described in the first aspect or any possible implementation thereof.

[0079] In a twelfth aspect, an embodiment of the present application further provides a chip comprising: an input interface, an output interface, and at least one processor. Optionally, the chip further comprises a memory. The at least one processor is configured to execute code in the memory. When the at least one processor executes the code, the chip implements the method described in the first aspect or any possible implementation thereof.

[0080] Optionally, the chip may also be an integrated circuit.

[0081] In a thirteenth aspect, an embodiment of the present application further provides a computer-readable storage medium for storing a computer program, which includes methods for implementing the above-mentioned first aspect or any possible implementation thereof.

[0082] In a fourteenth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to implement the method described in the above-mentioned first aspect or any possible implementation thereof.

[0083] The encoding and decoding device, computer storage medium, computer program product and chip provided in this embodiment are all used to execute the encoding and decoding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the encoding and decoding method provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0085] FIG1a is an exemplary block diagram of a decoding system provided in an embodiment of the present application;

[0086] FIG1b is an exemplary block diagram of a video decoding system provided in an embodiment of the present application;

[0087] FIG2 is an exemplary block diagram of a video encoder provided in an embodiment of the present application;

[0088] FIG3 is an exemplary block diagram of a video decoder provided in an embodiment of the present application;

[0089] FIG4 is an exemplary block diagram of a video decoding device provided in an embodiment of the present application;

[0090] FIG5 is an exemplary block diagram of a device provided in an embodiment of the present application;

[0091] FIG6 is a schematic diagram of a neural network-based image compression method provided in an embodiment of the present application;

[0092] FIG7 is a schematic diagram of an end-to-end image coding framework provided in an embodiment of the present application;

[0093] FIG8 is a schematic diagram of the structure of a neural network provided in an embodiment of the present application;

[0094] FIG9 is a schematic structural diagram of a video communication system provided in an embodiment of the present application;

[0095] FIG10 is a schematic diagram of a flow chart of an encoding method provided in an embodiment of the present application;

[0096] FIG11 is a schematic diagram of an encoding and decoding process provided in an embodiment of the present application;

[0097] FIG12 is a schematic diagram of another encoding and decoding process provided in an embodiment of the present application;

[0098] FIG13 is a schematic diagram of another encoding and decoding process provided in an embodiment of the present application;

[0099] FIG14 is a schematic diagram of a flow chart of a decoding method provided in an embodiment of the present application;

[0100] FIG15 is a schematic structural diagram of an encoding device provided in an embodiment of the present application;

[0101] FIG16 is a schematic structural diagram of a decoding device provided in an embodiment of the present application;

[0102] FIG17 is a schematic diagram of the structure of a chip provided in an embodiment of the present application;

[0103] FIG18 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0104] FIG19 is a schematic structural diagram of another electronic device provided in an embodiment of the present application;

[0105] FIG20 is a schematic diagram of the structure of a neural network provided in an embodiment of the present application;

[0106] FIG21 is a schematic diagram of the structure of a machine video encoding system provided in an embodiment of the present application;

[0107] FIG22 is a schematic diagram of the structure of a code stream provided in an embodiment of the present application. DETAILED DESCRIPTION

[0108] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the embodiments of this application.

[0109] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0110] The terms "first" and "second" and so on in the description and drawings of the embodiments of this application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.

[0111] Furthermore, the terms "including," "having," and any variations thereof, mentioned in the description of the embodiments of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0112] It should be noted that in the description of the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be interpreted as having priority or advantage over other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.

[0113] First, the terms involved in the embodiments of the present application are explained.

[0114] The Joint Photographic Experts Group (JPEG) artificial intelligence (AI) is a learning-based image coding standard that provides a single-stream, compact compressed domain representation, significantly improving compression efficiency over commonly used image coding standards while maintaining comparable subjective quality. JPEG AI is widely used in various fields. For example, JPEG AI can be applied to cloud storage, visual surveillance, autonomous vehicles and devices, image acquisition, storage, and management, real-time monitoring of visual data, and media distribution.

[0115] Information: data to be encoded / decoded, such as an image, a video, an audio clip, or a plain text file.

[0116] Feature map: Features extracted during neural network inference.

[0117] Symbol: The basic unit of information, such as a pixel in an image, a character in a plain text file, etc.

[0118] Encoding: The process of converting information into a string of 0s or 1s.

[0119] Code stream: A string of 0s and 1s obtained by encoding information.

[0120] Decoding: The process of restoring the code stream to information, which is the inverse process of encoding.

[0121] Prerequisite: The data is known during encoding and decoding. Using the premise during encoding and decoding can reduce the length of the code stream.

[0122] Entropy coding: Using probabilistic modeling to reduce the length of the bitstream to the theoretical minimum (Shannon entropy). Entropy coding can also reduce the length of the bitstream using the premise.

[0123] Entropy decoding: This method uses probabilistic modeling to decode data, which is the inverse of entropy coding. If the premise is used during entropy coding, the same premise must be used during entropy decoding.

[0124] Throughput: The number of symbols that can be encoded / decoded per second by entropy coding / decoding.

[0125] Register: A memory used by the processor to temporarily store instructions, data, etc. It has a very small capacity and very fast read and write speeds. Current computing devices typically have 8 to 32 16-bit, 32-bit, or 64-bit registers.

[0126] Memory: The main storage unit in a computing device, with a capacity much larger than a register, but the read and write speed is usually more than 10 times that of a register.

[0127] Feature map: The three-dimensional data output by the convolutional layer, activation layer, pooling layer, batch normalization layer, etc. in a convolutional neural network. The three dimensions are called width, height, and channel respectively.

[0128] Data encoding and decoding includes two parts: data encoding and data decoding. Data encoding is performed on the source side (or commonly referred to as the encoder side), and generally includes processing (e.g., compressing) the original data to reduce the amount of data required to represent the original data (thereby more efficiently storing and / or transmitting). Data decoding is performed on the destination side (or commonly referred to as the decoder side), and generally includes inverse processing relative to the encoder side to reconstruct the original data. The "encoding and decoding" of the data involved in the embodiments of the present application should be understood as the "encoding" or "decoding" of the data. The encoding part and the decoding part are also collectively referred to as encoding and decoding (encoding and decoding, CODEC).

[0129] In the case of lossless data encoding, the original data can be reconstructed, that is, the reconstructed original data has the same quality as the original data (assuming there is no transmission loss or other data loss during storage or transmission). In the case of lossy data encoding, further compression is performed through quantization, etc. to reduce the amount of data required to represent the original data, but the decoder side cannot fully reconstruct the original data, that is, the quality of the reconstructed original data is lower or worse than the quality of the original data.

[0130] The embodiments of the present application can be applied to video data and other data with compression / decompression requirements. The following uses the encoding of video data (referred to as video encoding) as an example to illustrate the embodiments of the present application. Other types of data (such as image data, audio data, integer data, and other data with compression / decompression requirements) can refer to the following description, and the embodiments of the present application will not be repeated here. It should be noted that, compared to video encoding, the encoding process of data such as audio data and integer data does not require data to be divided into blocks, but the data can be directly encoded.

[0131] Video coding generally refers to processing a sequence of images to form a video or video sequence. In the field of video coding, the terms "picture", "frame" or "image" can be used as synonyms.

[0132] Several video coding standards fall under the category of "lossy hybrid video codecs" (i.e., combining spatial and temporal prediction in the pixel domain with 2D transform coding in the transform domain for applying quantization). Each image in a video sequence is typically divided into a set of non-overlapping blocks, which are typically coded at the block level. In other words, the encoder typically processes, i.e., encodes, the video at the block (video block) level, for example, by generating a prediction block through spatial (intra-frame) and temporal (inter-frame) prediction; subtracting the prediction block from the current block (currently processed / to-be-processed block) to obtain a residual block; transforming and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compressed), while the decoder applies the inverse of the encoder's processing to the coded or compressed block to reconstruct the current block for representation. Furthermore, the encoder needs to repeat the decoder's processing steps so that the encoder and decoder generate the same predictions (e.g., intra-frame predictions and inter-frame predictions) and / or reconstructed pixels for processing, i.e., encoding, the subsequent block.

[0133] In the following embodiment of the decoding system 10 , the encoder 20 and the decoder 30 are described with reference to FIG. 1 a to FIG. 3 .

[0134] FIG1a is an exemplary block diagram of a decoding system 10 provided in an embodiment of the present application, such as a video decoding system 10 (or simply, decoding system 10) that can utilize the techniques of the embodiments of the present application. The video encoder 20 (or simply, encoder 20) and video decoder 30 (or simply, decoder 30) in the video decoding system 10 represent devices that can be used to perform various techniques according to the various examples described in the embodiments of the present application.

[0135] As shown in FIG. 1 a , a decoding system 10 includes a source device 12 for providing encoded image data 21 such as an encoded image to a destination device 14 for decoding the encoded image data 21 .

[0136] The source device 12 includes an encoder 20 , and optionally, may include an image source 16 , a preprocessor (or preprocessing unit) 18 such as an image preprocessor, and a communication interface (or communication unit) 22 .

[0137] The image source 16 may include or may be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer-animated images, or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage that stores any of the above images.

[0138] In order to distinguish the processing performed by the pre-processor (or pre-processing unit) 18 , the image (or image data) 17 may also be referred to as a raw image (or raw image data) 17 .

[0139] The preprocessor 18 is configured to receive raw image data 17 and preprocess the raw image data 17 to obtain a preprocessed image (or preprocessed image data) 19. For example, the preprocessing performed by the preprocessor 18 may include cropping, color format conversion (e.g., from RGB to YCbCr), color grading, or denoising. It will be appreciated that the preprocessor 18 may be an optional component.

[0140] The video encoder (or encoder) 20 is used to receive the pre-processed image data 19 and provide encoded image data 21 (which will be further described below with reference to FIG. 2 and the like).

[0141] The communication interface 22 in the source device 12 can be used to receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device through the communication channel 13 for storage or direct reconstruction.

[0142] The destination device 14 includes a decoder 30 and, in addition or alternatively, may include a communication interface (or communication unit) 28 , a post-processor (or post-processing unit) 32 , and a display device 34 .

[0143] The communication interface 28 in the destination device 14 is used to receive the encoded image data 21 (or any other processed version) directly from the source device 12 or from any other source device such as a storage device, for example, the storage device is a encoded image data storage device, and provide the encoded image data 21 to the decoder 30.

[0144] The communication interface 22 and the communication interface 28 can be used to send or receive encoded image data (or encoded data) 21 through a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any combination thereof.

[0145] For example, the communication interface 22 may be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transmission coding or processing for transmission over a communication link or network.

[0146] The communication interface 28 corresponds to the communication interface 22 , and can be used, for example, to receive transmission data and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain the encoded image data 21 .

[0147] Both the communication interface 22 and the communication interface 28 can be configured as a unidirectional communication interface as indicated by the arrow pointing from the source device 12 to the corresponding communication channel 13 of the destination device 14 in Figure 1a, or a bidirectional communication interface, and can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded image data transmission, etc.

[0148] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31 (which will be further described below with reference to FIG. 3 and the like).

[0149] The post-processor 32 is configured to post-process the decoded image data 31 (also referred to as reconstructed image data) such as the decoded image to obtain post-processed image data 33 such as the post-processed image. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color grading, cropping, or resampling, or any other processing for generating the decoded image data 31 for display on a display device 34 or the like.

[0150] The display device 34 is configured to receive the post-processed image data 33 and display the image to a user or viewer. The display device 34 may be or include any type of display for displaying the reconstructed image, such as an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display screen.

[0151] The decoding system 10 also includes a training engine 25, which is used to train the encoder 20 (especially the entropy encoding unit 270 in the encoder 20) or the decoder 30 (especially the entropy decoding unit 304 in the decoder 30) to perform entropy encoding on the image block to be encoded based on the estimated probability distribution. For a detailed description of the training engine 25, please refer to the following method test example.

[0152] Although FIG1a shows source device 12 and destination device 14 as separate devices, device embodiments may also include both source device 12 and destination device 14 or the functions of both source device 12 and destination device 14, that is, both source device 12 or the corresponding functions and destination device 14 or the corresponding functions. In these embodiments, source device 12 or the corresponding functions and destination device 14 or the corresponding functions may be implemented using the same hardware and / or software or through separate hardware and / or software or any combination thereof.

[0153] According to the description, the existence and (accurate) division of different units or functions in the source device 12 and / or the destination device 14 shown in FIG. 1 a may vary depending on actual devices and applications, which is obvious to those skilled in the art.

[0154] Please refer to Figure 1b, which is an exemplary block diagram of a video decoding system 40 provided in an embodiment of the present application. The encoder 20 (e.g., video encoder 20) or the decoder 30 (e.g., video decoder 30), or both, can be implemented by processing circuitry in the video decoding system 40 shown in Figure 1b, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding-specific processors, or any combination thereof. Please refer to Figures 2 and 3, Figure 2 is an exemplary block diagram of a video encoder provided in an embodiment of the present application, and Figure 3 is an exemplary block diagram of a video decoder provided in an embodiment of the present application. The encoder 20 can be implemented by processing circuitry 46 to include the various modules discussed with reference to the encoder 20 in Figure 2 and / or any other encoder systems or subsystems described herein. The decoder 30 can be implemented by processing circuitry 46 to include the various modules discussed with reference to the decoder 30 in Figure 3 and / or any other decoder systems or subsystems described herein. The processing circuitry 46 can be used to perform the various operations discussed below. As shown in FIG4 , if part of the technology is implemented in software, the device can store the software instructions in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware, thereby performing the technology of the embodiment of the present application. One of the video encoder 20 and the video decoder 30 can be integrated into a single device as part of a combined codec (encoder / decoder, CODEC), as shown in FIG1 b.

[0155] The source device 12 and the destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as a notebook computer or laptop, a mobile phone, a smart phone, a tablet or a tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, and a monitoring device, etc., and may not use or use any type of operating system. The source device 12 and the destination device 14 may also be devices in a cloud computing scenario, such as a virtual machine in a cloud computing scenario. In some cases, the source device 12 and the destination device 14 may be equipped with components for wireless communication. Therefore, the source device 12 and the destination device 14 may be wireless communication devices.

[0156] The source device 12 and the destination device 14 may be installed with virtual scene applications (APPs) such as virtual reality (VR), augmented reality (AR), or mixed reality (MR), and may run the VR, AR, or MR applications based on user operations (e.g., click, touch, slide, shake, voice control, etc.). The source device 12 and the destination device 14 may capture images / videos of any objects in the environment through cameras and / or sensors, and then display virtual objects on a display device based on the captured images / videos. The virtual objects may be virtual objects in the VR, AR, or MR scenes (i.e., objects in the virtual environment).

[0157] It should be noted that in the embodiment of the present application, the virtual scene application in the source device 12 and the destination device 14 can be an application built into the source device 12 and the destination device 14 themselves, or it can be an application provided by a third-party service provider and installed by the user. There is no specific limitation on this.

[0158] In addition, the source device 12 and the destination device 14 may be installed with a real-time video transmission application, such as a live broadcast application. The source device 12 and the destination device 14 may capture images / videos through cameras and then display the captured images / videos on a display device.

[0159] In some cases, the video decoding system 10 shown in FIG1a is merely exemplary, and the techniques provided in embodiments of the present application may be applicable to video encoding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from a local memory, sent over a network, and so on. A video encoding device may encode data and store the data in a memory, and / or a video decoding device may retrieve data from a memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data to a memory and / or retrieve and decode data from a memory.

[0160] Please refer to Figure 1b, which is an exemplary block diagram of a video decoding system 40 provided in an embodiment of the present application. As shown in Figure 1b, the video decoding system 40 may include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video encoder / decoder implemented by a processing circuit 46), an antenna 42, one or more processors 43, one or more memory storage devices 44 and / or a display device 45.

[0161] 1b, imaging device 41, antenna 42, processing circuit 46, video encoder 20, video decoder 30, processor 43, memory 44, and / or display device 45 are capable of communicating with one another. In different embodiments, video decoding system 40 may include only video encoder 20 or only video decoder 30.

[0162] In some instances, antenna 42 can be used to transmit or receive an encoded bitstream of video data. Additionally, in some instances, display device 45 can be used to present the video data. Processing circuitry 46 can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. Video decoding system 40 can also include an optional processor 43, which can similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. Furthermore, memory storage 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, memory storage 44 can be implemented as cache memory. In other instances, processing circuitry 46 can include memory (e.g., cache memory, etc.) for implementing an image buffer, etc.

[0163] In some examples, video encoder 20 implemented by logic circuitry may include an image buffer (e.g., implemented by processing circuitry 46 or memory storage 44) and a graphics processing unit (e.g., implemented by processing circuitry 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video encoder 20 implemented by processing circuitry 46 to implement the various modules discussed with reference to FIG. 2 and / or any other encoder systems or subsystems described herein. Logic circuitry may be used to perform the various operations discussed herein.

[0164] In some examples, the video decoder 30 may be implemented by processing circuitry 46 in a similar manner to implement the various modules discussed with reference to the video decoder 30 of FIG. 3 and / or any other decoder systems or subsystems described herein. In some examples, the logic circuit implementation of the video decoder 30 may include an image buffer (implemented by processing circuitry 46 or memory storage 44) and a graphics processing unit (e.g., implemented by processing circuitry 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 30 implemented by processing circuitry 46 to implement the various modules discussed with reference to FIG. 3 and / or any other decoder systems or subsystems described herein.

[0165] In some examples, antenna 42 may be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data related to the encoded video frames, indicators, index values, mode selection data, etc., as discussed herein, such as data related to encoding partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators (as discussed), and / or data defining encoding partitions). Video decoding system 40 may also include video decoder 30 coupled to antenna 42 and configured to decode the encoded bitstream. Display device 45 is configured to present the video frames.

[0166] It should be understood that for the examples described herein with reference to video encoder 20, video decoder 30 can be configured to perform the reverse process. With respect to signaling syntax elements, video decoder 30 can be configured to receive and parse such syntax elements and decode the associated video data accordingly. In some examples, video encoder 20 can entropy encode the syntax elements into an encoded video bitstream. In such examples, video decoder 30 can parse such syntax elements and decode the associated video data accordingly.

[0167] For ease of description, the embodiments of the present application are described with reference to the universal video coding (VVC) reference software or the high-efficiency video coding (HEVC) developed by the joint collaboration team on video coding (JCT-VC) of the ITU-T video coding experts group (VCEG) and the ISO / IEC motion picture experts group (MPEG). Those skilled in the art will understand that the embodiments of the present application are not limited to HEVC or VVC.

[0168] Encoders and encoding methods

[0169] As shown in FIG2 , the video encoder 20 includes an input terminal (or input interface) 201, a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output terminal (or output interface) 272. The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254, and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG2 may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.

[0170] Images and image segmentation (images and blocks)

[0171] Encoder 20 is operable to receive, via input 201 or the like, an image (or image data) 17, for example, an image from a sequence of images forming a video or video sequence. The received image or image data may also be a pre-processed image (or pre-processed image data) 19. For simplicity, the following description uses image 17. Image 17 may also be referred to as a current image or image to be encoded (particularly when distinguishing the current image from other images in video encoding, such as previously encoded and / or decoded images in the same video sequence, i.e., a video sequence that also includes the current image).

[0172] A (digital) image is, or can be considered to be, a two-dimensional array or matrix of pixels with intensity values. The pixels in the array are also referred to as pixels (or pels, short for picture elements). The number of pixels in the array or image in the horizontal and vertical directions (or axes) determines the image size and / or resolution. To represent color, three color components are typically used, meaning that an image can be represented as or include three pixel arrays. In the RBG format or color space, an image includes corresponding arrays of red, green, and blue pixels. However, in video coding, each pixel is typically represented in a luma / chroma format or color space, such as YCbCr, which includes a luma component indicated by Y (sometimes also indicated by L) and two chroma components, indicated by Cb and Cr. The luma component Y represents the brightness or grayscale level intensity (for example, in grayscale images, both are the same), while the two chroma components (abbreviated as chroma) Cb and Cr represent the chroma or color information components. Accordingly, an image in YCbCr format includes a luma pixel array of luma pixel values ​​(Y) and two chroma pixel arrays of chroma values ​​(Cb and Cr). An image in RGB format can be converted or transformed into YCbCr format, and vice versa, a process also known as color conversion or transformation. If the image is black and white, the image may include only a luma pixel array. Accordingly, the image may be, for example, a luma pixel array in monochrome format or a luma pixel array and two corresponding chroma pixel arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0173] In one embodiment, an embodiment of the video encoder 20 may include an image segmentation unit (not shown in FIG. 2 ) for segmenting the image 17 into a plurality of (typically non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC) or coding tree blocks (CTBs), or coding tree units (CTUs) in the H.265 / HEVC and VVC standards. The segmentation unit may be used to use the same block size for all images in a video sequence and a corresponding grid of defined block sizes, or to vary the block size between images or subsets or groups of images, and to segment each image into corresponding blocks.

[0174] In other embodiments, the video encoder may be configured to directly receive a block 203 of the image 17, for example, one, several or all blocks constituting the image 17. The image block 203 may also be referred to as a current image block or an image block to be encoded.

[0175] Like image 17, image block 203 is also or can be considered to be a two-dimensional array or matrix of pixels having intensity values ​​(pixel values), but image block 203 is smaller than image 17. In other words, block 203 may include one pixel array (e.g., a luminance array in the case of monochrome image 17, or a luminance array or chrominance array in the case of a color image), or three pixel arrays (e.g., one luminance array and two chrominance arrays in the case of a color image 17), or any other number and / or type of arrays depending on the color format used. The number of pixels in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Accordingly, a block may be an M×N (M columns×N rows) pixel array, or an M×N transform coefficient array, etc.

[0176] In one embodiment, the video encoder 20 shown in FIG. 2 is configured to encode the image 17 block by block, for example, performing encoding and prediction on each block 203 .

[0177] In one embodiment, the video encoder 20 shown in FIG2 may also be configured to partition and / or encode an image using slices (also referred to as video slices), where an image may be partitioned or encoded using one or more slices (typically non-overlapping). Each slice may include one or more blocks (e.g., coding tree units (CTUs)) or one or more groups of blocks (e.g., tiles in the H.265 / HEVC / VVC standard and bricks in the VVC standard).

[0178] In one embodiment, the video encoder 20 shown in Figure 2 can also be used to segment and / or encode an image using slices / coding block groups (also called video coding block groups) and / or coding blocks (also called video coding blocks), where the image can be segmented or encoded using one or more slices / coding block groups (usually non-overlapping), each slice / coding block group may include one or more blocks (e.g., CTUs) or one or more coding blocks, etc., where each coding block can be in a shape such as a rectangle and may include one or more complete or partial blocks (e.g., CTUs).

[0179] Residual calculation

[0180] The residual calculation unit 204 is used to calculate the residual block 205 (the prediction block 265 is described in detail later) based on the image block (or original block) 203 and the prediction block 265 in the following manner: for example, the pixel value of the prediction block 265 is subtracted from the pixel value of the image block 203 pixel by pixel (pixel by pixel) to obtain the residual block 205 in the pixel domain.

[0181] Quantification

[0182] The quantization unit 208 is configured to quantize the transform coefficients 207 by, for example, scalar quantization or vector quantization to obtain quantized transform coefficients 209 . The quantized transform coefficients 209 may also be referred to as quantized residual coefficients 209 .

[0183] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, during quantization, an n-bit transform coefficient may be rounded down to an m-bit transform coefficient, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, for scalar quantization, varying degrees of scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. The appropriate quantization step size may be indicated by a quantization parameter (QP). For example, the quantization parameter may be an index into a predefined set of appropriate quantization step sizes. For example, a smaller quantization parameter may correspond to fine quantization (a smaller quantization step size), while a larger quantization parameter may correspond to coarse quantization (a larger quantization step size), or vice versa. Quantization may include dividing by the quantization step size, while the corresponding or inverse dequantization performed by the inverse quantization unit 210, etc., may include multiplying by the quantization step size. Embodiments according to some standards, such as HEVC, may be used to determine the quantization step size using the quantization parameter. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation involving division. Additional scaling factors can be introduced for quantization and dequantization to restore the norm of the residual block that may have been modified by the scaling used in the fixed-point approximation of the equations for the quantization step size and quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream, etc. Quantization is a lossy operation, where larger quantization step sizes result in greater losses.

[0184] In one embodiment, the video encoder 20 (correspondingly, the quantization unit 208) may be configured to output a quantization parameter (QP), for example, directly or after being encoded or compressed by the entropy coding unit 270, such that the video decoder 30 may receive and use the quantization parameter for decoding.

[0185] Dequantization

[0186] The inverse quantization unit 210 is configured to perform inverse quantization performed by the quantization unit 208 on the quantized coefficients to obtain dequantized coefficients 211. For example, the inverse quantization scheme performed by the quantization unit 208 may be performed according to or using the same quantization step size as the quantization unit 208. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211, and correspond to the transform coefficients 207. However, due to the loss caused by quantization, the dequantized coefficients 211 are generally not identical to the transform coefficients.

[0187] reconstruction

[0188] The reconstruction unit 214 (e.g., the summer 214) is used to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the pixel domain, for example, by adding the pixel point values ​​of the reconstructed residual block 213 and the pixel point values ​​of the prediction block 265.

[0189] segmentation

[0190] The partitioning unit 262 can partition (or divide) an image block (or CTU) 203 into smaller parts, such as square or rectangular blocks. For an image with three pixel arrays, a CTU consists of N×N luminance pixel blocks and two corresponding chrominance pixel blocks. The maximum allowed size of a luminance block in a CTU is specified as 128×128 in the developing Universal Video Coding (VVC) standard, but may be specified to a value other than 128×128, such as 256×256, in the future. The CTUs of an image can be grouped / collected into slices / coding block groups, coding blocks, or bricks. A coding block covers a rectangular area of ​​an image and can be divided into one or more bricks. A brick consists of multiple CTU rows within a coding block. A coding block that is not partitioned into multiple bricks can be called a brick. However, a brick is a true subset of a coding block and is therefore not called a coding block. VVC supports two coding block group modes: raster scan slice / coding block group mode and rectangular slice mode. In raster scan CBG mode, a slice / CBG contains a sequence of CBs from a raster scan of the CBs of an image. In rectangular slice mode, a slice contains multiple bricks of an image that together form a rectangular region of the image. The bricks within a rectangular slice are arranged in the slice's brick raster scan order. These smaller blocks (also called sub-blocks) can be further split into smaller parts. This is also known as tree partitioning or hierarchical tree partitioning, where a root block at, for example, root tree level 0 (hierarchy level 0, depth 0) can be recursively split into two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchy level 1, depth 1). These blocks can be further split into two or more blocks at the next lower level, such as nodes at tree level 2 (hierarchy level 2, depth 2), and so on, until the partitioning is completed (because the end criteria are met, such as reaching the maximum tree depth or minimum block size). Blocks that are not further split are also called leaf blocks or leaf nodes of the tree. A tree divided into two parts is called a binary tree (BT), a tree divided into three parts is called a ternary tree (TT), and a tree divided into four parts is called a quadtree (QT).

[0191] Entropy Coding

[0192] The entropy coding unit 270 is configured to apply an entropy coding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC (CALVC) scheme, an arithmetic coding scheme, a binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements, resulting in coded image data 21 that can be output via an output terminal 272 in the form of a coded bitstream 21, etc., so that the video decoder 30, etc. can receive and use the parameters for decoding. The coded bitstream 21 can be transmitted to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.

[0193] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform-based encoder 20 may directly quantize the residual signal without a transform processing unit 206 for certain blocks or frames. In another implementation, the encoder 20 may have the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.

[0194] Decoder and decoding method

[0195] As shown in FIG3 , a video decoder 30 is configured to receive coded image data 21 (e.g., a coded bitstream 21) encoded by, for example, an encoder 20, and obtain a decoded image 331. The coded image data or bitstream includes information for decoding the coded image data, such as data representing image blocks of a coded video slice (and / or coding block group or coding block) and related syntax elements.

[0196] 3 , decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., summer 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter-prediction unit 344, and an intra-prediction unit 354. Inter-prediction unit 344 may be or include a motion compensation unit. In some examples, video decoder 30 may perform a decoding process that is generally the reverse of the encoding process described with reference to video encoder 100 of FIG. 2 .

[0197] As described with respect to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer DPB 230, inter-prediction unit 344, and intra-prediction unit 354 also constitute the "built-in decoder" of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 122, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Therefore, the explanations of the corresponding units and functions of video encoder 20 apply accordingly to the corresponding units and functions of video decoder 30.

[0198] Entropy decoding

[0199] The entropy decoding unit 304 is configured to parse the bitstream 21 (or generally, the encoded image data 21) and perform entropy decoding on the encoded image data 21 to obtain quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3 ), such as any or all of inter-frame prediction parameters (e.g., reference image indices and motion vectors), intra-frame prediction parameters (e.g., intra-frame prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the coding scheme of the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 may also be configured to provide inter-frame prediction parameters, intra-frame prediction parameters, and / or other syntax elements to the mode application unit 360, as well as to provide other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice and / or video block level. In addition to, or in lieu of, slices and corresponding syntax elements, coding block groups and / or coding blocks and corresponding syntax elements may also be received or used.

[0200] Dequantization

[0201] The inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally information related to inverse quantization) and quantization coefficients from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse quantize the decoded quantization coefficients 309 based on the quantization parameter to obtain inverse quantization coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include using the quantization parameter calculated by the video encoder 20 for each video block in the video slice to determine a degree of quantization, and thus a degree of inverse quantization to be performed.

[0202] reconstruction

[0203] The reconstruction unit 314 (eg, summer 314 ) is configured to add the reconstructed residual block 313 to the prediction block 365 to obtain the reconstructed block 315 in the pixel domain, eg, by adding the pixel values ​​of the reconstructed residual block 313 and the pixel values ​​of the prediction block 365 .

[0204] Other variations of the video decoder 30 may be used to decode the encoded image data 21. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 may directly inverse quantize the residual signal without the inverse transform processing unit 312 for certain blocks or frames. In another implementation, the video decoder 30 may have the inverse quantization unit 310 and the inverse transform processing unit 312 combined into a single unit.

[0205] It should be understood that the processing result of the current step can be further processed in the encoder 20 and the decoder 30 and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, the processing result of interpolation filtering, motion vector derivation, or loop filtering can be further operated on, such as clipping or shifting operations.

[0206] It should be noted that further operations can be performed on the derived motion vector of the current block (including but not limited to the control point motion vector of the affine mode, the sub-block motion vector of the affine, planar, ATMVP mode, the temporal motion vector, etc.). For example, the value of the motion vector is limited to a predefined range based on the representation bit of the motion vector. If the representation bit of the motion vector is bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" represents a power. For example, if bitDepth is set to 16, the range is -32768 to 32767; if bitDepth is set to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of four 4×4 sub-blocks in an 8×8 block) is limited so that the maximum difference between the integer parts of the four 4×4 sub-block MVs does not exceed N pixels, for example, not more than 1 pixel. Two methods of limiting the motion vector based on bitDepth are provided here.

[0207] Although the above embodiments primarily describe video coding, it should be noted that embodiments of the decoding system 10, encoder 20, and decoder 30, as well as other embodiments described herein, may also be used for still image processing or coding, i.e., processing or coding a single image in a video codec that is independent of any previous or subsequent images. In general, if image processing is limited to a single image 17, the inter-frame prediction unit 244 (encoder) and the inter-frame prediction unit 344 (decoder) may not be available. All other functionalities (also referred to as tools or techniques) of the video encoder 20 and video decoder 30, such as residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354, and / or loop filtering 220 / 320, entropy coding 270, and entropy decoding 304, may also be used for still image processing.

[0208] Please refer to Figure 4, which is an exemplary block diagram of a video decoding device 400 provided in an embodiment of the present application. Video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, video decoding device 400 can be a decoder, such as the video decoder 30 in Figure 1a, or an encoder, such as the video encoder 20 in Figure 1a.

[0209] Video decoding device 400 includes: an input port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; for example, processor 430 may be a neural network processor 430; a transmitter unit (Tx) 440 and an output port 450 (or output port 550) for transmitting data; and a memory 460 for storing data. Video decoding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to input port 410, receiver unit 420, transmitter unit 440, and output port 450 for outputting or receiving optical or electrical signals.

[0210] The processor 430 is implemented in hardware and software. The processor 430 can be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 communicates with the input port 410, the receiving unit 420, the transmitting unit 440, the output port 450, and the memory 460. The processor 430 includes a neural network-based codec 470. The neural network-based codec 470 implements the embodiments disclosed above. For example, the neural network-based codec 470 performs, processes, prepares, or provides various encoding operations. Therefore, the neural network-based codec 470 provides substantial improvements to the functionality of the video decoding device 400 and affects the switching of the video decoding device 400 to different states. Alternatively, the neural network-based codec 470 is implemented by instructions stored in the memory 460 and executed by the processor 430.

[0211] Memory 460 includes one or more disks, tape drives, and solid-state drives and can be used as overflow data storage for storing programs when such programs are selected for execution, and for storing instructions and data read during program execution. Memory 460 can be volatile and / or non-volatile and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0212] Please refer to FIG. 5 , which is an exemplary block diagram of an apparatus 500 provided in an embodiment of the present application. The apparatus 500 may be used as either or both of the source device 12 and the destination device 14 in FIG. 1 a .

[0213] The processor 502 in the apparatus 500 may be a central processing unit (CPU). Alternatively, the processor 502 may be any other type of device or devices, now available or developed in the future, capable of manipulating or processing information. While the disclosed implementations may be implemented using a single processor, such as the processor 502 shown, using more than one processor may provide greater speed and efficiency.

[0214] In one implementation, the memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that are accessed by the processor 502 via a bus 512. The memory 504 may also include an operating system 508 and application programs 510, which include at least one program that allows the processor 502 to perform the methods described herein. For example, the application programs 510 may include applications 1 through N, as well as a video decoding application that performs the methods described herein.

[0215] The apparatus 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with touch-sensitive elements that can be used to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.

[0216] Although bus 512 in device 500 is described herein as a single bus, bus 512 may include multiple buses. Furthermore, secondary storage may be directly coupled to other components of device 500 or accessed via a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, device 500 may have a variety of configurations.

[0217] Image encoding:

[0218] Today, multimedia data accounts for the majority of internet traffic. Image data compression plays a vital role in the storage and efficient transmission of multimedia data. Therefore, image coding technology is a highly practical technology. It is important to note that the English terms "image coding" and "image encoding" are often translated into "image coding" in Chinese. Image coding is a broad term that encompasses both the process of encoding an image into a bitstream and the process of decoding (decoding) the bitstream back into an image.

[0219] Image coding simply refers to the process of encoding an image into a bitstream. Image coding research has a long history. Researchers have proposed numerous methods and developed I-frame encoding schemes for various image and video coding standards, including JPEG, JPEG 2000, JPEG-XL, JPEG-XX, WebP, H.264 / AVC, H.264 / HEVC, H.26 / VVC, AVS3, and AV1. Most of these coding methods are based on transform, prediction, and entropy coding techniques. While these coding methods are currently widely used, the increase in image data volumes and the emergence of new media types have necessitated the need for coding methods with higher compression efficiency.

[0220] Image encoding based on deep learning:

[0221] In recent years, researchers have been studying deep learning-based image coding methods. Some have achieved promising results. For example, Balle et al. proposed an end-to-end optimized image coding method that outperforms the best existing image coding methods and even the best existing traditional coding standard, H.265 / HEVC.

[0222] Deep learning-based image coding relies on deep neural networks, typically convolutional neural networks. Several research works have proposed image coding methods based on Transformer networks. The architecture of a deep neural network can be designed manually or obtained through neural architecture search (NAS). The parameters of a deep neural network are obtained by using a loss function and a backpropagation algorithm.

[0223] Figure 6 shows a typical deep learning-based image compression method, also known as a neural network-based image compression method. Generally, a neural network-based image compression method includes the following components: a feature extraction module, a feature quantization module, an entropy coding module, an entropy decoding module, a feature dequantization module, and a feature decoding module. On the encoder side, the feature extraction module can use a nonlinear mapping activation function to obtain an extracted three-dimensional feature map through multi-layer convolution stacking. The feature quantization module quantizes the eigenvalues ​​of floating-point numbers through eigenvalue quantization to obtain quantized eigenvalues. The quantized eigenvalues ​​are losslessly entropy encoded to obtain an encoded bitstream. Upon receiving the entropy-coded bitstream, the decoder performs lossless entropy decoding to obtain the three-dimensional quantized eigenvalues. The feature decoding module decodes the features into a reconstructed image to achieve decoding.

[0224] After the image to be compressed passes through the feature extraction module and the feature quantization module, a three-dimensional feature quantization map is generated. When processing each eigenvalue in the three-dimensional feature quantization map, the entropy coding module uses the eigenvalues ​​in the processed neighborhood as context to estimate and obtain the probability distribution of the eigenvalue. Based on this probability distribution, subsequent encoding is performed to obtain the encoded bitstream.

[0225] With the outstanding performance of deep learning in various fields, researchers have proposed an end-to-end image coding solution based on deep learning. Figure 7 shows the coding framework. The specific technical solution is as follows: On the encoder side, the original image is input into the feature extraction module, and the feature map is output. The feature map passes through the side information extraction module, and the side information is output. On the encoder and decoder sides, it is input into the probability estimation module, and the probability distribution of each feature element is output to obtain the value of the feature element to be encoded. In addition, the feature map is input into the quantization module to obtain a quantized feature map. The entropy coding module performs entropy coding on each feature element in the quantized feature map based on the probability distribution of each feature element to obtain the encoded bit stream.

[0226] On the decoder side, the decoder parses the bitstream and outputs the probability distribution of the symbols to be encoded based on the accompanying information to obtain the value of the feature element to be decoded. The k-entropy decoding module performs arithmetic decoding on each feature element in the quantized feature map based on the probability distribution of each feature element to obtain the value of the feature element. The feature map is input to the image reconstruction module, and the reconstructed image is output.

[0227] Neural Networks

[0228] A neural network may include neurons. A neuron may be an operation unit that uses xs and an intercept 1 as input. The output of the operation unit may be:

[0229] Here, s=s=1,2,...,n, n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neuron, and f is the activation function of the neuron (activation function), which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neuron into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network constructed by connecting multiple single neurons together. Specifically, the output of a neuron can be the input to another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the characteristics of the local receptive field. The local receptive field can be an area including several neurons.

[0230] Convolutional Neural Networks:

[0231] A convolutional neural network (CNN) is a deep neural network with a convolutional architecture. A CNN includes a feature extractor, which consists of convolutional layers and subsampling layers. The feature extractor can be considered a filter. A convolutional layer is a layer of neurons in a CNN that performs convolution processing on the input signal. In a convolutional layer of a CNN, a neuron may only be connected to a subset of neurons in an adjacent layer. A convolutional layer typically includes several feature planes, each of which may include a number of neurons arranged in a rectangular shape. Neurons on the same feature plane share a weight, which is the convolution kernel. Shared weights can be understood as a position-independent way of extracting image information. The convolution kernel can be initialized as a matrix of random size. During the training process of a CNN, appropriate weights can be obtained for the convolution kernel through learning. In addition, shared weights directly reduce the number of connections between CNN layers and reduce the risk of overfitting.

[0232] Figure 8 schematically illustrates the general concept of processing by a neural network such as a CNN. A convolutional neural network consists of an input layer, an output layer, and multiple hidden layers. The input layer provides the input (e.g., a portion of an image, as shown in Figure 8) for processing. The hidden layers of a CNN typically consist of a series of convolutional layers, which perform convolutions with multiplications or other dot products. The result of these layers is one or more feature maps, sometimes also called channels. Subsampling may be involved in some or all layers. Therefore, as shown in Figure 8, the feature maps can become smaller. The activation function in a CNN is typically a Rectified Linear Unit (RELU) layer, followed by additional convolutions such as pooling, fully connected, and normalization layers. These are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. Although these layers are colloquially referred to as convolutions, this is just convention. Mathematically, it is technically a sliding dot product or cross-correlation. This has important implications for the indices in the matrix, as it affects how weights are determined at specific index points.

[0233] When programming a CNN for image processing, as shown in Figure 8, the input is a tensor with a shape of (number of images) x (image width) x (image height) x (image depth). Then, after passing through the convolutional layer, the image is abstracted into a feature map with a shape of (number of images) x (feature map width) x (feature map height) x (feature map channels). A convolutional layer in a neural network should have the following properties: A convolution kernel defined by width and height (hyperparameters). The number of input and output channels (hyperparameters). The depth of the convolution filter (input channels) should be equal to the number of channels in the input feature map (depth).

[0234] Traditional multilayer perceptron (MLP) models have been used for image recognition in the past. However, due to the full connectivity between nodes, their dimensionality is high and does not scale well with higher resolution images. A 1000×1000 pixel image with RGB color channels has 3 million weights, which is too high to be processed efficiently at scale with full connectivity. Furthermore, this network architecture does not account for the spatial structure of the data, treating input pixels that are far apart in the same way as pixels that are close together. This ignores the locality of reference in image data, both computationally and semantically. Therefore, full connectivity of neurons is wasteful for purposes such as image recognition, which is dominated by spatially local input patterns.

[0235] Convolutional neural networks are biologically inspired variants of multilayer perceptrons, specifically designed to mimic the behavior of the visual cortex. These models alleviate the challenges posed by MLP architectures by exploiting the strong spatially local correlations present in natural images. The convolutional layer is the core building block of a CNN. The parameters of this layer consist of a set of learnable filters (the kernels mentioned above) that have a small receptive field but extend to the full depth of the input volume. During the forward pass, each filter is convolved across the width and height of the input volume, computing the dot product between the entries of the filter and the input, and producing a two-dimensional activation map for that filter. Thus, the network learns filters that activate when it detects a certain type of feature at a certain spatial location in the input.

[0236] The activation maps of all filters stacked along the depth dimension form the complete output volume of the convolutional layer. Therefore, each entry in the output volume can also be interpreted as the output of a neuron that observes a small region in the input and shares parameters with neurons in the same activation map. A feature map, or activation map, is the output activation of a given filter. Feature map and activation have the same meaning. In some papers, it is called an activation map because it is a map corresponding to the activations of different parts of the image, and a feature map because it is also a map of certain features found in the image. High activation means that a certain feature has been found.

[0237] Another important concept in CNNs is pooling, a form of nonlinear downsampling. There are several nonlinear functions that implement pooling, with max pooling being the most common. It divides the input image into a set of non-overlapping rectangles and, for each such subregion, outputs the maximum value.

[0238] Intuitively, the precise location of a feature relative to other features is less important than its coarse location. This is the idea behind pooling in convolutional neural networks. Pooling layers are used to gradually reduce the spatial size of the representation, reducing the number of parameters, memory usage, and computational overhead in the network, thereby controlling overfitting. In CNN architectures, it is common to periodically insert pooling layers between consecutive convolutional layers. Pooling provides another form of translation invariance.

[0239] Pooling layers operate independently on each depth slice of the input and resize it spatially. The most common form is a pooling layer with filters of size 2×2, applying a downsampling stride of 2 along both width and height to each depth slice in the input, discarding 75% of the activations. In this case, each maximum operation exceeds 4 numbers. The depth dimension remains unchanged.

[0240] In addition to max pooling, the pooling unit can also use other functions such as average pooling or 2-norm pooling. Average pooling was often used historically but has recently fallen out of favor compared to max pooling, which performs better in practice. Due to the large reduction in representation size, there has been a recent trend to use smaller filters or to drop pooling layers entirely. "Region of interest" pooling (also known as ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important component of convolutional neural networks for object detection based on the Faster R-CNN architecture.

[0241] The aforementioned ReLU, short for Rectified Linear Unit, applies a non-saturating activation function. By setting negative values ​​to zero, it effectively removes negative values ​​from the activation map. This adds nonlinearity to the decision function and the entire network without affecting the receptive field of the convolutional layer. Other functions are also used to add nonlinearity, such as the saturated hyperbolic tangent and the sigmoid function. ReLU is often preferred over other functions because it trains neural networks several times faster without significantly affecting generalization accuracy.

[0242] After several convolutional and max-pooling layers, high-level reasoning in a neural network is done through fully connected layers. Neurons in a fully connected layer have connections to all activations in the previous layer, as seen in regular (non-convolutional) artificial neural networks. Therefore, their activations can be calculated as affine transformations, matrix multiplications followed by bias shifts (vector addition of learned or fixed bias terms).

[0243] The "loss layer" specifies how training penalizes deviations between predictions (outputs) and true labels and is typically the last layer in a neural network. A variety of loss functions can be used, appropriate for different tasks. Softmax loss is used to predict a single class among K mutually exclusive classes. Sigmoid cross-entropy loss is used to predict K independent probability values ​​in the range [0, 1]. Euclidean loss is used to regress to real-valued labels.

[0244] In summary, Figure 8 shows the data flow in a typical convolutional neural network. First, the input image passes through a convolutional layer and is abstracted into a feature map consisting of several channels, corresponding to the number of filters in the set of learnable filters in that layer (e.g., one channel per filter). The feature map is then subsampled using, for example, a pooling layer, which reduces the dimensionality of each channel in the feature map. The data then enters another convolutional layer, which may have a different number of output channels, resulting in a different number of channels in the feature map. As mentioned above, the number of input channels and output channels are hyperparameters of the layer. In order to establish the connectivity of the network, these parameters need to be synchronized between the two connected layers, for example, the number of input channels of the current layer should be equal to the number of output channels of the previous layer. For the first layer that processes the input data (e.g., an image), the number of input channels is typically equal to the number of channels of the data representation, such as 3 channels for RGB or YUV representation of images or videos, or 1 channel for grayscale image or video representation.

[0245] The present invention is applicable to video codecs, such as those in video communication systems, as shown in Figure 9: After video is captured using a video capture device, it undergoes a series of pre-processing steps and is then compressed and encoded to produce an encoded bitstream. A transmitting module transmits the bitstream via a transmission network to a receiving module, where it is decoded by a decoder and then rendered for display. Furthermore, the encoded bitstream can also be directly stored.

[0246] The present invention describes a neural network-based coding and decoding scheme for enhancement layers, which can be applied in a layered coding scheme.

[0247] The present invention can be used in devices or products containing video encoder and / or decoder functions, such as video processing hardware and software products, such as chips, as well as products or devices containing such chips, such as mobile phones and other media products.

[0248] FIG10 shows an encoding method provided by an embodiment of the present application. As shown in FIG10 , the method includes:

[0249] S1001. Perform AI encoding on an input image to obtain a first feature map of the input image.

[0250] Wherein, a reconstructed image can be obtained by decoding any one of the N groups of second feature maps, where N is a positive integer.

[0251] The first feature map is a three-dimensional data. As shown in Figure 8, the three-dimensional feature map is composed of multiple two-dimensional feature maps, each of which includes multiple elements (pixels). In the channel dimension, each two-dimensional feature map corresponds to a channel (such as an RPG channel).

[0252] In a possible implementation, a feature extraction module of an AI encoder may be used to perform feature extraction on an input image to obtain a first feature map of the input image.

[0253] Exemplarily, the feature extraction module of the JPEG AI encoder may be used to perform feature extraction on the input image to obtain a first feature map of the input image.

[0254] S1002: Determine N groups of second feature maps of the input image according to the first feature map of the input image.

[0255] Wherein, N is a positive integer. For example, N can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or other positive integers. Any of the N sets of second feature maps can be decoded to obtain a reconstructed image.

[0256] The second feature map is a two-dimensional data, as shown in Figure 8. As shown in Figure 8, the three-dimensional feature map is composed of multiple two-dimensional feature maps. Each two-dimensional feature map corresponds to a channel in the channel dimension, so the three-dimensional feature map can be divided into multiple two-dimensional feature maps in the channel dimension.

[0257] For example, a three-dimensional feature map with 192 channels can be channel-separated in the channel dimension to obtain 192 two-dimensional feature maps.

[0258] In one possible implementation, the first feature map may be subjected to channel separation to obtain M second feature maps of the input image. The M second feature maps are grouped according to channel information of the second feature map of the input image to obtain the N groups of second feature maps. M is a positive integer, M≤N. For example, M may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or other positive integers.

[0259] Here, M may be the number of channels of the first feature map.

[0260] In a possible implementation, the channel information includes channel entropy and / or channel variance, wherein the channel entropy is the entropy of all elements in the channel, and the channel variance is the variance of all elements in the channel.

[0261] For example, taking the example that each group of second feature maps includes K second feature maps, the K second feature maps with the highest channel variance can be divided into the first group, the K second feature maps with the second highest channel variance can be divided into the second group,..., the K second feature maps with the second lowest channel variance can be divided into the N-1th group, and the K second feature maps with the lowest channel variance can be divided into the Nth group.

[0262] For another example, taking the example that each group of second feature maps includes K second feature maps, the K second feature maps with the highest channel entropy can be divided into the first group, the K second feature maps with the second highest channel entropy can be divided into the second group,..., the K second feature maps with the second lowest channel entropy can be divided into the N-1th group, and the K second feature maps with the lowest channel entropy can be divided into the Nth group.

[0263] For another example, taking each group of second feature maps including K second feature maps, the second feature maps can be numbered in ascending order of channel entropy. Then, the K second feature maps with the highest numbers are grouped into the first group, the K second feature maps with the second highest numbers are grouped into the second group, ..., the K second feature maps with the second lowest numbers are grouped into the N-1th group, and the K second feature maps with the lowest numbers are grouped into the Nth group.

[0264] Exemplarily, if the first feature map has 384 channels, the first feature map can be subjected to channel separation to obtain 384 second feature maps of the input image. Then, according to the channel entropy of the 384 second feature maps, the 192 second feature maps with the highest channel entropy among the 384 second feature maps are divided into one group, and the 192 second feature maps with the lowest channel entropy are divided into another group.

[0265] As another example, as shown in FIG11 , the first feature map can be subjected to channel separation to obtain M second feature maps of the input image, and the M second feature maps can be divided into N groups in descending order of their channel entropy or channel variance.

[0266] As another example, as shown in FIG12 , the first feature map can be subjected to channel separation to obtain M second feature maps of the input image, and the M second feature maps can be divided into two groups in descending order of their channel entropy or channel variance.

[0267] S1003 , entropy encode the first group of second feature maps among the N groups of second feature maps of the input image into a bitstream.

[0268] In a possible implementation, the first group of second feature maps among the N groups of second feature maps of the input image may be entropy-coded into a bitstream using an asymmetric numeral system (ANS) algorithm.

[0269] In a possible implementation, the second group of second feature maps among the N groups of second feature maps may be entropy encoded into the bitstream.

[0270] In a possible implementation, the third group of second feature maps among the N groups of second feature maps may be entropy encoded into the bitstream.

[0271] In a possible implementation, at least one group of the N groups of second feature maps may be entropy encoded into the bitstream in descending order of channel entropy or channel variance.

[0272] Exemplarily, as shown in FIG11 , after the M second feature maps of the input image are divided into N groups, the N groups of second feature maps of the input image can be entropy encoded into the bitstream in descending order of channel entropy or channel variance.

[0273] As another example, as shown in FIG12 , after the M second feature maps of the input image are divided into 2 groups, the 2 groups of second feature maps of the input image can be entropy encoded into the bitstream in descending order of channel entropy or channel variance.

[0274] As another example, as shown in FIG13 , after the M second feature maps of the input image are divided into 2 groups, only the first group of second feature maps of the input image can be entropy encoded into the bitstream in descending order of channel entropy or channel variance.

[0275] It can be seen that in the method provided in the embodiment of the present application, a progressive encoding method is adopted to encode part of the feature map of the input image. Compared with encoding all the feature maps of the input image, encoding part of the feature map of the input image can reduce the image encoding delay. In addition, the first group of second feature maps of the input image is the feature map with the maximum channel entropy (variance) among the N groups of second feature maps of the input image, and the decoding end can obtain a complete reconstructed image based on the feature map.

[0276] For example, in a scenario with a poor network environment or insufficient decoding hardware resources, the embodiment of the present application adopts a progressive encoding method that does not require all feature maps. A complete reconstructed image can be obtained based only on decoding of the first group of second feature maps, thereby reducing the image encoding and decoding delay.

[0277] In one possible implementation, channel number information may be encoded into the bitstream. The channel number information indicates the number of channels of at least one set of second feature maps encoded into the bitstream. This allows a decoder to determine the number of channels of each set of second feature maps in the bitstream by decoding the channel number information, thereby facilitating recovery of the feature maps from the bitstream.

[0278] For example, the number of channels of the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps are 4, 3, and 3, respectively. Then, the channel number information of 4, 3, and 3 can be encoded into the above code stream.

[0279] In one possible implementation, first position information may be encoded into the bitstream, where the first position information indicates the position of at least one set of second feature maps encoded into the bitstream within the first feature map. This allows a decoder to determine the position of each set of second feature maps within the first feature map by decoding the first position information, thereby facilitating recovery of pre-group feature maps from the bitstream.

[0280] It should be noted that the decoding end needs to use the position of the second feature map in the above-mentioned first feature map to obtain the second feature map, and the position of the second feature map in the code stream may be different from the position of the second feature map in the above-mentioned first feature map. For this reason, it is necessary to encode the first position information of at least one group of second feature maps encoded into the above-mentioned code stream in the above-mentioned first feature map into the code stream.

[0281] For example, the first feature map is separated into 10 second feature maps through channels. The positions of these 10 second feature maps in the first feature map are recorded as 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 respectively. These 10 second feature maps are grouped and sorted and encoded into the bitstream in the order of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 respectively. Therefore, the first position information of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 can be encoded into the bitstream. After obtaining this information, the decoder can know that the position of the first second feature map obtained by decoding in the first feature map is 2, the position of the second second feature map is 9, and so on. The position of the tenth second feature map is 8 in the first feature map.

[0282] In one possible implementation, second position information may be encoded into the bitstream. The second position information indicates the position of at least one set of second feature maps encoded into the bitstream. This allows a decoder to determine the position of each set of second feature maps in the bitstream using the second position information, thereby facilitating parallel decoding.

[0283] In one possible implementation, the code stream length information may be added to the code stream, and the code stream length information is used to indicate the length of the code stream. In this way, the decoding end can determine the length of the code stream through the code stream length information, thereby facilitating parallel decoding at the decoding end.

[0284] FIG14 shows a decoding method provided by an embodiment of the present application. As shown in FIG14 , the method includes:

[0285] S1401: Perform entropy decoding on a code stream to obtain a first group of second feature maps of an input image.

[0286] Among them, the above-mentioned first group of second feature maps is a part of the feature map corresponding to the above-mentioned input image.

[0287] In a possible implementation, ANS entropy decoding may be performed on the code stream to obtain the first group of second feature maps of the input image.

[0288] In a possible implementation, the code stream may be entropy decoded to obtain a second group of second feature maps of the input image.

[0289] In a possible implementation, the code stream may be entropy decoded to obtain a third group of second feature maps of the input image.

[0290] In a possible implementation, the code stream may be entropy decoded to sequentially obtain L groups of second feature maps of the image, where L is a positive integer and is less than or equal to N.

[0291] Exemplarily, as shown in FIG11 , the code stream may be entropy decoded to sequentially obtain N groups of second feature maps of the image.

[0292] As another example, as shown in FIG12 , the bitstream is entropy decoded to obtain the first group of second feature maps of the input image, and then the bitstream is entropy decoded to obtain the second group of second feature maps of the input image.

[0293] As another example, as shown in FIG13 , the code stream may be entropy decoded to obtain only the first group of second feature maps of the input image.

[0294] S1402: Perform AI decoding on the first group of second feature maps of the input image to obtain a first reconstructed image.

[0295] In one possible implementation, the image restoration network of the AI ​​decoder may be used to perform AI decoding on the first group of second feature maps of the input image to obtain a first reconstructed image.

[0296] Exemplarily, the first reconstructed image can be obtained by performing AI decoding on the first group of second feature maps of the input image through the image restoration network of the JPEG AI decoder.

[0297] It should be noted that the first reconstructed image is not a partial image of the input image, but a reconstructed image having the same size as the input image but different image quality.

[0298] In a possible implementation, each time a set of second feature maps is solved, they can be merged with other solved second feature maps and sent to a subsequent image restoration network to restore the image.

[0299] Exemplarily, after AI decoding obtains the second group of second feature maps of the input image, AI decoding may be performed on the first group of second feature maps and the second group of second feature maps to obtain a second reconstructed image.

[0300] Exemplarily, after AI decoding obtains the third group of second feature maps of the input image, AI decoding can be performed on the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps to obtain a third reconstructed image.

[0301] In a possible implementation, the image quality of the second reconstructed image is better than that of the first reconstructed image, and the image quality of the third reconstructed image is better than that of the second reconstructed image.

[0302] The image quality is characterized by any one of the following variables: PSNR, MS-SSIM, or LPIPS.

[0303] Exemplarily, as shown in FIG11 , after decoding the first group of second feature maps of the input image, the first group of second feature maps of the input image are input into the image restoration network to obtain a first reconstructed image; after decoding the second group of second feature maps, the first group of second feature maps of the input image and the second group of second feature maps of the input image are input into the image restoration network to obtain a second reconstructed image; after decoding the third group of second feature maps of the input image, the first group of second feature maps of the input image, the second group of second feature maps of the input image and the third group of second feature maps of the input image are input into the image restoration network to obtain a third reconstructed image; after decoding the third group of second feature maps of the input image, the first group of second feature maps of the input image, the second group of second feature maps of the input image and the third group of second feature maps of the input image are input into the image restoration network to obtain a third reconstructed image; After the N-1th group of second feature maps of the input image are input, the 1st group of second feature maps of the input image, the 2nd group of second feature maps of the input image, the 3rd group of second feature maps of the input image, ..., the N-1th group of second feature maps of the input image are input into the image restoration network to obtain the N-1th reconstructed image; after decoding the Nth group of second feature maps of the input image, the 1st group of second feature maps of the input image, the 2nd group of second feature maps of the input image, the 3rd group of second feature maps of the input image, ..., the N-1th group of second feature maps of the input image, and the Nth group of second feature maps of the input image are input into the image restoration network to obtain the Nth reconstructed image.

[0304] It should be noted that the image quality of the first reconstructed image to the Nth reconstructed image increases sequentially. That is, the image quality of the first reconstructed image < the image quality of the second reconstructed image < the image quality of the third reconstructed image < ... < the image quality of the N-1th reconstructed image < the image quality of the Nth reconstructed image.

[0305] As another example, as shown in Figure 12, after decoding the first group of second feature maps of the input image, the first group of second feature maps of the input image can be input into the image restoration network to obtain a first reconstructed image; after decoding the second group of second feature maps of the input image, the first group of second feature maps of the input image and the second group of second feature maps of the input image are input into the image restoration network to obtain a second reconstructed image.

[0306] As another example, as shown in FIG13 , after decoding the first group of second feature maps of the input image, the first group of second feature maps of the input image may be input into an image restoration network to obtain a first reconstructed image.

[0307] It can be seen that in the method provided in the embodiment of the present application, a reconstructed image can be determined based on a partial feature map of the input image. Compared with determining the reconstructed image based on all the feature maps of the input image, the image decoding delay can be reduced by using a progressive decoding method to determine the reconstructed image based on a partial feature map of the input image. In addition, the first group of second feature maps of the input image is the feature map with the maximum channel entropy (variance) in the second feature map of the input image, and the decoding end can obtain a complete reconstructed image based on the feature map.

[0308] For example, in a scenario where the network environment is poor or decoding hardware resources are insufficient, the embodiment of the present application adopts a progressive decoding method to obtain a complete reconstructed image based only on the decoding of the first group of second feature maps.

[0309] In one possible implementation, the code stream may be decoded to obtain channel number information, where the channel number information is used to indicate the number of channels of at least one group of second feature maps obtained by decoding; and the first group of second feature maps may be AI decoded according to the channel number information to obtain a first reconstructed image.

[0310] For example, the number of channels of the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps are 4, 3, and 3, respectively. The channel number information of 4, 3, and 3 can be encoded into the above-mentioned code stream. Correspondingly, after the decoding end obtains the channel number information through decoding, it can know that the first decoding must obtain the second feature maps of 4 channels in order to decode the first group of second feature maps of the input image, the second decoding must obtain the second feature maps of 4 channels in order to decode the second group of second feature maps of the input image, and the third decoding must obtain the second feature maps of 3 channels in order to decode the third group of second feature maps of the input image.

[0311] In one possible implementation, the code stream may be decoded to obtain first position information, where the first position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the first feature map of the input image; and the first group of second feature maps may be AI decoded according to the first position information to obtain a first reconstructed image.

[0312] For example, the first feature map is separated into 10 second feature maps through channels. The positions of these 10 second feature maps in the first feature map are recorded as 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 respectively. These 10 second feature maps are grouped and sorted and encoded into the bitstream in the order of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 respectively. Therefore, the first position information of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 can be encoded into the bitstream. After obtaining this information, the decoder can know that the position of the first second feature map obtained by decoding in the first feature map is 2, the position of the second second feature map is 9, and so on. The position of the tenth second feature map is 8 in the first feature map.

[0313] The first second feature map obtained by decoding is placed at position 2 in the first feature map for image restoration, the second second feature map obtained by decoding is placed at position 9 in the first feature map for image restoration, ..., the tenth second feature map obtained by decoding is placed at position 8 in the first feature map for image restoration.

[0314] In one possible implementation, the code stream may be decoded to obtain second position information, where the second position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the code stream; the bit data corresponding to the first group of second feature maps in the code stream may be determined based on the second position information, and the corresponding bit data may be entropy decoded to obtain the first group of second feature maps.

[0315] For example, the corresponding positions of the bit data of the first, second, and third groups of second characteristic maps in the code stream are A1-A2, A3-A4, and A5-A6, respectively. The second position information of A1-A2, A3-A4, and A5-A6 can then be encoded into the code stream. Accordingly, after the decoder obtains the second position information through decoding, it can be determined that the bit data at positions A1-A2 in the decoded code stream can be used to obtain the first group of second characteristic maps, the bit data at positions A3-A4 in the decoded code stream can be used to obtain the second group of second characteristic maps, and the bit data at positions A5-A6 in the decoded code stream can be used to obtain the third group of second characteristic maps.

[0316] In one possible implementation, the above-mentioned code stream can be decoded to obtain code stream length information, where the above-mentioned code stream length information is used to indicate the length of the above-mentioned code stream; and the above-mentioned first group of second feature maps are AI decoded according to the above-mentioned code stream length information to obtain a first reconstructed image.

[0317] For example, the decoding time can be estimated based on the bitstream length. If the estimated decoding time exceeds a preset time, a progressive decoding method is used. Specifically, entropy decoding is performed on the bitstream to obtain a portion of the second feature map of the input image. AI decoding is then performed on this portion of the second feature map to obtain a reconstructed image of the input image.

[0318] The following will introduce an encoding device for executing the above encoding method in conjunction with FIG15 .

[0319] It is understandable that, in order to implement the above functions, the encoding device includes hardware and / or software modules that perform the corresponding functions. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0320] In the embodiment of the present application, the encoding device can be divided into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. In actual implementation, other division methods can be used.

[0321] In the case of dividing each functional module according to each function, Figure 15 shows a possible schematic diagram of the composition of the encoding device involved in the above embodiment. The device can be an electronic device, or a module applied to an electronic device (such as a processor, chip, or chip system, etc.), or a logical node, logical module, or software that can realize all or part of the functions of the electronic device. As shown in Figure 15, the encoding device 1500 may include: an encoding unit 1501 and a determining unit 1502.

[0322] The encoding unit 1501 is configured to perform AI encoding on the input image to obtain a first feature map of the input image.

[0323] The determining unit 1502 is configured to determine N groups of second feature maps of the input image based on the first feature map, where N is a positive integer.

[0324] The encoding unit 1501 is further configured to entropy encode the first group of second feature maps among the N groups of second feature maps into a bitstream.

[0325] In a possible implementation, the encoding unit 1501 is further configured to entropy encode a second group of second feature maps among the N groups of second feature maps into the bitstream.

[0326] In a possible implementation, the encoding unit 1501 is further configured to: entropy encode the third group of second feature maps among the N groups of second feature maps into the bitstream.

[0327] In one possible implementation, the determination unit 1502 is specifically used to: perform channel separation on the first feature map to obtain M second feature maps of the input image, where M is a positive integer; and group the M second feature maps according to the channel information of the second feature map of the input image to obtain the N groups of second feature maps.

[0328] In a possible implementation, the encoding unit 1501 is further configured to entropy encode at least one of the N groups of second feature maps into the bitstream in descending order of channel entropy or channel variance.

[0329] In a possible implementation, the encoding unit 1501 is further configured to encode channel number information into the bitstream, where the channel number information is used to indicate the number of channels of at least one set of second feature maps encoded into the bitstream.

[0330] In a possible implementation, the encoding unit 1501 is further configured to: encode first position information into the bitstream, where the first position information is used to indicate a position of at least one set of second feature maps encoded into the bitstream in the first feature map.

[0331] In a possible implementation, the encoding unit 1501 is further configured to: encode second position information into the code stream, where the second position information is used to indicate a position in the code stream of at least one set of second feature maps encoded into the code stream.

[0332] In a possible implementation, the encoding unit 1501 is further configured to: encode code stream length information into the code stream, where the code stream length information is used to indicate the length of the code stream.

[0333] The decoding device for executing the above decoding method will be introduced below with reference to FIG16 .

[0334] It is understandable that, in order to implement the above functions, the decoding device includes hardware and / or software modules that perform the corresponding functions. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0335] In the embodiments of the present application, the decoding device can be divided into functional modules according to the above-mentioned method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into a processing module. The above-mentioned integrated modules can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical functional division. In actual implementation, other division methods may be used.

[0336] FIG16 shows a possible schematic diagram of the composition of a decoding device involved in the above embodiment, where each functional module is divided according to its function. The device can be an electronic device, a module applied to an electronic device (such as a processor, chip, or chip system), or a logical node, logical module, or software that can implement all or part of the functions of the electronic device. As shown in FIG16 , the decoding device 1600 may include a decoding unit 1601 and a reconstruction unit 1602.

[0337] The decoding unit 1601 is configured to perform entropy decoding on the bitstream to obtain a first set of second feature maps of the input image, where N is a positive integer;

[0338] The reconstruction unit 1602 is configured to perform AI decoding on the first group of second feature maps to obtain a first reconstructed image of the input image.

[0339] In a possible implementation, the decoding unit 1601 is further configured to perform entropy decoding on the code stream to obtain a second group of second feature maps of the input image.

[0340] In a possible implementation, the reconstruction unit 1602 is further configured to perform AI decoding on the first group of second feature maps and the second group of second feature maps to obtain a second reconstructed image of the input image.

[0341] In a possible implementation, the decoding unit 1601 is further configured to perform entropy decoding on the bitstream to obtain a third group of second feature maps of the input image.

[0342] In a possible implementation, the reconstruction unit 1602 is further configured to perform AI decoding on the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps to obtain a third reconstructed image of the input image.

[0343] In one possible implementation, the image quality of the second reconstructed image is better than the image quality of the first reconstructed image, and the image quality of the third reconstructed image is better than the image quality of the second reconstructed image, and the image quality is characterized by any one of the following variables: PSNR, MS-SSIM or LPIPS.

[0344] In a possible implementation, the decoding unit 1601 is further configured to decode the code stream to obtain channel number information, where the channel number information is used to indicate the number of channels of at least one set of second feature maps obtained by decoding.

[0345] In a possible implementation, the reconstruction unit 1602 is further configured to perform AI decoding on the first group of second feature maps according to the channel number information to obtain a first reconstructed image.

[0346] In a possible implementation, the decoding unit 1601 is further configured to decode the code stream to obtain first position information, where the first position information is configured to indicate a position of at least one set of second feature maps obtained by decoding in the first feature map of the input image.

[0347] In a possible implementation, the reconstruction unit 1602 is further configured to perform AI decoding on the first group of second feature maps according to the first position information to obtain a first reconstructed image.

[0348] In a possible implementation, the decoding unit 1601 is further configured to decode the code stream to obtain second position information, where the second position information is used to indicate a position of the at least one set of second feature maps obtained by decoding in the code stream;

[0349] In a possible implementation, the decoding unit 1601 is further configured to determine the bit data corresponding to the first group of second feature maps in the code stream based on the second position information, and perform entropy decoding on the corresponding bit data to obtain the first group of second feature maps.

[0350] In a possible implementation, the decoding unit 1601 is further configured to decode the code stream to obtain code stream length information, where the code stream length information indicates the length of the code stream.

[0351] In a possible implementation, the reconstruction unit 1602 is further configured to perform AI decoding on the first group of second feature maps according to the bitstream length information to obtain a first reconstructed image.

[0352] The present application also provides a chip, which can be a chip for the encoding device or decoding device described above. FIG17 shows a schematic diagram of the structure of a chip 1700. Chip 1700 includes one or more processors 1701 and an interface circuit 1702. Optionally, chip 1700 may also include a bus 1703.

[0353] The processor 1701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above encoding and decoding method can be completed by a hardware integrated logic circuit in the processor 1701 or by software instructions.

[0354] Optionally, the processor 1701 may be a general-purpose processor, a digital signal processing (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods and steps disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor.

[0355] The interface circuit 1702 can be used to send or receive data, instructions or information. The processor 1701 can use the data, instructions or other information received by the interface circuit 1702 to process it, and can send the processing completion information through the interface circuit 1702.

[0356] Optionally, the chip also includes a memory, which may include a read-only memory and a random access memory, and provides operating instructions and data to the processor. Part of the memory may also include a non-volatile random access memory (NVRAM).

[0357] Optionally, the memory stores an executable software module or a data structure, and the processor can perform corresponding operations by calling an operation instruction stored in the memory (the operation instruction may be stored in an operating system).

[0358] Optionally, the chip can be used in the encoding device or decoding device involved in the embodiments of the present application. Optionally, the interface circuit 1702 can be used to output the execution result of the processor 1701. For the encoding and decoding methods provided in one or more embodiments of the present application, reference can be made to the aforementioned embodiments and will not be repeated here.

[0359] It should be noted that the corresponding functions of the processor 1701 and the interface circuit 1702 can be implemented through hardware design, software design, or a combination of hardware and software, and there is no limitation here.

[0360] Figure 18 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device can be an encoding device or a decoding device, a chip or functional module in an encoding device, or a chip or functional module in a decoding device. As shown in Figure 18, the electronic device 1800 includes a processor 1801, a transceiver 1802, and a communication circuit 1803.

[0361] Among them, the processor 1801 is used to execute any step in the encoding and decoding method provided in the embodiment of the present application, and in the process of executing any step in the encoding and decoding method provided in the embodiment of the present application, the transceiver 1802 and the communication line 1803 can be optionally called to complete the corresponding operation.

[0362] Furthermore, the electronic device 1800 may further include a memory 1804 . The processor 1801 , the memory 1804 and the transceiver 1802 may be connected via a communication line 1803 .

[0363] The processor 1801 is a processor, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 1801 may also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.

[0364] Transceiver 1802 is used to communicate with other devices or other communication networks, such as Ethernet, radio access networks (RAN), wireless local area networks (WLAN), etc. Transceiver 1802 can be a module, circuit, transceiver, or any device capable of implementing communication.

[0365] The transceiver 1802 is mainly used for sending and receiving commands and information, etc., and may include a transmitter and a receiver for sending and receiving commands and information, etc. respectively; operations other than sending and receiving commands and information, etc. are implemented by the processor.

[0366] The communication line 1803 is used to transmit information between the various components included in the electronic device 1800.

[0367] In one design, the processor can be considered as the logic circuit and the transceiver as the interface circuit.

[0368] The memory 1804 is used to store instructions, where the instructions may be computer programs.

[0369] Memory 1804 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM may be used, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). Memory 1804 may also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage media or other magnetic storage devices, etc. It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0370] It should be noted that memory 1804 can exist independently of processor 1801 or can be integrated with processor 1801. Memory 1804 can be used to store instructions, program code, or data. Memory 1804 can be located within or outside electronic device 1800, without limitation. Processor 1801 is configured to execute instructions stored in memory 1804 to implement the methods provided in the above embodiments of this application.

[0371] In one example, processor 1801 may include one or more processors, such as processor 0 and processor 1 in FIG. 18 .

[0372] As an optional implementation, the electronic device 1800 includes multiple processors. For example, in addition to the processor 1801 in FIG. 18 , it may also include a processor 1809 .

[0373] As an optional implementation, the electronic device 1800 further includes an output device 1805 and an input device 1806. For example, the input device 1806 is a keyboard, a mouse, a microphone, or a joystick, and the output device 1805 is a display screen, a speaker, or the like.

[0374] It should be pointed out that the electronic device 1800 can be a chip system or a device with a similar structure as shown in Figure 18. Among them, the chip system can be composed of chips, or it can include chips and other discrete devices. The actions, terms, etc. involved in the various embodiments of this application can refer to each other without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are only an example. Other names can also be used in the specific implementation without limitation. In addition, the component structure shown in Figure 18 does not constitute a limitation on the electronic device 1800. In addition to the components shown in Figure 18, the electronic device 1800 may include more or fewer components than those shown in Figure 18, or combine certain components, or arrange the components differently.

[0375] The processor and transceiver described in this application can be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit, a mixed-signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (NMOS), P-type metal oxide semiconductor (positive channel metal oxide semiconductor, PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0376] Figure 19 is a schematic diagram of the structure of another electronic device provided in an embodiment of the present application. The electronic device may be an encoding device or a decoding device, a chip or a functional module in an encoding device, or a chip or a functional module in a decoding device. For ease of explanation, Figure 19 only shows the main components of the electronic device, including a processor 1901, a memory 1902, a control circuit 1903, and an input / output device 1904. The processor 1901 is mainly used to process communication protocols and communication data, execute software programs, and process data of software programs. The memory 1902 is mainly used to store software programs and data. The control circuit 1903 is mainly used for power supply and transmission of various electrical signals. The input / output device 1904 is mainly used to receive data input by the user and output data to the user.

[0377] When the electronic device is a processor 1901, the control circuit 1903 may be a motherboard, the memory 1902 includes a hard disk, RAM, ROM and other media with storage functions, the processor 1901 may include a baseband processor 1901 and a central processing unit, the baseband processor is mainly used to process communication protocols and communication data, the central processing unit is mainly used to control the entire electronic device, execute software programs, and process software program data, the input and output devices 1904 include a display screen, a keyboard, and a mouse, etc.; the control circuit 1903 may further include or be connected to a transceiver circuit or transceiver, such as a network cable interface, etc., for sending or receiving data or signals, such as for data transmission and communication with other devices. Furthermore, it may also include an antenna for sending and receiving wireless signals for data / signal transmission with other devices.

[0378] An embodiment of the present application also provides an encoding device, which includes: at least one processor, when the at least one processor executes program code or instructions, it implements the above-mentioned related method steps to implement the encoding method in the above embodiment.

[0379] Optionally, the apparatus may further include at least one memory configured to store the program code or instruction.

[0380] An embodiment of the present application also provides a decoding device, which includes: at least one processor, when the at least one processor executes program code or instructions, it implements the above-mentioned related method steps to implement the decoding method in the above embodiment.

[0381] Optionally, the apparatus may further include at least one memory configured to store the program code or instruction.

[0382] An embodiment of the present application further provides a computer storage medium, which stores computer instructions. When the computer instructions are executed on a communication device, the communication device executes the above-mentioned related method steps to implement the encoding and decoding method in the above-mentioned embodiment.

[0383] The embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement the encoding and decoding method in the above-mentioned embodiment.

[0384] The present application also provides an encoding device, which may be a chip, integrated circuit, component, or module. Specifically, the device may include a processor and a memory for storing instructions, or the device may include at least one processor configured to retrieve instructions from an external memory. When the device is running, the processor may execute the instructions, causing the chip to perform the encoding method described in each of the above method embodiments.

[0385] The present application also provides a decoding device, which can be a chip, integrated circuit, component, or module. Specifically, the device can include a processor and a memory for storing instructions, or the device can include at least one processor for retrieving instructions from an external memory. When the device is running, the processor can execute the instructions, causing the chip to perform the decoding method described in each of the above method embodiments.

[0386] An embodiment of the present application further provides a code stream storage method, which includes: acquiring and storing the code stream obtained by the above encoding and decoding method.

[0387] The embodiment of the present application further provides a code stream storage device, which is used to obtain and store the code stream obtained by the above encoding and decoding method.

[0388] An embodiment of the present application further provides a code stream transmission method, which includes: acquiring and transmitting the code stream obtained by the above encoding and decoding method.

[0389] The embodiment of the present application further provides a bit stream transmission device, which is used to obtain and transmit the bit stream obtained by the above encoding and decoding method.

[0390] An embodiment of the present application further provides a computer-readable storage medium, on which a code stream obtained by the above encoding and decoding method is stored.

[0391] Please refer to Figure 20, which schematically illustrates the general concept of processing by a neural network such as a convolutional neural network (CNN). A convolutional neural network consists of input and output layers, as well as multiple hidden layers. The input layer is the layer that provides the input (a portion of the input image, as shown in Figure 20) for processing. The hidden layers of a CNN typically consist of a series of convolutional layers, which perform convolutions with multiplications or other dot products. The result of the layers is one or more feature maps (represented by empty solid rectangles), sometimes also called channels. Resampling (such as subsampling) may be involved in some or all layers. As a result, the feature maps may become smaller, as shown in Figure 20. Note that convolution with stride can also reduce the size of the input feature maps (resampling). The activation function in a CNN is typically a ReLU (rectified linear unit) layer, followed by additional convolutions such as pooling, fully connected, and normalization layers. These are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. Although these layers are colloquially referred to as convolutions, this is just convention. Mathematically speaking, it is technically a sliding dot product or cross correlation. The indexing in the matrix has important implications because it affects how the weight is determined at a particular index point.

[0392] When programming a CNN to process images, as shown in Figure 20, the input is a tensor of shape (number of images) x (image width) x (image height) x (image depth). It should be noted that the image depth can be composed of the number of channels in the image. After passing through the convolutional layer, the image is abstracted into a feature map with a shape of (number of images) x (feature map width) x (feature map height) x (feature map channels). The convolutional layer in a neural network should have the following properties: a convolution kernel defined by width and height (hyperparameters). The number of input and output channels (hyperparameters). The depth of the convolution filter (input channels) should be equal to the number of channels in the input feature map (depth).

[0393] Video Coding Machines (VCM) is another popular area of ​​computer science. The main idea behind this approach is to transmit coded representations of image or video information for further processing by computer vision (CV) algorithms, such as object segmentation, detection, and recognition. Compared to traditional image and video coding, which targets human perception, quality is characterized by performance on computer vision tasks, such as object detection accuracy, rather than reconstruction quality. This is illustrated in Figure 21.

[0394] Machine video coding, also known as collaborative intelligence, is a relatively new paradigm for efficiently deploying deep neural networks in mobile cloud infrastructure. By partitioning the network between the mobile side 2110 and the cloud side 2190 (e.g., cloud servers), the computational workload can be distributed, minimizing the system's overall energy and / or latency. Generally, collaborative intelligence is a paradigm in which neural network processing is distributed across two or more distinct computational nodes; for example, devices, but generally, any functionally defined node. The term "node" here does not refer to a neural network node as described above. Instead, a (computing) node here refers to a (physically or at least logically) independent device / module that implements a portion of a neural network. Such devices can be different servers, different end-user devices, a mixture of servers and / or user devices and / or the cloud and / or processors, etc. In other words, the computational nodes can be considered to belong to the same neural network and communicate with each other to transfer encoded data within / for the neural network. For example, to perform complex computations, one or more layers can be executed on a first device (such as a device on the mobile side 2110), while one or more layers can be executed on another device (such as a cloud server on the cloud side 2190). However, the distribution can also be more refined, and a single layer can be executed on multiple devices. In this disclosure, the term "multiple" means two or more. In some existing solutions, part of the neural network functionality is executed in a device (user device or edge device, etc.) or multiple such devices, and the output (feature map) is then passed to the cloud. The cloud is a collection of processing or computing systems located outside the device that is operating part of the neural network. The concept of collaborative intelligence has also been extended to model training. In this case, data flows in both directions: from the cloud to the mobile device during the backpropagation of training, and from the mobile device to the cloud during the forward pass and inference of training (as shown in Figure 21).

[0395] Some work has proposed semantic image compression by encoding deep features and then reconstructing the input image from them. Compression based on uniform quantization was demonstrated, followed by context-based adaptive arithmetic coding (CABAC) from H.264. In some scenarios, it may be more efficient to send the output of the hidden layer (deep feature map) from the mobile part 2110 to the cloud 2190, rather than sending compressed natural image data to the cloud and performing object detection using the reconstructed image. Therefore, it may be advantageous to compress the data (features) generated by the mobile side 2110, which may include a quantization layer 2120 for this purpose. Accordingly, the cloud side 2190 may include an inverse quantization layer 2160. Efficient compression of feature maps facilitates image and video compression and reconstruction for both human perception and machine vision. Entropy coding methods, such as arithmetic coding, are a popular method for compressing deep features (i.e., feature maps).

[0396] Please refer to Figure 22, which shows a code stream structure provided by an embodiment of the present application. As shown in Figure 22, the code stream includes: an image start point (Start of Image), a file header (File Header), entropy encoded data (Entropy Encoded Data) and an image end point (End of Image).

[0397] In a possible implementation, the channel number information, the first position information, the second position information, or the code stream length information may be stored in a file header.

[0398] In a possible implementation, the channel number information, the first position information, the second position information, or the code stream length information may be stored in entropy coded data.

[0399] Among them, the device, computer storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0400] It should be understood that in various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0401] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.

[0402] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0403] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0404] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0405] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0406] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods of each embodiment of the embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0407] The above description is merely a specific implementation of the embodiments of the present application, but the scope of protection of the embodiments of the present application is not limited thereto. Any person skilled in the art can easily conceive of changes or substitutions within the technical scope disclosed in the embodiments of the present application, and such changes or substitutions should be included in the scope of protection of the embodiments of the present application. Therefore, the scope of protection of the embodiments of the present application should be based on the scope of protection of the claims.

Claims

1. A coding method, characterized in that, Comprising: Performing artificial intelligence (AI) encoding on an input image to obtain a first feature map of the input image; Determining N groups of second feature maps of the input image according to the first feature map, where N is a positive integer; Entropy encoding the first group of second feature maps among the N groups of second feature maps into a bitstream.

2. The method according to claim 1, wherein The method further comprises: Entropy encoding the second group of second feature maps among the N groups of second feature maps into the bitstream.

3. The method according to claim 1 or 2, characterized in that, The method further comprises: Entropy encoding the third group of second feature maps among the N groups of second feature maps into the bitstream.

4. The method according to any one of claims 1 to 3, characterized in that, The determining the N groups of second features of the input image according to the first feature map includes: Performing channel separation on the first feature map to obtain M second feature maps of the input image, where M is a positive integer; Grouping the M second feature maps according to the channel information of the second feature maps of the input image to obtain the N groups of second feature maps.

5. The method according to claim 4, wherein The channel information includes channel entropy and / or channel variance, the channel entropy being the entropy of all elements within a channel, and the channel variance being the variance of all elements within a channel.

6. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: Entropy encoding at least one group of second feature maps among the N groups of second feature maps into the bitstream in descending order of channel entropy or channel variance.

7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: Encoding channel number information into the bitstream, the channel number information being used to indicate the number of channels of at least one group of second feature maps encoded into the bitstream.

8. The method according to any one of claims 1 to 7, characterized in that, The method further comprises: Encoding first position information into the bitstream, the first position information being used to indicate the position of at least one group of second feature maps encoded into the bitstream in the first feature map.

9. The method according to any one of claims 1 to 8, characterized in that, The method further comprises: Encoding second position information into the bitstream, the second position information being used to indicate the position of at least one group of second feature maps encoded into the bitstream in the bitstream.

10. A decoding method, characterized in that, Comprising: Performing entropy decoding on the bitstream to obtain the first group of second feature maps of the input image, the first group of second feature maps being a part of the feature maps corresponding to the input image; Performing AI decoding on the first group of second feature maps to obtain a first reconstructed image of the input image.

11. The method according to claim 10, wherein The method further comprises: Performing entropy decoding on the bitstream to obtain the second group of second feature maps, the second group of second feature maps being a part of the feature maps corresponding to the input image; Performing AI decoding on the first group of second feature maps and the second group of second feature maps to obtain a second reconstructed image of the input image.

12. The method according to claim 11, wherein The method further comprises: Performing entropy decoding on the bitstream to obtain the third group of second feature maps, the third group of second feature maps being a part of the feature maps corresponding to the input image; Performing AI decoding on the first group of second feature maps, the second group of second feature maps and the third group of second feature maps to obtain a third reconstructed image of the input image.

13. The method according to claim 12, characterized in that, The image quality of the second reconstructed image is better than the image quality of the first reconstructed image, and the image quality of the third reconstructed image is better than the image quality of the second reconstructed image, the image quality being characterized by any one of the following variables: peak signal-to-noise ratio (PSNR), multi-scale structural similarity index (MS-SSIM) or learned perceptual image patch similarity (LPIPS).

14. The method according to any one of claims 10 to 13, characterized in that, The method further includes: decoding the bitstream to obtain channel number information, where the channel number information is used to indicate the number of channels of at least one group of second feature maps; performing AI decoding on the first group of second feature maps according to the channel number information to obtain a first reconstructed image.

15. The method according to any one of claims 10 to 14, characterized in that, The method further includes: decoding the bitstream to obtain first position information, where the first position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the first feature map of the input image; performing AI decoding on the first group of second feature maps according to the first position information to obtain a first reconstructed image.

16. The method according to any one of claims 11 to 15, characterized in that, The method further includes: decoding the bitstream to obtain second position information, where the second position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the bitstream; determining the bit data corresponding to the first group of second feature maps in the bitstream according to the second position information, and performing entropy decoding on the corresponding bit data to obtain the first group of second feature maps.

17. An encoding device, characterized in that, It includes: an encoding unit and a determining unit; the encoding unit is configured to perform AI encoding on an input image to obtain a first feature map of the input image; the determining unit is configured to determine N groups of second feature maps of the input image according to the first feature map, where N is a positive integer; the encoding unit is further configured to entropy-encode the first group of second feature maps among the N groups of second feature maps into the bitstream.

18. The device according to claim 17, characterized in that, The encoding unit is further configured to: entropy-encode the second group of second feature maps among the N groups of second feature maps into the bitstream.

19. The device according to claim 17 or 18, characterized in that, The encoding unit is further configured to: entropy-encode the third group of second feature maps among the N groups of second feature maps into the bitstream.

20. The device according to any one of claims 17 to 19, characterized in that The determining unit is specifically configured to: perform channel separation on the first feature map to obtain M second feature maps of the input image, where M is a positive integer; group the M second feature maps according to the channel information of the second feature maps of the input image to obtain the N groups of second feature maps.

21. The device according to any one of claims 17 to 20, characterized in that, The encoding unit is further configured to: entropy-encode at least one group of second feature maps among the N groups of second feature maps into the bitstream in descending order of channel entropy or channel variance.

22. A decoding device, characterized in that, It includes: a decoding unit and a reconstruction unit; the decoding unit is configured to perform entropy decoding on the bitstream to obtain the first group of second feature maps of the input image, and the first group of second feature maps is a part of the feature map corresponding to the input image; the reconstruction unit is configured to perform AI decoding on the first group of second feature maps to obtain a first reconstructed image of the input image.

23. The device according to claim 22, characterized in that, the decoding unit is further configured to perform entropy decoding on the bitstream to obtain the second group of second feature maps, and the second group of second feature maps is a part of the feature map corresponding to the input image; the reconstruction unit is further configured to perform AI decoding on the first group of second feature maps and the second group of second feature maps to obtain a second reconstructed image of the input image.

24. The device according to claim 23, characterized in that, the decoding unit is further configured to perform entropy decoding on the bitstream to obtain the third group of second feature maps, and the third group of second feature maps is a part of the feature map corresponding to the input image; The reconstruction unit performs AI decoding on the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps to obtain a third reconstructed image of the input image.

25. The device according to claim 24, characterized in that, The image quality of the second reconstructed image is better than that of the first reconstructed image, and the image quality of the third reconstructed image is better than that of the second reconstructed image. The image quality is characterized by any one of the following variables: PSNR, MS-SSIM, or LPIPS.

26. An encoding device, characterized in that, Comprising at least one processor and a memory, the at least one processor executes a program or instructions stored in the memory to enable the encoding device to implement the method according to any one of claims 1 to 9 above.

27. A decoding device, characterized in that, Comprising at least one processor and a memory, the at least one processor executes a program or instructions stored in the memory to enable the decoding device to implement the method according to any one of claims 10 to 16 above.

28. A method for storing a bitstream, characterized in that, Comprising: Obtaining and storing a bitstream obtained by the method according to any one of claims 1 to 16.

29. A bitstream storage device, characterized in that The device is configured to obtain and store a bitstream obtained by the method according to any one of claims 1 to 16.

30. A bitstream transmission method, characterized in that, Comprising: Obtaining and transmitting a bitstream obtained by the method according to any one of claims 1 to 16.

31. A bitstream transmission device, characterized in that, The device is configured to obtain and transmit a bitstream obtained by the method according to any one of claims 1 to 16.

32. A computer-readable storage medium, characterized in that, Stored on the computer-readable storage medium is a bitstream obtained by the method according to any one of claims 1 to 16.

33. A computer-readable storage medium, characterized in that, For storing a computer program, when the computer program runs on a computer or a processor, enabling the computer or the processor to implement the method according to any one of claims 1 to 16 above.

34. A computer program product, characterized in that, The computer program product contains instructions, when the instructions run on a computer or a processor, enabling the computer or the processor to implement the method according to any one of claims 1 to 16 above.

Citation Information

Patent Citations

  • Coding and decoding method and device

    CN120343252A

  • Image coding and decoding method and device

    CN114554205A

  • Encoding and decoding method and electronic equipment

    CN116170596A

  • Inplausible neural network image compression method and system based on frequency decomposition, and storage medium

    CN117278757A

  • Method, apparatus, and storage medium for encoding / decoding multi-resolution feature map

    US20230342980A1