Coding and decoding method and device

Through the progressive encoding method, some feature maps are selected for encoding, which solves the problem of high encoding complexity of AI images and realizes efficient encoding and decoding under different conditions.

CN120343252APending Publication Date: 2025-07-18HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410176941.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2024-02-08
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing AI image encoding is complex, especially the entropy encoding and decoding process is time-consuming, which affects the implementation of the product.

Method used

The progressive encoding method is used to encode part of the feature map of the input image, select the feature map with maximum channel entropy or variance for preliminary encoding, and flexibly adjust the encoding process according to the network environment and hardware resources.

Benefits of technology

It reduces image encoding delay, adapts to different network environments and hardware resource conditions, and improves coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343252A_ABST
    Figure CN120343252A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a coding method, relates to the technical field of media, and is used for reducing image decoding time delay. The method comprises the following steps: performing feature extraction on an input image to obtain a first feature map of the input image; and determining N groups of second feature maps of the input image according to the first feature map. And carrying out entropy coding on the at least one group of second feature maps of the input image to obtain a code stream. Wherein N is a positive integer. It can be seen that in the method provided by the embodiment of the invention, the partial feature map of the input image can be coded by adopting the progressive coding method. Compared with encoding of all the feature maps of the input image, encoding of part of the feature maps of the input image can reduce image encoding time delay. In addition, the first group of second feature maps of the input image are feature maps with the maximum channel entropy (variance) in the N groups of second feature maps of the input image, and a decoding end can obtain a complete reconstructed image based on the feature maps.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application titled "Coding Method" with the application number 202410080945.2, filed with the Chinese Patent Office on January 18, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to the field of media technology, and in particular, to encoding and decoding methods and apparatuses. Background Art

[0003] Currently, significant breakthroughs have been made in artificial intelligence (AI) image coding. AI image coding significantly improves the compression efficiency compared to common image coding standards under the same subjective quality. AI image coding is widely used in various fields. For example, it can be applied to cloud storage, visual surveillance, autonomous driving vehicles and devices, image acquisition, storage and management, real-time monitoring of visual data, and media distribution.

[0004] The existing AI image coding has a high complexity, especially the entropy encoding and decoding process takes a long time, which has a certain impact on the implementation of products. Summary of the Invention

[0005] Embodiments of this application provide encoding and decoding methods and apparatuses for reducing the image encoding and decoding latency. To achieve the above objective, the embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, embodiments of this application provide an encoding method, which includes: performing AI encoding on an input image to obtain a first feature map of the input image; determining N groups of second feature maps of the input image according to the first feature map; and entropy encoding the first group of second feature maps among the N groups of second feature maps into a bitstream. Here, N is a positive integer.

[0007] It can be seen that in the method provided by the embodiments of this application, a progressive encoding method is adopted, and partial feature maps of the input image can be encoded. Compared with encoding all the feature maps of the input image, encoding partial feature maps of the input image can reduce the image encoding latency. In addition, the first group of second feature maps of the input image is the feature map with the largest channel entropy (variance) among the N groups of second feature maps of the input image, and the decoding end can obtain a complete reconstructed image based on this feature map.

[0008] For example, in a scenario where the network environment is poor or the decoding hardware resources are insufficient, the progressive encoding method adopted by the embodiments of this application does not require all the feature maps, and a complete reconstructed image can be obtained only by decoding based on the first group of second feature maps, thereby reducing the image encoding and decoding latency.

[0009] In a possible implementation, any one of the above N groups of second feature maps can be decoded to obtain a reconstructed image.

[0010] In a possible implementation, the second feature map of the 2nd group in the above N groups of second feature maps can be entropy-encoded into the above bitstream.

[0011] It can be seen that in the method provided by the embodiments of the present application, a progressive encoding method is adopted, and 2 groups of feature maps of the input image can be encoded. Compared with encoding all the feature maps of the input image, encoding some of the feature maps of the input image can reduce the image encoding delay.

[0012] In a possible implementation, the second feature map of the 3rd group in the above N groups of second feature maps can be entropy-encoded into the above bitstream.

[0013] It can be seen that in the method provided by the embodiments of the present application, a progressive encoding method is adopted, and 3 groups of feature maps of the input image can be encoded. Compared with encoding all the feature maps of the input image, encoding some of the feature maps of the input image can reduce the image encoding delay.

[0014] For example, in a scenario where the network environment is good or the decoding hardware resources are sufficient, the embodiments of the present application adopt a progressive encoding method. On the one hand, a complete reconstructed image can be obtained by decoding only based on the first group of second feature maps. On the other hand, a reconstructed image with better image quality can be obtained by decoding based on multiple groups of second feature maps. Therefore, it can be seen that the progressive encoding method provided by the embodiments of the present application can flexibly adapt to various network environments and hardware.

[0015] In a possible implementation, the above first feature map can be channel-separated to obtain M second feature maps of the above input image. The M second feature maps are grouped into the above N groups of second feature maps according to the channel information of the second feature maps of the above input image. Wherein, M is a positive integer.

[0016] It can be seen that in the method provided by the embodiments of the present application, after obtaining the feature maps of the input image, the feature maps of the input image can be channel-separated to divide the feature maps of the input image into multiple feature maps in the channel direction dimension, and then some of the feature maps in the channel dimension of the input image are encoded. Compared with encoding all the feature maps of the input image, encoding some of the feature maps of the input image by adopting a progressive encoding method can reduce the image encoding delay.

[0017] In a possible implementation, the above channel information includes channel entropy and / or channel variance. The channel entropy is the entropy of all elements in the channel, and the channel variance is the variance of all elements in the channel.

[0018] In a possible implementation, the above channel information may further include at least one of the channel number or the group number of channels. Among them, the channel number may be determined according to the channel entropy and / or the channel variance, and the group number of channels may be determined according to the channel entropy and / or the channel variance.

[0019] It can be seen that in the method provided by the embodiments of the present application, after obtaining the feature map of the input image, the feature map of the input image can be separated in channels by information such as channel entropy or channel variance to divide the feature map of the input image into multiple feature maps in the channel direction dimension, and then some feature maps in the channel dimension of the input image are encoded. Compared with encoding all the feature maps of the input image, by using the progressive encoding method, encoding some feature maps of the input image can reduce the image encoding delay.

[0020] In a possible implementation, at least one group of the second feature maps among the above N groups of second feature maps may be entropy encoded into the above bitstream in the order of decreasing channel entropy or channel variance.

[0021] It can be understood that the larger the channel entropy or channel variance of the feature map, the easier it is to recover and reconstruct the image. Therefore, at least one group of the second feature maps among the above N groups of second feature maps may be entropy encoded into the above bitstream in the order of decreasing channel entropy or channel variance, so that by using the progressive encoding method, encoding some feature maps of the input image can reduce the image encoding delay.

[0022] In a possible implementation, the channel number information may be encoded into the above bitstream, and the above channel number information is used to indicate the number of channels of at least one group of second feature maps encoded into the above bitstream. This can enable the decoding end to determine the number of channels of each group of second feature maps in the bitstream by decoding the channel number information, which helps the decoding end to recover the feature map from the bitstream.

[0023] For example, the number of channels of the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps are 4, 3, and 3 respectively. Then the channel number information of 4, 3, and 3 may be encoded into the above bitstream.

[0024] In a possible implementation, the first position information may be encoded into the above bitstream, and the above first position information is used to indicate the position of at least one group of second feature maps encoded into the above bitstream in the above first feature map. This can enable the decoding end to determine the position of each group of second feature maps in the first feature map by decoding the first position information, which helps the decoding end to recover the feature map before grouping from the bitstream.

[0025] It should be noted that the decoding end needs to use the position of the second feature map in the above-mentioned first feature map to decode the second feature map, and the position of the second feature map in the bitstream may be different from the position of the second feature map in the above-mentioned first feature map. Therefore, it is necessary to encode the first position information recording the positions of at least one group of second feature maps in the above-mentioned first feature map into the bitstream.

[0026] For example, the first feature map is separated into 10 second feature maps through channel separation. The positions of these 10 second feature maps in the first feature map are respectively recorded as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10. The grouping and sorting order of these 10 second feature maps into the bitstream is respectively 2, 9, 4, 7, 1, 5, 3, 6, 10, 8. Then, the first position information of 2, 9, 4, 7, 1, 5, 3, 6, 10, 8 can be encoded into the bitstream. After the decoding end obtains this information, it can know that the position of the first decoded second feature map in the first feature map is 2, the position of the second decoded second feature map in the first feature map is 9,..., and the position of the tenth decoded second feature map in the first feature map is 8.

[0027] In a possible implementation manner, the second position information can be encoded into the above-mentioned bitstream, and the second position information is used to indicate the positions of at least one group of second feature maps encoded into the above-mentioned bitstream in the bitstream. In this way, the decoding end can determine the positions of each group of second feature maps in the bitstream through the second position information, which helps the decoding end to perform parallel decoding.

[0028] In a possible implementation manner, the bitstream length information can be encoded into the above-mentioned bitstream, and the bitstream length information is used to indicate the length of the above-mentioned bitstream. In this way, the decoding end can determine the length of the bitstream through the bitstream length information, which helps the decoding end to perform parallel decoding.

[0029] In a second aspect, an embodiment of the present application provides a decoding method, and the method includes: performing entropy decoding on the bitstream to obtain the first group of second feature maps of the input image. Performing AI decoding on the above-mentioned first group of second feature maps to obtain the first reconstructed image of the above-mentioned input image. Wherein, the above-mentioned first group of second feature maps is a part of the feature maps corresponding to the above-mentioned input image.

[0030] It can be seen that in the method provided by the embodiment of the present application, the reconstructed image can be determined through partial feature maps of the input image. Compared with determining the reconstructed image through all the feature maps of the input image, using the progressive decoding method to determine the reconstructed image through partial feature maps of the input image can reduce the image decoding delay. In addition, the first group of second feature maps of the input image is the feature map with the largest channel entropy (variance) among the second feature maps of the input image, and the decoding end can obtain a complete reconstructed image based on this feature map.

[0031] For example, in a scenario where the network environment is poor or the decoding hardware resources are insufficient, the progressive decoding method adopted in the embodiments of the present application can obtain a complete reconstructed image only by decoding based on the first set of second feature maps.

[0032] It should be noted that the first reconstructed image is not a partial image of the input image, but a reconstructed image with the same size as the input image, but different image quality.

[0033] In a possible implementation, the bitstream can be entropy decoded to obtain the second set of second feature maps of the input image; the first set of second feature maps and the second set of second feature maps are AI decoded to obtain the second reconstructed image of the input image.

[0034] In a possible implementation, the bitstream can be entropy decoded to obtain the third set of second feature maps of the input image; the first set of second feature maps, the second set of second feature maps and the third set of second feature maps are AI decoded to obtain the third reconstructed image of the input image.

[0035] For example, in a scenario where the network environment is good or the decoding hardware resources are sufficient, the progressive decoding method adopted in the embodiments of the present application can, on the one hand, obtain a complete reconstructed image only by decoding based on the first set of second feature maps, and on the other hand, obtain a reconstructed image with better image quality by decoding based on multiple sets of second feature maps. Therefore, it can be seen that the progressive decoding method provided by the embodiments of the present application can flexibly adapt to various network environments and hardware.

[0036] In a possible implementation, the image quality of the second reconstructed image is better than the image quality of the first reconstructed image, and the image quality of the third reconstructed image is better than the image quality of the second reconstructed image.

[0037] Wherein, the image quality is characterized by any one of the following variables: peak signal to noise ratio (PSNR), multi-scale structural similarity (MS-SSIM) or learned perceptual image patch similarity (LPIPS).

[0038] In a possible implementation, the bitstream can be decoded to obtain channel number information, and the channel number information is used to indicate the channel number of at least one set of second feature maps obtained by decoding; the first set of second feature maps is AI decoded according to the channel number information to obtain the first reconstructed image.

[0039] For example, the number of channels of the second feature map of the first group, the second feature map of the second group, and the second feature map of the third group are 4, 3, and 3 respectively. Then, the channel number information of 4, 3, and 3 can be encoded into the above-mentioned bitstream. Correspondingly, after decoding to obtain the channel number information at the decoding end, it can be known that for the first decoding, 4 channels of the second feature map need to be decoded to decode the second feature map of the first group of the input image, for the second decoding, 4 channels of the second feature map need to be decoded to decode the second feature map of the second group of the input image, and for the third decoding, 3 channels of the second feature map need to be decoded to decode the second feature map of the third group of the input image.

[0040] In a possible implementation manner, the above-mentioned bitstream can be decoded to obtain first position information, and the first position information is used to indicate the position of at least one group of decoded second feature maps in the first feature map of the input image; the first group of second feature maps is subjected to AI decoding according to the first position information to obtain a first reconstructed image.

[0041] For example, 10 second feature maps are obtained by separating the channels of the first feature map, and the positions of these 10 second feature maps in the first feature map are respectively denoted as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10. The order of grouping and sorting these 10 second feature maps and encoding them into the bitstream is 2, 9, 4, 7, 1, 5, 3, 6, 10, 8. Then, the first position information of 2, 9, 4, 7, 1, 5, 3, 6, 10, 8 can be encoded into the bitstream. After the decoding end obtains this information, it can be known that the position of the first decoded second feature map in the first feature map is 2, the position of the second decoded second feature map in the first feature map is 9,..., and the position of the tenth decoded second feature map in the first feature map is 8.

[0042] The first decoded second feature map is placed at position 2 in the first feature map for image restoration, the second decoded second feature map is placed at position 9 in the first feature map for image restoration,..., and the tenth decoded second feature map is placed at position 8 in the first feature map for image restoration.

[0043] In a possible implementation manner, the above-mentioned bitstream can be decoded to obtain second position information, and the second position information is used to indicate the position of at least one group of decoded second feature maps in the above-mentioned bitstream; the bit data corresponding to the first group of second feature maps in the above-mentioned bitstream is determined according to the second position information, and the corresponding bit data is subjected to entropy decoding to obtain the first group of second feature maps.

[0044] For example, the corresponding positions of the bit data of the second feature maps of the first group, the second group, and the third group in the bitstream are A1 to A2, A3 to A4, and A5 to A6, respectively. Then, the second position information of A1 to A2, A3 to A4, and A5 to A6 can be encoded into the above bitstream. Correspondingly, after the decoding end decodes to obtain the second position information, it can be known that the bit data at the positions of A1 to A2 in the decoded bitstream can be used to obtain the second feature map of the first group, the bit data at the positions of A3 to A4 in the decoded bitstream can be used to obtain the second feature map of the second group, and the bit data at the positions of A5 to A6 in the decoded bitstream can be used to obtain the second feature map of the third group.

[0045] In a possible implementation manner, the above bitstream can be decoded to obtain bitstream length information of the bitstream. The bitstream length information is used to indicate the length of the bitstream. According to the bitstream length information, the second feature map of the first group is decoded by AI to obtain a first reconstructed image.

[0046] For example, the decoding duration can be estimated based on the bitstream length. In the case where the estimated decoding duration is greater than the preset duration, a progressive decoding method is adopted, that is, the bitstream is entropy decoded to obtain partial second feature maps of the input image. The partial second feature maps are decoded by AI to obtain a reconstructed image of the input image.

[0047] In a third aspect, an embodiment of the present application provides an encoding device, which includes: an encoding unit and a determining unit. The encoding unit is configured to perform AI encoding on an input image to obtain a first feature map of the input image. The determining unit is configured to determine N groups of second feature maps of the input image according to the first feature map, where N is a positive integer. The encoding unit is further configured to entropy encode the first group of second feature maps in the N groups of second feature maps into a bitstream.

[0048] In a possible implementation manner, the encoding unit is further configured to: entropy encode the second group of second feature maps in the N groups of second feature maps into the bitstream.

[0049] In a possible implementation manner, the encoding unit is further configured to: entropy encode the third group of second feature maps in the N groups of second feature maps into the bitstream.

[0050] In a possible implementation manner, the determining unit is specifically configured to: perform channel separation on the first feature map to obtain M second feature maps of the input image, where M is a positive integer; and group the M second feature maps according to the channel information of the second feature maps of the input image to obtain the N groups of second feature maps.

[0051] In a possible implementation, the above encoding unit is further configured to: entropy encode at least one group of the second feature maps among the above N groups of second feature maps into the above bitstream in descending order of channel entropy or channel variance.

[0052] In a possible implementation, the above encoding unit is further configured to: encode the channel number information into the above bitstream, where the channel number information is used to indicate the number of channels of at least one group of second feature maps encoded into the above bitstream.

[0053] In a possible implementation, the above encoding unit is further configured to: encode the first position information into the above bitstream, where the first position information is used to indicate the position of at least one group of second feature maps encoded into the above bitstream in the above first feature map.

[0054] In a possible implementation, the above encoding unit is further configured to: encode the second position information into the above bitstream, where the second position information is used to indicate the position of at least one group of second feature maps encoded into the above bitstream in the above bitstream.

[0055] In a possible implementation, the above encoding unit is further configured to: encode the bitstream length information into the above bitstream, where the bitstream length information is used to indicate the length of the above bitstream.

[0056] Fourthly, an embodiment of the present application provides a decoding device, which includes: a decoding unit and a reconstruction unit. The above decoding unit is configured to perform entropy decoding on the bitstream to obtain the first group of second feature maps of the input image, and the first group of second feature maps is a part of the feature maps corresponding to the input image. The above reconstruction unit is configured to perform AI decoding on the first group of second feature maps to obtain the first reconstructed image of the input image.

[0057] In a possible implementation, the above decoding unit is further configured to perform entropy decoding on the above bitstream to obtain the second group of second feature maps, and the second group of second feature maps is a part of the feature maps corresponding to the input image.

[0058] In a possible implementation, the above reconstruction unit is further configured to perform AI decoding on the first group of second feature maps and the second group of second feature maps to obtain the second reconstructed image of the input image.

[0059] In a possible implementation, the above decoding unit is further configured to perform entropy decoding on the above bitstream to obtain the third group of second feature maps, and the third group of second feature maps is a part of the feature maps corresponding to the input image.

[0060] In a possible implementation, the above reconstruction unit is further configured to perform AI decoding on the above first group of second feature maps, the above second group of second feature maps, and the above third group of second feature maps to obtain a third reconstructed image of the above input image.

[0061] In a possible implementation, the image quality of the above second reconstructed image is better than that of the above first reconstructed image, and the image quality of the above third reconstructed image is better than that of the above second reconstructed image. The above image quality is characterized by any one of the following variables: PSNR, MS-SSIM, or LPIPS.

[0062] In a possible implementation, the above decoding unit is further configured to decode the above bitstream to obtain channel number information, and the above channel number information is used to indicate the number of channels of at least one group of second feature maps obtained by decoding.

[0063] In a possible implementation, the above reconstruction unit is further configured to perform AI decoding on the above first group of second feature maps according to the above channel number information to obtain a first reconstructed image.

[0064] In a possible implementation, the above decoding unit is further configured to decode the above bitstream to obtain first position information, and the above first position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the first feature map of the above input image.

[0065] In a possible implementation, the above reconstruction unit is further configured to perform AI decoding on the above first group of second feature maps according to the above first position information to obtain a first reconstructed image.

[0066] In a possible implementation, the above decoding unit is further configured to decode the above bitstream to obtain second position information, and the above second position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the above bitstream.

[0067] In a possible implementation, the above decoding unit is further configured to determine the bit data corresponding to the above first group of second feature maps in the above bitstream according to the above second position information, and perform entropy decoding on the above corresponding bit data to obtain the above first group of second feature maps.

[0068] In a possible implementation, the above decoding unit is further configured to decode the above bitstream to obtain bitstream length information of the above bitstream, and the above bitstream length information is used to indicate the length of the above bitstream.

[0069] In a possible implementation, the above reconstruction unit is further configured to perform AI decoding on the above first group of second feature maps according to the above bitstream length information to obtain a first reconstructed image.

[0070] Fifth aspect, an embodiment of the present application further provides an encoding device, which includes: at least one processor, and when the at least one processor executes program code or instructions, the method described in the above first aspect or any possible implementation manner thereof is implemented.

[0071] Optionally, the device may further include at least one memory, and the at least one memory is used to store the program code or instructions.

[0072] Sixth aspect, an embodiment of the present application further provides a decoding device, which includes: at least one processor, and when the at least one processor executes program code or instructions, the method described in the above second aspect or any possible implementation manner thereof is implemented.

[0073] Seventh aspect, an embodiment of the present application further provides a method for storing a bitstream, which includes: acquiring and storing the bitstream obtained by the method described in the above first aspect or any possible implementation manner thereof.

[0074] Eighth aspect, an embodiment of the present application further provides a bitstream storage device, which is used to acquire and store the bitstream obtained by the method described in the above first aspect or any possible implementation manner thereof.

[0075] Ninth aspect, an embodiment of the present application further provides a method for transmitting a bitstream, which includes: acquiring and transmitting the bitstream obtained by the method described in the above first aspect or any possible implementation manner thereof.

[0076] Tenth aspect, an embodiment of the present application further provides a bitstream transmission device, which is used to acquire and transmit the bitstream obtained by the method described in the above first aspect or any possible implementation manner thereof.

[0077] Eleventh aspect, an embodiment of the present application further provides a computer-readable storage medium, on which the bitstream obtained by the method described in the above first aspect or any possible implementation manner thereof is stored

[0078] Twelfth aspect, an embodiment of the present application further provides a chip, which includes: an input interface, an output interface, and at least one processor. Optionally, the chip may further include a memory. The at least one processor is used to execute the code in the memory, and when the at least one processor executes the code, the chip implements the method described in the above first aspect or any possible implementation manner thereof.

[0079] Optionally, the above chip may also be an integrated circuit.

[0080] In a thirteenth aspect, an embodiment of the present application further provides a computer-readable storage medium for storing a computer program, where the computer program includes a method for implementing the method described in the first aspect or any possible implementation manner thereof above.

[0081] In a fourteenth aspect, an embodiment of the present application further provides a computer program product including instructions, which, when running on a computer, cause the computer to implement the method described in the first aspect or any possible implementation manner thereof above.

[0082] The encoding and decoding device, computer storage medium, computer program product, and chip provided in this embodiment are all used to execute the encoding and decoding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the encoding and decoding method provided above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0084] Figure 1a It is an exemplary block diagram of a decoding system provided in an embodiment of the present application;

[0085] Figure 1b It is an exemplary block diagram of a video decoding system provided in an embodiment of the present application;

[0086] Figure 2 It is an exemplary block diagram of a video encoder provided in an embodiment of the present application;

[0087] Figure 3 It is an exemplary block diagram of a video decoder provided in an embodiment of the present application;

[0088] Figure 4 It is an exemplary block diagram of a video decoding device provided in an embodiment of the present application;

[0089] Figure 5 It is an exemplary block diagram of a device provided in an embodiment of the present application;

[0090] Figure 6 It is a schematic diagram of a neural network-based image compression method provided in an embodiment of the present application;

[0091] Figure 7 It is a schematic diagram of an end-to-end image coding framework provided in an embodiment of the present application;

[0092] Figure 8A schematic structural diagram of a neural network provided by an embodiment of the present application;

[0093] Figure 9 A schematic structural diagram of a video communication system provided by an embodiment of the present application;

[0094] Figure 10 A schematic flow diagram of an encoding method provided by an embodiment of the present application;

[0095] Figure 11 A schematic diagram of an encoding and decoding process provided by an embodiment of the present application;

[0096] Figure 12 Another schematic diagram of an encoding and decoding process provided by an embodiment of the present application;

[0097] Figure 13 Another schematic diagram of an encoding and decoding process provided by an embodiment of the present application;

[0098] Figure 14 A schematic flow diagram of a decoding method provided by an embodiment of the present application;

[0099] Figure 15 A schematic structural diagram of an encoding device provided by an embodiment of the present application;

[0100] Figure 16 A schematic structural diagram of a decoding device provided by an embodiment of the present application;

[0101] Figure 17 A schematic structural diagram of a chip provided by an embodiment of the present application;

[0102] Figure 18 A schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0103] Figure 19 Another schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0104] Figure 20 A schematic structural diagram of a neural network provided by an embodiment of the present application;

[0105] Figure 21 A schematic structural diagram of a machine video encoding system provided by an embodiment of the present application;

[0106] Figure 22 A schematic structural diagram of a bitstream provided by an embodiment of the present application. Detailed implementation manners

[0107] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope protected by the embodiments of the present application.

[0108] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0109] The terms "first" and "second" in the description of the embodiments of the present application are used to distinguish different objects or different processes for the same object, rather than to describe the specific order of the objects.

[0110] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the embodiments of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0111] It should be noted that in the description of the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.

[0112] First, the terms related to the embodiments of the present application will be explained.

[0113] Joint Photographic Experts Group (JPEG) Artificial Intelligence (AI) is a learning-based image coding standard that provides a single-stream, compact compressed-domain representation, significantly improving compression efficiency compared to common image coding standards at the same subjective quality. JPEG AI is widely used in various fields. For example, JPEG AI can be applied to cloud storage, visual surveillance, autonomous vehicles and devices, image acquisition, storage and management, real-time monitoring of visual data, and media distribution.

[0114] Information: The data to be encoded / decoded. Such as an image, a video, an audio, a plain text file, etc.

[0115] Feature map: The features extracted during the neural network inference process.

[0116] Symbol: The basic unit composed of information. Such as a pixel point in an image, each character in a plain text file, etc.

[0117] Encoding: The process of converting information into a 0 / 1 string.

[0118] Bitstream: The 0 / 1 string obtained by encoding the information.

[0119] Decoding: The process of restoring the bitstream to information, which is the inverse process of encoding.

[0120] Premise: The data that is known during both encoding and decoding. Utilizing the premise during encoding and decoding can reduce the length of the bitstream.

[0121] Entropy encoding: A coding method using probability modeling, making the length of the bitstream approach the theoretical shortest (Shannon entropy). The premise can be utilized during entropy encoding to reduce the length of the bitstream.

[0122] Entropy decoding: A decoding method using probability modeling, which is the inverse process of entropy encoding. If the premise is utilized during entropy encoding, the exactly same premise needs to be used during entropy decoding.

[0123] Throughput: The number of symbols that entropy encoding / decoding can encode / decode per second.

[0124] Register: The memory used by the processor to temporarily store instructions, data, etc. It has a very small capacity and extremely fast read / write speed. There are usually 8 - 32 16-bit, 32-bit or 64-bit registers in current computing devices.

[0125] Memory: The main storage unit in a computing device, with a much larger capacity than registers, but the read / write speed is usually more than 10 times that of registers.

[0126] Feature map: The three-dimensional data output by convolutional layers, activation layers, pooling layers, batch normalization layers, etc. in a convolutional neural network. The three dimensions are respectively called width (Width), height (Height), and channel (Channel).

[0127] Data encoding and decoding include two parts: data encoding and data decoding. Data encoding is performed on the source side (or commonly referred to as the encoder side), and typically includes processing (e.g., compressing) the original data to reduce the amount of data required to represent the original data (thus enabling more efficient storage and / or transmission). Data decoding is performed on the destination side (or commonly referred to as the decoder side), and typically includes performing inverse processing relative to the encoder side to reconstruct the original data. The "encoding and decoding" of data involved in the embodiments of this application should be understood as "encoding" or "decoding" of data. The encoding part and the decoding part are also collectively referred to as encoding and decoding (encoding and decoding, CODEC).

[0128] In the case of lossless data encoding, the original data can be reconstructed, that is, the reconstructed original data has the same quality as the original data (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy data encoding, further compression is performed through quantization, etc., to reduce the amount of data required to represent the original data, and the decoder side cannot fully reconstruct the original data, that is, the quality of the reconstructed original data is lower or worse than the quality of the original data.

[0129] The embodiments of this application can be applied to video data and other data with compression / decompression requirements, etc. Hereinafter, the embodiments of this application will be described by taking the encoding of video data (abbreviated as video encoding) as an example. Other types of data (such as image data, audio data, integer data, and other data with compression / decompression requirements) can refer to the following description, and the embodiments of this application will not be elaborated herein. It should be noted that relative to video encoding, during the encoding process of data such as audio data and integer data, there is no need to divide the data into blocks, but the data can be directly encoded.

[0130] Video encoding generally refers to processing an image sequence that forms a video or a video sequence. In the field of video encoding, the terms "picture", "frame", or "image" can be used as synonyms.

[0131] Several video coding standards belong to "lossy hybrid video coding and decoding" (i.e., combining spatial and temporal prediction in the pixel domain with 2D transform coding for applying quantization in the transform domain). Each image in a video sequence is typically segmented into a set of non-overlapping blocks, and encoding is usually performed at the block level. In other words, an encoder typically processes, i.e., encodes video, at the block (video block) level. For example, by spatial (intra-frame) prediction and temporal (inter-frame) prediction to generate a predicted block; subtracting the predicted block from the current block (the currently processed / to-be-processed block) to obtain a residual block; transforming and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compressed), while the decoder side applies the inverse processing part relative to the encoder to the encoded or compressed block to reconstruct the current block for representation. Additionally, the encoder needs to repeat the processing steps of the decoder so that the encoder and the decoder generate the same predictions (e.g., intra-frame prediction and inter-frame prediction) and / or reconstructed pixels for processing, i.e., encoding subsequent blocks.

[0132] In the following embodiments of the decoding system 10, the encoder 20 and the decoder 30 are described according to Figures 1a to 3 this.

[0133] Figure 1a FIG. is an exemplary block diagram of a decoding system 10 provided by an embodiment of the present application. For example, a video decoding system 10 (or simply referred to as the decoding system 10) that can utilize the technology of the embodiment of the present application. The video encoder 20 (or simply referred to as the encoder 20) and the video decoder 30 (or simply referred to as the decoder 30) in the video decoding system 10 represent devices and the like that can be used to execute various techniques according to the various examples described in the embodiment of the present application.

[0134] As Figure 1a shown, the decoding system 10 includes a source device 12, and the source device 12 is configured to provide encoded image data 21 such as encoded images to a destination device 14 for decoding the encoded image data 21.

[0135] The source device 12 includes an encoder 20, and additionally, optionally, may include an image source 16, a pre-processor (or pre-processing unit) 18 such as an image pre-processor, and a communication interface (or communication unit) 22.

[0136] The image source 16 may include or may be any type of image capture device for capturing real-world images and the like, and / or any type of image generation device, such as a computer graphics processor for generating computer animated images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage device that stores any of the above images.

[0137] To distinguish the processing performed by the pre-processor (or pre-processing unit) 18, the image (or image data) 17 may also be referred to as the raw image (or raw image data) 17.

[0138] The pre-processor 18 is configured to receive the raw image data 17 and pre-process the raw image data 17 to obtain pre-processed image (or pre-processed image data) 19. For example, the pre-processing performed by the pre-processor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising. It should be understood that the pre-processing unit 18 may be an optional component.

[0139] The video encoder (or encoder) 20 is configured to receive the pre-processed image data 19 and provide encoded image data 21 (which will be further described below according to Figure 2 etc.).

[0140] The communication interface 22 in the source device 12 may be used to: receive the encoded image data 21 and transmit the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device via the communication channel 13 for storage or direct reconstruction.

[0141] The destination device 14 includes a decoder 30, and additionally, optionally, may include a communication interface (or communication unit) 28, a post-processor (or post-processing unit) 32, and a display device 34.

[0142] The communication interface 28 in the destination device 14 is configured to directly receive the encoded image data 21 (or any other processed version) from the source device 12 or from any other source device such as a storage device. For example, the storage device is an encoded image data storage device, and provide the encoded image data 21 to the decoder 30.

[0143] The communication interfaces 22 and 28 can be used to send or receive encoded image data (or encoded data) 21 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, etc., or via any type of network, such as a wired network, a wireless network, or any combination thereof, any type of private network and public network, or any combination of any type thereof.

[0144] For example, the communication interface 22 can be used to encapsulate the encoded image data 21 into a suitable format such as a packet, and / or use any type of transport encoding or processing to process the encoded image data for transmission over the communication link or communication network.

[0145] The communication interface 28 corresponds to the communication interface 22. For example, it can be used to receive the transmitted data and process the transmitted data using any type of corresponding transport decoding or processing and / or de-encapsulation to obtain the encoded image data 21.

[0146] Both the communication interface 22 and the communication interface 28 can be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by the arrow of the corresponding communication channel 13 pointing from the source device 12 to the destination device 14 in Figure 1a and can be used to send and receive messages, etc., to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as the transmission of encoded image data, etc.

[0147] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31 (which will be further described below according to Figure 3 etc.).

[0148] The post-processor 32 is used to post-process the decoded image, etc., the decoded image data 31 (also referred to as the reconstructed image data) to obtain post-processed image, etc., the post-processed image data 33. The post-processing performed by the post-processing unit 32 can include, for example, color format conversion (e.g., from YCbCr to RGB), color grading, cropping or resampling, or any other processing for generating the decoded image data 31 for display on a display device 34, etc.

[0149] The display device 34 is configured to receive the post - processed image data 33 to display an image to a user, viewer, etc. The display device 34 may be or include any type of display for presenting the reconstructed image. For example, an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro - LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.

[0150] The decoding system 10 further includes a training engine 25. The training engine 25 is configured to train the encoder 20 (especially the entropy coding unit 270 in the encoder 20) or the decoder 30 (especially the entropy decoding unit 304 in the decoder 30) to perform entropy coding on the image blocks to be encoded according to the estimated probability distribution. For a detailed description of the training engine 25, please refer to the following method test examples.

[0151] Although Figure 1a The source device 12 and the destination device 14 are shown as separate devices, but device embodiments may also include both the source device 12 and the destination device 14 or the functions of both the source device 12 and the destination device 14 at the same time, that is, include both the source device 12 or its corresponding function and the destination device 14 or its corresponding function at the same time. In these embodiments, the source device 12 or its corresponding function and the destination device 14 or its corresponding function may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0152] According to the description, Figure 1a The presence and (exact) division of different units or functions in the illustrated source device 12 and / or destination device 14 may vary according to the actual device and application, which is obvious to those skilled in the art.

[0153] Please refer to Figure 1b , Figure 1b FIG. is an exemplary block diagram of the video decoding system 40 provided by an embodiment of the present application. The encoder 20 (such as the video encoder 20) or the decoder 30 (such as the video decoder 30) or both can be implemented by, for example, Figure 1bThe processing circuitry in the illustrated video decoding system 40 is implemented by, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding-specific processors, or any combination thereof. Please refer to Figure 2 and Figure 3 , Figure 2 which is an exemplary block diagram of a video encoder provided by an embodiment of the present application, Figure 3 and Figure 3 is an exemplary block diagram of a video decoder provided by an embodiment of the present application. The encoder 20 may be implemented by the processing circuitry 46 to include various modules discussed with reference to Figure 2 the encoder 20 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented by the processing circuitry 46 to include various modules discussed with reference to Figure 3 the decoder 30 and / or any other decoder system or subsystem described herein. The processing circuitry 46 may be used to perform various operations discussed below. As Figure 4 shown, if part of the technology is implemented in software, the device may store the instructions of the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors, thereby implementing the technology of the embodiments of the present application. One of the video encoder 20 and the video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) in a single device, as Figure 1b

[0154] shown. The source device 12 and the destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as, for example, a laptop or notebook computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, and a monitoring device, etc., and may or may not use any type of operating system. The source device 12 and the destination device 14 may also be devices in a cloud computing scenario, such as virtual machines in a cloud computing scenario. In some cases, the source device 12 and the destination device 14 may be equipped with components for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.

[0155] The source device 12 and the destination device 14 can install virtual scene application programs (applications, APPs) such as virtual reality (VR) applications, augmented reality (AR) applications, or mixed reality (MR) applications, and can run VR applications, AR applications, or MR applications based on user operations (such as clicks, touches, swipes, shakes, voice controls, etc.). The source device 12 and the destination device 14 can collect images / videos of any object in the environment through a camera and / or a sensor, and then display virtual objects on a display device according to the collected images / videos. The virtual objects can be virtual objects in a VR scene, an AR scene, or an MR scene (i.e., objects in a virtual environment).

[0156] It should be noted that in the embodiments of the present application, the virtual scene application programs in the source device 12 and the destination device 14 can be application programs built into the source device 12 and the destination device 14 themselves, or can be application programs provided by a third-party service provider installed by the user, and no specific limitation is made thereto.

[0157] In addition, the source device 12 and the destination device 14 can install real-time video transmission applications, such as live broadcast applications. The source device 12 and the destination device 14 can collect images / videos through a camera, and then display the collected images / videos on a display device.

[0158] In some cases, Figure 1a the illustrated video decoding system 10 is merely exemplary, and the techniques provided by the embodiments of the present application are applicable to video coding settings (e.g., video encoding or video decoding), which may not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from a local memory, sent over a network, etc. A video encoding device can encode data and store the data in a memory, and / or a video decoding device can retrieve data from the memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data into a memory and / or retrieve and decode data from the memory.

[0159] Please refer to Figure 1b , Figure 1b which is an exemplary block diagram of a video decoding system 40 provided by an embodiment of the present application. As Figure 1b shown, the video decoding system 40 can include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video codec implemented by a processing circuit 46), an antenna 42, one or more processors 43, one or more memory memories 44, and / or a display device 45.

[0160] AsFigure 1b As shown, the imaging device 41, antenna 42, processing circuit 46, video encoder 20, video decoder 30, processor 43, memory 44, and / or display device 45 can communicate with each other. In different instances, the video decoding system 40 may include only the video encoder 20 or only the video decoder 30.

[0161] In some instances, the antenna 42 can be used to transmit or receive the encoded bitstream of video data. Additionally, in some instances, the display device 45 can be used to present video data. The processing circuit 46 can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. The video decoding system 40 can also include an optional processor 43, which can similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. Additionally, the memory 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting instance, the memory 44 can be implemented by cache memory. In other instances, the processing circuit 46 can include a memory (e.g., cache, etc.) for implementing an image buffer, etc.

[0162] In some instances, the video encoder 20 implemented by a logic circuit can include an image buffer (e.g., implemented by the processing circuit 46 or the memory 44) and a graphics processing unit (e.g., implemented by the processing circuit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit can include the video encoder 20 implemented by the processing circuit 46 to implement the various modules discussed with reference to Figure 2 and / or any other encoder system or subsystem described herein. The logic circuit can be used to perform the various operations discussed herein.

[0163] In some instances, the video decoder 30 can be implemented in a similar manner by the processing circuit 46 to implement the reference Figure 3The video decoder 30 and / or the various modules discussed for any other decoder system or subsystem described herein. In some examples, the video decoder 30 implemented by logic circuitry may include an image buffer (implemented by the processing circuitry 46 or the memory 44) and a graphics processing unit (e.g., implemented by the processing circuitry 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 30 implemented by the processing circuitry 46 to implement with reference to Figure 3 and / or the various modules discussed for any other decoder system or subsystem described herein.

[0164] In some examples, the antenna 42 may be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data, indicators, index values, mode selection data, etc. discussed herein related to the coded video frames, e.g., data related to coded partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators as discussed, and / or data defining the coded partitions). The video decoding system 40 may further include a video decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is used to present the video frames.

[0165] It should be understood that for the examples described with reference to the video encoder 20 in the embodiments of this application, the video decoder 30 may be used to perform the reverse process. Regarding the signaling syntax elements, the video decoder 30 may be used to receive and parse such syntax elements and accordingly decode the relevant video data. In some examples, the video encoder 20 may entropy code the syntax elements into an encoded video bitstream. In such examples, the video decoder 30 may parse such syntax elements and accordingly decode the relevant video data.

[0166] For ease of description, the embodiments of this application are described with reference to the Versatile Video Coding (VVC) reference software or the High-Efficiency Video Coding (HEVC) developed by the Video Coding Experts Group (VCEG) of ITU-T and the Joint Collaborative Team on Video Coding (JCT-VC) of ISO / IEC Moving Picture Experts Group (MPEG). Those of ordinary skill in the art understand that the embodiments of this application are not limited to HEVC or VVC.

[0167] Encoder and encoding method

[0168] As Figure 2As shown, the video encoder 20 includes an input end (or input interface) 201, a residual calculation unit 204, a transformation processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output end (or output interface) 272. The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a segmentation unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The video encoder 20 shown may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.

[0169] Image and image segmentation (image and block)

[0170] The encoder 20 can be used to receive an image (or image data) 17 through the input end 201, etc., for example, an image in an image sequence forming a video or a video sequence. The received image or image data may also be a preprocessed image (or preprocessed image data) 19. For simplicity, the following description uses the image 17. The image 17 may also be referred to as the current image or the image to be encoded (especially when distinguishing the current image from other images in video coding, other images such as previously encoded and / or decoded images in the same video sequence, i.e., the video sequence that also includes the current image).

[0171] (Digital) images are or can be regarded as two-dimensional arrays or matrices composed of pixel points with intensity values. Pixel points in the array can also be called pixels (pixel or pel, short for picture element). The number of pixel points in the horizontal and vertical directions (or axes) of the array or image determines the size and / or resolution of the image. To represent colors, usually three color components are adopted, that is, the image can be represented as or include three pixel point arrays. In the RGB format or color space, the image includes corresponding red, green, and blue pixel point arrays. However, in video coding, each pixel is usually represented in the luminance / chrominance format or color space, such as YCbCr, including the luminance component indicated by Y (sometimes also represented by L) and two chrominance components represented by Cb and Cr. The luminance component Y represents the luminance or gray level intensity (for example, they are the same in a grayscale image), while the two chrominance components Cb and Cr represent the chrominance or color information components. Accordingly, an image in the YCbCr format includes a luminance pixel point array of luminance pixel point values (Y) and two chrominance pixel point arrays of chrominance values (Cb and Cr). An image in the RGB format can be converted or transformed into the YCbCr format, and vice versa, and this process is also called color transformation or conversion. If the image is black and white, then the image can only include a luminance pixel point array. Accordingly, the image can be, for example, a luminance pixel point array in a monochrome format or a luminance pixel point array and two corresponding chrominance pixel point arrays in color formats such as 4:2:0, 4:2:2, and 4:4:4.

[0172] In one embodiment, an embodiment of the video encoder 20 may include an image segmentation unit ( Figure 2 not shown in the figure) for segmenting the image 17 into a plurality of (usually non-overlapping) image blocks 203. These blocks may also be called root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB), or coding tree units (CTU) in the H.265 / HEVC and VVC standards. The segmentation unit can be used to use the same block size for all images in the video sequence and use the corresponding grid defining the block size, or change the block size between images or subsets of images or groups of images, and segment each image into corresponding blocks.

[0173] In other embodiments, the video encoder can be used to directly receive the blocks 203 of the image 17, for example, one, several, or all of the blocks constituting the image 17. The image block 203 can also be called the current image block or the image block to be encoded.

[0174] Similar to Image 17, Image Block 203 is also or can be considered as a two-dimensional array or matrix composed of pixel points with intensity values (pixel point values), but Image Block 203 is smaller than Image 17. In other words, Block 203 can include an array of pixel points (e.g., a luminance array in the case of a monochrome image 17 or a luminance array or a chrominance array in the case of a color image) or three arrays of pixel points (e.g., one luminance array and two chrominance arrays in the case of a color image 17) or any other number and / or type of arrays according to the color format adopted. The number of pixel points in the horizontal and vertical directions (or axes) of Block 203 defines the size of Block 203. Accordingly, the block can be an array of M×N (M columns × N rows) pixel points, or an array of M×N transform coefficients, etc.

[0175] In one embodiment, Figure 2 The illustrated video encoder 20 is used to encode Image 17 block by block, for example, performing encoding and prediction on each block 203.

[0176] In one embodiment, Figure 2 The illustrated video encoder 20 can also be used to segment and / or encode an image using slices (also called video slices), where the image can be segmented or encoded using one or more slices (usually non-overlapping). Each slice can include one or more blocks (e.g., Coding Tree Unit CTU) or one or more groups of blocks (e.g., coding blocks (tile) in H.265 / HEVC / VVC standards and bricks in VVC standards).

[0177] In one embodiment, Figure 2 The illustrated video encoder 20 can also be used to segment and / or encode an image using slice / coding block group (also called video coding block group) and / or coding block (also called video coding block), where the image can be segmented or encoded using one or more slice / coding block groups (usually non-overlapping), each slice / coding block group can include one or more blocks (e.g., CTU) or one or more coding blocks, etc., where each coding block can be in a shape such as a rectangle and can include one or more complete or partial blocks (e.g., CTU).

[0178] Residual calculation

[0179] The residual calculation unit 204 is used to calculate the residual block 205 according to the image block (or original block) 203 and the prediction block 265 (the prediction block 265 is introduced in detail later) in the following way: for example, subtracting the pixel point values of the prediction block 265 from the pixel point values of the image block 203 pixel by pixel (pixel by pixel) to obtain the residual block 205 in the pixel domain.

[0180] Quantization

[0181] Quantization unit 208 is used to quantize transform coefficients 207 through, for example, scalar quantization or vector quantization to obtain quantized transform coefficients 209. The quantized transform coefficients 209 may also be referred to as quantized residual coefficients 209.

[0182] The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, during quantization, an n-bit transform coefficient can be rounded down to an m-bit transform coefficient, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different degrees of scaling can be applied to achieve finer or coarser quantization. A smaller quantization step corresponds to finer quantization, while a larger quantization step corresponds to coarser quantization. The appropriate quantization step can be indicated by the quantization parameter (QP). For example, the quantization parameter can be an index of a predefined set of appropriate quantization steps. For example, a smaller quantization parameter can correspond to fine quantization (smaller quantization step), and a larger quantization parameter can correspond to coarse quantization (larger quantization step), and vice versa. Quantization can include dividing by the quantization step, and the corresponding or inverse dequantization performed by dequantization unit 210 and the like can include multiplying by the quantization step. Embodiments according to some standards such as HEVC can be used to determine the quantization step using the quantization parameter. Generally, the quantization step can be calculated using a fixed-point approximation of an equation involving division according to the quantization parameter. Other scaling factors can be introduced for quantization and dequantization to recover the norm of the residual block that may be modified due to the scaling used in the fixed-point approximation of the equations for the quantization step and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream. Quantization is a lossy operation, where the larger the quantization step, the greater the loss.

[0183] In one embodiment, video encoder 20 (correspondingly, quantization unit 208) can be used to output the quantization parameter (QP), for example, directly output or output after being encoded or compressed by entropy coding unit 270, such that video decoder 30 can receive and use the quantization parameter for decoding.

[0184] Dequantization

[0185] Dequantization unit 210 is used to perform the inverse quantization of quantization unit 208 on the quantized coefficients to obtain dequantized coefficients 211. For example, it performs an inverse quantization scheme of the quantization scheme performed by quantization unit 208 according to or using the same quantization step as quantization unit 208. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211 and correspond to the transform coefficients 207. However, due to the loss caused by quantization, the dequantized coefficients 211 are usually not exactly the same as the transform coefficients.

[0186] Reconstruction

[0187] The reconstruction unit 214 (e.g., the adder 214) is used to add the transform block 213 (i.e., the reconstruction residual block 213) to the prediction block 265 to obtain the reconstruction block 215 in the pixel domain. For example, the pixel values of the reconstruction residual block 213 and the pixel values of the prediction block 265 are added together.

[0188] Segmentation

[0189] The segmentation unit 262 can segment (or partition) an image block (or CTU) 203 into smaller parts, such as small blocks in the shape of a square or a rectangle. For an image with an array of three pixels, a CTU consists of an N×N block of luminance pixels and two corresponding chrominance pixel blocks. The maximum allowed size of the luminance block in the currently developing Versatile Video Coding (VVC) standard is specified as 128×128, but it may be specified as a value different from 128×128 in the future, such as 256×256. The CTUs of an image can be grouped / collected into slices / coding tree units, coding blocks, or tiles. A coding block covers a rectangular area of an image, and a coding block can be divided into one or more tiles. A tile consists of multiple CTU rows within a coding block. A coding block that is not divided into multiple tiles can be called a tile. However, a tile is a proper subset of a coding block and thus is not called a coding block. VVC supports two coding tree unit modes, namely the raster scan slice / coding tree unit mode and the rectangular slice mode. In the raster scan coding tree unit mode, a slice / coding tree unit contains a sequence of coding blocks in the raster scan of the coding blocks of an image. In the rectangular slice mode, a slice contains multiple tiles of an image, and these tiles together form a rectangular area of the image. The tiles within a rectangular slice are arranged in the tile raster scan order of the slice. These smaller blocks (which can also be called sub-blocks) can be further segmented into even smaller parts. This is also called tree segmentation or hierarchical tree segmentation, where the root block at the root tree level 0 (hierarchical level 0, depth 0), etc., can be recursively segmented into two or more blocks at the next lower tree level, such as the nodes at tree level 1 (hierarchical level 1, depth 1). These blocks can in turn be segmented into two or more blocks at the next lower level, such as tree level 2 (hierarchical level 2, depth 2), etc., until the segmentation ends (because an end criterion is met, such as reaching the maximum tree depth or the minimum block size). Blocks that are not further segmented are also called leaf blocks or leaf nodes of the tree. A tree segmented into two parts is called a binary tree (BT), a tree segmented into three parts is called a ternary tree (TT), and a tree segmented into four parts is called a quadtree (QT).

[0190] Entropy coding

[0191] The entropy coding unit 270 is used to apply an entropy coding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CALVC), arithmetic coding scheme, binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements, to obtain encoded image data 21 that can be output in the form of an encoded bitstream 21 etc. through the output terminal 272, so that a video decoder 30 etc. can receive and use the parameters for decoding. The encoded bitstream 21 can be transmitted to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.

[0192] Other structural variants of the video encoder 20 can be used to encode a video stream. For example, a non-transform-based encoder 20 can directly quantize the residual signal when some blocks or frames do not have a transform processing unit 206. In another implementation, the encoder 20 can have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.

[0193] Decoder and decoding method

[0194] As Figure 3 shown, the video decoder 30 is used to receive encoded image data 21 (e.g., encoded bitstream 21) encoded by, for example, the encoder 20, to obtain a decoded image 331. The encoded image data or bitstream includes information for decoding the encoded image data, such as data representing image blocks of an encoded video slice (and / or encoded group of blocks or encoded block) and related syntax elements.

[0195] In Figure 3In the example of, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (such as an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. In some examples, the video decoder 30 may perform a decoding process that is generally the reverse of the encoding process described with reference to Figure 2 the video encoder 100.

[0196] As described for the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer DPB 230, the inter prediction unit 344, and the intra prediction unit 354 also constitute the "built-in decoder" of the video encoder 20. Accordingly, the inverse quantization unit 310 may be functionally the same as the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally the same as the inverse transform processing unit 122, the reconstruction unit 314 may be functionally the same as the reconstruction unit 214, the loop filter 320 may be functionally the same as the loop filter 220, and the decoded picture buffer 330 may be functionally the same as the decoded picture buffer 230. Therefore, the explanations of the corresponding units and functions of the video encoder 20 correspondingly apply to the corresponding units and functions of the video decoder 30.

[0197] Entropy Decoding

[0198] The entropy decoding unit 304 is used to parse the bitstream 21 (or generally the encoded picture data 21) and perform entropy decoding on the encoded picture data 21 to obtain quantization coefficients 309 and / or decoded encoded parameters ( Figure 3 not shown in), such as any one or all of inter prediction parameters (such as reference picture indices and motion vectors), intra prediction parameters (such as intra prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements, etc. The entropy decoding unit 304 may be used to apply a decoding algorithm or scheme corresponding to the encoding scheme of the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may also be used to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode application unit 360, and other parameters to other units of the decoder 30. The video decoder 30 may receive video slice and / or video block-level syntax elements. Additionally, or as an alternative to slices and corresponding syntax elements, encoded block groups and / or encoded blocks and corresponding syntax elements may be received or used.

[0199] Inverse Quantization

[0200] The inverse quantization unit 310 may be used to receive a quantization parameter (QP) (or generally information related to inverse quantization) and quantized coefficients from the encoded picture data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse quantize the decoded quantized coefficients 309 based on the quantization parameter to obtain inverse quantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include determining the degree of quantization using the quantization parameter calculated by the video encoder 20 for each video block in the video slice, and also determining the degree of inverse quantization to be performed.

[0201] Reconstruction

[0202] The reconstruction unit 314 (e.g., adder 314) is used to add the reconstructed residual block 313 to the prediction block 365 to obtain a reconstructed block 315 in the pixel domain. For example, the pixel values of the reconstructed residual block 313 and the pixel values of the prediction block 365 are added together.

[0203] Other variants of the video decoder 30 may be used to decode the encoded picture data 21. For example, the decoder 30 may produce an output video stream without the loop filter unit 320. For example, a non-transform based decoder 30 may directly inverse quantize the residual signal without the inverse transform processing unit 312 for some blocks or frames. In another implementation, the video decoder 30 may have the inverse quantization unit 310 and the inverse transform processing unit 312 combined into a single unit.

[0204] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clip or shift operations may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.

[0205] It should be noted that further operations can be performed on the derived motion vectors of the current block (including but not limited to the control point motion vectors in the affine mode, the sub-block motion vectors in the affine, planar, and ATMVP modes, the temporal motion vectors, etc.). For example, the value of the motion vector is restricted to a predefined range according to the representation bits of the motion vector. If the representation bits of the motion vector are bitDepth, the range is from -2^(bitDepth - 1) to 2^(bitDepth - 1) - 1, where "^" represents exponentiation. For example, if bitDepth is set to 16, the range is from -32768 to 32767; if bitDepth is set to 18, the range is from -131072 to 131071. For example, the value of the derived motion vector (such as the MV of 4 4×4 sub-blocks in an 8×8 block) is restricted such that the maximum difference between the integer parts of the 4 4×4 sub-block MVs does not exceed N pixels, for example, does not exceed 1 pixel. Two methods for restricting the motion vector according to bitDepth are provided here.

[0206] Although the above embodiments mainly describe video coding and decoding, it should be noted that the embodiments of the decoding system 10, the encoder 20, and the decoder 30, as well as other embodiments described herein, can also be used for still image processing or coding and decoding, that is, the processing or coding and decoding of a single image independent of any previous or consecutive images in video coding and decoding. Generally, if the image processing is limited to a single image 17, the inter-frame prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 can equally be used for static image processing, such as residual calculation 204 / 304, transformation 206, quantization 208, dequantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354, and / or loop filtering 220 / 320, entropy coding 270, and entropy decoding 304.

[0207] Please refer to Figure 4 , Figure 4 which is an exemplary block diagram of the video decoding device 400 provided by the embodiments of the present application. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 400 can be a decoder, such as Figure 1a the video decoder 30 in Figure 1a or an encoder, such as

[0208] Video decoding device 400 includes: an input port 410 (or input port 410) for receiving data and a receiver unit (Rx) 420; a processor, logic unit, or central processing unit (CPU) 430 for processing data; for example, the processor 430 here can be a neural network processor 430; a transmitter unit (Tx) 440 and an output port 450 (or output port 550) for transmitting data; and a memory 460 for storing data. Video decoding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 410, the receiving unit 420, the transmitting unit 440, and the output port 450 for the exit or entry of optical or electrical signals.

[0209] Processor 430 is implemented by hardware and software. Processor 430 can be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with the input port 410, the receiving unit 420, the transmitting unit 440, the output port 450, and the memory 460. Processor 430 includes a neural network-based codec 470. The neural network-based codec 470 implements the embodiments disclosed above. For example, the neural network-based codec 470 performs, processes, prepares, or provides various encoding operations. Therefore, the neural network-based codec 470 provides a substantial improvement to the functions of video decoding device 400 and affects the switching of video decoding device 400 to different states. Alternatively, the neural network-based codec 470 is implemented by instructions stored in the memory 460 and executed by the processor 430.

[0210] Memory 460 includes one or more disks, tape drives, and solid-state drives and can be used as an overflow data storage device for storing such programs when a selected program is to be executed and for storing instructions and data read during program execution. Memory 460 can be volatile and / or non-volatile and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0211] Please refer to Figure 5 , Figure 5An exemplary block diagram of a device 500 provided in an embodiment of the present application, the device 500 can be used as Figure 1a Either or both of the source device 12 and the destination device 14 in .

[0212] The processor 502 in the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices that are currently available or will be developed in the future and are capable of manipulating or processing information. Although a single processor such as the processor 502 shown in the figure may be used to implement the disclosed implementation, using more than one processor is faster and more efficient.

[0213] In one implementation, the memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessed by the processor 502 via a bus 512. The memory 504 may also include an operating system 508 and an application 510, which includes at least one program that allows the processor 502 to perform the methods described herein. For example, the application 510 may include applications 1 to N, and also include a video decoding application that performs the methods described herein.

[0214] The apparatus 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element that may be used to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.

[0215] Although bus 512 in device 500 is described herein as a single bus, bus 512 may include multiple buses. In addition, auxiliary storage may be directly coupled to other components of device 500 or accessed through a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Therefore, device 500 may have a variety of configurations.

[0216] Image encoding:

[0217] Nowadays, multimedia data occupies a large part of Internet traffic. The compression of image data plays an important role in the storage and efficient transmission of multimedia data. Therefore, image coding technology is a technology with great practical value. It should be noted that the Chinese translation of the English terms "image coding" and "image encoding" is usually "image coding". Image coding is broad and includes the process of encoding an image into a bit stream and the process of decoding (decoding) a bit stream into an image.

[0218] Image encoding only refers to the process of encoding an image into a bitstream. The research on image encoding has a long history. Researchers have proposed a large number of methods and developed I-frame encoding methods for various image encoding standards and video encoding standards such as JPEG, JPEG2000, JPEG-XL, JPEG-XX, WebP, H.264 / AVC, H.264 / HEVC, H.26 / VVC, AVS3, and AV1. Most of these encoding methods are based on transform, prediction, and entropy encoding techniques. Although these encoding methods are currently widely used, due to the increase in image data volume and the emergence of new media types, encoding methods with higher compression efficiency are needed.

[0219] Image Encoding Based on Deep Learning:

[0220] In recent years, researchers have studied image encoding methods based on deep learning. Some researchers have achieved good results. For example, Balle et al. proposed an end-to-end optimized image encoding method that is superior to existing state-of-the-art image encodings and even superior to the existing best traditional encoding standard H.265 / HEVC.

[0221] Image encoding based on deep learning is based on deep neural networks, usually convolutional neural networks. Some research works have proposed an image encoding method based on the Transformer network. The structure of the deep neural network can be designed manually or obtained through neural architecture search (NAS). The parameters of the deep neural network are obtained by using the loss function and the backpropagation algorithm.

[0222] Figure 6 Shows a typical deep learning-based image compression method, also known as neural network-based image compression. Generally, neural network-based image compression methods include the following parts: feature extraction module, feature quantization module, entropy encoding module, entropy decoding module, feature dequantization module, and feature decoding module. On the encoder side, the feature extraction module can use a non-linear mapping activation function to obtain the extracted three-dimensional feature map through multi-layer convolution stacking. The feature quantization module quantizes the floating-point feature values through eigenvalue quantization to obtain the quantized feature values. The quantized feature values are subjected to lossless entropy encoding to obtain the encoded bitstream. When receiving the bitstream of entropy encoding, the decoder performs lossless entropy decoding to obtain the three-dimensional quantized feature values. The feature decoding module decodes the features into a reconstructed image to achieve decoding.

[0223] After the image to be compressed passes through the feature extraction module and the feature quantization module, a three-dimensional feature quantization map is obtained. When processing each eigenvalue in the three-dimensional feature quantization map, the entropy coding module can estimate and obtain the probability distribution of the eigenvalue by using the eigenvalues in the processed neighborhood as context, and perform subsequent coding based on the probability distribution to obtain a coded bitstream.

[0224] With the excellent performance of deep learning in various fields, researchers have proposed an end-to-end image coding solution based on deep learning. Figure 7 The coding framework is shown. The specific technical solution is as follows: On the encoder side, the original image is input into the feature extraction module, and a feature map is output. The feature map passes through the side information extraction module, and side information is output. On both the encoder and decoder sides, it is input into the probability estimation module, and the probability distribution of each feature element is output to obtain the value of the feature element to be coded. In addition, the feature map is input into the quantization module to obtain a quantized feature map. The entropy coding module performs entropy coding on each feature element in the quantized feature map based on the probability distribution of each feature element to obtain a coded bitstream.

[0225] On the decoder side, the decoder parses the bitstream and outputs the probability distribution of the symbol to be coded based on the side information to obtain the value of the feature element to be decoded. The entropy decoding module performs arithmetic decoding on each feature element in the quantized feature map based on the probability distribution of each feature element to obtain the value of the feature element. The feature map is input into the image reconstruction module, and the reconstructed image is output.

[0226] Neural network

[0227] A neural network can include neurons. A neuron can be an operation unit that uses \(x_s\) and an intercept \(1\) as inputs. The output of the operation unit can be:

[0228]

[0229] Here, \(s = 1,2,\cdots,n\), \(n\) is a natural number greater than \(1\), \(W_s\) is the weight of \(x_s\), \(b\) is the bias of the neuron, and \(f\) is the activation function of the neuron (activation function), which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting multiple individual neurons together. Specifically, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be a region including several neurons.

[0230] Convolutional neural network:

[0231] A Convolutional Neural Network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network includes a feature extractor, which includes a convolutional layer and a subsampling layer. The feature extractor can be regarded as a filter. The convolutional layer is a layer of neurons in the convolutional neural network that performs convolutional processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can be connected to only a part of the neurons in the adjacent layer. The convolutional layer usually includes several feature planes, and each feature plane can include some neurons arranged in a rectangle. Neurons on the same feature plane share a weight, and the shared weight here is the convolutional kernel. The shared weight can be understood as a way of extracting image information that is independent of position. The convolutional kernel can be initialized in the form of a matrix of random size. During the training process of the convolutional neural network, appropriate weights can be obtained for the convolutional kernel through learning. In addition, the shared weight directly reduces the connections between the layers of the convolutional neural network and reduces the risk of overfitting.

[0232] Figure 8 Schematically shows the general concept of processing by a neural network such as a CNN. A convolutional neural network consists of an input layer, an output layer, and multiple hidden layers. The input layer is the layer that provides the input (e.g., Figure 8 a part of the image shown) for processing. The hidden layers of a CNN usually consist of a series of convolutional layers that perform convolution with multiplication or other dot products. The result of the layer is one or more feature maps, sometimes also called channels. Subsampling may be involved in some or all of the layers. Thus, as Figure 8 shown, the feature maps can become smaller. The activation function in a CNN is usually a ReLU (Rectified Linear Unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, and normalization layers, which are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. Although these layers are popularly called convolutions, this is just a convention. Mathematically, it is technically a sliding dot product or cross-correlation. This has important implications for the indices in the matrix because it affects how the weights are determined at a specific index point.

[0233] When programming a CNN for processing images, as Figure 8 shown, the input is a tensor with the shape (number of images) x (image width) x (image height) x (image depth). Then, after passing through the convolutional layer, the image is abstracted into a feature map with the shape (number of images) x (feature map width) x (feature map height) x (feature map channels). The convolutional layer in a neural network should have the following properties. A convolutional kernel (hyperparameter) defined by width and height. The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) should be equal to the number of channels (depth) of the input feature map.

[0234] In the past, traditional multi-layer perceptron (MLP) models have been used for image recognition. However, due to their full connectivity between nodes, they have a high dimensionality and do not scale well with higher-resolution images. An image of 1000×1000 pixels with RGB color channels has 3 million weights, which is too high to be effectively processed at scale in a fully connected manner. Additionally, this network architecture does not consider the spatial structure of the data, treating input pixels that are far apart in the same way as those that are close. This ignores the locality of reference in image data both computationally and semantically. Therefore, the full connectivity of neurons is wasteful for purposes such as image recognition, which is dominated by spatially local input patterns.

[0235] Convolutional neural networks are biologically inspired variants of multi-layer perceptrons, specifically designed to mimic the behavior of the visual cortex. These models alleviate the challenges posed by the MLP architecture by exploiting the strong spatial local correlations present in natural images. The convolutional layer is the core building block of a CNN. The parameters of this layer consist of a set of learnable filters (the kernels mentioned above), which have small receptive fields but extend across the entire depth of the input volume. During the forward pass, each filter is convolved over the width and height of the input volume, computing the dot product between the entries of the filter and the input, and producing a two-dimensional activation map for that filter. Thus, the network learns filters that activate when it detects a particular type of feature at a certain spatial location in the input.

[0236] Stacking the activation maps of all filters along the depth dimension forms the complete output volume of the convolutional layer. Thus, each entry in the output volume can also be interpreted as the output of a neuron that observes a small region in the input and shares parameters with neurons in the same activation map. A feature map or activation map is the output activation of a given filter. Feature map and activation have the same meaning. In some papers, it is called an activation map because it is a mapping corresponding to the activation of different parts of the image, and it is also a feature map because it is also a mapping where a certain feature is found in the image. High activation means that a certain feature has been found.

[0237] Another important concept in convolutional neural networks is pooling, which is a form of non-linear downsampling. There are several non-linear functions that can implement pooling, among which max pooling is the most common. It divides the input image into a set of non-overlapping rectangles, and for each such sub-region, it outputs the maximum value.

[0238] Intuitively, the exact location of a feature is less important than its rough location relative to other features. This is the idea behind using pooling in convolutional neural networks. Pooling layers are used to gradually reduce the spatial size of the representation, reducing the number of parameters, memory footprint, and computational load in the network, and thus also controlling overfitting. In CNN architectures, it is common to periodically insert pooling layers between successive convolutional layers. Pooling operations provide another form of translational invariance.

[0239] Pooling layers operate independently on each depth slice of the input and resize it spatially. The most common form is a pooling layer with a filter of size 2×2, which applies a downsampling stride of 2 along the width and height on each depth slice of the input, discarding 75% of the activations. In this case, each max operation is over 4 numbers. The depth dimension remains unchanged.

[0240] In addition to max pooling, pooling units can also use other functions, such as average pooling or l2 - norm pooling. Average pooling has been used frequently historically but has fallen out of favor recently compared to max pooling, which performs better in practice. Due to the significant reduction in representation size, there has been a recent trend to use smaller filters or to discard pooling layers altogether. "Region of interest" pooling (also known as ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important part of convolutional neural networks for object detection based on the fast R - CNN architecture.

[0241] The above ReLU is an abbreviation for rectified linear unit, which applies a non - saturated activation function. By setting negative values to zero, it effectively removes negative values from the activation map. It increases the non - linearity of the decision function and the entire network without affecting the receptive field of the convolutional layer. Other functions are also used to increase non - linearity, such as the saturated hyperbolic tangent and the sigmoid function. ReLU is generally more popular than other functions because it trains neural networks several times faster without significantly affecting the generalization accuracy.

[0242] After several convolutional layers and max - pooling layers, high - level inference in the neural network is done through fully - connected layers. Neurons in fully - connected layers are connected to all activations in the previous layer, as seen in regular (non - convolutional) artificial neural networks. Thus, their activations can be computed as an affine transformation, followed by a bias offset (a vector addition of learned or fixed bias terms) after matrix multiplication.

[0243] The "loss layer" specifies how training penalizes the deviation between predictions (outputs) and true labels, typically the last layer of a neural network. Various loss functions suitable for different tasks can be used. Softmax loss is used to predict a single class out of K mutually exclusive classes. Sigmoid cross-entropy loss is used to predict K independent probability values in the range [0, 1]. Euclidean loss is used for regression to real-valued labels.

[0244] In summary, Figure 8 shows the data flow in a typical convolutional neural network. First, the input image passes through a convolutional layer and is abstracted into a feature map consisting of several channels, corresponding to the number of filters in a set of learnable filters for that layer (e.g., one channel per filter). Then the feature map is subsampled using, for example, a pooling layer, which reduces the dimensionality of each channel in the feature map. The subsequent data enters another convolutional layer, which may have a different number of output channels, resulting in a different number of channels in the feature map. As mentioned above, the number of input channels and output channels are hyperparameters of the layer. To establish the connectivity of the network, these parameters need to be synchronized between two connected layers, e.g., the number of input channels of the current layer should be equal to the number of output channels of the previous layer. For the first layer that processes input data (e.g., an image), the number of input channels is usually equal to the number of channels of the data representation, e.g., 3 channels for an RGB or YUV representation of an image or video, or 1 channel for a grayscale image or video representation.

[0245] The present invention is applied to a video codec, such as a video codec in a video communication system, such as Figure 9 shown: After a video is captured using a video capture device, it undergoes a series of preprocessing, and then the processed video is compressed and encoded to obtain an encoded bitstream. The bitstream is sent to a receiving module via a transmission network using a sending module, and after being decoded by a decoder, it can be rendered and displayed. In addition, the bitstream after video encoding can also be directly stored.

[0246] The present invention describes a neural network-based encoding and decoding scheme for enhancement layers, and its application scenario can be a hierarchical coding scheme.

[0247] The present invention can be used in devices or products with video encoder and / or decoder functions, such as video processing software and hardware products, such as chips. And in products or devices containing such chips, such as media products like mobile phones.

[0248] Figure 10 Shows an encoding method provided by an embodiment of the present application, such as Figure 10 shown, the method includes:

[0249] S1001. Perform AI encoding on the input image to obtain a first feature map of the input image.

[0250] Among them, any one of the N groups of second feature maps can be decoded to obtain a reconstructed image, where N is a positive integer.

[0251] The first feature map is a three-dimensional data. As Figure 8 shown, the three-dimensional feature map is composed of multiple two-dimensional feature maps, and each two-dimensional feature map includes multiple elements (pixel points). Each two-dimensional feature map corresponds to one channel (such as the RPG channel) in the channel dimension

[0252] In a possible implementation, the first feature map of the input image can be obtained by extracting features from the input image through the feature extraction module of the AI encoder.

[0253] Exemplarily, the first feature map of the input image can be obtained by extracting features from the input image through the feature extraction module of the JPEG AI encoder.

[0254] S1002. Determine N groups of second feature maps of the input image according to the first feature map of the input image.

[0255] Among them, N is a positive integer. For example, N can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or other positive integers. Any one of the N groups of second feature maps can be decoded to obtain a reconstructed image.

[0256] The second feature map is a two-dimensional data. As Figure 8 shown, as Figure 8 shown, the three-dimensional feature map is composed of multiple two-dimensional feature maps, and each two-dimensional feature map corresponds to one channel in the channel dimension, so the three-dimensional feature map can be divided into multiple two-dimensional feature maps in the channel dimension.

[0257] For example, for a three-dimensional feature map with 192 channels, the three-dimensional feature map can be separated into 192 two-dimensional feature maps in the channel dimension.

[0258] In a possible implementation, the above first feature map can be separated into M second feature maps of the above input image. The M second feature maps are grouped according to the channel information of the second feature map of the above input image to obtain the above N groups of second feature maps. M is a positive integer, and M ≤ N. For example, M can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or other positive integers.

[0259] Among them, M can be the number of channels of the first feature map.

[0260] In a possible implementation, the channel information includes channel entropy and / or channel variance. Among them, the above channel entropy is the entropy of all elements in the channel, and the above channel variance is the variance of all elements in the channel.

[0261] For example, taking the case where each group of second feature maps includes K second feature maps, the K second feature maps with the highest channel variances can be assigned to group 1, the K second feature maps with the second highest channel variances can be assigned to group 2, …, the K second feature maps with the second lowest channel variances can be assigned to group N - 1, and the K second feature maps with the lowest channel variances can be assigned to group N.

[0262] For another example, taking the case where each group of second feature maps includes K second feature maps, the K second feature maps with the highest channel entropies can be assigned to group 1, the K second feature maps with the second highest channel entropies can be assigned to group 2, …, the K second feature maps with the second lowest channel entropies can be assigned to group N - 1, and the K second feature maps with the lowest channel entropies can be assigned to group N.

[0263] For another example, taking the case where each group of second feature maps includes K second feature maps, the second feature maps can be numbered in ascending order of channel entropy. Then the K second feature maps with the highest numbers can be assigned to group 1, the K second feature maps with the second highest numbers can be assigned to group 2, …, the K second feature maps with the second lowest numbers can be assigned to group N - 1, and the K second feature maps with the lowest numbers can be assigned to group N.

[0264] Exemplarily, if the first feature map has 384 channels, then the above first feature map can be separated into 384 second feature maps of the above input image, and then 192 second feature maps with the highest channel entropies among the 384 second feature maps can be grouped into one group according to the channel entropies of the 384 second feature maps, and 192 second feature maps with the lowest channel entropies can be grouped into another group.

[0265] For another example, as Figure 11 shown, the above first feature map can be separated into M second feature maps of the above input image, and the M second feature maps can be divided into N groups in descending order of the channel entropy or channel variance of the M second feature maps.

[0266] For another example, as Figure 12 shown, the above first feature map can be separated into M second feature maps of the above input image, and the M second feature maps can be divided into 2 groups in descending order of the channel entropy or channel variance of the M second feature maps.

[0267] S1003. Entropy encode the first group of second feature maps among the N groups of second feature maps of the input image into the bitstream.

[0268] In a possible implementation, the first set of second feature maps among N sets of second feature maps of the input image can be entropy encoded into a bitstream using the asymmetric numeral system (ANS).

[0269] In a possible implementation, the second set of second feature maps among the above-mentioned N sets of second feature maps can also be entropy encoded into the above-mentioned bitstream.

[0270] In a possible implementation, the third set of second feature maps among the above-mentioned N sets of second feature maps can also be entropy encoded into the above-mentioned bitstream.

[0271] In a possible implementation, at least one set of second feature maps among the above-mentioned N sets of second feature maps can be entropy encoded into the above-mentioned bitstream in the order from large to small according to channel entropy or channel variance.

[0272] Exemplarily, as Figure 11 shown, after dividing M second feature maps of the input image into N groups, the N groups of second feature maps of the input image can be entropy encoded into a bitstream in the order from large to small according to channel entropy or channel variance.

[0273] Another exemplarily, as Figure 12 shown, after dividing M second feature maps of the input image into 2 groups, the 2 groups of second feature maps of the input image can be entropy encoded into a bitstream in the order from large to small according to channel entropy or channel variance.

[0274] Another exemplarily, as Figure 13 shown, after dividing M second feature maps of the input image into 2 groups, only the first set of second feature maps of the input image can be entropy encoded into a bitstream in the order from large to small according to channel entropy or channel variance.

[0275] It can be seen that in the method provided by the embodiments of the present application, a progressive encoding method is adopted, and partial feature maps of the input image can be encoded. Compared with encoding all feature maps of the input image, encoding partial feature maps of the input image can reduce the image encoding delay. In addition, the first set of second feature maps of the input image is the feature map with the largest channel entropy (variance) among N sets of second feature maps of the input image, and the decoding end can obtain a complete reconstructed image based on this feature map.

[0276] For example, in a scenario where the network environment is poor or the decoding hardware resources are insufficient, the progressive encoding method adopted by the embodiments of the present application does not require all feature maps, and a complete reconstructed image can be obtained only by decoding based on the first set of second feature maps, thereby reducing the image encoding and decoding delay.

[0277] In a possible implementation, the channel number information can be encoded into the above-mentioned bitstream, and the channel number information is used to indicate the number of channels of at least one group of second feature maps encoded into the above-mentioned bitstream. In this way, the decoding end can determine the number of channels of each group of second feature maps in the bitstream by decoding the channel number information, which helps the decoding end to recover the feature maps from the bitstream.

[0278] For example, the number of channels of the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps are 4, 3, and 3 respectively. Then, the channel number information of 4, 3, and 3 can be encoded into the above-mentioned bitstream.

[0279] In a possible implementation, the first position information can be encoded into the above-mentioned bitstream, and the first position information is used to indicate the position of at least one group of second feature maps encoded into the above-mentioned bitstream in the above-mentioned first feature map. In this way, the decoding end can determine the position of each group of second feature maps in the first feature map by decoding the first position information, which helps the decoding end to recover the feature maps before grouping from the bitstream.

[0280] It should be noted that the decoding end needs to use the position of the second feature map in the above-mentioned first feature map to obtain the second feature map, and the position of the second feature map in the bitstream may be different from the position of the second feature map in the above-mentioned first feature map. Therefore, it is necessary to encode the first position information recording the position of at least one group of second feature maps encoded into the above-mentioned bitstream in the above-mentioned first feature map into the bitstream.

[0281] For example, the first feature map is separated into 10 second feature maps through channel separation, and the positions of these 10 second feature maps in the first feature map are respectively recorded as 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10. The grouping and sorting order of these 10 second feature maps into the bitstream are 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 respectively. Then, the first position information of 2, 9, 4, 7, 1, 5, 3, 6, 10, and 8 can be encoded into the bitstream. After obtaining this information, the decoding end can know that the position of the first decoded second feature map in the first feature map is 2, the position of the second decoded second feature map in the first feature map is 9,..., and the position of the tenth decoded second feature map in the first feature map is 8.

[0282] In a possible implementation, the second position information can be encoded into the above-mentioned bitstream, and the second position information is used to indicate the position of at least one group of second feature maps encoded into the above-mentioned bitstream in the above-mentioned bitstream. In this way, the decoding end can determine the position of each group of second feature maps in the bitstream by the second position information, which helps the decoding end to perform parallel decoding.

[0283] In a possible implementation, the bitstream length information can be included in the above-mentioned bitstream, and the above-mentioned bitstream length information is used to indicate the length of the above-mentioned bitstream. In this way, the decoding end can determine the length of the bitstream through the bitstream length information, which helps the decoding end to perform parallel decoding.

[0284] Figure 14 The following shows a decoding method provided by an embodiment of the present application. As Figure 14 shown, the method includes:

[0285] S1401. Perform entropy decoding on the bitstream to obtain the first set of second feature maps of the input image.

[0286] Wherein, the above-mentioned first set of second feature maps is a part of the feature maps corresponding to the above-mentioned input image.

[0287] In a possible implementation, ANS entropy decoding can be performed on the bitstream to obtain the first set of second feature maps of the input image.

[0288] In a possible implementation, entropy decoding can also be performed on the bitstream to obtain the second set of second feature maps of the input image.

[0289] In a possible implementation, entropy decoding can also be performed on the bitstream to obtain the third set of second feature maps of the input image.

[0290] In a possible implementation, entropy decoding can also be performed on the bitstream to sequentially obtain L sets of second feature maps of the image. Wherein, L is a positive integer, and L is less than or equal to N.

[0291] Exemplarily, as Figure 11 shown, entropy decoding can be performed on the bitstream to sequentially obtain N sets of second feature maps of the image.

[0292] Another example is, as Figure 12 shown, perform entropy decoding on the bitstream to obtain the first set of second feature maps of the input image, and then perform entropy decoding on the bitstream to obtain the second set of second feature maps of the input image.

[0293] Another example is, as Figure 13 shown, entropy decoding can be performed on the bitstream to obtain only the first set of second feature maps of the input image.

[0294] S1402. Perform AI decoding on the first set of second feature maps of the input image to obtain a first reconstructed image.

[0295] In a possible implementation, the first set of second feature maps of the input image can be subjected to AI decoding through the image restoration network of the AI decoder to obtain a first reconstructed image.

[0296] Exemplarily, the first reconstructed image can be obtained by performing AI decoding on the first set of second feature maps of the input image through the image restoration network of the JPEG AI decoder.

[0297] It should be noted that the first reconstructed image is not a partial image of the input image, but a reconstructed image with the same size as the input image but different image quality.

[0298] In a possible implementation, each time a set of second feature maps is solved, it can be merged with other already solved second feature maps and sent to the subsequent image restoration network to restore the image.

[0299] Exemplarily, after performing AI decoding on the second set of second feature maps of the input image, AI decoding can be performed on the first set of second feature maps and the second set of second feature maps to obtain the second reconstructed image.

[0300] Exemplarily, after performing AI decoding on the third set of second feature maps of the input image, AI decoding can be performed on the first set of second feature maps, the second set of second feature maps, and the third set of second feature maps to obtain the third reconstructed image.

[0301] In a possible implementation, the image quality of the second reconstructed image is better than that of the first reconstructed image, and the image quality of the third reconstructed image is better than that of the second reconstructed image.

[0302] Among them, the above image quality is characterized by any one of the following variables: PSNR, MS-SSIM, or LPIPS.

[0303] Exemplarily, such as Figure 11As shown, after decoding the first set of second feature maps of the input image, the first set of second feature maps of the input image can be input into the image restoration network to obtain the first reconstructed image; after decoding the second set of second feature maps, the first set of second feature maps of the input image and the second set of second feature maps of the input image can be input into the image restoration network to obtain the second reconstructed image; after decoding the third set of second feature maps of the input image, the first set of second feature maps of the input image, the second set of second feature maps of the input image, and the third set of second feature maps of the input image can be input into the image restoration network to obtain the third reconstructed image; after decoding the (N - 1)th set of second feature maps of the input image, the first set of second feature maps of the input image, the second set of second feature maps of the input image, the third set of second feature maps of the input image, ……, the (N - 1)th set of second feature maps of the input image can be input into the image restoration network to obtain the (N - 1)th reconstructed image; after decoding the Nth set of second feature maps of the input image, the first set of second feature maps of the input image, the second set of second feature maps of the input image, the third set of second feature maps of the input image, ……, the (N - 1)th set of second feature maps of the input image, and the Nth set of second feature maps of the input image can be input into the image restoration network to obtain the Nth reconstructed image.

[0304] It should be noted that the image quality of the above first reconstructed image to the above Nth reconstructed image increases in sequence. That is, the image quality of the first reconstructed image < the image quality of the second reconstructed image < the image quality of the third reconstructed image < …… < the image quality of the (N - 1)th reconstructed image < the image quality of the Nth reconstructed image.

[0305] Also exemplarily, as Figure 12 shown, after decoding the first set of second feature maps of the input image, the first set of second feature maps of the input image can be input into the image restoration network to obtain the first reconstructed image; after decoding the second set of second feature maps of the input image, the first set of second feature maps of the input image and the second set of second feature maps of the input image can be input into the image restoration network to obtain the second reconstructed image.

[0306] Also exemplarily, as Figure 13 shown, after decoding the first set of second feature maps of the input image, the first set of second feature maps of the input image can be input into the image restoration network to obtain the first reconstructed image.

[0307] It can be seen that in the method provided by the embodiments of the present application, the reconstructed image can be determined through partial feature maps of the input image. Compared with determining the reconstructed image through all the feature maps of the input image, using the progressive decoding method to determine the reconstructed image through partial feature maps of the input image can reduce the image decoding delay. In addition, the first set of second feature maps of the input image is the feature map with the largest channel entropy (variance) among the second feature maps of the input image, and the decoding end can obtain a complete reconstructed image based on this feature map.

[0308] For example, in a scenario where the network environment is poor or the decoding hardware resources are insufficient, the embodiments of the present application can obtain a complete reconstructed image by using the progressive decoding method and decoding only based on the first group of second feature maps.

[0309] In a possible implementation, the above-mentioned bitstream can be decoded to obtain channel number information, and the above-mentioned channel number information is used to indicate the number of channels of at least one group of second feature maps obtained by decoding; the first group of second feature maps is subjected to AI decoding according to the above-mentioned channel number information to obtain a first reconstructed image.

[0310] For example, the number of channels of the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps are 4, 3, and 3 respectively. Then, the channel number information of 4, 3, and 3 can be encoded into the above-mentioned bitstream. Correspondingly, after the decoding end decodes and obtains the channel number information, it can know that the first group of second feature maps of the input image can be decoded only by decoding 4-channel second feature maps for the first time, the second group of second feature maps of the input image can be decoded only by decoding 4-channel second feature maps for the second time, and the third group of second feature maps of the input image can be decoded only by decoding 3-channel second feature maps for the third time.

[0311] In a possible implementation, the above-mentioned bitstream can be decoded to obtain first position information, and the above-mentioned first position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the first feature map of the above-mentioned input image; the first group of second feature maps is subjected to AI decoding according to the above-mentioned first position information to obtain a first reconstructed image.

[0312] For example, 10 second feature maps are obtained by channel separation from the first feature map, and the positions of these 10 second feature maps in the first feature map are respectively recorded as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10. The grouping and sorting order of these 10 second feature maps and their incorporation into the bitstream are 2, 9, 4, 7, 1, 5, 3, 6, 10, 8 respectively. Then, the first position information of 2, 9, 4, 7, 1, 5, 3, 6, 10, 8 can be incorporated into the bitstream. After the decoding end obtains this information, it can know that the position of the first decoded second feature map in the first feature map is 2, the position of the second decoded second feature map in the first feature map is 9, ……, and the position of the tenth decoded second feature map in the first feature map is 8.

[0313] The first decoded second feature map is placed at position 2 in the first feature map for image restoration, the second decoded second feature map is placed at position 9 in the first feature map for image restoration, ……, and the tenth decoded second feature map is placed at position 8 in the first feature map for image restoration.

[0314] In a possible implementation, the above-mentioned bitstream can be decoded to obtain second position information, which is used to indicate the position of at least one group of decoded second feature maps in the above-mentioned bitstream; the bit data corresponding to the first group of second feature maps in the above-mentioned bitstream is determined according to the above-mentioned second position information, and the corresponding bit data is entropy decoded to obtain the first group of second feature maps.

[0315] For example, the corresponding positions of the bit data of the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps in the bitstream are A1 to A2, A3 to A4, and A5 to A6, respectively. Then, the second position information of A1 to A2, A3 to A4, and A5 to A6 can be encoded into the above-mentioned bitstream. Correspondingly, after the decoding end decodes to obtain the second position information, it can be known that the bit data at the positions of A1 to A2 in the decoded bitstream can obtain the first group of second feature maps, the bit data at the positions of A3 to A4 in the decoded bitstream can obtain the second group of second feature maps, and the bit data at the positions of A5 to A6 in the decoded bitstream can obtain the third group of second feature maps.

[0316] In a possible implementation, the above-mentioned bitstream can be decoded to obtain bitstream length information of the above-mentioned bitstream, and the bitstream length information is used to indicate the length of the above-mentioned bitstream; the first group of second feature maps is AI decoded according to the above-mentioned bitstream length information to obtain a first reconstructed image.

[0317] For example, the decoding duration can be estimated through the bitstream length. When the estimated decoding duration is greater than the preset duration, a progressive decoding method is adopted, that is, the bitstream is entropy decoded to obtain partial second feature maps of the input image. The partial second feature maps are AI decoded to obtain the reconstructed image of the input image.

[0318] Next, a coding device for performing the above coding method will be introduced in combination with Figure 15 the above.

[0319] It can be understood that in order to implement the above functions, the coding device includes corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0320] Embodiments of the present application can divide the encoding device into functional modules according to the above method examples. For example, each functional module can be corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0321] In the case of dividing each functional module corresponding to each function, Figure 15 FIG. shows a possible schematic composition of the encoding device involved in the above embodiment. This device can be an electronic device, or a module applied to an electronic device (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of the electronic device. As Figure 15 shown, the encoding device 1500 can include: an encoding unit 1501 and a determination unit 1502.

[0322] The above encoding unit 1501 is used to perform AI encoding on the input image to obtain the first feature map of the input image.

[0323] The above determination unit 1502 is used to determine N groups of second feature maps of the input image according to the first feature map, where N is a positive integer.

[0324] The above encoding unit 1501 is further used to entropy-encode the first group of second feature maps among the N groups of second feature maps into the bitstream.

[0325] In a possible implementation manner, the above encoding unit 1501 is further used to: entropy-encode the second group of second feature maps among the N groups of second feature maps into the bitstream.

[0326] In a possible implementation manner, the above encoding unit 1501 is further used to: entropy-encode the third group of second feature maps among the N groups of second feature maps into the bitstream.

[0327] In a possible implementation manner, the above determination unit 1502 is specifically used to: perform channel separation on the first feature map to obtain M second feature maps of the input image, where M is a positive integer; group the M second feature maps according to the channel information of the second feature maps of the input image to obtain the N groups of second feature maps.

[0328] In a possible implementation manner, the above encoding unit 1501 is further used to: entropy-encode at least one group of second feature maps among the N groups of second feature maps into the bitstream in the order from largest to smallest according to channel entropy or channel variance.

[0329] In a possible implementation, the above-mentioned encoding unit 1501 is further configured to: encode the number of channels information into the above-mentioned bitstream, and the number of channels information is used to indicate the number of channels of at least one group of second feature maps encoded into the above-mentioned bitstream.

[0330] In a possible implementation, the above-mentioned encoding unit 1501 is further configured to: encode the first position information into the above-mentioned bitstream, and the first position information is used to indicate the position of at least one group of second feature maps encoded into the above-mentioned bitstream in the above-mentioned first feature map.

[0331] In a possible implementation, the above-mentioned encoding unit 1501 is further configured to: encode the second position information into the above-mentioned bitstream, and the second position information is used to indicate the position of at least one group of second feature maps encoded into the above-mentioned bitstream in the above-mentioned bitstream.

[0332] In a possible implementation, the above-mentioned encoding unit 1501 is further configured to: encode the bitstream length information into the above-mentioned bitstream, and the bitstream length information is used to indicate the length of the above-mentioned bitstream.

[0333] Next, a decoding device for performing the above decoding method will be described in conjunction with Figure 16 introduce the decoding device for performing the above decoding method.

[0334] It can be understood that, in order to implement the above functions, the decoding device includes corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0335] The embodiments of the present application can divide the function modules of the decoding device according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0336] In the case of dividing each function module corresponding to each function, Figure 16FIG. 0 shows a possible schematic composition of the decoding device involved in the above embodiments. The device can be an electronic device, or a module applied to an electronic device (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of the electronic device. As Figure 16 shown, the decoding device 1600 may include: a decoding unit 1601 and a reconstruction unit 1602.

[0337] The above decoding unit 1601 is used to perform entropy decoding on the bitstream to obtain the first set of second feature maps of the input image, where N is a positive integer;

[0338] The above reconstruction unit 1602 is used to perform AI decoding on the above first set of second feature maps to obtain the first reconstructed image of the input image.

[0339] In a possible implementation manner, the above decoding unit 1601 is further used to perform entropy decoding on the bitstream to obtain the second set of second feature maps of the input image.

[0340] In a possible implementation manner, the above reconstruction unit 1602 is further used to perform AI decoding on the above first set of second feature maps and the above second set of second feature maps to obtain the second reconstructed image of the input image.

[0341] In a possible implementation manner, the above decoding unit 1601 is further used to perform entropy decoding on the bitstream to obtain the third set of second feature maps of the input image.

[0342] In a possible implementation manner, the above reconstruction unit 1602 is further used to perform AI decoding on the above first set of second feature maps, the above second set of second feature maps, and the above third set of second feature maps to obtain the third reconstructed image of the input image.

[0343] In a possible implementation manner, the image quality of the above second reconstructed image is better than the image quality of the above first reconstructed image, and the image quality of the above third reconstructed image is better than the image quality of the above second reconstructed image. The above image quality is characterized by any one of the following variables: PSNR, MS-SSIM, or LPIPS.

[0344] In a possible implementation manner, the above decoding unit 1601 is further used to decode the above bitstream to obtain channel number information, and the above channel number information is used to indicate the number of channels of at least one set of second feature maps obtained by decoding.

[0345] In a possible implementation manner, the above reconstruction unit 1602 is further used to perform AI decoding on the above first set of second feature maps according to the above channel number information to obtain the first reconstructed image.

[0346] In a possible implementation, the above decoding unit 1601 is further configured to decode the above bitstream to obtain first position information, where the first position information is used to indicate the position of at least one set of second feature maps obtained by decoding in the first feature map of the above input image.

[0347] In a possible implementation, the above reconstruction unit 1602 is further configured to perform AI decoding on the above first set of second feature maps according to the above first position information to obtain a first reconstructed image.

[0348] In a possible implementation, the above decoding unit 1601 is further configured to decode the above bitstream to obtain second position information, where the second position information is used to indicate the position of at least one set of second feature maps obtained by decoding in the above bitstream;

[0349] In a possible implementation, the above decoding unit 1601 is further configured to determine the bit data corresponding to the above first set of second feature maps in the above bitstream according to the above second position information, and perform entropy decoding on the corresponding bit data to obtain the above first set of second feature maps.

[0350] In a possible implementation, the above decoding unit 1601 is further configured to decode the above bitstream to obtain bitstream length information of the above bitstream, where the bitstream length information is used to indicate the length of the above bitstream.

[0351] In a possible implementation, the above reconstruction unit 1602 is further configured to perform AI decoding on the above first set of second feature maps according to the above bitstream length information to obtain a first reconstructed image.

[0352] An embodiment of this application further provides a chip, and the chip may be a chip of the above encoding device or decoding device. Figure 17 A schematic structural diagram of a chip 1700 is shown. The chip 1700 includes one or more processors 1701 and an interface circuit 1702. Optionally, the above chip 1700 may further include a bus 1703.

[0353] The processor 1701 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above encoding and decoding method may be completed by the integrated logic circuit in the hardware of the processor 1701 or instructions in software form.

[0354] Optionally, the above-mentioned processor 1701 may be a general-purpose processor, a digital signal processing (DSP) processor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0355] The interface circuit 1702 can be used for sending or receiving data, instructions, or information. The processor 1701 can process the data, instructions, or other information received by the interface circuit 1702 and send the processed information through the interface circuit 1702.

[0356] Optionally, the chip further includes a memory. The memory may include a read-only memory and a random access memory, and provide operation instructions and data to the processor. A part of the memory may also include a non-volatile random access memory (NVRAM).

[0357] Optionally, the memory stores executable software modules or data structures. The processor can execute corresponding operations by calling the operation instructions stored in the memory (the operation instructions may be stored in the operating system).

[0358] Optionally, the chip can be used in the encoding device or decoding device involved in the embodiments of the present application. Optionally, the interface circuit 1702 can be used to output the execution result of the processor 1701. For the encoding and decoding methods provided by one or more embodiments of the embodiments of the present application, reference can be made to the foregoing respective embodiments, which will not be elaborated here.

[0359] It should be noted that the respective functions corresponding to the processor 1701 and the interface circuit 1702 can be implemented through hardware design, can also be implemented through software design, or can be implemented through a combination of software and hardware, and no limitation is made here.

[0360] Figure 18 It is a schematic structural diagram of an electronic device provided by the embodiments of the present application. The electronic device can be an encoding device or a decoding device, a chip or a functional module in the encoding device, or a chip or a functional module in the decoding device. As Figure 18 shown, the electronic device 1800 includes a processor 1801, a transceiver 1802, and a communication line 1803.

[0361] Among them, the processor 1801 is used to execute any step of the encoding and decoding method provided in the embodiments of the present application. During the execution of any step of the encoding and decoding method provided in the embodiments of the present application, the transceiver 1802 and the communication line 1803 can be selectively called to complete the corresponding operations.

[0362] Further, the electronic device 1800 may further include a memory 1804. Among them, the processor 1801, the memory 1804 and the transceiver 1802 can be connected through the communication line 1803.

[0363] Among them, the processor 1801 is a processor, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 1801 may also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.

[0364] The transceiver 1802 is used to communicate with other devices or other communication networks. The other communication networks may be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The transceiver 1802 may be a module, a circuit, a transceiver, or any device capable of implementing communication.

[0365] The transceiver 1802 is mainly used for sending and receiving commands, information, etc. It may include a transmitter and a receiver, which are respectively used for sending and receiving commands, information, etc. Operations other than sending and receiving commands, information, etc. are implemented by the processor.

[0366] The communication line 1803 is used to transmit information between the components included in the electronic device 1800.

[0367] In one design, the processor can be regarded as a logic circuit, and the transceiver can be regarded as an interface circuit.

[0368] The memory 1804 is used to store instructions. Among them, the instructions may be computer programs.

[0369] Among them, the memory 1804 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). The memory 1804 can also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, etc. It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0370] It should be noted that the memory 1804 can exist independently of the processor 1801 or can be integrated with the processor 1801. The memory 1804 can be used to store instructions, program codes, or some data, etc. The memory 1804 can be located inside the electronic device 1800 or outside the electronic device 1800, without limitation. The processor 1801 is configured to execute the instructions stored in the memory 1804 to implement the method provided in the foregoing embodiments of the present application.

[0371] In one example, the processor 1801 can include one or more processors, such as Figure 18 processor 0 and processor 1 in

[0372] As an alternative implementation, the electronic device 1800 includes multiple processors. For example, in addition to the processor 1801 in Figure 18 , the processor 1809 may also be included.

[0373] As an alternative implementation, the electronic device 1800 further includes an output device 1805 and an input device 1806. Exemplarily, the input device 1806 is a device such as a keyboard, a mouse, a microphone, or a joystick, and the output device 1805 is a device such as a display screen or a speaker.

[0374] It should be noted that the electronic device 1800 may be a chip system or a device with a Figure 18 similar structure. Among them, the chip system may be composed of chips or may include chips and other discrete devices. Actions, terms, etc. involved among the embodiments of the present application may refer to each other without limitation. In the embodiments of the present application, the message names or parameter names in the messages exchanged between various devices are only examples, and other names may also be adopted in specific implementations without limitation. In addition, Figure 18 the component structure shown in Figure 18 does not constitute a limitation on the electronic device 1800. In addition to the Figure 18 shown components, the electronic device 1800 may include more or fewer components than

[0375] shown, or combine some components, or have different component arrangements.

[0376] Figure 19Schematic diagram of another electronic device provided by an embodiment of the present application. The electronic device may be an encoding device or a decoding device, a chip or a functional module in the encoding device, or a chip or a functional module in the decoding device. For ease of explanation, Figure 19 only the main components of the electronic device are shown, including a processor 1901, a memory 1902, a control circuit 1903, and an input / output device 1904. The processor 1901 is mainly used for processing communication protocols and communication data, executing software programs, and processing data of software programs. The memory 1902 is mainly used for storing software programs and data. The control circuit 1903 is mainly used for power supply and transmission of various electrical signals. The input / output device 1904 is mainly used for receiving data input by the user and outputting data to the user.

[0377] When the electronic device is the processor 1901, the control circuit 1903 may be a main board, the memory 1902 includes media with storage functions such as a hard disk, RAM, and ROM. The processor 1901 may include a baseband processor 1901 and a central processor. The baseband processor is mainly used for processing communication protocols and communication data, and the central processor is mainly used for controlling the entire electronic device, executing software programs, and processing data of software programs. The input / output device 1904 includes a display screen, a keyboard, a mouse, etc.; the control circuit 1903 may further include or be connected to a transceiver circuit or a transceiver, such as: a network interface, etc., for sending or receiving data or signals, such as for data transmission and communication with other devices. Further, an antenna may also be included for wireless signal transceiver for data / signal transmission with other devices.

[0378] An embodiment of the present application also provides an encoding device, the device includes: at least one processor, when the at least one processor executes program code or instructions, implementing the above-mentioned related method steps to implement the encoding method in the above embodiment.

[0379] Optionally, the device may further include at least one memory for storing the program code or instructions.

[0380] An embodiment of the present application also provides a decoding device, the device includes: at least one processor, when the at least one processor executes program code or instructions, implementing the above-mentioned related method steps to implement the decoding method in the above embodiment.

[0381] Optionally, the device may further include at least one memory for storing the program code or instructions.

[0382] An embodiment of the present application further provides a computer storage medium, in which computer instructions are stored. When the computer instructions run on a communication device, the communication device is caused to execute the above-related method steps to implement the encoding and decoding method in the above embodiment.

[0383] An embodiment of the present application further provides a computer program product. When the computer program product runs on a computer, the computer is caused to execute the above-related steps to implement the encoding and decoding method in the above embodiment.

[0384] An embodiment of the present application further provides an encoding device, which may specifically be a chip, an integrated circuit, a component, or a module. Specifically, the device may include a processor connected to a memory for storing instructions, or the device includes at least one processor for obtaining instructions from an external memory. When the device runs, the processor can execute the instructions to cause the chip to execute the encoding method in each of the above method embodiments.

[0385] An embodiment of the present application further provides a decoding device, which may specifically be a chip, an integrated circuit, a component, or a module. Specifically, the device may include a processor connected to a memory for storing instructions, or the device includes at least one processor for obtaining instructions from an external memory. When the device runs, the processor can execute the instructions to cause the chip to execute the decoding method in each of the above method embodiments.

[0386] An embodiment of the present application further provides a method for storing a bitstream, the method including: obtaining and storing the bitstream obtained by the above encoding and decoding method.

[0387] An embodiment of the present application further provides a bitstream storage device, which is used to obtain and store the bitstream obtained by the above encoding and decoding method.

[0388] An embodiment of the present application further provides a method for transmitting a bitstream, the method including: obtaining and transmitting the bitstream obtained by the above encoding and decoding method.

[0389] An embodiment of the present application further provides a bitstream transmission device, which is used to obtain and transmit the bitstream obtained by the above encoding and decoding method.

[0390] An embodiment of the present application further provides a computer-readable storage medium, on which the bitstream obtained by the above encoding and decoding method is stored.

[0391] Please refer to Figure 20 , Figure 20 which schematically illustrates the general concept of processing through a neural network such as a convolutional neural network (CNN). A convolutional neural network consists of an input layer, an output layer, and multiple hidden layers. The input layer provides inputs (such asFigure 20 a layer that processes a part of the input image as shown). The hidden layers of a CNN typically consist of a series of convolutional layers that are convolved with multiplications or other dot products. The result of the layer is one or more feature maps (represented by empty solid rectangles), sometimes also called channels. Resampling (such as subsampling) may be involved in some or all of the layers. As a result, the feature maps may become smaller, as Figure 20 shown. Note that convolution with a stride can also reduce the size of the input feature map (resampling). The activation function in a CNN is typically a ReLU (Rectified Linear Unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, and normalization layers, which are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. Although these layers are popularly called convolutions, this is just a convention. Mathematically, it is technically a sliding dot product or cross-correlation. This has important implications for the indexing in the matrix because it affects how the weights are determined at a specific index point.

[0392] When programming a CNN for processing images, as Figure 20 shown, the input is a tensor of shape (number of images) x (image width) x (image height) x (image depth). It should be noted that the image depth can be composed of the channels of the image. After passing through the convolutional layer, the image is abstracted into a feature map with the shape (number of images) x (feature map width) x (feature map height) x (feature map channels). The convolutional layer in a neural network should have the following properties. A convolutional kernel (hyperparameter) defined by width and height. The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) should be equal to the number of channels (depth) of the input feature map.

[0393] Machine video coding (VCM) is another popular direction in computer science today. The main idea behind this approach is to transmit an encoded representation of image or video information for further processing by computer vision (CV) algorithms such as object segmentation, detection, and recognition. Compared with traditional image and video coding that aims at human perception, the quality feature is the performance of computer vision tasks, such as object detection accuracy, rather than the reconstruction quality. As Figure 21 shown.

[0394] Machine video coding, also known as collaborative intelligence, is a relatively new paradigm for the efficient deployment of deep neural networks in a mobile cloud infrastructure. By partitioning the network between the mobile side 2110 and the cloud side 2190 (e.g., a cloud server), the computational workload can be distributed such that the total energy and / or latency of the system is minimized. Generally, collaborative intelligence is a paradigm in which the processing of a neural network is distributed between two or more different computing nodes; e.g., a device, but in general, any functionally defined node. Here, the term "node" does not refer to the above-mentioned neural network nodes. Instead, the (computing) nodes here refer to (physically or at least logically) independent devices / modules that implement parts of the neural network. Such devices can be different servers, different end-user devices, a mixture of servers and / or user devices and / or clouds and / or processors, etc. In other words, the computing nodes can be considered as nodes belonging to the same neural network and communicate with each other to transmit encoded data within / for the neural network. For example, in order to be able to perform complex computations, one or more layers can be executed on a first device (such as a device on the mobile side 2110), and one or more layers can be executed in another device (such as a cloud server on the cloud side 2190). However, the distribution can also be finer, and a single layer can be executed on multiple devices. In the present disclosure, the term "plurality" means two or more. In some existing solutions, a part of the neural network function is executed in a device (a user device or an edge device, etc.) or multiple such devices, and then the output (feature map) is passed to the cloud. The cloud is a collection of processing or computing systems located outside the device that is operating a part of the neural network. The concept of collaborative intelligence has also been extended to model training. In this case, data flows bidirectionally: from the cloud to the mobile device during the backpropagation of training, and from the mobile device to the cloud during the forward pass and inference of training (as Figure 21 shown).

[0395] Some work has proposed semantic image compression by encoding deep features and then reconstructing the input image from them. Compression based on uniform quantization followed by context-based adaptive arithmetic coding (CABAC) from H.264 is shown. In some scenarios, it may be more efficient to send the output of the hidden layer (deep feature map) from the mobile part 2110 to the cloud 2190 rather than sending compressed natural image data to the cloud and performing object detection using the reconstructed image. Therefore, it may be advantageous to compress the data (features) generated by the mobile side 2110, and the mobile side 2110 can include a quantization layer 2120 for this purpose. Accordingly, the cloud side 2190 can include an inverse quantization layer 2160. The efficient compression of feature maps is beneficial for image and video compression and reconstruction for human perception and machine vision. Entropy coding methods, such as arithmetic coding, are a popular method for compressing deep features (i.e., feature maps).

[0396] Please refer to Figure 22 , Figure 22 which shows a bitstream structure provided by an embodiment of the present application. As Figure 22 shown, the bitstream includes: Start of Image, File Header, Entropy Encoded Data, and End of Image.

[0397] In a possible implementation manner, the above channel number information, the above first position information, the above second position information, or the bitstream length information may be stored in the file header.

[0398] In a possible implementation manner, the above channel number information, the above first position information, the above second position information, or the bitstream length information may be stored in the entropy encoded data.

[0399] Wherein, the device, computer storage medium, computer program product, or chip provided by this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0400] It should be understood that in various embodiments of the embodiments of the present application, the magnitude of the sequence numbers of the above processes does not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0401] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0402] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0403] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0404] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0405] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0406] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0407] The above is only the specific implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the embodiments of the present application, and all of them should be covered by the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.

Claims

1. A coding method, characterized in that, Including: Performing artificial intelligence (AI) encoding on an input image to obtain a first feature map of the input image; Determining N groups of second feature maps of the input image according to the first feature map, where N is a positive integer; Entropy encoding the first group of second feature maps among the N groups of second feature maps into a bitstream.

2. The method according to claim 1, wherein The method further includes: Entropy encoding the second group of second feature maps among the N groups of second feature maps into the bitstream.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Entropy encoding the third group of second feature maps among the N groups of second feature maps into the bitstream.

4. The method according to any one of claims 1 to 3, characterized in that The determining N groups of second features of the input image according to the first feature map includes: Performing channel separation on the first feature map to obtain M second feature maps of the input image, where M is a positive integer; Grouping the M second feature maps according to the channel information of the second feature maps of the input image to obtain the N groups of second feature maps.

5. The method according to claim 4, characterized in that, The channel information includes channel entropy and / or channel variance, the channel entropy is the entropy of all elements in the channel, and the channel variance is the variance of all elements in the channel.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Entropy encoding at least one group of second feature maps among the N groups of second feature maps into the bitstream in descending order of channel entropy or channel variance.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Encoding channel number information into the bitstream, where the channel number information is used to indicate the number of channels of at least one group of second feature maps encoded into the bitstream.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Encoding first position information into the bitstream, where the first position information is used to indicate the position of at least one group of second feature maps encoded into the bitstream in the first feature map.

9. The method according to any one of claims 1 to 8, characterized in that The method further includes: Encoding second position information into the bitstream, where the second position information is used to indicate the position of at least one group of second feature maps encoded into the bitstream in the bitstream.

10. A decoding method, characterized in that, Including: Performing entropy decoding on the bitstream to obtain the first group of second feature maps of the input image, where the first group of second feature maps is a part of the feature maps corresponding to the input image; Performing AI decoding on the first group of second feature maps to obtain a first reconstructed image of the input image.

11. The method according to claim 10, characterized in that, The method further includes: Performing entropy decoding on the bitstream to obtain the second group of second feature maps, where the second group of second feature maps is a part of the feature maps corresponding to the input image; Performing AI decoding on the first group of second feature maps and the second group of second feature maps to obtain a second reconstructed image of the input image.

12. The method according to claim 11, wherein The method further includes: Performing entropy decoding on the bitstream to obtain the third group of second feature maps, where the third group of second feature maps is a part of the feature maps corresponding to the input image; Performing AI decoding on the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps to obtain a third reconstructed image of the input image.

13. The method according to claim 12, wherein The image quality of the second reconstructed image is better than that of the first reconstructed image, and the image quality of the third reconstructed image is better than that of the second reconstructed image. The image quality is characterized by any one of the following variables: peak signal-to-noise ratio (PSNR), multi-scale structural similarity index (MS-SSIM), or learned perceptual image patch similarity (LPIPS).

14. The method according to any one of claims 10 to 13, characterized in that, The method further includes: Decoding the bitstream to obtain channel number information, where the channel number information is used to indicate the number of channels of at least one group of second feature maps obtained by decoding; Performing AI decoding on the first group of second feature maps according to the channel number information to obtain a first reconstructed image.

15. The method according to any one of claims 10 to 14, characterized in that The method further includes: Decoding the bitstream to obtain first position information, where the first position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the first feature map of the input image; Performing AI decoding on the first group of second feature maps according to the first position information to obtain a first reconstructed image.

16. The method according to any one of claims 11 to 15, characterized in that The method further includes: Decoding the bitstream to obtain second position information, where the second position information is used to indicate the position of at least one group of second feature maps obtained by decoding in the bitstream; Determining the bit data corresponding to the first group of second feature maps in the bitstream according to the second position information, and performing entropy decoding on the corresponding bit data to obtain the first group of second feature maps.

17. An encoding device, characterized in that, It includes: An encoding unit and a determining unit; The encoding unit is used to perform AI encoding on the input image to obtain the first feature map of the input image; The determining unit is used to determine N groups of second feature maps of the input image according to the first feature map, where N is a positive integer; The encoding unit is further used to entropy encode the first group of second feature maps among the N groups of second feature maps into the bitstream.

18. The device according to claim 17, characterized in that, The encoding unit is further used for: Entropy encoding the second group of second feature maps among the N groups of second feature maps into the bitstream.

19. The device according to claim 17 or 18, characterized in that, The encoding unit is further used for: Entropy encoding the third group of second feature maps among the N groups of second feature maps into the bitstream.

20. The device according to any one of claims 17 to 19, characterized in that, The determining unit is specifically used for: Performing channel separation on the first feature map to obtain M second feature maps of the input image, where M is a positive integer; Grouping the M second feature maps according to the channel information of the second feature maps of the input image to obtain the N groups of second feature maps.

21. The device according to any one of claims 17 to 20, characterized in that, The encoding unit is further used for: Entropy encoding at least one group of second feature maps among the N groups of second feature maps into the bitstream in the order of decreasing channel entropy or channel variance.

22. A decoding device, characterized in that, It includes: A decoding unit and a reconstruction unit; The decoding unit is used to perform entropy decoding on the bitstream to obtain the first group of second feature maps of the input image, where the first group of second feature maps is a part of the feature map corresponding to the input image; The reconstruction unit is used to perform AI decoding on the first group of second feature maps to obtain the first reconstructed image of the input image.

23. The device according to claim 22, wherein The decoding unit is further used to perform entropy decoding on the bitstream to obtain the second group of second feature maps, where the second group of second feature maps is a part of the feature map corresponding to the input image; The reconstruction unit is further used to perform AI decoding on the first group of second feature maps and the second group of second feature maps to obtain the second reconstructed image of the input image.

24. The device according to claim 23, characterized in that, The decoding unit is further used to perform entropy decoding on the bitstream to obtain the third group of second feature maps, where the third group of second feature maps is a part of the feature map corresponding to the input image; The reconstruction unit performs AI decoding on the first group of second feature maps, the second group of second feature maps, and the third group of second feature maps to obtain a third reconstructed image of the input image.

25. The device according to claim 24, characterized in that, The image quality of the second reconstructed image is better than that of the first reconstructed image, and the image quality of the third reconstructed image is better than that of the second reconstructed image. The image quality is characterized by any one of the following variables: PSNR, MS-SSIM, or LPIPS.

26. An encoding device, characterized in that, Comprising at least one processor and a memory, the at least one processor executes programs or instructions stored in the memory to enable the encoding device to implement the method according to any one of claims 1 to 9 above.

27. A decoding device, characterized in that, Comprising at least one processor and a memory, the at least one processor executes programs or instructions stored in the memory to enable the decoding device to implement the method according to any one of claims 10 to 16 above.

28. A method for storing a bitstream, characterized in that, Comprising: Obtaining and storing a bitstream obtained by the method according to any one of claims 1 to 16.

29. A bitstream storage device, characterized in that, The device is configured to obtain and store a bitstream obtained by the method according to any one of claims 1 to 16.

30. A bitstream transmission method, characterized in that, Comprising: Obtaining and transmitting a bitstream obtained by the method according to any one of claims 1 to 16.

31. A bitstream transmission device, characterized in that, The device is configured to obtain and transmit a bitstream obtained by the method according to any one of claims 1 to 16.

32. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a bitstream obtained by the method according to any one of claims 1 to 16.

33. A computer-readable storage medium, characterized in that, For storing a computer program, when the computer program runs on a computer or a processor, enabling the computer or the processor to implement the method according to any one of claims 1 to 16 above.

34. A computer program product, characterized in that, The computer program product contains instructions, when the instructions run on a computer or a processor, enabling the computer or the processor to implement the method according to any one of claims 1 to 16 above.

Citation Information

Cited By

  • Encoding method and apparatus, and decoding method and apparatus

    WO2025152469A1