Coding and decoding method and device

By performing transformation operations and entropy encoding of the local areas of AI image encoding, the problem of poor image quality at extremely high code rates is solved, and higher image quality and compression efficiency are achieved.

CN120343261APending Publication Date: 2025-07-18HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410166019.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2024-02-05
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In extremely high code rate or lossless near vision, the image quality of AI image encoding is poor, and due to network capacity and generalization, there is greater local distortion.

Method used

The local areas of AI image encoding are enhanced by traditional coding schemes. By dividing the input image into multiple image blocks for transformation operations, and encoding the transformation coefficient and position information into the code stream, the image quality is improved by combining entropy coding technology.

Benefits of technology

The overall image quality of AI image encoding is improved, and the local distortion problem caused by the capacity and generalization of AI encoding networks is compensated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343261A_ABST
    Figure CN120343261A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a coding method, relates to the technical field of media, and is used for improving the image quality of AI image coding. The method comprises the following steps: carrying out AI coding on an input image to obtain a feature map; entropy coding is carried out on the feature map to obtain a code stream; performing transformation operation on a target area of the input image to obtain a transformation coefficient; encoding the transformation coefficient to the code stream; and encoding the position information to the code stream. Wherein the position information is used for indicating the position of the target area in the input image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application titled "Coding Method" with the application number 202410077119.2 filed with the Chinese Patent Office on January 18, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to the field of media technology, and in particular, to coding and decoding methods and devices. Background Art

[0003] Currently, significant breakthroughs have been made in Artificial Intelligence (AI) image coding. AI image coding significantly improves the compression efficiency compared to common image coding standards under the same subjective quality. AI image coding is widely applied in various fields. For example, it can be applied to cloud storage, visual surveillance, autonomous driving vehicles and devices, image acquisition, storage and management, real-time monitoring of visual data, and media distribution.

[0004] In the case of lossy coding at medium and low bitrates, under the same quality, the compression ratio of the AI image coding scheme is significantly better than that of traditional image coding schemes; however, in the case of extremely high bitrates or near visually lossless conditions, due to factors such as network capacity and network generalization, the image quality of AI image coding is poor. Summary of the Invention

[0005] Embodiments of this application provide a coding method for improving the image quality of AI image coding. To achieve the above objective, the embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, embodiments of this application provide a coding method, which includes: performing AI coding on an input image to obtain a feature map; performing entropy coding on the feature map to obtain a bitstream; performing a transformation operation on a target region of the input image to obtain transformation coefficients; encoding the transformation coefficients into the bitstream; and encoding position information into the bitstream. Wherein, the position information is used to indicate the position of the target region in the input image.

[0007] The method provided by the embodiments of this application additionally uses a traditional coding scheme for quality enhancement on local regions of the image, making up for the deficiency that the local distortion of the image quality caused by the network capacity and generalization of AI coding is relatively large. Compared with only performing AI coding on the input image, additionally using a traditional coding scheme for quality enhancement on local regions of the image can improve the overall image quality of AI image coding.

[0008] In a possible implementation, the target region can be divided into multiple image blocks; performing the transformation operation on the multiple image blocks to obtain transformation coefficients.

[0009] The method provided by the embodiments of the present application additionally divides the local region of the input image into multiple image blocks and performs a transformation operation on the local region of the input image, thereby improving the image quality of the local region of the input image. Compared with only performing AI encoding on the input image, additionally performing a transformation operation on the local region of the input image can improve the image quality of AI image encoding.

[0010] In a possible implementation manner, the above position information can be entropy encoded into the above bitstream.

[0011] It can be seen that the position information indicating the position of the target region in the above input image can be encoded into the bitstream by means of entropy encoding.

[0012] In a possible implementation manner, the above transform coefficients can be entropy encoded into the above bitstream.

[0013] It can be seen that the transform coefficients indicating the coefficients used for the transformation operation on the above target region can be encoded into the bitstream by means of entropy encoding.

[0014] In a possible implementation manner, the above transform coefficients are quantized; the quantized transform coefficients are encoded into the above bitstream. Quantization can reduce the bit depth related to the transform coefficients.

[0015] In a possible implementation manner, the transform parameters can be encoded (entropy encoded) into the above bitstream. Among them, the transform parameters are the transform parameters used for the above transformation operation.

[0016] In a possible implementation manner, the transformation operation can include at least one of discrete cosine transform (DCT), discrete Fourier transform (DFT), or discrete wavelet transform (DWT).

[0017] The method provided by the embodiments of the present application additionally performs transformation operations such as DCT, DFT, or DWT on the local region of the input image, thereby improving the image quality of the local region of the input image. Compared with only performing AI encoding on the input image, additionally performing a transformation operation on the local region of the input image can improve the image quality of AI image encoding.

[0018] In a possible implementation manner, the above position information is mask map information, and the mask map information includes a first region, and the first region is used to indicate the target region. For example, the above mask map can be composed of 0 and 1, and the region where 1 is located is the target region.

[0019] It can be seen that the position information indicating the position of the target region in the above input image can be saved in the form of a mask map.

[0020] In a possible implementation, the above transformation coefficient can be a residual.

[0021] In a second aspect, an embodiment of the present application provides a decoding method, which includes: performing entropy decoding on a bitstream to obtain a feature map; performing AI decoding on the above feature map to obtain a reconstructed image; performing entropy decoding on the above bitstream to obtain transformation coefficients; performing an inverse transformation operation on the above transformation coefficients to obtain the reconstructed pixel values of the target region; performing entropy decoding on the above bitstream to obtain position information; and determining a fused image according to the above position information, the reconstructed pixel values of the above target region, and the above reconstructed image. Wherein, the above position information is used to indicate the position of the target region of the above reconstructed image.

[0022] The method provided by the embodiment of the present application additionally uses a traditional decoding scheme to enhance the quality of a local region of an image, making up for the deficiency that the capacity and generalization of the AI decoding network cause a large local distortion in the image quality. Compared with only performing AI encoding on a bitstream to obtain a reconstructed image, additionally using a traditional decoding scheme to enhance the quality of a local region of an image can improve the overall image quality of AI image decoding.

[0023] In a possible implementation, the target region in the above reconstructed image can be determined according to the above position information; and the pixel values of the target region in the above reconstructed image are updated according to the above reconstructed pixel values to obtain the above fused image.

[0024] In a possible implementation, entropy decoding is performed on the above bitstream to obtain quantized transformation coefficients; and the above quantized transformation coefficients are inverse quantized to obtain the above transformation coefficients.

[0025] In a possible implementation, the above transformation coefficients can be inverse quantized.

[0026] In a third aspect, an embodiment of the present application provides an encoding device, which includes: an encoding unit and a transformation unit. The above encoding unit is used to perform AI encoding on an input image to obtain a feature map. The above encoding unit is further used to perform entropy encoding on the above feature map to obtain a bitstream. The above transformation unit is used to perform a transformation operation on the target region of the above input image to obtain transformation coefficients. The above encoding unit is further used to encode the above transformation coefficients into the above bitstream. The above encoding unit is further used to encode position information into the above bitstream, and the above position information is used to indicate the position of the above target region in the above input image.

[0027] In a possible implementation, the above transformation unit is specifically configured to: divide the above target region into multiple image blocks; perform the above transformation operation on the above multiple image blocks to obtain transformation coefficients.

[0028] In a possible implementation, quantize the above transformation coefficients; encode the quantized transformation coefficients into the above bitstream.

[0029] In a fourth aspect, an embodiment of the present application provides a decoding device, which includes: a decoding unit, a transformation unit, and a fusion unit. The above decoding unit is configured to perform entropy decoding on the bitstream to obtain a feature map. The above decoding unit is further configured to perform AI decoding on the above feature map to obtain a reconstructed image. The above decoding unit is further configured to perform entropy decoding on the above bitstream to obtain transformation coefficients. The above transformation unit is configured to perform an inverse transformation operation on the above transformation coefficients to obtain the reconstructed pixel values of the target region. The above decoding unit is further configured to perform entropy decoding on the above bitstream to obtain position information, and the above position information is used to indicate the position of the above target region. The above fusion unit is configured to determine a fused image according to the above position information, the reconstructed pixel values of the above target region, and the above reconstructed image.

[0030] In a possible implementation, the above fusion unit is specifically configured to: determine the target region in the above reconstructed image according to the above position information; update the pixel values of the target region in the above reconstructed image according to the above reconstructed pixel values to obtain the above fused image.

[0031] In a possible implementation, the above decoding unit is specifically configured to: perform entropy decoding on the above bitstream to obtain quantized transformation coefficients; perform inverse quantization on the above quantized transformation coefficients to obtain the above transformation coefficients.

[0032] In a fifth aspect, an embodiment of the present application further provides an encoding device, which includes: at least one processor, and when the at least one processor executes program code or instructions, the method described in the above first aspect or any of its possible implementations is implemented.

[0033] Optionally, the device may further include at least one memory, and the at least one memory is used to store the program code or instructions.

[0034] In a sixth aspect, an embodiment of the present application further provides a decoding device, which includes: at least one processor, and when the at least one processor executes program code or instructions, the method described in the above second aspect or any of its possible implementations is implemented.

[0035] In a seventh aspect, an embodiment of the present application further provides a method for storing a bitstream, which includes: acquiring and storing the bitstream obtained by the method described in the above first aspect or any of its possible implementations.

[0036] In an eighth aspect, an embodiment of the present application further provides a bitstream storage device, which is used to acquire and store the bitstream obtained by the method described in the first aspect or any possible implementation manner thereof.

[0037] In a ninth aspect, an embodiment of the present application further provides a bitstream transmission method, which includes: acquiring and transmitting the bitstream obtained by the method described in the first aspect or any possible implementation manner thereof.

[0038] In a tenth aspect, an embodiment of the present application further provides a bitstream transmission device, which is used to acquire and transmit the bitstream obtained by the method described in the first aspect or any possible implementation manner thereof.

[0039] In an eleventh aspect, an embodiment of the present application further provides a computer-readable storage medium, on which the bitstream obtained by the method described in the first aspect or any possible implementation manner thereof is stored

[0040] In a twelfth aspect, an embodiment of the present application further provides a chip, including: an input interface, an output interface, and at least one processor. Optionally, the chip further includes a memory. The at least one processor is used to execute the code in the memory, and when the at least one processor executes the code, the chip implements the method described in the first aspect or any possible implementation manner thereof.

[0041] Optionally, the above chip may also be an integrated circuit.

[0042] In a thirteenth aspect, an embodiment of the present application further provides a computer-readable storage medium, which is used to store a computer program, and the computer program includes a method for implementing the method described in the first aspect or any possible implementation manner thereof.

[0043] In a fourteenth aspect, an embodiment of the present application further provides a computer program product containing instructions, which, when running on a computer, enables the computer to implement the method described in the first aspect or any possible implementation manner thereof.

[0044] The encoding and decoding device, computer storage medium, computer program product, and chip provided in this embodiment are all used to execute the encoding and decoding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the encoding and decoding method provided above, and will not be elaborated here. Description of the Drawings

[0045] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0046] Figure 1a An exemplary block diagram of a decoding system provided by an embodiment of the present application;

[0047] Figure 1b An exemplary block diagram of a video decoding system provided by an embodiment of the present application;

[0048] Figure 2 An exemplary block diagram of a video encoder provided by an embodiment of the present application;

[0049] Figure 3 An exemplary block diagram of a video decoder provided by an embodiment of the present application;

[0050] Figure 4 An exemplary block diagram of a video decoding device provided by an embodiment of the present application;

[0051] Figure 5 An exemplary block diagram of a device provided by an embodiment of the present application;

[0052] Figure 6 A schematic diagram of an image compression method based on a neural network provided by an embodiment of the present application;

[0053] Figure 7 A schematic diagram of an end-to-end image coding framework provided by an embodiment of the present application;

[0054] Figure 8 A schematic diagram of the structure of a neural network provided by an embodiment of the present application;

[0055] Figure 9 A schematic diagram of the structure of a video communication system provided by an embodiment of the present application;

[0056] Figure 10 A schematic diagram of the flow of an encoding method provided by an embodiment of the present application;

[0057] Figure 11 A schematic diagram of the flow of a decoding method provided by an embodiment of the present application;

[0058] Figure 12 A schematic diagram of the structure of an encoding device provided by an embodiment of the present application;

[0059] Figure 13Schematic diagram of a decoding device provided by an embodiment of the present application;

[0060] Figure 14 Schematic diagram of a chip provided by an embodiment of the present application;

[0061] Figure 15 Schematic diagram of an electronic device provided by an embodiment of the present application;

[0062] Figure 16 Schematic diagram of another electronic device provided by an embodiment of the present application;

[0063] Figure 17 Schematic diagram of a neural network provided by an embodiment of the present application;

[0064] Figure 18 Schematic diagram of a machine video coding system provided by an embodiment of the present application;

[0065] Figure 19 Schematic diagram of a bitstream provided by an embodiment of the present application;

[0066] Figure 20 Schematic diagram of hierarchical coding provided by an embodiment of the present application;

[0067] Figure 21 Schematic diagram of end-to-end image coding based on a hyperprior structure provided by an embodiment of the present application. Detailed implementation manners

[0068] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the embodiments of the present application.

[0069] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.

[0070] The terms "first" and "second" in the specification and drawings of the embodiments of the present application are used to distinguish different objects or different processes for the same object, rather than to describe the specific order of the objects.

[0071] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the embodiments of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include other steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.

[0072] It should be noted that in the description of the embodiments of the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.

[0073] First, the terms related to the embodiments of the present application are explained.

[0074] Joint Photographic Experts Group (JPEG) Artificial Intelligence (AI) is a learning-based image coding standard that provides a single-stream, compact representation in the compression domain, significantly improving compression efficiency compared to common image coding standards at the same subjective quality. JPEG AI is widely used in various fields. For example, JPEG AI can be applied to cloud storage, visual surveillance, autonomous vehicles and devices, image acquisition, storage and management, real-time monitoring of visual data, and media distribution.

[0075] Convolutional Neural Network (CNN): A neural network that includes convolutional layers, and may also include modules such as activation layers (such as ReLU, PReLU, etc.), pooling layers, batch normalization layers (BN layers), and fully connected layers. Typical convolutional neural networks include LeNet, AlexNet, VGGNet, ResNet, etc. A basic CNN can be composed of a backbone network and a head network; a complex CNN is composed of a backbone, a neck, and a head network.

[0076] Feature Map: Three-dimensional data output by convolutional layers, activation layers, pooling layers, batch normalization layers, etc. in a convolutional neural network. The three dimensions are respectively called width (Width), height (Height), and channel (Channel).

[0077] Backbone network: The first part of a convolutional neural network, which functions to extract feature maps of multiple scales from the input image. It is usually composed of convolutional layers, pooling layers, activation layers, etc., and does not contain fully connected layers. Generally, the feature maps output by the layers closer to the input image in the backbone network have a larger resolution (width and height) but fewer channels. Typical backbone networks include VGG-16, ResNet-50, ResNeXt-101, etc.

[0078] Head network: The last part of a convolutional neural network, which functions to process the feature maps to obtain the prediction results output by the neural network. Common head networks include fully connected layers, softmax modules, etc.

[0079] Neck network: The middle part of a convolutional neural network, which functions to further integrate and process the feature maps generated by the backbone to obtain new feature maps. Common networks such as the Feature Pyramid Network (FPN) in Faster RCNN

[0080] Bottleneck structure: A multi-layer network structure where the input data of the network first passes through one or more neural network layers to obtain intermediate data, and the intermediate data then passes through one or more neural network layers to obtain the output data. The amount of intermediate data (i.e., the product of width, height, and number of channels) is lower than the amount of input data and output data.

[0081] Data encoding and decoding include two parts: data encoding and data decoding. Data encoding is performed on the source side (or usually referred to as the encoder side), and generally includes processing (e.g., compressing) the original data to reduce the amount of data required to represent the original data (thereby enabling more efficient storage and / or transmission). Data decoding is performed on the destination side (or usually referred to as the decoder side), and generally includes performing inverse processing relative to the encoder side to reconstruct the original data. The "encoding and decoding" of the data involved in the embodiments of this application should be understood as the "encoding" or "decoding" of the data. The encoding part and the decoding part are also collectively referred to as encoding and decoding (encoding and decoding, CODEC).

[0082] In the case of lossless data encoding, the original data can be reconstructed, that is, the reconstructed original data has the same quality as the original data (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy data encoding, further compression is performed through quantization, etc., to reduce the amount of data required to represent the original data, and the decoder side cannot fully reconstruct the original data, that is, the quality of the reconstructed original data is lower or worse than the quality of the original data.

[0083] Embodiments of the present application can be applied to video data and other data with compression / decompression requirements, etc. The following takes video data encoding (hereinafter referred to as video encoding) as an example to illustrate the embodiments of the present application. Other types of data (such as image data, audio data, integer data, and other data with compression / decompression requirements) can refer to the following description, and the embodiments of the present application will not be elaborated further. It should be noted that compared with video encoding, during the encoding process of data such as audio data and integer data, there is no need to divide the data into blocks, but the data can be directly encoded.

[0084] Video encoding generally refers to processing an image sequence that forms a video or video sequence. In the field of video encoding, the terms "picture", "frame", or "image" can be used as synonyms.

[0085] Several video encoding standards belong to "lossy hybrid video codecs" (i.e., combining spatial and temporal prediction in the pixel domain with 2D transform coding for applying quantization in the transform domain). Each image in a video sequence is usually divided into a set of non-overlapping blocks, and encoding is usually performed at the block level. In other words, the encoder usually processes and encodes the video at the block (video block) level. For example, prediction blocks are generated through spatial (intra-frame) prediction and temporal (inter-frame) prediction; the prediction blocks are subtracted from the current block (the currently processed / block to be processed) to obtain a residual block; the residual block is transformed and quantized in the transform domain to reduce the amount of data to be transmitted (compressed), and the decoder side applies the inverse processing part relative to the encoder to the encoded or compressed block to reconstruct the current block for representation. In addition, the encoder needs to repeat the processing steps of the decoder so that the encoder and the decoder generate the same predictions (e.g., intra-frame prediction and inter-frame prediction) and / or reconstructed pixels for processing, i.e., encoding subsequent blocks.

[0086] In the following embodiments of the decoding system 10, the encoder 20 and the decoder 30 are described according to Figures 1a to 3 this.

[0087] Figure 1a FIG. 15 is an exemplary block diagram of a decoding system 10 provided by an embodiment of the present application. For example, a video decoding system 10 (or simply referred to as the decoding system 10) that can utilize the technology of the embodiments of the present application. The video encoder 20 (or simply referred to as the encoder 20) and the video decoder 30 (or simply referred to as the decoder 30) in the video decoding system 10 represent devices that can be used to execute various techniques according to the various examples described in the embodiments of the present application.

[0088] As Figure 1a shown, the decoding system 10 includes a source device 12, and the source device 12 is used to provide encoded image data 21 such as encoded images to a destination device 14 for decoding the encoded image data 21.

[0089] The source device 12 includes an encoder 20, and additionally, optionally, may include an image source 16, a pre-processor (or pre-processing unit) 18 such as an image pre-processor, and a communication interface (or communication unit) 22.

[0090] The image source 16 may include or may be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer animation images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage device that stores any of the above images.

[0091] To distinguish the processing performed by the pre-processor (or pre-processing unit) 18, the image (or image data) 17 may also be referred to as the original image (or original image data) 17.

[0092] The pre-processor 18 is configured to receive the original image data 17 and pre-process the original image data 17 to obtain pre-processed image (or pre-processed image data) 19. For example, the pre-processing performed by the pre-processor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising. It can be understood that the pre-processing unit 18 may be an optional component.

[0093] The video encoder (or encoder) 20 is configured to receive the pre-processed image data 19 and provide encoded image data 21 (which will be further described below according to Figure 2 etc.).

[0094] The communication interface 22 in the source device 12 can be used to: receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device via the communication channel 13 for storage or direct reconstruction.

[0095] The destination device 14 includes a decoder 30, and additionally, optionally, may include a communication interface (or communication unit) 28, a post-processor (or post-processing unit) 32, and a display device 34.

[0096] The communication interface 28 in the destination device 14 is configured to directly receive the encoded image data 21 (or any other processed version) from the source device 12 or from any other source device such as a storage device. For example, the storage device is an encoded image data storage device, and provide the encoded image data 21 to the decoder 30.

[0097] The communication interfaces 22 and 28 can be used to send or receive encoded image data (or encoded data) 21 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, etc., or via any type of network, such as a wired network, a wireless network, or any combination thereof, any type of private network and public network, or any combination of any type thereof.

[0098] For example, the communication interface 22 can be used to encapsulate the encoded image data 21 into a suitable format such as a packet, and / or use any type of transmission encoding or processing to process the encoded image data for transmission over the communication link or communication network.

[0099] The communication interface 28 corresponds to the communication interface 22. For example, it can be used to receive the transmitted data and process the transmitted data using any type of corresponding transmission decoding or processing and / or de-encapsulation to obtain the encoded image data 21.

[0100] Both the communication interface 22 and the communication interface 28 can be configured as Figure 1a a unidirectional communication interface as indicated by the arrow of the corresponding communication channel 13 pointing from the source device 12 to the destination device 14 as shown in, or a bidirectional communication interface, and can be used to send and receive messages, etc., to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as the transmission of the encoded image data, etc.

[0101] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31 (which will be further described below according to Figure 3 etc.).

[0102] The post-processor 32 is used to post-process the decoded image, etc., the decoded image data 31 (also referred to as the reconstructed image data) to obtain post-processed image, etc., the post-processed image data 33. The post-processing performed by the post-processing unit 32 can include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, cropping or resampling, or any other processing for generating the decoded image data 31 for display on a display device 34, etc.

[0103] The display device 34 is configured to receive the post - processed image data 33 to display an image to a user, viewer, etc. The display device 34 may be or include any type of display for representing the reconstructed image. For example, an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro - LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.

[0104] The decoding system 10 further includes a training engine 25. The training engine 25 is configured to train the encoder 20 (especially the entropy encoding unit 270 in the encoder 20) or the decoder 30 (especially the entropy decoding unit 304 in the decoder 30) to perform entropy encoding on the image block to be encoded according to the estimated probability distribution. For a detailed description of the training engine 25, please refer to the following method test examples.

[0105] Although Figure 1a The source device 12 and the destination device 14 are shown as separate devices, but the device embodiments may also include both the source device 12 and the destination device 14 or the functions of both the source device 12 and the destination device 14 at the same time, that is, including both the source device 12 or the corresponding function and the destination device 14 or the corresponding function at the same time. In these embodiments, the source device 12 or the corresponding function and the destination device 14 or the corresponding function may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0106] According to the description, Figure 1a The presence and (exact) division of different units or functions in the shown source device 12 and / or destination device 14 may vary according to the actual device and application, which is obvious to those skilled in the art.

[0107] Please refer to Figure 1b , Figure 1b FIG. is an exemplary block diagram of a video decoding system 40 provided by an embodiment of the present application. The encoder 20 (such as a video encoder 20) or the decoder 30 (such as a video decoder 30), or both, may be implemented by, for example, Figure 1bThe processing circuitry in the illustrated video decoding system 40 is implemented by, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding dedicated processors, or any combination thereof. Refer to Figure 2 and Figure 3 , Figure 2 which is an exemplary block diagram of a video encoder provided by an embodiment of the present application, Figure 3 and which is an exemplary block diagram of a video decoder provided by an embodiment of the present application. The encoder 20 may be implemented by the processing circuitry 46 to include various modules discussed with reference to Figure 2 the encoder 20 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented by the processing circuitry 46 to include various modules discussed with reference to Figure 3 the decoder 30 and / or any other decoder system or subsystem described herein. The processing circuitry 46 may be used to perform various operations discussed below. As Figure 4 illustrated, if some techniques are implemented in software, the device may store the instructions of the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors, thereby implementing the techniques of the embodiments of the present application. One of the video encoder 20 and the video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) in a single device, as Figure 1b illustrated.

[0108] The source device 12 and the destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, for example, a laptop or notebook computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, and a monitoring device, etc., and may or may not use any type of operating system. The source device 12 and the destination device 14 may also be devices in a cloud computing scenario, such as virtual machines in a cloud computing scenario. In some cases, the source device 12 and the destination device 14 may be equipped with components for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.

[0109] The source device 12 and the destination device 14 can install virtual scene application programs (applications, APPs) such as virtual reality (VR) applications, augmented reality (AR) applications, or mixed reality (MR) applications, and can run VR applications, AR applications, or MR applications based on user operations (such as clicking, touching, swiping, shaking, voice control, etc.). The source device 12 and the destination device 14 can collect images / videos of any object in the environment through a camera and / or a sensor, and then display virtual objects on a display device according to the collected images / videos. The virtual objects can be virtual objects in a VR scene, an AR scene, or an MR scene (i.e., objects in a virtual environment).

[0110] It should be noted that in the embodiments of the present application, the virtual scene application programs in the source device 12 and the destination device 14 can be application programs built into the source device 12 and the destination device 14 themselves, or can be application programs provided by third-party service providers installed by the user, and no specific limitation is made thereto.

[0111] In addition, the source device 12 and the destination device 14 can install real-time video transmission applications, such as live broadcast applications. The source device 12 and the destination device 14 can collect images / videos through a camera, and then display the collected images / videos on a display device.

[0112] In some cases, Figure 1a the illustrated video decoding system 10 is merely exemplary, and the technology provided by the embodiments of the present application is applicable to video coding settings (e.g., video encoding or video decoding), which may not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from a local memory, sent over a network, etc. A video encoding device can encode data and store the data in a memory, and / or a video decoding device can retrieve data from the memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data into a memory and / or retrieve and decode data from the memory.

[0113] Please refer to Figure 1b , Figure 1b which is an exemplary block diagram of the video decoding system 40 provided by the embodiments of the present application. As Figure 1b shown, the video decoding system 40 can include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video codec implemented by a processing circuit 46), an antenna 42, one or more processors 43, one or more memory memories 44, and / or a display device 45.

[0114] AsFigure 1b As shown, the imaging device 41, antenna 42, processing circuit 46, video encoder 20, video decoder 30, processor 43, memory 44, and / or display device 45 can communicate with each other. In different instances, the video decoding system 40 may include only the video encoder 20 or only the video decoder 30.

[0115] In some instances, the antenna 42 can be used to transmit or receive an encoded bitstream of video data. Additionally, in some instances, the display device 45 can be used to present video data. The processing circuit 46 can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. The video decoding system 40 can also include an optional processor 43, which can similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. Additionally, the memory 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting instance, the memory 44 can be implemented by cache memory. In other instances, the processing circuit 46 can include a memory (e.g., a cache, etc.) for implementing an image buffer, etc.

[0116] In some instances, the video encoder 20 implemented by logic circuits can include an image buffer (e.g., implemented by the processing circuit 46 or the memory 44) and a graphics processing unit (e.g., implemented by the processing circuit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit can include the video encoder 20 implemented by the processing circuit 46 to implement various modules discussed with reference to Figure 2 and / or any other encoder system or subsystem described herein. The logic circuits can be used to perform various operations discussed herein.

[0117] In some instances, the video decoder 30 can be implemented in a similar manner by the processing circuit 46 to implement with reference to Figure 3The video decoder 30 and / or the various modules discussed for any other decoder system or subsystem described herein. In some examples, the logic circuit-implemented video decoder 30 may include an image buffer (implemented by the processing circuit 46 or the memory 44) and a graphics processing unit (e.g., implemented by the processing circuit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 30 implemented by the processing circuit 46 to implement with reference to Figure 3 and / or the various modules discussed for any other decoder system or subsystem described herein.

[0118] In some examples, the antenna 42 may be used to receive the encoded bitstream of video data. As discussed, the encoded bitstream may include data, indicators, index values, mode selection data, etc. discussed herein related to the encoded video frames, e.g., data related to the encoded partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators as discussed, and / or data defining the encoded partitions). The video decoding system 40 may also include a video decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is used to present the video frames.

[0119] It should be understood that for the examples described with reference to the video encoder 20 in the embodiments of the present application, the video decoder 30 may be used to perform the opposite process. Regarding the signaling syntax elements, the video decoder 30 may be used to receive and parse such syntax elements and accordingly decode the relevant video data. In some examples, the video encoder 20 may entropy encode the syntax elements into an encoded video bitstream. In such examples, the video decoder 30 may parse such syntax elements and accordingly decode the relevant video data.

[0120] For ease of description, the embodiments of the present application are described with reference to the versatile video coding (VVC) reference software or the high-efficiency video coding (HEVC) developed by the video coding experts group (VCEG) of ITU-T and the joint collaboration team on video coding (JCT-VC) of ISO / IEC motion picture experts group (MPEG). Those of ordinary skill in the art understand that the embodiments of the present application are not limited to HEVC or VVC.

[0121] Encoder and encoding method

[0122] As Figure 2As shown, the video encoder 20 includes an input end (or input interface) 201, a residual calculation unit 204, a transformation processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output end (or output interface) 272. The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a segmentation unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The video encoder 20 shown may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.

[0123] Image and image segmentation (image and block)

[0124] The encoder 20 can be used to receive an image (or image data) 17 through the input end 201, etc., for example, an image in an image sequence forming a video or a video sequence. The received image or image data can also be a preprocessed image (or preprocessed image data) 19. For simplicity, the following description uses the image 17. The image 17 can also be referred to as the current image or the image to be encoded (especially when distinguishing the current image from other images in video coding, such as the previously encoded and / or decoded images in the same video sequence, i.e., the video sequence that also includes the current image).

[0125] (Digital) images are or can be regarded as two-dimensional arrays or matrices composed of pixel points with intensity values. Pixel points in the array can also be referred to as pixels (pixel or pel, short for picture element). The number of pixel points in the horizontal and vertical directions (or axes) of the array or image determines the size and / or resolution of the image. To represent colors, usually three color components are used, that is, the image can be represented as or include three pixel point arrays. In the RGB format or color space, the image includes corresponding red, green, and blue pixel point arrays. However, in video coding, each pixel is usually represented in a luminance / chrominance format or color space, such as YCbCr, including a luminance component indicated by Y (sometimes also denoted by L) and two chrominance components denoted by Cb and Cr. The luminance (luma) component Y represents the luminance or gray-level intensity (for example, they are the same in a grayscale image), while the two chrominance components Cb and Cr represent the chrominance or color information components. Correspondingly, an image in the YCbCr format includes a luminance pixel point array of luminance pixel point values (Y) and two chrominance pixel point arrays of chrominance values (Cb and Cr). An image in the RGB format can be converted or transformed into the YCbCr format, and vice versa, and this process is also called color transformation or conversion. If the image is black and white, then the image can only include a luminance pixel point array. Correspondingly, the image can be, for example, a luminance pixel point array in a monochrome format or a luminance pixel point array and two corresponding chrominance pixel point arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0126] In one embodiment, an embodiment of the video encoder 20 may include an image segmentation unit ( Figure 2 not shown in the figure) for segmenting the image 17 into a plurality of (usually non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB), or coding tree units (CTU) in the H.265 / HEVC and VVC standards. The segmentation unit can be used to use the same block size for all images in a video sequence and a corresponding grid defining the block size, or to change the block size between images or subsets or groups of images, and segment each image into corresponding blocks.

[0127] In other embodiments, the video encoder can be used to directly receive the blocks 203 of the image 17, for example, one, several, or all of the blocks constituting the image 17. The image block 203 can also be referred to as the current image block or the image block to be encoded.

[0128] Similar to Image 17, Image Block 203 is also or can be considered as a two-dimensional array or matrix composed of pixel points with intensity values (pixel values), but Image Block 203 is smaller than Image 17. In other words, Block 203 can include an array of pixel points (e.g., a luminance array in the case of a monochrome image 17 or a luminance array and two chrominance arrays in the case of a color image) or three arrays of pixel points (e.g., one luminance array and two chrominance arrays in the case of a color image 17) or any other number and / or type of arrays according to the color format adopted. The number of pixel points in the horizontal and vertical directions (or axes) of Block 203 defines the size of Block 203. Accordingly, the block can be an array of M×N (M columns × N rows) pixel points, or an array of M×N transform coefficients, etc.

[0129] In one embodiment, Figure 2 the illustrated video encoder 20 is used to encode Image 17 block by block, for example, performing encoding and prediction on each block 203.

[0130] In one embodiment, Figure 2 the illustrated video encoder 20 can also be used to segment and / or encode an image using slices (also known as video slices), where the image can be segmented or encoded using one or more slices (usually non-overlapping). Each slice can include one or more blocks (e.g., Coding Tree Units CTUs) or one or more groups of blocks (e.g., coding blocks (tiles) in the H.265 / HEVC / VVC standards and bricks in the VVC standard).

[0131] In one embodiment, Figure 2 the illustrated video encoder 20 can also be used to segment and / or encode an image using slice / coding block groups (also known as video coding block groups) and / or coding blocks (also known as video coding blocks), where the image can be segmented or encoded using one or more slice / coding block groups (usually non-overlapping), each slice / coding block group can include one or more blocks (e.g., CTUs) or one or more coding blocks, etc., where each coding block can be in a shape such as a rectangle and can include one or more complete or partial blocks (e.g., CTUs).

[0132] Residual calculation

[0133] The residual calculation unit 204 is used to calculate the residual block 205 based on the image block (or original block) 203 and the prediction block 265 (the prediction block 265 is introduced in detail later) in the following way: for example, subtracting the pixel value of the prediction block 265 from the pixel value of the image block 203 pixel by pixel (pixel by pixel) to obtain the residual block 205 in the pixel domain.

[0134] Quantization

[0135] Quantization unit 208 is used to quantize transform coefficients 207 through, for example, scalar quantization or vector quantization to obtain quantized transform coefficients 209. The quantized transform coefficients 209 may also be referred to as quantized residual coefficients 209.

[0136] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded down to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different degrees of scaling can be applied to achieve finer or coarser quantization. A smaller quantization step corresponds to finer quantization, while a larger quantization step corresponds to coarser quantization. The appropriate quantization step can be indicated by the quantization parameter (QP). For example, the quantization parameter can be an index of a predefined set of appropriate quantization steps. For example, a smaller quantization parameter may correspond to fine quantization (smaller quantization step), and a larger quantization parameter may correspond to coarse quantization (larger quantization step), and vice versa. Quantization may include dividing by the quantization step, and the corresponding or inverse dequantization performed by dequantization unit 210 and the like may include multiplying by the quantization step. Embodiments according to some standards such as HEVC can be used to determine the quantization step using the quantization parameter. Generally, the quantization step can be calculated using a fixed-point approximation of an equation involving division according to the quantization parameter. Other scaling factors may be introduced for quantization and dequantization to restore the norm of the residual block that may be modified due to the scaling used in the fixed-point approximation of the equations for the quantization step and quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream. Quantization is a lossy operation, where the larger the quantization step, the greater the loss.

[0137] In one embodiment, video encoder 20 (correspondingly, quantization unit 208) may be used to output a quantization parameter (QP), for example, directly output or output after being encoded or compressed by entropy coding unit 270, such that video decoder 30 can receive and use the quantization parameter for decoding.

[0138] Dequantization

[0139] Dequantization unit 210 is used to perform the inverse quantization of quantization unit 208 on the quantized coefficients to obtain dequantized coefficients 211, for example, perform an inverse quantization scheme of the quantization scheme performed by quantization unit 208 according to or using the same quantization step as quantization unit 208. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211, corresponding to transform coefficients 207, but due to the loss caused by quantization, the dequantized coefficients 211 are usually not exactly the same as the transform coefficients.

[0140] Reconstruction

[0141] The reconstruction unit 214 (e.g., the adder 214) is used to add the transform block 213 (i.e., the reconstruction residual block 213) to the prediction block 265 to obtain a reconstructed block 215 in the pixel domain. For example, the pixel values of the reconstruction residual block 213 and the pixel values of the prediction block 265 are added together.

[0142] Segmentation

[0143] The segmentation unit 262 can segment (or partition) an image block (or CTU) 203 into smaller parts, such as small blocks in the shape of a square or a rectangle. For an image with a three-pixel array, a CTU consists of an N×N block of luminance pixels and two corresponding chrominance pixel blocks. The maximum allowed size of the luminance block in the currently developing Versatile Video Coding (VVC) standard is specified as 128×128, but it may be specified as a value different from 128×128 in the future, such as 256×256. The CTUs of an image can be grouped / concentrated into slices / Coding Tree Units (CTUs), Coding Blocks, or tiles. A Coding Block covers a rectangular area of an image, and a Coding Block can be divided into one or more tiles. A tile consists of multiple CTU rows within a Coding Block. A Coding Block that is not divided into multiple tiles can be called a tile. However, a tile is a proper subset of a Coding Block and thus is not called a Coding Block. VVC supports two Coding Tree Unit modes, namely the raster scan slice / Coding Tree Unit mode and the rectangular slice mode. In the raster scan Coding Tree Unit mode, a slice / Coding Tree Unit contains a sequence of Coding Blocks in the raster scan of the Coding Blocks of an image. In the rectangular slice mode, a slice contains multiple tiles of an image, and these tiles together form a rectangular area of the image. The tiles within a rectangular slice are arranged in the tile raster scan order of the slice. These smaller blocks (which can also be called sub-blocks) can be further divided into even smaller parts. This is also called tree segmentation or hierarchical tree segmentation, where a root block at the root tree level 0 (hierarchical level 0, depth 0), etc., can be recursively divided into two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchical level 1, depth 1). These blocks can in turn be divided into two or more blocks at the next lower level, such as tree level 2 (hierarchical level 2, depth 2), etc., until the segmentation ends (because an end criterion is met, such as reaching the maximum tree depth or the minimum block size). Blocks that are not further divided are also called leaf blocks or leaf nodes of the tree. A tree divided into two parts is called a binary tree (BT), a tree divided into three parts is called a ternary tree (TT), and a tree divided into four parts is called a quadtree (QT).

[0144] Entropy coding

[0145] The entropy coding unit 270 is used to apply an entropy coding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CALVC), arithmetic coding scheme, binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements, to obtain coded image data 21 that can be output in the form of a coded bitstream 21 etc. through the output terminal 272, such that a video decoder 30 etc. can receive and use the parameters for decoding. The coded bitstream 21 can be transmitted to the video decoder 30, or stored in a memory for later transmission or retrieval by the video decoder 30.

[0146] Other structural variants of the video encoder 20 can be used to encode a video stream. For example, a non-transform-based encoder 20 can directly quantize the residual signal in the case where some blocks or frames do not have a transform processing unit 206. In another implementation, the encoder 20 can have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.

[0147] Decoder and decoding method

[0148] As Figure 3 shown, the video decoder 30 is used to receive, for example, the coded image data 21 (e.g., the coded bitstream 21) encoded by the encoder 20, to obtain a decoded image 331. The coded image data or bitstream includes information for decoding the coded image data, such as data representing image blocks of a coded video slice (and / or coded group of blocks or coded block) and related syntax elements.

[0149] In Figure 3In the example of, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (such as an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. In some examples, the video decoder 30 may perform a decoding process that is generally opposite to the encoding process described with reference to Figure 2 the video encoder 100.

[0150] As described for the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer DPB 230, the inter prediction unit 344, and the intra prediction unit 354 also constitute the "built-in decoder" of the video encoder 20. Accordingly, the inverse quantization unit 310 may be functionally the same as the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally the same as the inverse transform processing unit 122, the reconstruction unit 314 may be functionally the same as the reconstruction unit 214, the loop filter 320 may be functionally the same as the loop filter 220, and the decoded picture buffer 330 may be functionally the same as the decoded picture buffer 230. Therefore, the explanations of the corresponding units and functions of the video encoder 20 correspondingly apply to the corresponding units and functions of the video decoder 30.

[0151] Entropy Decoding

[0152] The entropy decoding unit 304 is used to parse the bitstream 21 (or generally the encoded picture data 21) and perform entropy decoding on the encoded picture data 21 to obtain quantization coefficients 309 and / or decoded encoded parameters ( Figure 3 not shown in), such as any one or all of inter prediction parameters (such as reference picture indices and motion vectors), intra prediction parameters (such as intra prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be used to apply a decoding algorithm or scheme corresponding to the encoding scheme of the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may also be used to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode application unit 360, and other parameters to other units of the decoder 30. The video decoder 30 may receive video slice and / or video block-level syntax elements. Additionally, or as an alternative to slices and corresponding syntax elements, encoded block groups and / or encoded blocks and corresponding syntax elements may be received or used.

[0153] Inverse Quantization

[0154] The inverse quantization unit 310 can be used to receive a quantization parameter (QP) (or generally information related to inverse quantization) and quantized coefficients from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse-quantize the decoded quantized coefficients 309 based on the quantization parameter to obtain inverse-quantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include determining the degree of quantization using the quantization parameter calculated by the video encoder 20 for each video block in the video slice, and also determining the degree of inverse quantization to be performed.

[0155] Reconstruction

[0156] The reconstruction unit 314 (e.g., adder 314) is used to add the reconstructed residual block 313 to the prediction block 365 to obtain a reconstructed block 315 in the pixel domain. For example, the pixel values of the reconstructed residual block 313 and the pixel values of the prediction block 365 are added together.

[0157] Other variants of the video decoder 30 can be used to decode the encoded image data 21. For example, the decoder 30 can produce an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 can directly inverse-quantize the residual signal without the inverse transform processing unit 312 for some blocks or frames. In another implementation, the video decoder 30 can have the inverse quantization unit 310 and the inverse transform processing unit 312 combined into a single unit.

[0158] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step can be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clip or shift operations can be performed on the processing results of interpolation filtering, motion vector derivation, or loop filtering.

[0159] It should be noted that further operations can be performed on the derived motion vectors of the current block (including but not limited to the control point motion vectors in affine mode, the sub-block motion vectors in affine, planar, and ATMVP modes, the temporal motion vectors, etc.). For example, the value of the motion vector is restricted to a predefined range according to the representation bits of the motion vector. If the representation bits of the motion vector are bitDepth, the range is from -2^(bitDepth - 1) to 2^(bitDepth - 1) - 1, where "^" represents the power. For example, if bitDepth is set to 16, the range is from -32768 to 32767; if bitDepth is set to 18, the range is from -131072 to 131071. For example, the values of the derived motion vectors (such as the MVs of 4 4×4 sub-blocks in an 8×8 block) are restricted such that the maximum difference between the integer parts of the 4 4×4 sub-block MVs does not exceed N pixels, for example, does not exceed 1 pixel. Two methods for restricting the motion vector according to bitDepth are provided here.

[0160] Although the above embodiments mainly describe video coding and decoding, it should be noted that the embodiments of the decoding system 10, the encoder 20, and the decoder 30, as well as other embodiments described herein, can also be used for still image processing or coding and decoding, that is, the processing or coding and decoding of a single image independent of any previous or consecutive images in video coding and decoding. Generally, if the image processing is limited to a single image 17, the inter-frame prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 can equally be used for static image processing, such as residual calculation 204 / 304, transformation 206, quantization 208, dequantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354, and / or loop filtering 220 / 320, entropy coding 270, and entropy decoding 304.

[0161] Please refer to Figure 4 , Figure 4 which is an exemplary block diagram of the video decoding device 400 provided by the embodiments of the present application. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 400 can be a decoder, such as Figure 1a the video decoder 30 in Figure 1a or an encoder, such as

[0162] Video decoding device 400 includes: an input port 410 (or input port 410) for receiving data and a receiver unit (Rx) 420; a processor, logic unit, or central processing unit (CPU) 430 for processing data; for example, the processor 430 here can be a neural network processor 430; a transmitter unit (Tx) 440 and an output port 450 (or output port 550) for transmitting data; and a memory 460 for storing data. Video decoding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 410, the receiving unit 420, the transmitting unit 440, and the output port 450 for the exit or entry of optical or electrical signals.

[0163] Processor 430 is implemented by hardware and software. Processor 430 can be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with the input port 410, the receiving unit 420, the transmitting unit 440, the output port 450, and the memory 460. Processor 430 includes a neural network-based codec 470. The neural network-based codec 470 implements the embodiments disclosed above. For example, the neural network-based codec 470 performs, processes, prepares, or provides various encoding operations. Therefore, the neural network-based codec 470 provides a substantial improvement to the functions of video decoding device 400 and affects the switching of video decoding device 400 to different states. Alternatively, the neural network-based codec 470 is implemented by instructions stored in the memory 460 and executed by the processor 430.

[0164] Memory 460 includes one or more disks, tape drives, and solid-state drives, and can be used as an overflow data storage device for storing such programs when a selected program is executed, and storing instructions and data read during program execution. Memory 460 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0165] Please refer to Figure 5 , Figure 5An exemplary block diagram of a device 500 provided in an embodiment of the present application, the device 500 can be used as Figure 1a Either or both of the source device 12 and the destination device 14 in .

[0166] The processor 502 in the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices that are currently available or will be developed in the future and are capable of manipulating or processing information. Although a single processor such as the processor 502 shown in the figure may be used to implement the disclosed implementation, using more than one processor is faster and more efficient.

[0167] In one implementation, the memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessed by the processor 502 via a bus 512. The memory 504 may also include an operating system 508 and an application 510, which includes at least one program that allows the processor 502 to perform the methods described herein. For example, the application 510 may include applications 1 to N, and also include a video decoding application that performs the methods described herein.

[0168] The apparatus 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element that may be used to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.

[0169] Although bus 512 in device 500 is described herein as a single bus, bus 512 may include multiple buses. In addition, auxiliary storage may be directly coupled to other components of device 500 or accessed through a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Therefore, device 500 may have a variety of configurations.

[0170] Image encoding:

[0171] Nowadays, multimedia data occupies a large part of Internet traffic. The compression of image data plays an important role in the storage and efficient transmission of multimedia data. Therefore, image coding technology is a technology with great practical value. It should be noted that the Chinese translation of the English terms "image coding" and "image encoding" is usually "image coding". Image coding is broad and includes the process of encoding an image into a bit stream and the process of decoding (decoding) a bit stream into an image.

[0172] Image coding only refers to the process of encoding an image into a bitstream. The research on image coding has a long history. Researchers have proposed a large number of methods and developed I-frame coding methods for various image coding standards and video coding standards such as JPEG, JPEG2000, JPEG-XL, JPEG-XX, WebP, H.264 / AVC, H.264 / HEVC, H.26 / VVC, AVS3, and AV1. Most of these coding methods are based on transform, prediction, and entropy coding techniques. Although these coding methods are currently widely used, due to the increase in image data volume and the emergence of new media types, coding methods with higher compression efficiency are needed.

[0173] Image Coding Based on Deep Learning:

[0174] In recent years, researchers have studied image coding methods based on deep learning. Some researchers have achieved good results. For example, Balle et al. proposed an end-to-end optimized image coding method that is better than the existing best image coding and even better than the existing best traditional coding standard H.265 / HEVC.

[0175] Image coding based on deep learning is based on deep neural networks, usually convolutional neural networks. Some research works have proposed an image coding method based on the Transformer network. The structure of the deep neural network can be designed manually or obtained through neural architecture search (NAS). The parameters of the deep neural network are obtained by using the loss function and the backpropagation algorithm.

[0176] Figure 6 Shows a typical deep learning-based image compression method, also known as neural network-based image compression method. Generally, the neural network-based image compression method includes the following parts: feature extraction module, feature quantization module, entropy coding module, entropy decoding module, feature dequantization module, and feature decoding module. On the encoder side, the feature extraction module can use a non-linear mapping activation function to obtain the extracted three-dimensional feature map through multi-layer convolution stacking. The feature quantization module quantizes the floating-point feature values through eigenvalue quantization to obtain the quantized feature values. The quantized feature values are subjected to lossless entropy coding to obtain the encoded bitstream. When receiving the bitstream of entropy coding, the decoder performs lossless entropy decoding to obtain the three-dimensional quantized feature values. The feature decoding module decodes the features into a reconstructed image to achieve decoding.

[0177] After the image to be compressed passes through the feature extraction module and the feature quantization module, a three-dimensional feature quantization map is obtained. When processing each eigenvalue in the three-dimensional feature quantization map, the entropy coding module can estimate and obtain the probability distribution of the eigenvalue by using the eigenvalues in the processed neighborhood as context, and perform subsequent coding based on the probability distribution to obtain a coded bitstream.

[0178] With the excellent performance of deep learning in various fields, researchers have proposed an end-to-end image coding solution based on deep learning. Figure 7 The coding framework is shown. The specific technical solution is as follows: On the encoder side, the original image is input into the feature extraction module, and a feature map is output. The feature map passes through the side information extraction module, and side information is output. On both the encoder and decoder sides, it is input into the probability estimation module, and the probability distribution of each feature element is output to obtain the value of the feature element to be coded. In addition, the feature map is input into the quantization module to obtain a quantized feature map. The entropy coding module performs entropy coding on each feature element in the quantized feature map based on the probability distribution of each feature element to obtain a coded bitstream.

[0179] On the decoder side, the decoder parses the bitstream and outputs the probability distribution of the symbol to be coded based on the side information to obtain the value of the feature element to be decoded as k. The entropy decoding module performs arithmetic decoding on each feature element in the quantized feature map based on the probability distribution of each feature element to obtain the value of the feature element. The feature map is input into the image reconstruction module, and a reconstructed image is output.

[0180] Neural network

[0181] A neural network can include neurons. A neuron can be an operation unit that uses xs and an intercept 1 as inputs. The output of the operation unit can be:

[0182]

[0183] Here, s = s = 1, 2,..., n, n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neuron, and f is the activation function (activation function) of the neuron, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting multiple individual neurons together. Specifically, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be a region including several neurons.

[0184] Convolutional neural network:

[0185] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network includes a feature extractor, which includes a convolutional layer and a subsampling layer. The feature extractor can be regarded as a filter. The convolutional layer is a layer of neurons in the convolutional neural network that performs convolutional processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can be connected to only a part of the neurons in the adjacent layer. The convolutional layer usually includes several feature planes, and each feature plane can include some neurons arranged in a rectangle. Neurons on the same feature plane share a weight, and the shared weight here is the convolutional kernel. The shared weight can be understood as a way of extracting image information that is independent of position. The convolutional kernel can be initialized in the form of a matrix of random size. During the training process of the convolutional neural network, appropriate weights can be obtained for the convolutional kernel through learning. In addition, the shared weight directly reduces the connections between the layers of the convolutional neural network and reduces the risk of overfitting.

[0186] Figure 8 Schematically shows the general concept of processing by a neural network such as a CNN. A convolutional neural network consists of an input layer, an output layer, and multiple hidden layers. The input layer is the layer that provides an input (such as Figure 8 a part of the image shown) for processing. The hidden layers of a CNN usually consist of a series of convolutional layers that are convolved with multiplication or other dot products. The result of the layer is one or more feature maps, sometimes also called channels. Subsampling may be involved in some or all layers. Thus, as Figure 8 shown, the feature maps can become smaller. The activation function in a CNN is usually a ReLU (Rectified Linear Unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, and normalization layers, which are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. Although these layers are popularly called convolutions, this is just a convention. Mathematically, it is technically a sliding dot product or cross-correlation. This has important implications for the indices in the matrix because it affects how the weights are determined at a specific index point.

[0187] When programming a CNN for processing images, as Figure 8 shown, the input is a tensor with the shape (number of images) x (image width) x (image height) x (image depth). Then, after passing through the convolutional layer, the image is abstracted into a feature map with the shape (number of images) x (feature map width) x (feature map height) x (feature map channels). The convolutional layer in a neural network should have the following properties. A convolutional kernel (hyperparameter) defined by width and height. The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) should be equal to the number of channels (depth) of the input feature map.

[0188] In the past, traditional multi-layer perceptron (MLP) models have been used for image recognition. However, due to their full connectivity between nodes, they have high dimensionality and do not scale well with higher-resolution images. An image of 1000×1000 pixels with RGB color channels has 3 million weights, which is too high to be effectively processed at scale in a fully connected manner. Additionally, this network architecture does not consider the spatial structure of the data, treating input pixels that are far apart in the same way as those that are close. This ignores the locality of reference in the image data both computationally and semantically. Therefore, the full connectivity of neurons is wasteful for purposes such as image recognition, which is dominated by spatially local input patterns.

[0189] Convolutional neural networks are biologically inspired variants of multi-layer perceptrons, specifically designed to mimic the behavior of the visual cortex. These models alleviate the challenges posed by the MLP architecture by exploiting the strong spatial local correlations present in natural images. The convolutional layer is the core building block of a CNN. The parameters of this layer consist of a set of learnable filters (the kernels mentioned above), which have small receptive fields but extend across the entire depth of the input volume. During the forward pass, each filter is convolved over the width and height of the input volume, computing the dot product between the entries of the filter and the input, and producing a two-dimensional activation map for that filter. Thus, the network learns filters that activate when it detects a particular type of feature at a certain spatial location in the input.

[0190] Stacking the activation maps of all filters along the depth dimension forms the complete output volume of the convolutional layer. Thus, each entry in the output volume can also be interpreted as the output of a neuron that observes a small region in the input and shares parameters with neurons in the same activation map. A feature map or activation map is the output activation of a given filter. Feature map and activation have the same meaning. In some papers, it is called an activation map because it is a mapping corresponding to the activation of different parts of the image, and it is also a feature map because it is also a mapping where a certain feature is found in the image. High activation means that a certain function has been found.

[0191] Another important concept in convolutional neural networks is pooling, which is a form of non-linear downsampling. There are several non-linear functions that can implement pooling, and max pooling is the most common. It divides the input image into a set of non-overlapping rectangles, and for each such sub-region, it outputs the maximum value.

[0192] Intuitively, the exact location of a feature is less important than its rough location relative to other features. This is the idea behind using pooling in convolutional neural networks. Pooling layers are used to gradually reduce the spatial size of the representation, decreasing the number of parameters, memory footprint, and computational load in the network, and thus also controlling overfitting. In a CNN architecture, it is common to periodically insert pooling layers between successive convolutional layers. Pooling operations provide another form of translational invariance.

[0193] Pooling layers operate independently on each depth slice of the input and resize it spatially. The most common form is a pooling layer with a filter of size 2×2, which applies 2 downsampling strides along the width and height on each depth slice in the input, discarding 75% of the activations. In this case, each max operation is over 4 numbers. The depth dimension remains unchanged.

[0194] In addition to max pooling, pool units can also use other functions, such as average pooling or l2-norm pooling. Average pooling was historically often used but has fallen out of favor recently compared to max pooling, which performs better in practice. Due to the significant reduction in representation size, there has been a recent trend to use smaller filters or to discard pooling layers altogether. "Region of interest" pooling (also known as ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important part of convolutional neural networks for object detection based on the fast R-CNN architecture.

[0195] The above ReLU is an abbreviation for rectified linear unit, which applies a non-saturating activation function. By setting negative values to zero, it effectively removes negative values from the activation map. It increases the non-linearity of the decision function and the entire network without affecting the receptive field of the convolutional layer. Other functions are also used to increase non-linearity, such as the saturated hyperbolic tangent and the sigmoid function. ReLU is generally more popular than other functions because it trains neural networks several times faster without significantly affecting the generalization accuracy.

[0196] After several convolutional layers and max pooling layers, high-level inference in a neural network is done through fully connected layers. Neurons in fully connected layers are connected to all activations in the previous layer, as seen in regular (non-convolutional) artificial neural networks. Thus, their activations can be computed as an affine transformation, followed by a bias offset (a vector addition of learned or fixed bias terms) after matrix multiplication.

[0197] The "loss layer" specifies how training penalizes the deviation between the predictions (outputs) and the true labels, typically the last layer of a neural network. Various loss functions suitable for different tasks can be used. Softmax loss is used to predict a single class out of K mutually exclusive classes. Sigmoid cross-entropy loss is used to predict K independent probability values in the range [0, 1]. Euclidean loss is used for regression to real-valued labels.

[0198] In summary, Figure 8 shows the data flow in a typical convolutional neural network. First, the input image passes through the convolutional layer and is abstracted into a feature map consisting of several channels, corresponding to the number of filters in a set of learnable filters for that layer (e.g., one channel per filter). Then the feature map is subsampled using, for example, a pooling layer, which reduces the dimension of each channel in the feature map. The subsequent data enters another convolutional layer, which may have a different number of output channels, resulting in a different number of channels in the feature map. As described above, the number of input channels and output channels are hyperparameters of the layer. To establish the connectivity of the network, these parameters need to be synchronized between two connected layers, e.g., the number of input channels of the current layer should be equal to the number of output channels of the previous layer. For the first layer that processes the input data (e.g., an image), the number of input channels is usually equal to the number of channels of the data representation, e.g., 3 channels for the RGB or YUV representation of an image or video, or 1 channel for the grayscale image or video representation.

[0199] The method provided by the embodiments of this application is applied to a video codec, such as a video codec in a video communication system, such as Figure 9 shown: After the video is captured using a video capture device, it undergoes a series of preprocessing, and then the processed video is compressed and encoded to obtain an encoded bitstream. The bitstream is sent to the receiving module via the transmission network using the sending module, and after being decoded by the decoder, it can be rendered and displayed. In addition, the bitstream after video encoding can also be directly stored.

[0200] The embodiments of this application describe a neural network-based encoding and decoding scheme for enhancement layers, and its application scenario can be a hierarchical coding scheme.

[0201] The method provided by the embodiments of this application can be used in devices or products with video encoder and / or decoder functions, such as video processing software and hardware products, such as chips. And in products or devices containing such chips, such as media products like mobile phones.

[0202] Figure 10 shows a flowchart of an encoding method provided by the embodiments of this application. As Figure 10 shown, the encoding method provided by the embodiments of this application includes:

[0203] S1001. Perform AI encoding on the input image to obtain a feature map.

[0204] Among them, the above AI encoding can be any AI encoder.

[0205] For example, the input image can be input into the JPEG AI encoder for AI encoding to obtain a feature map.

[0206] S1002. Perform entropy encoding on the feature map to obtain a bitstream.

[0207] For example, asymmetric numeral system (ANS) entropy encoding can be performed on the feature map to obtain a bitstream.

[0208] S1003. Perform a transformation operation on the target region of the input image to obtain transformation coefficients.

[0209] In a possible implementation, the above transformation coefficients can be residuals.

[0210] In a possible implementation, one or more regions of the input image can be selected as the target region.

[0211] Exemplarily, one or more regions of the input image can be selected as the target region according to the image distortion (image distortion degree) of the region.

[0212] For example, the region in the input image with an image distortion (image distortion degree) greater than the image distortion threshold can be used as the target region.

[0213] Also exemplarily, one or more regions of interest to the human eye in the input image can be selected as the target region.

[0214] Also exemplarily, one or more regions of the input image can be randomly selected as the target region.

[0215] In a possible implementation, the above transformation operation includes at least one of DCT, DFT, or DWT.

[0216] For example, DCT transformation can be performed on the target region of the input image.

[0217] In a possible implementation, the above target region can be divided into multiple image blocks. Perform the above transformation operation on the multiple image blocks.

[0218] Exemplarily, the above target region can be divided into multiple image blocks according to the granularity of HxW. Perform the above transformation operation on the multiple image blocks. Among them, H is a positive integer, W is a positive integer, and H and W can be the same.

[0219] For example, the above target region can be divided into multiple image blocks at a granularity of 8x8 based on a binary mask image. Perform DCT operations on the above multiple image blocks.

[0220] Among them, the binary mask image represents all 8x8 image blocks in the original image that need to be subjected to DCT transformation. The 8x8 block corresponding to the point with a value of 1 in the mask image represents the image block that needs to be subjected to DCT transformation.

[0221] S1004. Encode the transformation coefficients into a bitstream.

[0222] In a possible implementation, the above transformation coefficients can be quantized; the quantized transformation coefficients are encoded into the above bitstream.

[0223] In a possible implementation, the above transformation coefficients can be entropy-encoded into the above bitstream.

[0224] Exemplarily, the direct current (DC) transformation coefficients and alternating current (AC) transformation coefficients of the input image can be encoded separately. Among them, NxN transformation coefficients are obtained by performing a transformation on an NxN block. The transformation coefficient at the (0,0) position is the DC transformation coefficient, and the transformation coefficients at other positions are AC transformation coefficients.

[0225] For example, differential coding can be used for the DC transformation coefficients of the input image, that is, the difference from the DC component of the previous block is encoded. According to the magnitude of the difference, it is divided into multiple groups. The group number of each group is entropy-encoded using the tabled asymmetric numeral systems (tANS) algorithm, and the magnitude within each group is encoded using a fixed-length code.

[0226] For example, the DC transformation coefficients of the input image can be differentially encoded according to the table shown in Table 1. Table 1 shows the group number, the corresponding value range for each group, and the number of bits required for the value. When performing tANS encoding on the group number, first count the probabilities P0 to P11 of each group number, represent the probability values using 8-bit unsigned integers, write the probabilities into the bitstream, and at the same time, perform entropy encoding on the group number according to the probabilities P0 to P11 and write it into the bitstream.

[0227] Table 1

[0228] Group number Numerical range Number of bits required for amplitude 0 0 - 1 -1,1 1 2 -3,-2,2,-3 2 3 -7,…,-4,4,…7 3 4 -15,…,-8,8,…-15 4 5 -31,…,-16,16,…,31 5 6 -63,…,-32,32,…,63 6 7 -127,…,-64,64,…,127 7 8 -255,…,-128,128,255 8 9 -511,…,-256,256,…,511 9 10 -1023,…,-512,512,…,1023 10 11 -2047,…,-1024,1024,…,2047 11

[0229] For example, run-length encoding can be used for the AC transformation coefficients of the input image. Among them, the encoding scheme for non-zero coefficients can be the same as the encoding method for the differential magnitude of the DC transformation coefficients.

[0230] Run - length coding: Encode the number R of zeros between the current non - zero coefficient and the previous non - zero coefficient. The maximum value of this number is 15, and then encode the current non - zero coefficient S. If the number of zeros is greater than 15, then encode (15, 0). When R is 16, it means that all the alternating transform coefficients from the current position to the last one are zeros. Entropy - encode the number R of consecutive zeros and the group number of coefficient S using tANS: Count the probabilities Pr of different values of R from 0 to 16, and represent the 1x17 probability vector with 17 8 - bit unsigned integers and write them into the bitstream; Count the probabilities Ps of different values of the group number of S from 0 to 11, and represent the 1x12 probability vector with 12 8 - bit unsigned integers and write them into the bitstream; Use ANS to perform entropy encoding on R and the group number of coefficient S according to Pr and Ps. The amplitude of coefficient S is represented by a fixed - length code with the number of bits in its group and written into the bitstream.

[0231] For example, the run - length coding can be performed on the alternating transform coefficients of the input image according to the table shown in Table 2. Table 2 shows the group number, the corresponding value range for each group, and the number of bits required for the values.

[0232] Table 2

[0233]

[0234]

[0235] S1005. Encode the position information into the bitstream.

[0236] Among them, the above - mentioned position information is used to indicate the position of the target area in the above - mentioned input image.

[0237] In a possible implementation, the position information can be mask map information (such as a binary mask map). Among them, the above - mentioned mask map information includes a first area, and the first area is used to indicate the target area. For example, the above - mentioned mask map can be composed of 0s and 1s, and the area where 1 is located is the target area.

[0238] In a possible implementation, the above - mentioned position information can be entropy - encoded into the above - mentioned bitstream.

[0239] Exemplarily, for an image of HxW, divided into 8x8 blocks, the height of the obtained binary mask map is ((H + 7) >> 3), and the width is ((W + 7) >> 3). Convert the mask map into a 1 - D 01 vector by raster scanning, and then perform entropy encoding on the 01 vector generated by the mask map to write it into the above - mentioned bitstream.

[0240] For example, the entropy coding can be performed on the 01 vector generated from the mask image by using the run length coding method. The number of 0 values before the non-zero value 1 is represented by an unsigned 8-bit number; a bitstream is generated from the 01 string organized by the number of 0s in 8 bits and 0 / 1 in 1 bit. (When the number of 0 values is greater than 255, 255 is binarized and written into the bitstream, and the 256th 0 is also written into the bitstream).

[0241] For another example, ANS can be used to perform entropy coding on the above 01 vector. The frequency P of 0 occurrences in the 01 vector is counted, and this frequency is used as the probability of 0. The probability P is represented by an 8-bit integer A and written into the bitstream. The probability of 1 is (256 - A) / 256. Based on the probability P, ANS is used to perform entropy coding on the 01 vector.

[0242] The method provided by the embodiments of the present application additionally performs quality enhancement on the local area of the image by using a traditional coding scheme, making up for the deficiency that the local distortion of the image quality is relatively large due to the capacity and generalization of the AI coding network. Compared with only performing AI coding on the input image, additionally performing quality enhancement on the local area of the image by using a traditional coding scheme can improve the overall image quality of AI image coding.

[0243] In a possible implementation manner, the transformation parameters can be encoded (entropy encoded) into the above bitstream. The transformation parameters are the transformation parameters used for the above transformation operation.

[0244] Figure 11 The flowchart of a coding method provided by the embodiments of the present application is shown. As Figure 11 shown, the coding method provided by the embodiments of the present application includes:

[0245] S1101. Perform entropy decoding on the bitstream to obtain a feature map.

[0246] For example, the bitstream can be subjected to ANS entropy decoding to obtain a feature map.

[0247] S1102. Perform AI decoding on the feature map to obtain a reconstructed image.

[0248] Among them, the above AI decoding can be any AI decoder.

[0249] For example, the feature map can be input into a JPEG AI decoder for AI decoding to obtain a reconstructed image.

[0250] S1103. Perform entropy decoding on the bitstream to obtain transformation coefficients.

[0251] Exemplarily, the direct current transformation coefficients and alternating current transformation coefficients in the bitstream can be entropy decoded respectively.

[0252] For example, for DC coefficient decoding, the probabilities P0 to P11 of the DC differential values in each group can be decoded from the bitstream. The numerical grouping is shown in Table 1. Using tANS entropy decoding on the group numbers according to the probabilities P0 to P11, the group number where the current DC differential value is located is obtained from the bitstream. According to the number of bits required for the values within this group number, the corresponding bits are read from the bitstream to obtain the differential value. The sum of this differential value and the DC component of the previous block gives the current DC coefficient value. The DC coefficients of all DCT blocks are obtained according to this process.

[0253] Also, for example, for AC coefficient decoding, the probabilities Pr of the number of 0s R between adjacent non-zero coefficients taking values from 0 to 16 can be decoded from the bitstream. This 1x17 probability vector is represented by 17 8-bit unsigned integers; the probabilities Ps of the AC coefficient S taking different values for group numbers from 0 to 11 are decoded from the bitstream. This 1x12 probability vector is represented by 12 8-bit unsigned integers; using ANS, R is entropy decoded from the bitstream according to Pr. When R is 16, all the AC coefficients from the current position to the 63rd one are 0; otherwise, using ANS, the group number of the coefficient S is entropy decoded from the bitstream according to Ps. According to the group where the coefficient S is located, the number of bits required for the amplitude of the coefficient S is obtained, and the corresponding bits are read from the bitstream to obtain the value of the AC coefficient S, and the first R positions before the current coefficient S are filled with 0. According to this process, the AC coefficients of all DCT blocks are obtained.

[0254] In a possible implementation, the bitstream can be entropy decoded to obtain the quantized transform coefficients; the quantized transform coefficients are inverse quantized.

[0255] S1104. Perform an inverse transform operation on the transform coefficients to obtain the reconstructed pixel values of the target region.

[0256] S1105. Entropy decode the bitstream to obtain the position information.

[0257] Among them, the above position information is used to indicate the position of the above target region, and the above transform coefficients are the coefficients used for the transform operation on the above target region.

[0258] In a possible implementation, the bitstream can be entropy decoded to obtain a binary mask map representing the position information.

[0259] Exemplarily, for an HxW image divided into 8x8 blocks, the height of the obtained binary mask map is ((H + 7) >> 3), and the width is ((W + 7) >> 3). The mask map is converted into a 1D 01 vector by raster scanning.

[0260] For example, the above 01 vector is decoded using the run length coding method. An unsigned 8-bit number represents the number of 0s before the non-zero value 1; the number of 0s and the 0 / 1 value are generated from the 01 string in the code stream in the order of (the number of 8-bit 0s, 1-bit 0 / 1), thereby obtaining the entire binary mask image.

[0261] For another example, ANS is used to perform entropy decoding on the above 01 vector. The frequency P of the occurrence of 0s in the 01 vector is decoded from the code stream, and this probability P is represented by an 8-bit integer A. The probability of 1 is (256 - A) / 256. Based on the probability P, ANS is used to perform entropy decoding on the 01 vector to obtain the binary mask image.

[0262] S1106. Determine the fused image according to the position information, the reconstructed pixel values of the target region, and the reconstructed image.

[0263] In a possible implementation manner, the target region in the reconstructed image may be determined according to the above position information; the pixel values of the target region in the reconstructed image are updated according to the above reconstructed pixel values to obtain the fused image.

[0264] Exemplarily, according to the binary mask image, the region corresponding to the enhancement layer image block in the reconstructed image may be determined, and the pixel values of the region corresponding to the enhancement layer image block in the reconstructed image are set to the reconstructed pixel values of the target region of the reconstructed image.

[0265] Next, the Figure 12 encoding apparatus for performing the above encoding method will be introduced.

[0266] It can be understood that, in order to implement the above functions, the encoding apparatus includes corresponding hardware and / or software modules for performing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0267] The embodiments of the present application may divide the functional modules of the encoding apparatus according to the above method examples. For example, each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module. The above integrated module may be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0268] In the case where each functional module is divided according to each function, Figure 12 FIG. shows a possible schematic composition of the encoding device involved in the above embodiment. The device may be an electronic device, or a module applied to an electronic device (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of the electronic device. As Figure 12 shown, the encoding device 1200 may include: an encoding unit 1201 and a transformation unit 1202.

[0269] The above encoding unit 1201 is configured to perform AI encoding on the input image to obtain a feature map.

[0270] The above encoding unit 1201 is further configured to perform entropy encoding on the above feature map to obtain a bitstream.

[0271] The above transformation unit 1202 is configured to perform a transformation operation on the target region of the above input image to obtain transformation coefficients.

[0272] The above encoding unit 1201 is further configured to encode the above transformation coefficients into the above bitstream.

[0273] The above encoding unit 1201 is further configured to encode position information into the above bitstream, and the position information is used to indicate the position number of the above target region in the above input image.

[0274] In a possible implementation manner, the above transformation unit 1202 is specifically configured to: divide the above target region into a plurality of image blocks; perform the above transformation operation on the above plurality of image blocks to obtain transformation coefficients.

[0275] In a possible implementation manner, the above encoding unit 1201 is specifically configured to: quantize the above transformation coefficients; encode the quantized transformation coefficients into the above bitstream.

[0276] Next, a decoding device for performing the above decoding method will be described in conjunction with Figure 13 introduce the decoding device for performing the above decoding method.

[0277] It can be understood that in order to implement the above functions, the decoding device includes corresponding hardware and / or software modules for performing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0278] The embodiments of the present application can divide the decoding device into functional modules according to the above method examples. For example, each functional module can be corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0279] In the case of dividing each functional module corresponding to each function, Figure 13 FIG. shows a possible schematic composition diagram of the decoding device involved in the above embodiment. This device can be an electronic device, or a module applied to an electronic device (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of the electronic device. As Figure 13 shown, the decoding device 1300 can include: a decoding unit 1301, a transformation unit 1302, and a fusion unit 1303.

[0280] The above decoding unit 1301 is used to perform entropy decoding on the bitstream to obtain a feature map.

[0281] The above decoding unit 1301 is further used to perform AI decoding on the above feature map to obtain a reconstructed image.

[0282] The above decoding unit 1301 is further used to perform entropy decoding on the above bitstream to obtain transform coefficients.

[0283] The above transformation unit 1302 is used to perform an inverse transformation operation on the above transform coefficients to obtain the reconstructed pixel values of the target region.

[0284] The above decoding unit 1301 is further used to perform entropy decoding on the above bitstream to obtain position information, and the above position information is used to indicate the position of the above target region.

[0285] The above fusion unit 1303 is used to determine a fused image according to the above position information, the reconstructed pixel values of the above target region, and the above reconstructed image.

[0286] In a possible implementation manner, the above fusion unit 1303 is specifically used to: determine the target region in the above reconstructed image according to the above position information; update the pixel values of the target region in the above reconstructed image according to the above reconstructed pixel values to obtain the above fused image.

[0287] In a possible implementation manner, the above decoding unit 1301 is specifically used to: perform entropy decoding on the above bitstream to obtain quantized transform coefficients; perform inverse quantization on the above quantized transform coefficients to obtain the above transform coefficients.

[0288] The embodiment of the present application further provides a chip, which can be the chip of the above encoding device or decoding device. Figure 14 Fig. 1400 shows a schematic structural diagram of a chip 1400. The chip 1400 includes one or more processors 1401 and an interface circuit 1402. Optionally, the above chip 1400 may further include a bus 1403.

[0289] The processor 1401 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above encoding and decoding method can be completed by the integrated logic circuit in the hardware of the processor 1401 or instructions in the form of software.

[0290] Optionally, the above processor 1401 may be a general-purpose processor, a digital signal processing (DSP) processor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0291] The interface circuit 1402 can be used for sending or receiving data, instructions, or information. The processor 1401 can use the data, instructions, or other information received by the interface circuit 1402 for processing, and can send the processed information through the interface circuit 1402.

[0292] Optionally, the chip further includes a memory, which can include a read-only memory and a random access memory, and provides operation instructions and data to the processor. A part of the memory may also include a non-volatile random access memory (NVRAM).

[0293] Optionally, the memory stores executable software modules or data structures, and the processor can execute corresponding operations by calling the operation instructions stored in the memory (the operation instructions can be stored in the operating system).

[0294] Optionally, the chip can be used in the encoding device or decoding device involved in the embodiment of the present application. Optionally, the interface circuit 1402 can be used to output the execution result of the processor 1401. For the encoding and decoding method provided by one or more embodiments of the present application, reference can be made to the foregoing respective embodiments, which will not be elaborated here.

[0295] It should be noted that the functions corresponding to the processor 1401 and the interface circuit 1402 can be implemented by hardware design, software design, or a combination of both, and there is no limitation here.

[0296] Figure 15 FIG. is a schematic structural diagram of an electronic device provided in an embodiment of the present application. The electronic device can be an encoding device or a decoding device, a chip or a functional module in the encoding device, or a chip or a functional module in the decoding device. As Figure 15 shown, the electronic device 1500 includes a processor 1501, a transceiver 1502, and a communication line 1503.

[0297] Among them, the processor 1501 is used to execute any step in the encoding and decoding method provided in the embodiment of the present application, and during the execution of any step in the encoding and decoding method provided in the embodiment of the present application, the transceiver 1502 and the communication line 1503 can be selectively called to complete the corresponding operations.

[0298] Furthermore, the electronic device 1500 may further include a memory 1504. Among them, the processor 1501, the memory 1504, and the transceiver 1502 can be connected through the communication line 1503.

[0299] Among them, the processor 1501 is a processor, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 1501 may also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.

[0300] The transceiver 1502 is used to communicate with other devices or other communication networks. The other communication networks can be Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc. The transceiver 1502 can be a module, a circuit, a transceiver, or any device capable of implementing communication.

[0301] The transceiver 1502 is mainly used for the transmission and reception of commands and information, etc., and may include a transmitter and a receiver, which are respectively used for the transmission and reception of commands and information, etc.; operations other than the transmission and reception of commands and information, etc. are implemented by the processor.

[0302] A communication line 1503 is used to transmit information between components included in the electronic device 1500.

[0303] In one design, the processor can be regarded as a logic circuit and the transceiver as an interface circuit.

[0304] A memory 1504 is used to store instructions. Among them, the instructions can be computer programs.

[0305] Among them, the memory 1504 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). The memory 1504 can also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, etc. It should be noted that the memories of the systems and methods described herein are intended to include but not be limited to these and any other suitable types of memories.

[0306] It should be noted that the memory 1504 can exist independently of the processor 1501 or be integrated with the processor 1501. The memory 1504 can be used to store instructions, program codes, or some data, etc. The memory 1504 can be located inside the electronic device 1500 or outside the electronic device 1500, without limitation. The processor 1501 is used to execute the instructions stored in the memory 1504 to implement the method provided in the above embodiments of the present application.

[0307] In one example, the processor 1501 can include one or more processors, such as Figure 15 the processor 0 and the processor 1 in

[0308] As an alternative implementation, the electronic device 1500 includes multiple processors. For example, in addition to Figure 15 the processor 1501 in

[0309] As an alternative implementation, the electronic device 1500 further includes an output device 1505 and an input device 1506. Exemplarily, the input device 1506 is a device such as a keyboard, a mouse, a microphone, or a joystick, and the output device 1505 is a device such as a display screen or a speaker.

[0310] It should be noted that the electronic device 1500 can be a chip system or a device with a Figure 15 similar structure in Figure 15 Among them, the chip system can be composed of chips or can include chips and other discrete devices. The actions, terms, etc. involved among the embodiments of the present application can be referred to each other without limitation. The message names or parameter names in the messages exchanged between the devices in the embodiments of the present application are only examples, and other names can also be adopted in specific implementations without limitation. In addition, Figure 15 the shown component structure does not constitute a limitation on the electronic device 1500. Except for Figure 15 the components shown, the electronic device 1500 can include more or fewer components than

[0311] The processors and transceivers described in this application can be implemented on integrated circuits (ICs), analog ICs, radio frequency integrated circuits, mixed-signal ICs, application specific integrated circuits (ASICs), printed circuit boards (PCBs), electronic devices, etc. The processors and transceivers can also be fabricated using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal-oxide-semiconductor (NMOS), positive channel metal oxide semiconductor (PMOS), bipolar junction transistor (BJT), BiCMOS, silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0312] Figure 16 FIG. 4 is a schematic structural diagram of another electronic device provided by an embodiment of this application. The electronic device can be an encoding device or a decoding device, a chip or a functional module in the encoding device, or a chip or a functional module in the decoding device. For ease of description, Figure 16 only the main components of the electronic device are shown, including a processor 1601, a memory 1602, a control circuit 1603, and an input / output device 1604. The processor 1601 is mainly used to process communication protocols and communication data, execute software programs, and process data of the software programs. The memory 1602 is mainly used to store software programs and data. The control circuit 1603 is mainly used for power supply and transmission of various electrical signals. The input / output device 1604 is mainly used to receive data input by the user and output data to the user.

[0313] When the electronic device is the processor 1601, the control circuit 1603 can be the motherboard. The memory 1602 includes media with storage functions such as hard disks, RAM, and ROM. The processor 1601 can include a baseband processor 1601 and a central processing unit. The baseband processor is mainly used to process communication protocols and communication data, and the central processing unit is mainly used to control the entire electronic device, execute software programs, and process data of software programs. The input / output device 1604 includes a display screen, a keyboard, a mouse, etc. The control circuit 1603 can further include or be connected to a transceiver circuit or transceiver, such as a network interface, etc., for sending or receiving data or signals, such as performing data transmission and communication with other devices. Further, it can also include an antenna for wireless signal transceiver for data / signal transmission with other devices.

[0314] An embodiment of the present application further provides an encoding device, which includes: at least one processor. When the at least one processor executes program code or instructions, the relevant method steps are implemented to realize the encoding method in the above embodiment.

[0315] Optionally, the device can further include at least one memory for storing the program code or instructions.

[0316] An embodiment of the present application further provides a decoding device, which includes: at least one processor. When the at least one processor executes program code or instructions, the relevant method steps are implemented to realize the decoding method in the above embodiment.

[0317] Optionally, the device can further include at least one memory for storing the program code or instructions.

[0318] An embodiment of the present application further provides a computer storage medium, in which computer instructions are stored. When the computer instructions run on a communication device, the communication device is caused to execute the relevant method steps to realize the encoding and decoding methods in the above embodiment.

[0319] An embodiment of the present application further provides a computer program product. When the computer program product runs on a computer, the computer is caused to execute the above relevant steps to realize the encoding and decoding methods in the above embodiment.

[0320] An embodiment of the present application further provides an encoding device, which can specifically be a chip, an integrated circuit, a component, or a module. Specifically, the device can include a processor connected to a memory for storing instructions, or the device includes at least one processor for obtaining instructions from an external memory. When the device runs, the processor can execute the instructions to cause the chip to execute the encoding method in each of the above method embodiments.

[0321] An embodiment of the present application further provides a decoding device, which may specifically be a chip, an integrated circuit, a component, or a module. Specifically, the device may include a processor connected to a memory for storing instructions, or the device includes at least one processor for obtaining instructions from an external memory. When the device runs, the processor may execute the instructions to cause the chip to execute the decoding method in each of the above method embodiments.

[0322] An embodiment of the present application further provides a method for storing a bitstream, the method including: obtaining and storing the bitstream obtained by the above encoding and decoding method.

[0323] An embodiment of the present application further provides a bitstream storage device for obtaining and storing the bitstream obtained by the above encoding and decoding method.

[0324] An embodiment of the present application further provides a method for transmitting a bitstream, the method including: obtaining and transmitting the bitstream obtained by the above encoding and decoding method.

[0325] An embodiment of the present application further provides a bitstream transmission device for obtaining and transmitting the bitstream obtained by the above encoding and decoding method.

[0326] An embodiment of the present application further provides a computer-readable storage medium having stored thereon the bitstream obtained by the above encoding and decoding method.

[0327] Please refer to Figure 17 , Figure 17 which schematically illustrates the general concept of processing by a neural network such as a convolutional neural network (CNN). A convolutional neural network consists of an input layer, an output layer, and multiple hidden layers. The input layer is the layer that provides the input (such as Figure 17 a part of the input image shown) for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that are convolved with multiplications or other dot products. The result of a layer is one or more feature maps (represented by empty solid rectangles), sometimes also referred to as channels. Resampling (such as subsampling) may be involved in some or all of the layers. As a result, the feature maps may become smaller, as Figure 17 shown. Note that convolution with a stride may also reduce the size of the input feature map (resampling). The activation function in a CNN is typically a ReLU (rectified linear unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, and normalization layers, which are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. Although these layers are commonly referred to as convolutions, this is just a convention. Mathematically, it is technically a sliding dot product or cross-correlation. This has important implications for the indices in the matrix as it affects how the weights are determined at a specific index point.

[0328] When programming a CNN for processing images, as Figure 17 shown, the input is a tensor of shape (number of images) x (image width) x (image height) x (image depth). It should be known that the image depth can be composed of the channels of the image. After passing through the convolutional layer, the image is abstracted into a feature map with the shape (number of images) x (feature map width) x (feature map height) x (feature map channels). The convolutional layer in the neural network should have the following properties. A convolutional kernel (hyperparameter) defined by the width and height. The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) should be equal to the number of channels (depth) of the input feature map.

[0329] Machine video coding (VCM) is another popular direction in computer science today. The main idea behind this approach is to transmit an encoded representation of the image or video information for further processing by computer vision (CV) algorithms, such as object segmentation, detection, and recognition. Compared with traditional image and video coding that aims at human perception, the quality feature is the performance of computer vision tasks, such as object detection accuracy, rather than the reconstruction quality. As Figure 18 shown.

[0330] Machine video coding, also known as collaborative intelligence, is a relatively new paradigm for the efficient deployment of deep neural networks in mobile cloud infrastructure. By partitioning the network between the mobile side 1810 and the cloud side 1890 (e.g., a cloud server), the computational workload can be distributed to minimize the total energy and / or latency of the system. Generally, collaborative intelligence is a paradigm in which the processing of a neural network is distributed among two or more different computing nodes; e.g., devices, but in general, any functionally defined node. Here, the term "node" does not refer to the above-mentioned neural network nodes. Instead, the (computing) nodes here refer to (physically or at least logically) independent devices / modules that implement parts of the neural network. Such devices can be different servers, different end-user devices, mixtures of servers and / or user devices and / or clouds and / or processors, etc. In other words, the computing nodes can be considered as nodes belonging to the same neural network and communicate with each other to transfer encoded data within / for the neural network. For example, in order to be able to perform complex computations, one or more layers can be executed on a first device (such as a device on the mobile side 1810), and one or more layers can be executed in another device (such as a cloud server on the cloud side 1890). However, the distribution can also be finer, and a single layer can be executed on multiple devices. In the present disclosure, the term "plurality" refers to two or more. In some existing solutions, a part of the neural network function is executed in a device (user device or edge device, etc.) or multiple such devices, and then the output (feature map) is passed to the cloud. The cloud is a collection of processing or computing systems located outside the device that is operating a part of the neural network. The concept of collaborative intelligence has also been extended to model training. In this case, data flows bidirectionally: from the cloud to the mobile device during the backpropagation of training, and from the mobile device to the cloud during the forward pass and inference of training (as Figure 18 shown).

[0331] Some work has proposed semantic image compression by encoding deep features and then reconstructing the input image from them. Compression based on uniform quantization followed by context-based adaptive arithmetic coding (CABAC) from H.264 is shown. In some scenarios, it may be more efficient to send the output of the hidden layer (deep feature map) from the mobile part 1810 to the cloud 1890 rather than sending compressed natural image data to the cloud and performing object detection using the reconstructed image. Therefore, it may be advantageous to compress the data (features) generated by the mobile side 1810, and the mobile side 1810 can include a quantization layer 1820 for this purpose. Accordingly, the cloud side 1890 can include an inverse quantization layer 1860. The efficient compression of feature maps is beneficial for image and video compression and reconstruction for human perception and machine vision. Entropy coding methods, such as arithmetic coding, are a popular method for compressing deep features (i.e., feature maps).

[0332] Please refer to Figure 19 , Figure 19 which shows a bitstream structure provided by an embodiment of the present application. As Figure 19 shown, the bitstream includes: Start of Image, File Header, Entropy Encoded Data, and End of Image.

[0333] In a possible implementation, the above position information, the above transformation coefficients, Gaussian distribution parameter index numbers, Gaussian distribution parameters, target probability distributions, and target probability distribution index numbers may be stored in the file header.

[0334] In a possible implementation, the target quality residual matrix or the target quality matrix may be stored in the entropy encoded data.

[0335] Please refer to Figure 20 , Figure 20 which shows a schematic diagram of hierarchical coding provided by an embodiment of the present application.

[0336] During the video network transmission process, especially in the process of real-time multi-user communication over the Internet, there are certain differences in the network bandwidth and device processing capabilities among multiple users. This differential network bandwidth resource poses a need for adaptive bitrate regulation for different users. Hierarchical coding proposes the concepts of temporal, spatial, and quality scalability. By adding the information of the enhancement layer on the basis of the base layer, video content with higher frame rate, resolution, and quality can be obtained. Different users can select whether to require the bitstream of the enhancement layer to match their respective electronic device processing capabilities and network bandwidths.

[0337] Figure 20 is a schematic diagram of a hierarchical coding scheme provided by an embodiment of the present application. As Figure 20 shown. The hierarchical coding scheme includes: performing AI coding on the input image to obtain a feature map; performing entropy coding on the feature map to obtain a bitstream; writing the position information for recording the target area of the input image into the bitstream; performing a transformation operation on the target area of the input image; and writing the transformation coefficients used in the transformation operation into the bitstream.

[0338] Please refer to Figure 21 , Figure 21 which shows an end-to-end image coding based on a hyperprior structure provided by an embodiment of the present application.

[0339] As deep learning has shown excellent performance in various fields, researchers have proposed an end-to-end image coding scheme based on deep learning. A typical coding framework is as Figure 21 shown, and the specific scheme is as follows:

[0340] At the encoding end, the original image input feature extraction module, i.e., the Ga module, outputs a feature map, which is input into the quantization module to obtain a quantized feature map. On the one hand, the feature map passes through the side information extraction module, i.e., the Ha module, to output side information z. The side information z is input into the quantization module for quantization to obtain [z'], which will be entropy encoded and written into the bitstream. The encoding end performs entropy decoding operation to obtain the decoded [z'], which is input into the probability estimation module, i.e., the Hs module, to output the probability distribution of each feature element of the feature map (entropy encoding and then decoding is to ensure encoding and decoding synchronization). On the other hand, the entropy encoding module performs entropy encoding on each feature element in the feature map according to the probability distribution of each feature element of [z'], to obtain a compressed bitstream. Among them, the side information is also a kind of feature information, represented as a three-dimensional feature map, and the number of feature elements it contains is less than that of the feature map. Both the feature map [x] and the feature map [y] are MxWxH, where M is the number of channels, the same as the number of channels of the last convolution layer, and WxH is related to the width and height of the input image and the downsampling ratio (i.e., stride) set for each convolution operation in the network. In this figure, it means that both the width and height of the input image are downsampled 4 times with a ratio of 2.

[0341] At the decoding end, first decode to obtain the side information, and then input the side information into the Hs module to output the probability distribution of the symbol to be decoded. The entropy decoding module performs arithmetic decoding on each feature element in the feature map according to the probability distribution of each feature element in [z'], to obtain the value of each feature element in the feature map, so as to obtain the decoded feature map. The decoded feature map is input into the image reconstruction module, i.e., the Gs module, to output the reconstructed image.

[0342] It should be noted that, in Figure 2 the scheme shown in Figure 21 the input image contains three RGB channels, and the R, G, and B components have the same W and H. The image data with a size of 3xWxH can be input into the network for encoding and decoding. When the input image is in YUV444 format, the network shown in

[0343] can also be used for encoding and decoding.

[0343] Among them, the device, computer storage medium, computer program product, or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects they can achieve can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0344] It should be understood that in various embodiments of this application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of this application.

[0345] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of this application.

[0346] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0347] In several embodiments provided in the embodiments of this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0348] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0349] In addition, the functional units in the various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0350] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application, in essence, or the parts that contribute to the prior art, or parts of the technical solutions, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0351] As described above, the above is only the specific implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the embodiments of the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.

Claims

1. A coding method, characterized in that, Comprising: Performing artificial intelligence (AI) encoding on an input image to obtain a feature map; Performing entropy encoding on the feature map to obtain a bitstream; Performing a transformation operation on a target region of the input image to obtain transformation coefficients; Encoding the transformation coefficients into the bitstream; Encoding position information into the bitstream, where the position information is used to indicate the position of the target region in the input image.

2. The method according to claim 1, wherein The performing a transformation operation on a target region of the input image to obtain transformation coefficients includes: Dividing the target region into a plurality of image blocks; Performing the transformation operation on the plurality of image blocks to obtain transformation coefficients.

3. The method according to claim 1 or 2, characterized in that, The encoding the transformation coefficients into the bitstream includes: Quantizing the transformation coefficients; Encoding the quantized transformation coefficients into the bitstream.

4. The method according to any one of claims 1 to 3, characterized in that The transformation operation includes at least one of a discrete cosine transform (DCT), a discrete Fourier transform (DFT), or a discrete wavelet transform (DWT).

5. The method according to any one of claims 1 to 4, characterized in that The position information is mask map information, and the mask map information includes a first region, where the first region is used to indicate the target region.

6. A decoding method, characterized in that, Comprising: Performing entropy decoding on the bitstream to obtain a feature map; Performing AI decoding on the feature map to obtain a reconstructed image; Performing entropy decoding on the bitstream to obtain transformation coefficients; Performing an inverse transformation operation on the transformation coefficients to obtain reconstructed pixel values of the target region; Performing entropy decoding on the bitstream to obtain position information, where the position information is used to indicate the position of the target region; Determining a fused image based on the position information, the reconstructed pixel values of the target region, and the reconstructed image.

7. The method according to claim 6, wherein The determining a fused image based on the position information, the reconstructed pixel values of the target region, and the reconstructed image includes: Determining the target region in the reconstructed image according to the position information; Updating the pixel values of the target region in the reconstructed image according to the reconstructed pixel values to obtain the fused image.

8. The method according to claim 6 or 7, characterized in that, The performing entropy decoding on the bitstream to obtain transformation coefficients includes: Performing entropy decoding on the bitstream to obtain quantized transformation coefficients; Performing inverse quantization on the quantized transformation coefficients to obtain the transformation coefficients.

9. An encoding device, characterized in that, Comprising an encoding unit and a transformation unit; The encoding unit is configured to perform AI encoding on an input image to obtain a feature map; The encoding unit is further configured to perform entropy encoding on the feature map to obtain a bitstream; The transformation unit is configured to perform a transformation operation on a target region of the input image to obtain transformation coefficients; The encoding unit is further configured to encode the transformation coefficients into the bitstream; The encoding unit is further configured to encode position information into the bitstream, where the position information is used to indicate the position of the target region in the input image.

10. The device according to claim 9, characterized in that, The transformation unit is specifically configured to: Divide the target region into a plurality of image blocks; Perform the transformation operation on the plurality of image blocks to obtain transformation coefficients.

11. The device according to claim 9 or 10, characterized in that, The encoding unit is specifically configured to: Quantize the transformation coefficients; Encode the quantized transformation coefficients into the bitstream.

12. A decoding device, characterized in that, Comprising a decoding unit, a transformation unit, and a fusion unit; The decoding unit is configured to perform entropy decoding on the bitstream to obtain a feature map; The decoding unit is further configured to perform AI decoding on the feature map to obtain a reconstructed image; The decoding unit is further configured to perform entropy decoding on the bitstream to obtain transform coefficients; The transformation unit is configured to perform an inverse transformation operation on the transform coefficients to obtain the reconstructed pixel values of the target region; The decoding unit is further configured to perform entropy decoding on the bitstream to obtain position information, where the position information is used to indicate the position of the target region; The fusion unit is configured to determine a fused image according to the position information, the reconstructed pixel values of the target region, and the reconstructed image.

13. The device according to claim 12, characterized in that Specifically, the fusion unit is configured to: Determine the target region in the reconstructed image according to the position information; Update the pixel values of the target region in the reconstructed image according to the reconstructed pixel values to obtain the fused image.

14. The device according to claim 12 or 13, characterized in that, Specifically, the decoding unit is configured to: Perform entropy decoding on the bitstream to obtain quantized transform coefficients; Perform inverse quantization on the quantized transform coefficients to obtain the transform coefficients.

15. An encoding device, characterized in that, Comprising at least one processor and a memory, the at least one processor executes a program or instructions stored in the memory, so that the encoding device implements the method according to any one of claims 1 to 5 above.

16. A decoding device, characterized in that, Comprising at least one processor and a memory, the at least one processor executes a program or instructions stored in the memory, so that the decoding device implements the method according to any one of claims 6 to 8 above.

17. A method for storing a bitstream, characterized in that, Comprising: Obtain and store the bitstream obtained by the method according to any one of claims 1 to 8.

18. A bitstream storage device, characterized in that, The device is configured to obtain and store the bitstream obtained by the method according to any one of claims 1 to 8.

19. A bitstream transmission method, characterized in that, Comprising: Obtain and transmit the bitstream obtained by the method according to any one of claims 1 to 8.

20. A bitstream transmission device, characterized in that, The device is configured to obtain and transmit the bitstream obtained by the method according to any one of claims 1 to 8.

21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores the bitstream obtained by the method according to any one of claims 1 to 8.

22. A computer-readable storage medium, characterized in that, For storing a computer program, when the computer program runs on a computer or a processor, the computer or the processor implements the method according to any one of claims 1 to 8 above.

23. A computer program product, characterized in that, The computer program product contains instructions, when the instructions run on a computer or a processor, the computer or the processor implements the method according to any one of claims 1 to 8 above.

Citation Information

Cited By

  • Encoding method and device, and decoding method and device

    WO2025152470A1