Coding method and apparatus and decoding method and apparatus

By using the indication of reference image queue information in hierarchical encoding, the encoding and codec performance is optimized, and the problem of insufficient performance in video encoding is solved, and more efficient video encoding adaptability is achieved.

WO2025148890A1PCT designated stage expired Publication Date: 2025-07-17HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/071116
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2025-01-07
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The existing airspace hierarchical encoding scheme has not been fully optimized in video encoding, resulting in poor encoding and decoding performance and it is difficult to adapt to the differences between diverse network conditions and terminal equipment.

Method used

By using the first indication information in the reference image queue information in the hierarchical encoding scene, the reference image decoding order index of the current hierarchical encoding image is determined, the number of bits is saved, and the encoding and decoding performance is improved.

Benefits of technology

The encoding and codec performance in hierarchical encoding scenarios is improved, and the processing capabilities of different network conditions and terminal devices are adapted to achieve more efficient video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071116_17072025_PF_FP_ABST
    Figure CN2025071116_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a coding method and apparatus and a decoding method and apparatus, relating to the technical field of media, capable of improving the coding and decoding performance in layered coding scenarios. The decoding method comprises: parsing a code stream to obtain a current layered coded image and reference image queue information of the current layered coded image, wherein the reference image queue information comprises first indication information, and the first indication information is used for indicating a decoding sequence index of a reference image of the current layered coded image; when the decoding sequence index of the reference image is equal to the decoding sequence index of the current layered coded image, adding an image in a decoded image buffer area that has the same reference layer identifier as the current layered coded image and the same first indication information as the current layered coded image to a reference image queue of the current layered coded image; and decoding the current layered coded image on the basis of the reference image queue.
Need to check novelty before this filing date? Find Prior Art

Description

A coding and decoding method and device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 10, 2024, with application number 202410041913.1 and application name “A coding and decoding method and device”. This application also claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 8, 2024, with application number 202410178188.2 and application name “A coding and decoding method and device”. The entire contents of these patent applications are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of media technology, and in particular to a coding and decoding method and device. Background Art

[0003] With the rapid development of network infrastructure, storage capacity, and computing power, the scope of video applications has gradually expanded, and the processing capabilities and resolutions of connected terminal devices vary greatly. Traditional video encoding generally uses transcoding and simulcasting to improve the compatibility of the bitstream with network bandwidth, terminal processing capabilities, and user needs.

[0004] Currently, a design scheme for spatial layered coding (SVC) has been proposed. This allows a single encoding to generate multiple video streams with varying resolutions, adapting to diverse network conditions, terminal devices, and user needs. In security applications, spatial layering not only enables video sources to be viewed on multiple devices, from high-definition monitoring equipment to mobile phones, but also allows for the selection of appropriate resolutions for storage and archiving.

[0005] There are still technical points in the design of spatial layered coding that need to be optimized and improved. Summary of the Invention

[0006] The present application provides a coding and decoding method and apparatus, which can improve coding and decoding performance in layered coding scenarios.

[0007] This application adopts the following technical solutions:

[0008] In a first aspect, the present application provides a decoding method, comprising: parsing a bitstream to obtain a current layered coded image and reference image queue information of the current layered coded image; the reference image queue information includes first indication information; the first indication information is used to indicate a decoding order index DOI of a reference image of the current layered coded image; wherein, when the decoding order index of the reference image is equal to the decoding order index of the current layered coded image, an image in a decoding image buffer having the same reference layer identifier as the current layered coded image and the same first indication information as the current layered coded image is added to the reference image queue of the current layered coded image; and based on the reference image queue, the current layered coded image is decoded.

[0009] In this application, in the layered coding scenario, the inter-layer reference frame does not need to be specifically identified. Instead, the inter-layer reference frame is located by using the fact that the reference frame DOI is the same as the current image DOI, plus the reference layer number transmitted in the sequence parameter set SPS. This saves bits and improves the encoding and decoding performance in the layered coding scenario.

[0010] In a possible implementation, the code stream includes multiple layered coded images of the current image.

[0011] In one possible implementation, the distance index of the 0th image in the reference image queue of the current layered coded image satisfies:

[0012] DistanceIndex=2×(POI_Currlayer-1), where DistanceIndex represents the distance index, and POI_Currlayer represents the display order index of the current layered coded image.

[0013] In one possible implementation, if the current image uses a knowledge image, each of the multiple layered coded images has a knowledge image.

[0014] In a possible implementation, when the knowledge image is a non-display knowledge image, the indexes of the image blocks of the multiple layered coded images of the current image included in each access unit in the code stream are the same.

[0015] In a possible implementation, when the knowledge image is a non-display knowledge image, each access unit in the code stream includes the same number of image blocks of multiple layered coded images of the current image.

[0016] In one possible implementation, the decoding method provided in the present application further includes: parsing the code stream to obtain second indication information and third indication information; wherein the second indication information is used to indicate whether the current image uses rights protection; the third indication information is used to indicate the rights protection mode of the current image, and the rights protection mode includes any one of the following: layered independent coding mode, layered reference coding mode or single-layer coding mode.

[0017] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights and the third indication information indicates that the rights protection mode is a layered independent coding mode or a layered reference coding mode, the number of multiple layered coded images of the current image is 2 or 4.

[0018] In a possible implementation, when the second indication information indicates that the current image is protected by usage rights, and the third indication information indicates that the rights protection mode is a layered independent coding mode, there is no inter-layer dependency among the multiple layered coded images.

[0019] In a possible implementation, the value of the fourth indication information in the code stream is a first value, and the first value is used to indicate that there is no inter-layer dependency among the multiple layered coded images.

[0020] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights and the third indication information indicates that the permission protection mode is a layered independent coding mode or a layered reference coding mode, the layered coded image with an even layer identifier is a low-authority layer, and the layered coded image with an odd layer identifier is a high-authority layer.

[0021] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights, the third indication information indicates that the rights protection mode is a layered reference coding mode, and the number of multiple layered coded images is 2, the layered coded images with even-numbered layer identifiers in the multiple layered coded images are independently encoded, and the layered coded images with odd-numbered layer identifiers are not independently encoded.

[0022] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights, the third indication information indicates that the rights protection mode is a layered reference coding mode, and the number of multiple layered coded images is 4, the layered coded images with even-numbered layer identifiers among the multiple layered coded images are independently encoded; the layered coded images with odd-numbered layer identifiers among the multiple layered coded images are not independently encoded, and the layered coded images with odd-numbered layer identifiers are dependent on the layered coded images with even-numbered layer identifiers.

[0023] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights and the third indication information indicates that the rights protection mode is a layered reference coding mode, the resolution of the current layered coded image among multiple layered coded images is the same as the resolution of the reference image of the current layered coded image.

[0024] In a second aspect, the present application provides a coding method, comprising: encoding multiple layered images of a current image to obtain multiple layered coded images; writing the multiple layered coded images and reference image queue information into a bitstream; the reference image queue information includes first indication information; the first indication information is used to indicate a decoding order index of a reference image of a layered coded image, and when the reference image is an inter-layer reference image of the current layered coded image, the decoding order index of the reference image is equal to the decoding order index of the layered coded image.

[0025] In this application, in the layered coding scenario, the inter-layer reference frame does not need to be specifically identified. Instead, the inter-layer reference frame is located by using the fact that the reference frame DOI is the same as the current image DOI, plus the reference layer number transmitted in the SPS. This saves bits and improves the encoding and decoding performance in the layered coding scenario.

[0026] In one possible implementation, the distance index of the 0th image in the reference image queue of the layered coded image satisfies:

[0027] DistanceIndex=2×(POI_Currlayer-1), where DistanceIndex represents the distance index, and POI_Currlayer represents the display order index of the layered coded image.

[0028] In one possible implementation, if the current image uses a knowledge image, each of the multiple layered coded images has a knowledge image.

[0029] In a possible implementation, when the knowledge image is a non-display knowledge image, the indexes of the image blocks of the multiple layered coded images of the current image included in each access unit in the code stream are the same.

[0030] In a possible implementation, when the knowledge image is a non-display knowledge image, each access unit in the code stream includes the same number of image blocks of multiple layered coded images of the current image.

[0031] In one possible implementation, the code stream also includes: second indication information and / or third indication information; wherein the second indication information is used to indicate whether the current image uses rights protection; the third indication information is used to indicate the rights protection mode of the current image, and the rights protection mode includes any one of the following: layered independent coding mode, layered reference coding mode or single-layer coding mode.

[0032] In one possible implementation, the code stream also includes fourth indication information and / or fifth indication information; the fourth indication information is used to indicate whether multiple layered coded images of the current image have inter-layer dependencies; the fifth indication information is used to indicate whether a layered coded image is independently coded.

[0033] When the second indication information indicates that the current image is protected by usage rights and the third indication information indicates that the rights protection mode is layered independent coding mode or layered reference coding mode, the number of multiple layered coded images of the current image is 2 or 4.

[0034] In a possible implementation, when the second indication information indicates that the current image is protected by usage rights and the third indication information indicates that the rights protection mode is a layered independent coding mode, the fourth indication information indicates that there is no inter-layer dependency among the multiple layered coded images.

[0035] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights and the third indication information indicates that the permission protection mode is a layered independent coding mode or a layered reference coding mode, the layered coded image with an even layer identifier is a low-authority layer, and the layered coded image with an odd layer identifier is a high-authority layer.

[0036] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights, the third indication information indicates that the rights protection mode is a layered reference coding mode, and the number of multiple layered coded images is 2, the fifth indication information indicates that the layered coded images with an even layer identifier among the multiple layered coded images are independently encoded.

[0037] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights, the third indication information indicates that the rights protection mode is a layered reference coding mode, and the number of multiple layered coded images is 4, the fifth indication information indicates that the layered coded images with even-numbered layer identifiers among the multiple layered coded images are independently encoded, and the layered coded images with odd-numbered layer identifiers among the multiple layered coded images are not independently encoded, and the layered coded images with odd-numbered layer identifiers are dependent on the layered coded images with even-numbered layer identifiers.

[0038] In one possible implementation, when the second indication information indicates that the current image is protected by usage rights and the third indication information indicates that the rights protection mode is a layered reference coding mode, a layered coded image among multiple layered coded images has the same resolution as a reference image of the layered coded image.

[0039] In a third aspect, the present application provides a decoding device, comprising modules for implementing the method described in the first aspect and any one of its possible implementations. The encoding device has the functionality to implement the behaviors described in the method examples of any one of the first aspect and its possible implementations. The functionality may be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the functionality described above.

[0040] In a fourth aspect, the present application provides an encoding device, comprising modules for implementing the method described in the second aspect and any one of its possible implementations. The encoding device has the functionality to implement the behaviors described in the method examples of any one of the second aspect and its possible implementations. The functionality may be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the functionality described above.

[0041] In a fifth aspect, the present application provides an image processing device comprising at least one processor and a memory, wherein the at least one processor executes a program or instruction stored in the memory so that the image processing device implements the method described in the first aspect, the second aspect or any possible implementation thereof.

[0042] Optionally, in the present application, the above-mentioned image processing device can be an encoding device or a decoding device, or the image processing device is a part of the encoding device or the decoding device, which is not specifically limited.

[0043] In a sixth aspect, the present application also provides a computer-readable storage medium for storing a computer program, which includes methods for implementing the above-mentioned first aspect, second aspect or any possible implementation thereof.

[0044] In a seventh aspect, the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to implement the method described in the first aspect, the second aspect or any possible implementation thereof.

[0045] In an eighth aspect, the present application further provides a chip comprising: an input interface, an output interface, and at least one processor. Optionally, the chip further comprises a memory. The at least one processor is configured to execute code in the memory. When the at least one processor executes the code, the chip implements the method described in the first aspect, the second aspect, or any possible implementation thereof.

[0046] Optionally, the chip may also be an integrated circuit.

[0047] The encoding device, decoding device, image processing device, computer storage medium, computer program product and chip provided in this application are all used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the method provided above and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] FIG1a is an exemplary block diagram of a decoding system provided in an embodiment of the present application;

[0049] FIG1b is an exemplary block diagram of a video decoding system provided in an embodiment of the present application;

[0050] FIG2 is an exemplary block diagram of a video encoder provided in an embodiment of the present application;

[0051] FIG3 is an exemplary block diagram of a video decoder provided in an embodiment of the present application;

[0052] FIG4 is an exemplary schematic diagram of candidate image blocks provided in an embodiment of the present application;

[0053] FIG5 is an exemplary block diagram of a video decoding device provided in an embodiment of the present application;

[0054] FIG6 is an exemplary block diagram of a device provided in an embodiment of the present application;

[0055] FIG7 is a schematic diagram of a layered coding framework provided in an embodiment of the present application;

[0056] FIG8 is a schematic diagram of a flow chart of an encoding method provided in an embodiment of the present application;

[0057] FIG9 is a schematic diagram of a coding structure of an LTR mode and a CRR slice mode provided in an embodiment of the present application;

[0058] FIG10 is a schematic diagram of a coding structure of a dual-layer LTR and CRR slicing mode provided in an embodiment of the present application;

[0059] FIG11 is a schematic diagram of a code stream structure in an LTR mode and a CRR slice mode provided in an embodiment of the present application;

[0060] FIG12 is a flowchart of an algorithm for a privacy protection mode of mode 0 provided in an embodiment of the present application;

[0061] FIG13 is an algorithm flow chart of a privacy protection mode 1 provided in an embodiment of the present application;

[0062] FIG14 is a schematic diagram of a flow chart of a decoding method provided in an embodiment of the present application;

[0063] FIG15 is a schematic structural diagram of an encoding device provided in an embodiment of the present application;

[0064] FIG16 is a schematic structural diagram of a decoding device provided in an embodiment of the present application;

[0065] FIG17 is a schematic diagram of the structure of a chip provided in an embodiment of the present application;

[0066] FIG18 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0067] FIG19 is a schematic structural diagram of an image processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0068] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0069] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0070] In the description of the embodiments of the present application, unless otherwise specified, “plurality” means two or more.

[0071] Data encoding and decoding includes two parts: data encoding and data decoding. Data encoding is performed on the source side (or commonly referred to as the encoder side), and generally includes processing (e.g., compressing) the original data to reduce the amount of data required to represent the original data (thereby more efficiently storing and / or transmitting). Data decoding is performed on the destination side (or commonly referred to as the decoder side), and generally includes inverse processing relative to the encoder side to reconstruct the original data. The "encoding and decoding" of the data involved in the embodiments of the present application should be understood as the "encoding" or "decoding" of the data. The encoding part and the decoding part are also collectively referred to as encoding and decoding (encoding and decoding, CODEC).

[0072] In the case of lossless data encoding, the original data can be reconstructed, that is, the reconstructed original data has the same quality as the original data (assuming there is no transmission loss or other data loss during storage or transmission). In the case of lossy data encoding, further compression is performed through quantization, etc. to reduce the amount of data required to represent the original data, but the decoder side cannot fully reconstruct the original data, that is, the quality of the reconstructed original data is lower or worse than the quality of the original data.

[0073] The embodiments of the present application can be applied to video data and other data with compression / decompression requirements. The following uses the encoding of video data (referred to as video encoding) as an example to illustrate the embodiments of the present application. Other types of data (such as image data, audio data, integer data, and other data with compression / decompression requirements) can refer to the following description, and the embodiments of the present application will not be repeated here. It should be noted that, compared to video encoding, the encoding process of data such as audio data and integer data does not require data to be divided into blocks, but the data can be directly encoded.

[0074] Video coding generally refers to processing a sequence of images to form a video or video sequence. In the field of video coding, the terms "picture", "frame" or "image" can be used as synonyms.

[0075] Several video coding standards fall under the category of "lossy hybrid video codecs" (i.e., combining spatial and temporal prediction in the pixel domain with 2D transform coding in the transform domain for applying quantization). Each image in a video sequence is typically divided into a set of non-overlapping blocks, which are typically coded at the block level. In other words, the encoder typically processes, i.e., encodes, the video at the block (video block) level, for example, by generating a prediction block through spatial (intra-frame) and temporal (inter-frame) prediction; subtracting the prediction block from the current block (currently processed / to-be-processed block) to obtain a residual block; transforming and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compressed), while the decoder applies the inverse of the encoder's processing to the coded or compressed block to reconstruct the current block for representation. Furthermore, the encoder needs to repeat the decoder's processing steps so that the encoder and decoder generate the same predictions (e.g., intra-frame predictions and inter-frame predictions) and / or reconstructed pixels for processing, i.e., encoding, the subsequent block.

[0076] In the following embodiment of the decoding system 10 , the encoder 20 and the decoder 30 are described with reference to FIG. 1 a to FIG. 3 .

[0077] FIG1a is an exemplary block diagram of a decoding system 10 provided in an embodiment of the present application, such as a video decoding system 10 (or simply, decoding system 10) that can utilize the techniques of the embodiments of the present application. The video encoder 20 (or simply, encoder 20) and video decoder 30 (or simply, decoder 30) in the video decoding system 10 represent devices that can be used to perform various techniques according to the various examples described in the embodiments of the present application.

[0078] As shown in FIG. 1 a , a decoding system 10 includes a source device 12 for providing encoded image data 21 such as an encoded image to a destination device 14 for decoding the encoded image data 21 .

[0079] The source device 12 includes an encoder 20 , and optionally, may include an image source 16 , a preprocessor (or preprocessing unit) 18 such as an image preprocessor, and a communication interface (or communication unit) 22 .

[0080] Image source 16 may include or may be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer-animated images, or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage for storing any of the above images.

[0081] In order to distinguish the processing performed by the pre-processor (or pre-processing unit) 18 , the image (or image data) 17 may also be referred to as a raw image (or raw image data) 17 .

[0082] The preprocessor 18 is configured to receive raw image data 17 and preprocess the raw image data 17 to obtain a preprocessed image (or preprocessed image data) 19. For example, the preprocessing performed by the preprocessor 18 may include cropping, color format conversion (e.g., from RGB to YCbCr), color grading, or denoising. It will be appreciated that the preprocessor 18 may be an optional component.

[0083] The video encoder (or encoder) 20 is used to receive the pre-processed image data 19 and provide encoded image data 21 (which will be further described below with reference to FIG. 2 and the like).

[0084] The communication interface 22 in the source device 12 can be used to receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device through the communication channel 13 for storage or direct reconstruction.

[0085] The destination device 14 includes a decoder 30 and, in addition or alternatively, may include a communication interface (or communication unit) 28 , a post-processor (or post-processing unit) 32 , and a display device 34 .

[0086] The communication interface 28 in the destination device 14 is used to receive the encoded image data 21 (or any other processed version) directly from the source device 12 or from any other source device such as a storage device, for example, the storage device is a encoded image data storage device, and provide the encoded image data 21 to the decoder 30.

[0087] The communication interface 22 and the communication interface 28 can be used to send or receive encoded image data (or encoded data) 21 through a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any combination thereof.

[0088] For example, the communication interface 22 may be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transmission coding or processing for transmission over a communication link or network.

[0089] The communication interface 28 corresponds to the communication interface 22 , and can be used, for example, to receive transmission data and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain the encoded image data 21 .

[0090] Both the communication interface 22 and the communication interface 28 can be configured as a unidirectional communication interface as indicated by the arrow pointing from the source device 12 to the corresponding communication channel 13 of the destination device 14 in Figure 1a, or a bidirectional communication interface, and can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded image data transmission, etc.

[0091] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31 (which will be further described below with reference to FIG. 3 and the like).

[0092] The post-processor 32 is configured to post-process the decoded image data 31 (also referred to as reconstructed image data) such as the decoded image to obtain post-processed image data 33 such as the post-processed image. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color grading, cropping, or resampling, or any other processing for generating the decoded image data 31 for display on a display device 34 or the like.

[0093] The display device 34 is configured to receive the post-processed image data 33 and display the image to a user or viewer. The display device 34 may be or include any type of display for displaying the reconstructed image, such as an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display screen.

[0094] The decoding system 10 also includes a training engine 25, which is used to train the encoder 20 (especially the entropy encoding unit 270 in the encoder 20) or the decoder 30 (especially the entropy decoding unit 304 in the decoder 30) to perform entropy encoding on the image block to be encoded based on the estimated probability distribution. For a detailed description of the training engine 25, please refer to the following method test example.

[0095] Although FIG1a shows source device 12 and destination device 14 as separate devices, device embodiments may also include both source device 12 and destination device 14 or the functions of both source device 12 and destination device 14, that is, both source device 12 or the corresponding functions and destination device 14 or the corresponding functions. In these embodiments, source device 12 or the corresponding functions and destination device 14 or the corresponding functions may be implemented using the same hardware and / or software or through separate hardware and / or software or any combination thereof.

[0096] According to the description, the existence and (accurate) division of different units or functions in the source device 12 and / or the destination device 14 shown in FIG. 1 a may vary depending on actual devices and applications, which is obvious to those skilled in the art.

[0097] Please refer to Figure 1b, which is an exemplary block diagram of a video decoding system 40 provided in an embodiment of the present application. The encoder 20 (e.g., video encoder 20) or the decoder 30 (e.g., video decoder 30), or both, can be implemented by processing circuitry in the video decoding system 40 shown in Figure 1b, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding-specific processors, or any combination thereof. Please refer to Figures 2 and 3, Figure 2 is an exemplary block diagram of a video encoder provided in an embodiment of the present application, and Figure 3 is an exemplary block diagram of a video decoder provided in an embodiment of the present application. The encoder 20 can be implemented by processing circuitry 46 to include the various modules discussed with reference to the encoder 20 in Figure 2 and / or any other encoder systems or subsystems described herein. The decoder 30 can be implemented by processing circuitry 46 to include the various modules discussed with reference to the decoder 30 in Figure 3 and / or any other decoder systems or subsystems described herein. The processing circuitry 46 described above can be used to perform the various operations discussed below. As shown in Figure 5, if part of the technology is implemented in software, the device can store the software instructions in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware, thereby performing the technology of the embodiment of the present application. One of the video encoder 20 and the video decoder 30 can be integrated into a single device as part of a combined codec (encoder / decoder, CODEC), as shown in Figure 1b.

[0098] The source device 12 and the destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as a notebook computer or laptop, a mobile phone, a smart phone, a tablet or a tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, and a monitoring device, etc., and may not use or use any type of operating system. The source device 12 and the destination device 14 may also be devices in a cloud computing scenario, such as a virtual machine in a cloud computing scenario. In some cases, the source device 12 and the destination device 14 may be equipped with components for wireless communication. Therefore, the source device 12 and the destination device 14 may be wireless communication devices.

[0099] The source device 12 and the destination device 14 may be installed with virtual scene applications (APPs) such as virtual reality (VR), augmented reality (AR), or mixed reality (MR), and may run the VR, AR, or MR applications based on user operations (e.g., click, touch, slide, shake, voice control, etc.). The source device 12 and the destination device 14 may capture images / videos of any objects in the environment through cameras and / or sensors, and then display virtual objects on a display device based on the captured images / videos. The virtual objects may be virtual objects in the VR, AR, or MR scenes (i.e., objects in the virtual environment).

[0100] It should be noted that in the embodiment of the present application, the virtual scene application in the source device 12 and the destination device 14 can be an application built into the source device 12 and the destination device 14 themselves, or it can be an application provided by a third-party service provider and installed by the user. There is no specific limitation on this.

[0101] In addition, the source device 12 and the destination device 14 may be installed with a real-time video transmission application, such as a live broadcast application. The source device 12 and the destination device 14 may capture images / videos through cameras and then display the captured images / videos on a display device.

[0102] In some cases, the video decoding system 10 shown in FIG1a is merely exemplary, and the techniques provided in embodiments of the present application may be applicable to video encoding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between the encoding device and the decoding device. In other examples, data is retrieved from local memory, sent over a network, and so on. The video encoding device can encode the data and store the data in memory, and / or the video decoding device can retrieve the data from memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to memory and / or retrieve and decode data from memory.

[0103] Please refer to Figure 1b, which is an exemplary block diagram of a video decoding system 40 provided in an embodiment of the present application. As shown in Figure 1b, the video decoding system 40 may include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video encoder / decoder implemented by a processing circuit 46), an antenna 42, one or more processors 43, one or more memory storage devices 44 and / or a display device 45.

[0104] 1b, imaging device 41, antenna 42, processing circuit 46, video encoder 20, video decoder 30, processor 43, memory 44, and / or display device 45 are capable of communicating with one another. In different embodiments, video decoding system 40 may include only video encoder 20 or only video decoder 30.

[0105] In some instances, antenna 42 can be used to transmit or receive an encoded bitstream of video data. Additionally, in some instances, display device 45 can be used to present the video data. Processing circuitry 46 can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. Video decoding system 40 can also include an optional processor 43, which can similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. Furthermore, memory storage 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, memory storage 44 can be implemented as cache memory. In other instances, processing circuitry 46 can include memory (e.g., cache memory, etc.) for implementing an image buffer, etc.

[0106] In some examples, video encoder 20 implemented by logic circuitry may include an image buffer (e.g., implemented by processing circuitry 46 or memory storage 44) and a graphics processing unit (e.g., implemented by processing circuitry 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video encoder 20 implemented by processing circuitry 46 to implement the various modules discussed with reference to FIG. 2 and / or any other encoder systems or subsystems described herein. Logic circuitry may be used to perform the various operations discussed herein.

[0107] In some examples, the video decoder 30 may be implemented by processing circuitry 46 in a similar manner to implement the various modules discussed with reference to the video decoder 30 of FIG. 3 and / or any other decoder systems or subsystems described herein. In some examples, the logic circuit implementation of the video decoder 30 may include an image buffer (implemented by processing circuitry 46 or memory storage 44) and a graphics processing unit (e.g., implemented by processing circuitry 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 30 implemented by processing circuitry 46 to implement the various modules discussed with reference to FIG. 3 and / or any other decoder systems or subsystems described herein.

[0108] In some examples, antenna 42 may be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data related to the encoded video frames, indicators, index values, mode selection data, etc., as discussed herein, such as data related to encoding partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators (as discussed), and / or data defining encoding partitions). Video decoding system 40 may also include video decoder 30 coupled to antenna 42 and configured to decode the encoded bitstream. Display device 45 is configured to present the video frames.

[0109] It should be understood that for the examples described herein with reference to video encoder 20, video decoder 30 can be configured to perform the reverse process. With respect to signaling syntax elements, video decoder 30 can be configured to receive and parse such syntax elements and decode the associated video data accordingly. In some examples, video encoder 20 can entropy encode the syntax elements into an encoded video bitstream. In such examples, video decoder 30 can parse such syntax elements and decode the associated video data accordingly.

[0110] For ease of description, the embodiments of the present application are described with reference to the universal video coding (VVC) reference software or the high-efficiency video coding (HEVC) developed by the joint collaboration team on video coding (JCT-VC) of the ITU-T video coding experts group (VCEG) and the ISO / IEC motion picture experts group (MPEG). Those skilled in the art will understand that the embodiments of the present application are not limited to HEVC or VVC.

[0111] Encoders and encoding methods

[0112] As shown in FIG2 , the video encoder 20 includes an input terminal (or input interface) 201, a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output terminal (or output interface) 272. The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254, and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG2 may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.

[0113] 2 , the inter-frame prediction unit is a trained target model (also known as a neural network) that processes an input image, image region, or image block to generate a prediction value for the input image block. For example, the neural network for inter-frame prediction receives an input image, image region, or image block and generates a prediction value for the input image, image region, or image block.

[0114] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 constitute the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-frame prediction unit 244, and the intra-frame prediction unit 254 constitute the backward signal path of the encoder, where the backward signal path of the encoder 20 corresponds to the signal path of the decoder (see decoder 30 in Figure 3). The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer 230, the inter-frame prediction unit 244, and the intra-frame prediction unit 254 also constitute the "internal decoder" of the video encoder 20.

[0115] Images and image segmentation (images and blocks)

[0116] Encoder 20 is operable to receive, via input 201 or the like, an image (or image data) 17, for example, an image from a sequence of images forming a video or video sequence. The received image or image data may also be a pre-processed image (or pre-processed image data) 19. For simplicity, the following description uses image 17. Image 17 may also be referred to as a current image or image to be encoded (particularly when distinguishing the current image from other images in video encoding, such as previously encoded and / or decoded images in the same video sequence, i.e., a video sequence that also includes the current image).

[0117] A (digital) image is, or can be considered to be, a two-dimensional array or matrix of pixels with intensity values. The pixels in the array are also referred to as pixels (or pels, short for picture elements). The number of pixels in the array or image in the horizontal and vertical directions (or axes) determines the image size and / or resolution. To represent color, three color components are typically used, meaning that an image can be represented as or include three pixel arrays. In the RBG format or color space, an image includes corresponding arrays of red, green, and blue pixels. However, in video coding, each pixel is typically represented in a luma / chroma format or color space, such as YCbCr, which includes a luma component indicated by Y (sometimes also indicated by L) and two chroma components, indicated by Cb and Cr. The luma component Y represents the brightness or grayscale level intensity (for example, in grayscale images, both are the same), while the two chroma components (abbreviated as chroma) Cb and Cr represent the chroma or color information components. Accordingly, an image in YCbCr format includes a luma pixel array of luma pixel values ​​(Y) and two chroma pixel arrays of chroma values ​​(Cb and Cr). An image in RGB format can be converted or transformed into YCbCr format, and vice versa, a process also known as color conversion or transformation. If the image is black and white, the image may include only a luma pixel array. Accordingly, the image may be, for example, a luma pixel array in monochrome format or a luma pixel array and two corresponding chroma pixel arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0118] In one embodiment, an embodiment of the video encoder 20 may include an image segmentation unit (not shown in FIG. 2 ) for segmenting the image 17 into a plurality of (typically non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC) or coding tree blocks (CTBs), or coding tree units (CTUs) in the H.265 / HEVC and VVC standards. The segmentation unit may be used to use the same block size for all images in a video sequence and a corresponding grid of defined block sizes, or to vary the block size between images or subsets or groups of images, and to segment each image into corresponding blocks.

[0119] In other embodiments, the video encoder may be configured to directly receive a block 203 of the image 17, for example, one, several or all blocks constituting the image 17. The image block 203 may also be referred to as a current image block or an image block to be encoded.

[0120] Like image 17, image block 203 is also or can be considered to be a two-dimensional array or matrix of pixels having intensity values ​​(pixel values), but image block 203 is smaller than image 17. In other words, block 203 may include one pixel array (e.g., a luminance array in the case of monochrome image 17, or a luminance array or chrominance array in the case of a color image), or three pixel arrays (e.g., one luminance array and two chrominance arrays in the case of a color image 17), or any other number and / or type of arrays depending on the color format used. The number of pixels in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Accordingly, a block may be an M×N (M columns×N rows) pixel array, or an M×N transform coefficient array, etc.

[0121] In one embodiment, the video encoder 20 shown in FIG. 2 is configured to encode the image 17 block by block, for example, performing encoding and prediction on each block 203 .

[0122] In one embodiment, the video encoder 20 shown in FIG2 may also be configured to partition and / or encode an image using slices (also referred to as video slices), where an image may be partitioned or encoded using one or more slices (typically non-overlapping). Each slice may include one or more blocks (e.g., coding tree units (CTUs)) or one or more groups of blocks (e.g., tiles in the H.265 / HEVC / VVC standard and bricks in the VVC standard).

[0123] In one embodiment, the video encoder 20 shown in Figure 2 can also be used to segment and / or encode an image using slices / coding block groups (also called video coding block groups) and / or coding blocks (also called video coding blocks), where the image can be segmented or encoded using one or more slices / coding block groups (usually non-overlapping), each slice / coding block group may include one or more blocks (e.g., CTUs) or one or more coding blocks, etc., where each coding block can be in a shape such as a rectangle and may include one or more complete or partial blocks (e.g., CTUs).

[0124] Residual calculation

[0125] The residual calculation unit 204 is used to calculate the residual block 205 (the prediction block 265 is described in detail later) based on the image block (or original block) 203 and the prediction block 265 in the following manner: for example, the pixel value of the prediction block 265 is subtracted from the pixel value of the image block 203 pixel by pixel (pixel by pixel) to obtain the residual block 205 in the pixel domain.

[0126] Transform

[0127] The transform processing unit 206 is configured to perform a discrete cosine transform (DCT) or a discrete sine transform (DST) on the pixel values ​​of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207 may also be referred to as transform residual coefficients, representing the residual block 205 in the transform domain.

[0128] The transform processing unit 206 may be used to apply an integerized approximation of the DCT / DST, such as the transform specified for H.265 / HEVC. This integerized approximation is typically scaled by a factor compared to the orthogonal DCT transform. In order to maintain the norm of the residual block after the forward and inverse transforms, additional scaling factors are used as part of the transform process. The scaling factors are typically selected based on certain constraints, such as whether the scaling factor is a power of 2 for the shift operation, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. For example, a specific scaling factor is specified for the inverse transform on the encoder 20 side by the inverse transform processing unit 212 (and for the corresponding inverse transform on the decoder 30 side by, for example, the inverse transform processing unit 312), and correspondingly, a corresponding scaling factor may be specified for the forward transform on the encoder 20 side by the transform processing unit 206.

[0129] In one embodiment, the video encoder 20 (correspondingly, the transform processing unit 206) may be configured to output transform parameters such as one or more transform types, for example, directly output or output after being encoded or compressed by the entropy coding unit 270, such that the video decoder 30 may receive and use the transform parameters for decoding.

[0130] Quantification

[0131] The quantization unit 208 is configured to quantize the transform coefficients 207 by, for example, scalar quantization or vector quantization to obtain quantized transform coefficients 209 . The quantized transform coefficients 209 may also be referred to as quantized residual coefficients 209 .

[0132] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, during quantization, an n-bit transform coefficient may be rounded down to an m-bit transform coefficient, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, for scalar quantization, varying degrees of scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. The appropriate quantization step size may be indicated by a quantization parameter (QP). For example, the quantization parameter may be an index into a predefined set of appropriate quantization step sizes. For example, a smaller quantization parameter may correspond to fine quantization (a smaller quantization step size), while a larger quantization parameter may correspond to coarse quantization (a larger quantization step size), or vice versa. Quantization may include dividing by the quantization step size, while the corresponding or inverse dequantization performed by the inverse quantization unit 210, etc., may include multiplying by the quantization step size. Embodiments according to some standards, such as HEVC, may be used to determine the quantization step size using the quantization parameter. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation involving division. Additional scaling factors can be introduced for quantization and dequantization to restore the norm of the residual block that may have been modified by the scaling used in the fixed-point approximation of the equations for the quantization step size and quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream, etc. Quantization is a lossy operation, where larger quantization step sizes result in greater losses.

[0133] In one embodiment, the video encoder 20 (correspondingly, the quantization unit 208) may be configured to output a quantization parameter (QP), for example, directly or after being encoded or compressed by the entropy coding unit 270, such that the video decoder 30 may receive and use the quantization parameter for decoding.

[0134] Dequantization

[0135] The inverse quantization unit 210 is configured to perform inverse quantization performed by the quantization unit 208 on the quantized coefficients to obtain dequantized coefficients 211. For example, the inverse quantization scheme performed by the quantization unit 208 may be performed according to or using the same quantization step size as the quantization unit 208. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211, which correspond to the transform coefficients 207. However, due to the loss caused by quantization, the dequantized coefficients 211 are generally not identical to the transform coefficients.

[0136] Inverse transform

[0137] The inverse transform processing unit 212 is configured to perform the inverse transform of the transform performed by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the pixel domain. The reconstructed residual block 213 may also be referred to as a transform block 213.

[0138] reconstruction

[0139] The reconstruction unit 214 (e.g., the summer 214) is used to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the pixel domain, for example, by adding the pixel point values ​​of the reconstructed residual block 213 and the pixel point values ​​of the prediction block 265.

[0140] Filtering

[0141] The loop filter unit 220 (or simply "loop filter" 220) is used to filter the reconstructed block 215 to obtain a filter block 221, or generally to filter the reconstructed pixels to obtain filtered pixel values. For example, the loop filter unit is used to smoothly perform pixel conversion or improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process can be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop reshaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process can also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Although the loop filter unit 220 is shown as a loop filter in FIG2 , in other configurations, the loop filter unit 220 can be implemented as a post-loop filter. The filter block 221 can also be referred to as a filter reconstruction block 221.

[0142] In one embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) may be configured to output loop filter parameters (e.g., SAO filter parameters, ALF filter parameters, or LMCS parameters), for example, directly or after being entropy-encoded by the entropy coding unit 270, such that the decoder 30 may receive and use the same or different loop filter parameters for decoding.

[0143] Decoded Image Buffer

[0144] The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data for use by the video encoder 20 when encoding video data. The DPB 230 may be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer 230 may be used to store one or more filter blocks 221. The decoded picture buffer 230 may also be used to store other previously filtered blocks, such as previously reconstructed and filtered blocks 221, for the same current picture or a different picture, such as a previously reconstructed picture, and may provide a complete previously reconstructed, i.e., decoded picture (and corresponding reference blocks and pixels) and / or a partially reconstructed current picture (and corresponding reference blocks and pixels), for example, for inter-frame prediction. The decoded image buffer 230 may also be used to store one or more unfiltered reconstructed blocks 215, or generally to store unfiltered reconstructed pixels, for example, reconstructed blocks 215 that have not been filtered by the loop filtering unit 220, or reconstructed blocks or reconstructed pixels that have not undergone any other processing.

[0145] Mode selection (segmentation and prediction)

[0146] The mode selection unit 260 includes a segmentation unit 262, an inter-frame prediction unit 244, and an intra-frame prediction unit 254, which are configured to receive or obtain original image data, such as an original block 203 (current block 203 of the current image 17) and reconstructed image data, such as filtered and / or unfiltered reconstructed pixels or reconstructed blocks of the same (current) image and / or one or more previously decoded images, from the decoded image buffer 230 or other buffer (e.g., a column buffer, not shown in FIG. 2). The reconstructed image data is used as reference image data required for prediction, such as inter-frame prediction or intra-frame prediction, to obtain a prediction block 265 or a prediction value 265.

[0147] The mode selection unit 260 may be used to determine or select a partitioning for the current block (including no partitioning) and prediction mode (eg, intra-frame or inter-frame prediction mode), generate a corresponding prediction block 265 , and calculate the residual block 205 and reconstruct the reconstruction block 215 .

[0148] In one embodiment, the mode selection unit 260 may be configured to select a segmentation and prediction mode (e.g., from prediction modes supported or available by the mode selection unit 260) that provides the best match or minimum residual (minimum residual means better compression during transmission or storage), or provides minimum signaling overhead (minimum signaling overhead means better compression during transmission or storage), or simultaneously considers or balances both. The mode selection unit 260 may be configured to determine the segmentation and prediction mode based on rate distortion optimization (RDO), i.e., select the prediction mode that provides the minimum rate distortion optimization. Terms such as "best," "lowest," and "optimal" herein do not necessarily refer to "best," "lowest," or "optimal" overall, but may also refer to situations where termination or selection criteria are met, e.g., values ​​exceeding or falling below a threshold or other limit may result in a "suboptimal selection" but reduce complexity and processing time.

[0149] In other words, the partitioning unit 262 may be configured to partition an image in a video sequence into a sequence of coding tree units (CTUs), the CTU 203 being further partitioned into smaller block portions or sub-blocks (again forming blocks), e.g., by iteratively using quad-tree partitioning (QT), binary-tree partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, and for, e.g., performing prediction on each of the block portions or sub-blocks, wherein the mode selection comprises selecting a tree structure for partitioning the block 203 and selecting a prediction mode to be applied to each of the block portions or sub-blocks.

[0150] The segmentation (eg, performed by segmentation unit 262) and prediction processes (eg, performed by inter-prediction unit 244 and intra-prediction unit 254) performed by video encoder 20 are described in detail below.

[0151] segmentation

[0152] The partitioning unit 262 can partition (or divide) an image block (or CTU) 203 into smaller parts, such as square or rectangular blocks. For an image with three pixel arrays, a CTU consists of N×N luminance pixel blocks and two corresponding chrominance pixel blocks. The maximum allowed size of a luminance block in a CTU is specified as 128×128 in the developing Universal Video Coding (VVC) standard, but may be specified to a value other than 128×128, such as 256×256, in the future. The CTUs of an image can be grouped / collected into slices / coding block groups, coding blocks, or bricks. A coding block covers a rectangular area of ​​an image and can be divided into one or more bricks. A brick consists of multiple CTU rows within a coding block. A coding block that is not partitioned into multiple bricks can be called a brick. However, a brick is a true subset of a coding block and is therefore not called a coding block. VVC supports two coding block group modes: raster scan slice / coding block group mode and rectangular slice mode. In raster scan CBG mode, a slice / CBG contains a sequence of CBs from a raster scan of the CBs of an image. In rectangular slice mode, a slice contains multiple bricks of an image that together form a rectangular region of the image. The bricks within a rectangular slice are arranged in the slice's brick raster scan order. These smaller blocks (also called sub-blocks) can be further split into smaller parts. This is also known as tree partitioning or hierarchical tree partitioning, where a root block at, for example, root tree level 0 (hierarchy level 0, depth 0) can be recursively split into two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchy level 1, depth 1). These blocks can be further split into two or more blocks at the next lower level, such as nodes at tree level 2 (hierarchy level 2, depth 2), and so on, until the partitioning is completed (because the end criteria are met, such as reaching the maximum tree depth or minimum block size). Blocks that are not further split are also called leaf blocks or leaf nodes of the tree. A tree divided into two parts is called a binary tree (BT), a tree divided into three parts is called a ternary tree (TT), and a tree divided into four parts is called a quadtree (QT).

[0153] For example, a coding tree unit (CTU) may be or include a CTB for luma pixels, two corresponding CTBs for chroma pixels of an image with a three-pixel array, or a CTB for pixels of a monochrome image, or a CTB for pixels of an image encoded using three independent color planes and syntax structures for encoding pixels. Accordingly, a coding tree block (CTB) may be an N×N block of pixels, where N may be set to a value such that a component is divided into CTBs, which is known as partitioning. A coding unit (CU) may be or include a coding block of luma pixels, two corresponding coding blocks for chroma pixels of an image with a three-pixel array, or a coding block of pixels of a monochrome image, or a coding block of pixels of an image encoded using three independent color planes and syntax structures for encoding pixels. Accordingly, a coding block (CB) may be an M×N block of pixels, where M and N may be set to a value such that a CTB is divided into coding blocks, which is known as partitioning.

[0154] For example, in an embodiment, according to HEVC, a coding tree unit (CTU) can be divided into multiple CUs using a quadtree structure represented as a coding tree. A decision is made at the leaf-CU level whether to use inter-frame (temporal) prediction or intra-frame (spatial) prediction to encode an image region. Each leaf-CU can be further divided into one, two, or four PUs according to the PU partition type. The same prediction process is used within a PU, and relevant information is transmitted to the decoder in units of PUs. After applying the prediction process according to the PU partition type to obtain a residual block, the leaf-CU can be divided into transform units (TUs) according to other quadtree structures similar to the coding tree for the CU.

[0155] For example, in an embodiment, according to the latest video coding standard currently under development (called Versatile Video Coding (VVC), a combined quadtree of nested multi-type trees (such as binary trees and ternary trees) is used to divide the segment structure for partitioning the coding tree unit. In the coding tree structure within the coding tree unit, the CU can be square or rectangular. For example, the coding tree unit (CTU) is first partitioned by the quadtree structure. The quadtree leaf nodes are further partitioned by the multi-type tree structure. The multi-type tree structure has four partition types: vertical binary tree partition (SPLIT_BT_VER), horizontal binary tree partition (SPLIT_BT_HOR), vertical ternary tree partition The tree nodes of the multi-type tree are called coding units (CUs), unless the CU is too large for the maximum transform length, in which case such segmentation is used for prediction and transform processing without any other splitting. In most cases, this means that the block sizes of CUs, PUs, and TUs in the coding block structure of the quadtree nested multi-type tree are the same. This exception occurs when the maximum supported transform length is less than the width or height of the color components of the CU. VVC has developed a unique signaling mechanism for the split partitioning information in the coding structure with quadtree nested multi-type trees. In the signaling mechanism, the coding The tree unit (CTU) as the root of the quadtree is first split by the quadtree structure. Then each quadtree leaf node (when large enough) is further split into a multi-type tree structure. In the multi-type tree structure, the first flag (mtt_split_cu_flag) is used to indicate whether the node is further split. When the node is further split, the second flag (mtt_split_cu_vertical_flag) is used to indicate the division direction, and the third flag (mtt_split_cu_binary_flag) is used to indicate whether the division is a binary tree division or a ternary tree division. According to mtt_split_c The values ​​of u_vertical_flag and mtt_split_cu_binary_flag allow the decoder to derive the multi-type tree split mode (MttSplitMode) of the CU based on predefined rules or tables. It should be noted that for certain designs, such as the 64×64 luma block and 32×32 chroma pipeline design in the VVC hardware decoder, TT splitting is not allowed when the width or height of the luma coding block is greater than 64. TT splitting is also not allowed when the width or height of the chroma coding block is greater than 32. The pipeline design divides the image into multiple virtual pipeline data units (VPDUs), each of which is defined as a non-overlapping unit in the image. In the hardware decoder, consecutive VPDUs are processed simultaneously in multiple pipeline stages. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is desirable to keep the VPDU small.In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the ternary tree (TT) and binary tree (BT) partitioning may increase the VPDU size.

[0156] In addition, it should be noted that when a part of the tree node block exceeds the bottom or the right boundary of the image, the tree node block is forcibly divided until all pixels of each coding CU are located within the image boundary.

[0157] For example, the intra sub-partitions (ISP) tool may vertically or horizontally divide the luma intra prediction block into two or four sub-partitions according to the block size.

[0158] In one example, mode select unit 260 of video encoder 20 may be used to perform any combination of the segmentation techniques described above.

[0159] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a (predetermined) prediction mode set. The prediction mode set may include, for example, an intra-frame prediction mode and / or an inter-frame prediction mode.

[0160] Intra-frame prediction

[0161] The intra prediction mode set may include 35 different intra prediction modes, for example, non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in HEVC, or may include 67 different intra prediction modes, for example, non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in VVC. For example, several traditional angle intra prediction modes are adaptively replaced by wide-angle intra prediction modes for non-square blocks defined in VVC. For another example, in order to avoid the division operation of DC prediction, only the longer side is used to calculate the average value of the non-square block. In addition, the intra prediction result of the planar mode can also be modified using the position dependent intra prediction combination (PDPC) method.

[0162] The intra prediction unit 254 is configured to generate an intra prediction block 265 by reconstructing pixels in adjacent blocks of the same current image according to an intra prediction mode in the intra prediction mode set.

[0163] The intra-frame prediction unit 254 (or generally the mode selection unit 260) is also used to output intra-frame prediction parameters (or generally information indicating the selected intra-frame prediction mode for the block) in the form of syntax elements 266 to the entropy coding unit 270 for inclusion in the encoded image data 21, so that the video decoder 30 can perform operations, such as receiving and using the prediction parameters for decoding.

[0164] The intra prediction modes in HEVC include DC prediction mode, plane prediction mode and 33 angular prediction modes, with a total of 35 candidate prediction modes. The current block can use the pixels of the reconstructed image blocks on the left and above as references for intra prediction. The image blocks in the surrounding area of ​​the current block used for intra prediction of the current block are called reference blocks, and the pixels in the reference blocks are called reference pixels. Among the 35 candidate prediction modes, the DC prediction mode is applicable to areas with flat textures in the current block. All pixels in this area use the average value of the reference pixels in the reference block as prediction; the plane prediction mode is applicable to image blocks with smoothly changing textures. The current block that meets this condition uses the reference pixels in the reference block for bilinear interpolation as the prediction of all pixels in the current block; the angular prediction mode uses the characteristic that the texture of the current block is highly correlated with the texture of the adjacent reconstructed image blocks, and copies the values ​​of the reference pixels in the corresponding reference block along a certain angle as the prediction of all pixels in the current block.

[0165] The HEVC encoder selects an optimal intra-frame prediction mode for the current block from 35 candidate prediction modes and writes this optimal intra-frame prediction mode into the video bitstream. To improve the coding efficiency of intra-frame prediction, the encoder / decoder derives three most probable modes from the optimal intra-frame prediction modes of the reconstructed image blocks in the surrounding area using intra-frame prediction. If the optimal intra-frame prediction mode selected for the current block is one of these three most probable modes, a first index is encoded to indicate that the selected optimal intra-frame prediction mode is one of these three most probable modes; if the selected optimal intra-frame prediction mode is not one of these three most probable modes, a second index is encoded to indicate that the selected optimal intra-frame prediction mode is one of the other 32 modes (other than the aforementioned three most probable modes among the 35 candidate prediction modes). The HEVC standard uses a 5-bit fixed-length code as the aforementioned second index.

[0166] The HEVC encoder derives the three most probable modes by selecting the optimal intra-frame prediction mode of the image block to the left and the image block above the current block and adding them to a set. If these two optimal intra-frame prediction modes are the same, only one is retained in the set. If these two optimal intra-frame prediction modes are the same and both are angular prediction modes, two angular prediction modes adjacent to the angular direction are selected and added to the set. Otherwise, the planar prediction mode, the DC mode, and the vertical prediction mode are selected and added to the set in sequence until the number of modes in the set reaches three.

[0167] After the HEVC decoder performs entropy decoding on the bitstream, it obtains the mode information of the current block, which includes an indicator indicating whether the optimal intra-frame prediction mode of the current block is among the three most probable modes, and the index of the optimal intra-frame prediction mode of the current block among the three most probable modes or the index of the optimal intra-frame prediction mode of the current block among the other 32 modes.

[0168] Inter-frame prediction

[0169] In a possible implementation, the set of inter-prediction modes depends on the available reference picture (i.e., at least part of the previously decoded picture stored in the DBP 230 as mentioned above) and other inter-prediction parameters, e.g., on whether the entire reference picture is used or only a part of the reference picture is used, e.g., a search window area around the area of ​​the current block, to search for the best matching reference block, and / or on whether pixel interpolation such as half-pixel, quarter-pixel and / or 1 / 16 interpolation is performed, for example.

[0170] In addition to the above prediction modes, skip mode and / or direct mode may also be employed.

[0171] For example, in extended merge prediction, the merge candidate list of this mode consists of the following five candidate types in order: spatial MVP from spatially neighboring CUs, temporal MVP from collocated CUs, history-based MVP from the FIFO table, pairwise average MVP, and zero MV. Decoder-side motion vector refinement (DMVR) based on bilateral matching can be used to increase the accuracy of the MV in merge mode. Merge mode with MVD (MMVD) is derived from merge mode with motion vector difference. The MMVD flag is sent immediately after the skip flag and merge flag to specify whether the CU uses MMVD mode. A CU-level adaptive motion vector resolution (AMVR) scheme can be used. AMVR supports encoding the CU's MVD with different precisions. The MVD of the current CU is adaptively selected based on the prediction mode of the current CU. When the CU is encoded in merge mode, the combined inter / intra prediction (CIIP) mode can be applied to the current CU. The CIIP prediction is obtained by weighted averaging the inter and intra prediction signals. For affine motion compensation prediction, the affine motion field of the block is described by the motion information of 2 control points (4 parameters) or 3 control points (6 parameters) motion vectors. Subblock-based temporal motion vector prediction (SbTMVP) is similar to the temporal motion vector prediction (TMVP) in HEVC, but it predicts the motion vector of the sub-CU within the current CU. Bidirectional optical flow (BDOF), formerly known as BIO, is a simplified version that reduces calculations, especially in terms of the number of multiplications and the size of the multipliers. In the triangle partitioning mode, the CU is evenly divided into two triangular parts using diagonal partitioning and anti-diagonal partitioning. In addition, the bidirectional prediction mode is extended based on the simple average to support the weighted average of the two prediction signals.

[0172] The inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2 ). The motion estimation unit may be configured to receive or obtain an image block 203 (the current image block 203 of the current image 17 ) and a decoded image 231 , or at least one or more previously reconstructed blocks, e.g., reconstructed blocks of one or more other / different previously decoded images 231 , to perform motion estimation. For example, a video sequence may include the current image and the previously decoded image 231 , or in other words, the current image and the previously decoded image 231 may be part of or form a sequence of images forming the video sequence.

[0173] For example, the encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different images among a plurality of other images, and provide the reference image (or reference image index) and / or the offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. This offset is also referred to as a motion vector (MV).

[0174] The motion compensation unit is configured to obtain, for example, receive, inter-frame prediction parameters and perform inter-frame prediction based on or using the inter-frame prediction parameters to obtain an inter-frame prediction block 246. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction block based on a motion / block vector determined by motion estimation, and may also include performing interpolation with sub-pixel accuracy. Interpolation filtering can generate pixel points of other pixels from pixel points of known pixels, thereby potentially increasing the number of candidate prediction blocks that can be used to encode the image block. Upon receiving a motion vector corresponding to a PU of the current image block, the motion compensation unit may locate the prediction block pointed to by the motion vector in one of the reference picture lists.

[0175] The motion compensation unit may also generate syntax elements associated with blocks and video slices for use by video decoder 30 when decoding image blocks of a video slice. In addition to, or in lieu of, slices and corresponding syntax elements, coding block groups and / or coding blocks and corresponding syntax elements may be generated or used.

[0176] In the process of obtaining a candidate motion vector list in an advanced motion vector prediction (AMVP) mode, motion vectors (MVs) that can be added to the candidate motion vector list as alternatives include MVs of spatially and temporally adjacent image blocks of a current block, wherein the MVs of spatially adjacent image blocks can further include the MVs of a left candidate image block located to the left of the current block and the MVs of an upper candidate image block located above the current block. For example, please refer to FIG. 4 , which is an exemplary schematic diagram of candidate image blocks provided in an embodiment of the present application. As shown in FIG. 4 , the set of left candidate image blocks includes {A0, A1}, the set of upper candidate image blocks includes {B0, B1, B2}, and the set of temporally adjacent candidate image blocks includes {C, T}. All three sets can be added to the candidate motion vector list as alternatives. However, according to existing coding standards, the maximum length of the candidate motion vector list for AMVP is 2. Therefore, it is necessary to determine the MVs of up to two image blocks to be added to the candidate motion vector list from the three sets according to a prescribed order. The order may be to give priority to the set of candidate image blocks {A0, A1} to the left of the current block (consider A0 first, and then consider A1 if A0 is not available), then consider the set of candidate image blocks {B0, B1, B2} above the current block (consider B0 first, and then consider B1 if B0 is not available, and then consider B2 if B1 is not available), and finally consider the set of candidate image blocks {C, T} that are adjacent to the current block in the time domain (consider T first, and then consider C if T is not available).

[0177] After obtaining the candidate motion vector list, the optimal MV is determined from the candidate motion vector list using the rate distortion cost (RD cost). The candidate motion vector with the smallest RD cost is used as the motion vector predictor (MVP) for the current block. The rate distortion cost is calculated using the following formula: J = SAD + λR

[0178] Wherein, J represents RD cost, SAD is the sum of absolute differences (SAD) between the pixel values ​​of the predicted block obtained after motion estimation using the candidate motion vector and the pixel values ​​of the current block, R represents the bit rate, and λ represents the Lagrange multiplier.

[0179] The encoder passes the index of the determined MVP in the candidate motion vector list to the decoder. Furthermore, a motion search can be performed within a neighborhood centered on the MVP to obtain the actual motion vector of the current block. The encoder calculates the motion vector difference (MVD) between the MVP and the actual motion vector and also passes the MVD to the decoder. The decoder parses the index, finds the corresponding MVP in the candidate motion vector list based on the index, parses the MVD, and adds the MVD to the MVP to obtain the actual motion vector of the current block.

[0180] When obtaining the candidate motion information list in Merge mode, the motion information that can be added to the candidate motion information list includes the motion information of spatially or temporally adjacent image blocks to the current block. For spatially and temporally adjacent image blocks, refer to Figure 4. The spatially adjacent candidate motion information in the candidate motion information list comes from five spatially adjacent blocks (A0, A1, B0, B1, and B2). If the spatially adjacent blocks are unavailable or are intra-predicted, their motion information is not added to the candidate motion information list. The temporal candidate motion information for the current block is obtained by scaling the MV of the corresponding block in the reference frame based on the picture order count (POC) of the reference frame and the current frame. The block at position T in the reference frame is first determined to be available. If not, the block at position C is selected. After obtaining the candidate motion information list, the optimal motion information from the candidate motion information list is determined using the RD cost as the motion information for the current block. The encoder transmits the index of the optimal motion information in the candidate motion information list (denoted as the merge index) to the decoder.

[0181] Entropy Coding

[0182] The entropy coding unit 270 is configured to apply an entropy coding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC (CALVC) scheme, an arithmetic coding scheme, a binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements, resulting in coded image data 21 that can be output via an output terminal 272 in the form of a coded bitstream 21, etc., so that the video decoder 30, etc. can receive and use the parameters for decoding. The coded bitstream 21 can be transmitted to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.

[0183] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform-based encoder 20 may directly quantize the residual signal without a transform processing unit 206 for certain blocks or frames. In another implementation, the encoder 20 may have the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.

[0184] Decoder and decoding method

[0185] As shown in FIG3 , a video decoder 30 is configured to receive coded image data 21 (e.g., coded bitstream 21) encoded by, for example, an encoder 20, and generate a decoded image 331. The coded image data or bitstream includes information used to decode the coded image data, such as data representing image blocks of a coded video slice (and / or coding block group or coding block) and related syntax elements.

[0186] 3 , decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., summer 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter-prediction unit 344, and an intra-prediction unit 354. Inter-prediction unit 344 may be or include a motion compensation unit. In some examples, video decoder 30 may perform a decoding process that is generally the reverse of the encoding process described with reference to video encoder 20 of FIG. 2 .

[0187] As described above with respect to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer DPB 230, inter-frame prediction unit 344, and intra-frame prediction unit 354 also constitute the "built-in decoder" of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 122, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Therefore, the explanations of the corresponding units and functions of video encoder 20 apply accordingly to the corresponding units and functions of video decoder 30.

[0188] Entropy decoding

[0189] The entropy decoding unit 304 is configured to parse the bitstream 21 (or generally, the encoded image data 21) and perform entropy decoding on the encoded image data 21 to obtain quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3 ), such as any or all of inter-frame prediction parameters (e.g., reference image indices and motion vectors), intra-frame prediction parameters (e.g., intra-frame prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the coding scheme of the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 may also be configured to provide inter-frame prediction parameters, intra-frame prediction parameters, and / or other syntax elements to the mode application unit 360, as well as to provide other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice and / or video block level. In addition to, or in lieu of, slices and corresponding syntax elements, coding block groups and / or coding blocks and corresponding syntax elements may also be received or used.

[0190] Dequantization

[0191] The inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally information related to inverse quantization) and quantization coefficients from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse quantize the decoded quantization coefficients 309 based on the quantization parameter to obtain inverse quantization coefficients 311. The inverse quantization coefficients 311 may also be referred to as transform coefficients 311. The inverse quantization process may include using the quantization parameter calculated by the video encoder 20 for each video block in the video slice to determine a degree of quantization, and thus a degree of inverse quantization to be performed.

[0192] Inverse transform

[0193] The inverse transform processing unit 312 may be configured to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain a reconstructed residual block 213 in the pixel domain. The reconstructed residual block 213 may also be referred to as a transform block 313. The transform may be an inverse transform, such as an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may also be configured to receive transform parameters or corresponding information from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304) to determine the transform to be applied to the dequantized coefficients 311.

[0194] reconstruction

[0195] The reconstruction unit 314 (eg, summer 314 ) is configured to add the reconstructed residual block 313 to the prediction block 365 to obtain the reconstructed block 315 in the pixel domain, eg, by adding the pixel values ​​of the reconstructed residual block 313 and the pixel values ​​of the prediction block 365 .

[0196] Filtering

[0197] The loop filter unit 320 (in or after the encoding loop) is used to filter the reconstructed block 315 to obtain a filter block 321, thereby smoothly performing pixel conversion or improving video quality. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process can be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop reshaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process can also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Although the loop filter unit 320 is shown as a loop filter in Figure 3, in other configurations, the loop filter unit 320 can be implemented as a post-loop filter.

[0198] Decoded Image Buffer

[0199] The decoded video blocks 321 of one picture are then stored in a decoded picture buffer 330 which stores the decoded picture 331 as a reference picture for subsequent motion compensation of other pictures and / or for respective output displays.

[0200] The decoder 30 is used to output the decoded image 311 through the output terminal 312, etc., for display to the user or for the user to view.

[0201] predict

[0202] The inter-frame prediction unit 344 may be functionally identical to the inter-frame prediction unit 244 (particularly the motion compensation unit), and the intra-frame prediction unit 354 may be functionally identical to the inter-frame prediction unit 254 and may determine the partitioning or segmentation and perform prediction based on the segmentation and / or prediction parameters or corresponding information received from the coded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304). The mode application unit 360 may be configured to perform prediction (intra-frame or inter-frame prediction) for each block based on the reconstructed image, block, or corresponding pixel point (filtered or unfiltered), resulting in a prediction block 365.

[0203] When the video slice is encoded as an intra-coded (I) slice, the intra-prediction unit 354 in the mode application unit 360 is configured to generate a prediction block 365 for the image block of the current video slice based on the indicated intra-prediction mode and data from previously decoded blocks of the current image. When the video image is encoded as an inter-coded (i.e., B or P) slice, the inter-prediction unit 344 (e.g., a motion compensation unit) in the mode application unit 360 is configured to generate a prediction block 365 for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. For inter-prediction, these prediction blocks can be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 can construct reference frame list 0 and list 1 using a default construction technique based on the reference pictures stored in DPB 330. The same or similar processes may be applied to embodiments of coding block groups (e.g., video coding block groups) and / or coding blocks (e.g., video coding blocks) in addition to or instead of slices (e.g., video slices), e.g., video may be encoded using I, P, or B coding block groups and / or coding blocks.

[0204] Mode application unit 360 is configured to determine prediction information for video blocks of a current video slice by parsing motion vectors and other syntax elements, and to use the prediction information to generate a prediction block for the current video block being decoded. For example, mode application unit 360 uses received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used to encode the video blocks of the video slice, an inter-prediction slice type (e.g., a B slice, a P slice, or a GPB slice), construction information for one or more reference picture lists for the slice, a motion vector for each inter-coded video block in the slice, an inter-prediction state for each inter-coded video block in the slice, and other information to decode the video blocks within the current video slice. In addition to or in lieu of slices (e.g., video slices), the same or similar processes may be applied to embodiments of coding block groups (e.g., video coding block groups) and / or coding blocks (e.g., video coding blocks), for example, video may be encoded using I, P, or B coding block groups and / or coding blocks.

[0205] In one embodiment, the video encoder 30 of FIG3 may also be configured to partition and / or decode an image using slices (also referred to as video slices), where an image may be partitioned or decoded using one or more (typically non-overlapping) slices. Each slice may include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., coding blocks in the H.265 / HEVC / VVC standard and bricks in the VVC standard).

[0206] In one embodiment, the video decoder 30 shown in Figure 3 can also be used to segment and / or decode an image using slices / coding block groups (also called video coding block groups) and / or coding blocks (also called video coding blocks), where the image can be segmented or decoded using one or more slices / coding block groups (usually non-overlapping), each slice / coding block group may include one or more blocks (e.g., CTUs) or one or more coding blocks, etc., where each coding block can be in a shape such as a rectangle and may include one or more complete or partial blocks (e.g., CTUs).

[0207] Other variations of the video decoder 30 may be used to decode the encoded image data 21. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 may directly inverse quantize the residual signal without the inverse transform processing unit 312 for certain blocks or frames. In another implementation, the video decoder 30 may have the inverse quantization unit 310 and the inverse transform processing unit 312 combined into a single unit.

[0208] It should be understood that the processing result of the current step can be further processed in the encoder 20 and the decoder 30 and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, the processing result of interpolation filtering, motion vector derivation, or loop filtering can be further operated on, such as clipping or shifting operations.

[0209] It should be noted that further operations can be performed on the derived motion vector of the current block (including but not limited to control point motion vectors in affine mode, sub-block motion vectors in affine, planar, ATMVP modes, temporal motion vectors, etc.). For example, the value of the motion vector can be limited to a predefined range based on the representation bit of the motion vector. If the representation bit of the motion vector is bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" represents a power. For example, if bitDepth is set to 16, the range is -32768 to 32767; if bitDepth is set to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MVs of four 4×4 sub-blocks in an 8×8 block) is limited so that the maximum difference between the integer parts of the above four 4×4 sub-block MVs does not exceed N pixels, for example, not more than 1 pixel. Two methods of limiting motion vectors based on bitDepth are provided here.

[0210] Although the above embodiments primarily describe video coding, it should be noted that embodiments of the decoding system 10, encoder 20, and decoder 30, as well as other embodiments described herein, may also be used for still image processing or coding, i.e., processing or coding a single image in a video codec that is independent of any previous or subsequent images. In general, if image processing is limited to a single image 17, the inter-frame prediction unit 244 (encoder) and the inter-frame prediction unit 344 (decoder) may not be available. All other functionalities (also referred to as tools or techniques) of the video encoder 20 and video decoder 30, such as residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354, and / or loop filtering 220 / 320, entropy coding 270, and entropy decoding 304, may also be used for still image processing.

[0211] Please refer to Figure 5, which is an exemplary block diagram of a video decoding device 500 provided in an embodiment of the present application. Video decoding device 500 is suitable for implementing the disclosed embodiments described herein. In one embodiment, video decoding device 500 can be a decoder, such as the video decoder 30 in Figure 1a, or an encoder, such as the video encoder 20 in Figure 1a.

[0212] Video decoding device 500 includes: an input port 510 (or input port 510) and a receiver unit (Rx) 520 for receiving data; a processor, logic unit, or central processing unit (CPU) 530 for processing data; for example, processor 530 may be a neural network processor 530; a transmitter unit (Tx) 540 and an output port 550 (or output port 550) for transmitting data; and a memory 560 for storing data. Video decoding device 500 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to input port 510, receiver unit 520, transmitter unit 540, and output port 550 for outputting or transmitting optical or electrical signals.

[0213] The processor 530 is implemented in hardware and software. The processor 530 can be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 530 communicates with the input port 510, the receiving unit 520, the transmitting unit 540, the output port 550, and the memory 560. The processor 530 includes a decoding module 570 (e.g., a neural network-based decoding module 570). The decoding module 570 implements the embodiments disclosed above. For example, the decoding module 570 performs, processes, prepares, or provides various encoding operations. Therefore, the decoding module 570 provides substantial improvements to the functionality of the video decoding device 500 and affects the switching of the video decoding device 500 to different states. Alternatively, the decoding module 570 is implemented by instructions stored in the memory 560 and executed by the processor 530.

[0214] Memory 560 includes one or more disks, tape drives, and solid-state drives and can be used as overflow data storage for storing programs when such programs are selected for execution, and for storing instructions and data read during program execution. Memory 560 can be volatile and / or non-volatile and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0215] Please refer to FIG. 6 , which is an exemplary block diagram of an apparatus 600 provided in an embodiment of the present application. The apparatus 600 may be used as either or both of the source device 12 and the destination device 14 in FIG. 1 a .

[0216] The processor 602 in the apparatus 600 may be a central processing unit. Alternatively, the processor 602 may be any other type of device or devices, now available or developed in the future, capable of manipulating or processing information. While the disclosed implementations may be implemented using a single processor, such as the processor 602 shown, using more than one processor may provide greater speed and efficiency.

[0217] In one implementation, the memory 604 in the apparatus 600 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 604. The memory 604 may include code and data 606 accessed by the processor 602 via a bus 612. The memory 604 may also include an operating system 608 and application programs 610, which include at least one program that allows the processor 602 to perform the methods described herein. For example, the application programs 610 may include applications 1 through N, as well as a video decoding application that performs the methods described herein.

[0218] The apparatus 600 may also include one or more output devices, such as a display 618. In one example, the display 618 may be a touch-sensitive display that combines a display with touch-sensitive elements that can be used to sense touch input. The display 618 may be coupled to the processor 602 via the bus 612.

[0219] Although bus 612 in device 600 is described herein as a single bus, bus 612 may include multiple buses. Furthermore, secondary storage may be directly coupled to other components of device 600 or accessed via a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, device 600 may have a variety of configurations.

[0220] The coding and decoding method provided in the embodiment of the present application involves layered coding and decoding. Referring to the schematic diagram of the layered coding framework shown in Figure 7, a low-resolution image is obtained by downsampling the input image. For example, the input image is downsampled to obtain Layer i (i.e., layer i) and Layer j (i.e., layer j). Images in the same access unit (access unit, AU) are encoded from low to high according to the number of layers. If layer i is the inter-layer reference layer of layer j (i.e., Ref(j) == i, Ref is the abbreviation of reference), the reconstructed image of layer i is added to the reference frame library as a reference frame. In the output code stream, the image bit streams of different layers of the same access unit are arranged and interleaved in the order of increasing layer number.

[0221] It should be understood that the bitstream (i.e., the bitstream) after encoding the image includes one or more consecutive AUs, and an AU includes a group of network abstraction layer (NAL) units that are associated with each other according to specified rules. According to the content contained in the NAL unit, the types of NAL units may include access unit boundary NAL units (if present), security parameter set NAL units, sequence parameter set NAL units, picture parameter set NAL units, picture header NAL units, extension data NAL units, supplemental enhancement information NAL units (if present), coded slice NAL units, coded video sequence end NAL units (if present), and stream end NAL units (if present).

[0222] In order to facilitate understanding of the technical solution of this application, some technical terms involved in this method are explained below.

[0223] 1. Reference picture: An image used for inter-frame prediction of subsequent images during the decoding process.

[0224] 2. Reference picture list: A picture list consisting of reference pictures of the current picture. The reference picture list can include reference picture list 0 and reference picture list 1.

[0225] Reference picture list 0 is a picture list consisting of reference pictures for the current picture. Reference picture list 0 may include past and future pictures in the display order. Reference picture list 1 is a picture list consisting of reference pictures for the current picture. Reference picture list 1 may include past and future pictures in the display order.

[0226] 3. Reference index: The number of the reference image in the reference image queue, where the number in reference image queue 0 is called the L0 reference index, and the number in reference image queue 1 is called the L1 reference index.

[0227] 4. Layer Identifier (layer_id): A 3-bit unsigned integer. Indicates the layer identifier of the current image. The layer identifier value ranges from 0 to MAX_LAYERS. When the layer identifier is encoded starting from 0, the layer identifier value ranges from 0 to MAX_LAYERS-1. The layer_id of the sequence parameter set and picture parameter set NALU (NAL unit) is 0. The layer_id of the picture header NAL unit and all coded slice NAL units of a coded image should be the same. The value of layerId is equal to the value of layer_id.

[0228] 5. Number of Spatial Layers (num_of_layers_minus1): A 2-bit unsigned integer. This value, plus 1, indicates the number of layers in the spatial hierarchy. Num_of_layers_minus1 is transmitted in the bitstream only if the profile is 0xXX. For example, if the profile is 0x40 or 0x42, the number of spatial layers is transmitted in the bitstream. The value of NumOfLayers is equal to num_of_layers_minus1 + 1. If num_of_layers_minus1 is not present in the bitstream, the value of NumOfLayers is 1. The value of NumOfLayers should range from 1 to MAX_LAYER.

[0229] 6. Privacy_enable_flag: A binary variable. A value of '1' indicates that privacy protection can be used; a value of '0' indicates that privacy protection should not be used. The value of PrivacyEnableFlag is equal to the value of privacy_enable_flag. If privacy_enable_flag is not present in the bitstream, the value of PrivacyEnableFlag is 0.

[0230] 7. Rights protection mode index (privacy_mode_index): 2-bit unsigned integer, ranging from 0 to 3. Indicates the rights protection mode. The value of '0' indicates the privacy protection layered independent coding mode, the value of '1' indicates the privacy protection layered reference coding mode, the value of '3' indicates the privacy protection single layer coding mode, and the value of '2' is reserved. For example, the value of '2' can be reserved to indicate a coding mode that supports masking privacy. The value of PrivacyModeIndex is equal to the value of privacy_mode_index. If privacy_mode_index does not exist in the bitstream, the value of PrivacyModeIndex is equal to 0.

[0231] 8. Spatial dependency layer number of layer i (ref_layer_id[i]): 2-bit unsigned integer. Specifies the LayerId value of the reference layer for inter-layer dependency coding of spatial layer i whose LayerId is equal to ref_layer_id[i]. The value of ref_layer_id[i] ranges from 0 to NumOfLayers-1. The value of refLayerId[i] is equal to ref_layer_id[i].

[0232] 9. Horizontal size (horizontal_size[i]): A 14-bit unsigned integer. Specifies the width of the displayable area (aligned with the left edge of the image) of the luma component of the image whose LayerId is i, i.e., the number of horizontal samples. The value of horizontalSize[i] is equal to horizontal_size.

[0233] HorizontalSize[LayerId] shall not be '0'. The unit of HorizontalSize[LayerId] shall be the number of samples per line of the image. The upper left sample of the displayable area shall be aligned with the upper left sample of the decoded image.

[0234] 10. Vertical size (vertical_size[i]): 14-bit unsigned integer. Specifies the height of the displayable area (aligned with the top edge of the image) of the luma component of an image with LayerId equal to i, i.e., the number of vertical scan lines. The value of verticalSize[i] is equal to vertical_size.

[0235] VerticalSize[LayerId] shall not be 0. The unit of VerticalSize[LayerId] shall be the number of rows of image samples.

[0236] It can be understood that the above horizontal size and vertical size are the sizes of the input image.

[0237] 11. Horizontal coded image size (pic_width_in_luma[i]): 14-bit unsigned integer. Specifies the width of the coded image with LayerId equal to i. The value of PicWidthInLuma[i] is equal to pic_width_in_luma[i]. The value of PicWidthInLuma[i] cannot be 0 and must be an integer multiple of MinCuSize.

[0238] The value of HorizontalSize[i] must be less than or equal to the value of PicWidthInLuma[i].

[0239] 12. Vertical coded image size (pic_height_in_luma[i]): 14-bit unsigned integer. Specifies the coded image height for LayerId i. The value of PicHeightInLuma[i] is equal to pic_height_in_luma[i]. The value of PicHeightInLuma[i] cannot be 0 and must be an integer multiple of MinCuSize.

[0240] The value of VerticalSize[i] must be less than or equal to the value of PicHeightInLuma[i].

[0241] It can be understood that the above horizontal coding size and vertical coding size are the coding sizes of the image in the actual coding process.

[0242] 13. Number of layers supported by layered encoding and decoding

[0243] When the coding and decoding method provided in the embodiment of the present application performs layered coding and decoding, the number of supported layers is multiple. In one case, considering the complexity of coding and decoding, this solution supports a spatial domain layered structure of up to 4 layers. A 2-bit syntax num_of_layers_minus1 is added to the sequence parameter set (SPS) to identify the number of layers in the current layered structure. Frames in different layers of the same AU share the same POC number, and the frame decoding order index DOI of different layers in the same AU is constrained to be the same.

[0244] 14. Inter-layer dependencies

[0245] Add syntax for inter-layer dependency in SPS, constraining the enhancement layer cross-layer reference to only refer to a lower-layer frame with the same resolution, and using inter-layer pixel prediction. The syntax for inter-layer dependency includes:

[0246] Spatial Layer Independent Coding Flag (sps_all_independent_layers_flag): A binary variable. This flag indicates whether each layer in the current hierarchical structure is independently coded, without inter-layer prediction. A value of '1' indicates that all spatial layers are independently coded, without inter-layer dependency coding. A value of '0' indicates that inter-layer dependency coding is present. The value of spsAllIndependentLayersFlag is equal to the value of sps_all_independent_layers_flag. If sps_all_independent_layers_flag is not present in the bitstream, the value of SpsAllIndependentLayersFlag is 1.

[0247] Spatial layer i independent coding flag (sps_independent_layer_flag[i]): A binary variable that identifies whether layer i is independently coded. A value of '1' indicates that the spatial layer with LayerId equal to i is independently coded and does not use inter-layer dependency coding. A value of '0' indicates that the spatial layer with LayerId equal to i can use inter-layer dependency coding. The value of spsIndependentLayerFlag[i] is equal to the value of sps_independent_layer_flag[i]. If sps_independent_layer_flag[i] does not exist in the bitstream, the value of SpsIndependentLayerFlag[i] is equal to 1.

[0248] Reference layer identifier (ref_layer_id[i]): When layer i is not independently coded, it identifies the layer number of the layer referenced by layer i. The value range of ref_layer_id[i] is 0 to i-1.

[0249] 15. Resolution of each layer

[0250] In SPS, the image width and height as well as the encoded width and height data are transmitted separately for each layer.

[0251] horizontal_size[i], vertical_size[i]: identifies the width and height data of the input image at layer i.

[0252] pic_width_in_luma[i], pic_height_in_luma[i]: identifies the width and height data of the coded image of the i-th layer.

[0253] As you can understand, the resolution of an image is the product of its width and height.

[0254] 16. Reference frame list management

[0255] Reference frame management: In SPS, a dedicated reference picture queue configuration set is transmitted for each layer, namely reference_picture_list_set, abbreviated as rpls.

[0256] rpl1_index_exist_flag[i]: Indicates whether the reference picture queue 1 index of layer i exists in the codestream.

[0257] rpl1_same_as_rpl0_flag[i]: identifies whether the reference picture queue 1 of the i-th layer is consistent with the reference picture queue 0.

[0258] num_ref_pic_list_set[i][0]: num_ref_pic_list_set[i][1]: indicates the number of reference picture queue configuration sets.

[0259] num_ref_default_active_minus1[i][0], num_ref_default_active_minus1[i][1]: indicates the default maximum value of the reference index in the reference picture queue when decoding the picture.

[0260] 17. Patch division

[0261] The syntax for transmitting each layer of patch partitioning information independently in SPS.

[0262] stable_patch_flag[i]: patch partition consistency flag

[0263] uniform_patch_flag[i]: uniform patch size flag

[0264] patch_width_minus1[i],patch_height_minus1[i]: identifies the width and height of the patch, in LCU units.

[0265] 18. Knowledge stream flag (library_stream_flag): A binary variable. A value of '1' indicates that the current sequence parameter set is the sequence parameter set corresponding to the knowledge image; a value of '0' indicates that the current sequence parameter set is the sequence parameter set corresponding to the display image. The value of LibraryStreamFlag is equal to the value of library_stream_flag.

[0266] 19. Knowledge Picture Enable Flag (library_picture_enable_flag): A binary variable. A value of '1' indicates that inter-frame prediction pictures using knowledge pictures as reference pictures can exist in the coded video sequence; a value of '0' indicates that inter-frame prediction pictures using knowledge pictures as reference pictures should not exist in the coded video sequence. The value of LibraryPictureEnableFlag is equal to the value of library_picture_enable_flag.

[0267] 20. Knowledge picture mode index (library_picture_mode_index): A 2-bit unsigned integer ranging from 0 to 3, indicating the knowledge picture mode. A value of '0' indicates knowledge picture single-stream mode, a value of '1' indicates long-term reference picture mode, a value of '2' indicates knowledge picture dual-stream mode, and a value of '3' is reserved. The value of LibraryPictureModeIndex is equal to the value of library_picture_mode_index. If library_picture_mode_index does not exist in the bitstream, the value of LibraryPictureModeIndex is 0.

[0268] It should be noted that in the knowledge image single-stream mode, only non-display knowledge images are allowed in the coded video sequence, and display knowledge images are not allowed. The NAL units of the non-display knowledge image coding slices and the NAL units of the display image coding slices are interwoven to form a coded video sequence.

[0269] In long-term reference mode, only display knowledge images are allowed in the coded video sequence, and non-display knowledge images are not allowed. All coded slice NAL units of display knowledge images are continuous in the order of increasing coded slice index and should not be interleaved with the access units of display images.

[0270] In the knowledge image dual-stream mode, only display image coding NAL units are allowed in the main bitstream coded video sequence, and only knowledge image coding slice NAL units are allowed in the knowledge bitstream coded image sequence. The LibraryStreamFlag is used to determine whether the current coded video sequence is in the main bitstream or the knowledge bitstream.

[0271] 21. Patch width (patch_width_minus1[i]) and patch height (patch_height_minus1[i]): The width and height of the patch in the image with LayerId equal to i, in LCUs. The value of patch_width_minus1 should be less than 256, and the value of patch_height_minus1 should be less than 144. The value of patchWidth[i] is equal to patch_width_minus1[i] plus 1, and the value of patchHeight[i] is equal to patch_height_minus1[i] plus 1.

[0272] The width, height and position of each slice in the image in LCU units are obtained as follows: tempW = PatchWidth[LayerId] tempH = PatchHeight[LayerId] LcuSize = 1< <LcuSizeInBit PictureWidthInLcu[LayerId]=(HorizontalSize[LayerId]+LcuSize-1) / LcuSize PictureHeightInLcu[LayerId]=(VerticalSize[LayerId]+LcuSize-1) / LcuSize if(tempW> PictureWidthInLcu[LayerId]){ tempW=PictureWidthInLcu[LayerId]} if(tempH>PictureHeightInLcu[LayerId]){ tempH=PictureHeightInLcu[LayerId]} if(tempW <Min(2,PictureWidthInLcu[LayerId])){ tempW=Min(2,PictureWidthInLcu[LayerId])} numPatchColumns[LayerId]=PictureWidthInLcu[LayerId] / tempW numPatchRows[LayerId]=PictureHeightInLcu[LayerId] / tempH for(i=0;i<numPatchRows[LayerId];i++){ for(j=0;j<numPatchColumns[LayerId];j++){ PatchSizeInLcu[LayerId][i][j]-> Width=PatchWidth[LayerId] PatchSizeInLcu[LayerId][i][j]->Height=PatchHeight[LayerId] PatchSizeInLcu[LayerId][i][j]->PosX=(PatchWidth[LayerId]+1)*j PatchSizeInLcu[i][j]->PosY=(PatchHeight[LayerId]+1)*i}} temp=PictureHeightInLcu[LayerId]%tempH if(temp!=0){ for(j=0;j <numPatchColumns[LayerId];j++){ PatchSizeInLcu[LayerId][numPatchRows[LayerId]][j]->Width=tempW PatchSizeInLcu[LayerId][numPatchRows[LayerId]][j]->Height=temp PatchSizeInLcu[LayerId][numPatchRows[LayerId]][j]->PosX=tempW*j PatchSizeInLcu[LayerId][numPatchRows[LayerId]][j]->PosY=tempH*numPatchRows[LayerId]} numPatchRows[LayerId]++} temp=PictureWidthInLcu[LayerId]%tempW if(temp!=0){ for(j=0;j<numPatchRows[LayerId];j++){ PatchSizeInLcu[LayerId][j][numPatchColums[LayerId]-1]-> Width += temp}};

[0273] In this embodiment of the present application, 3 bits are used in the reserved bits of nal_unit to identify the spatial layer layer_id. If ssvc is not enabled and the base layer data is not enabled, the flag bit is set to 0. The layer_id of NAL units with nal_unit_type of 7, 8, 9, 11, and 16 is constrained to be 0.

[0274] At the same time, in order to constrain the complexity of codec products, the required maximum layer level can be set in different profiles. For example, referring to the NAL unit syntax definition shown in Table 1, the spatial layer identifier layer_id is included.

[0275] Table 1. NAL unit syntax definition

[0276] In one implementation, all layer-specific information (common to all layers) is placed in a sequence parameter set (SPS), including the number of supported layers, inter-layer dependencies, resolution of each layer, RPLs, and patch divisions. Only one SPS is transmitted (i.e., one SPS carries all layer-specific information). For example, Table 2 below illustrates the raw byte sequence payload (RSRP) of a sequence parameter set.

[0277] Table 2. Sequence parameter set RBSP definition

[0278] Based on the above, an embodiment of the present application provides a coding method for layered encoding of a current image. The current image includes multiple layered images (i.e., the current image has multiple layers). The encoder encodes each layered image to obtain multiple layered coded images (i.e., the coded image after encoding a layered image is called a layered coded image). The encoder encodes images in the same AU from lowest to highest layer identifier layer_id. This solution does not require modifications to block-level encoding and decoding operations, but supports flexible layered design through only high-level syntax changes.

[0279] As shown in FIG8 , the encoding method provided in this embodiment of the present application includes S801 - S802 .

[0280] S801: Encode multiple layered images of a current image to obtain multiple layered encoded images.

[0281] That is, construct the coded image of each layer, downsample the input image to the set resolution, that is, downsample the input image so that the horizontal size of the input image is horizontal_size[CurrLayerId] and the vertical size is vertical_size[CurrLayerId], and encode the downsampled image to obtain the coded image of the current layer. Among the multiple layers, the resolution of the highest layer should be consistent with the input image, and the resolution of the higher layer should be greater than or equal to the resolution of the lower layer.

[0282] For example, the current image is downsampled to obtain images with three resolutions, the three resolutions being the first resolution, the second resolution, and the third resolution, from smallest to largest. The third resolution is equal to the resolution of the input image (i.e., the original resolution before downsampling). The current image is layered encoded, the number of layered images is 3, and the layer number i is 0, 1, and 2. The resolution corresponding to layer 0 is the first resolution, the resolution corresponding to layer 1 is the second resolution, and the resolution corresponding to layer 2 is the third resolution.

[0283] It is understandable that during the encoding of multiple layered images of a current image, there may or may not be inter-layer dependencies between the multiple layered images. The existence of inter-layer dependencies between multiple layered images means that encoding one or more layered images of the current image requires reliance on other layered images (i.e., encoding the other layered images as reference images). The absence of inter-layer dependencies between multiple layered images means that the multiple layered images are encoded independently without using the layered images as reference images.

[0284] S802. Write multiple layered coded images and reference image queue information into a bitstream, where the reference image queue information includes first indication information; the first indication information is used to indicate a decoding order index (DOI) of a reference image of a layered coded image, where the decoding order index of the reference image is equal to the decoding order index of the layered coded image.

[0285] When the reference picture of a layered coded image is an inter-layer reference picture of the layered coded image, the decoding order index of the reference picture of the layered coded image is equal to the decoding order index of the layered coded image. The DOIs of different images in the temporal domain are different, while the DOIs of different layers of the same image in the spatial domain are the same.

[0286] After the encoder completes encoding of multiple layered images, it constructs a reference frame list, i.e., a reference picture list (RefPicDoiList) for the layered images, which includes the DOIs of the reference images. Thus, constructing the reference list is the first indication information for generating the layered coded images.

[0287] It is understood that during the encoding process of a layered image of a current image (referred to as the current layered image or the current layer), the reference image of the current layered image may include a temporal reference frame and / or a spatial reference frame. For ease of description, in the following embodiments, the temporal reference frame is referred to as an inter-frame reference frame, and the spatial reference frame is referred to as an inter-layer reference frame.

[0288] If sps_independent_layer_flag[CurrLayerId] is 1, that is, when the current layer (the layer indicated by CurrLayerId) does not use cross-layer reference (that is, the current layer is independently coded in the spatial domain), the frame coding of this layer can only use the coded frames of the same layer as the reference frames for inter-frame prediction. That is, the reference picture of the current layer is the layer picture with the same layer number as the current layer in multiple layers of the coded frame in the temporal domain. In other words, when sps_independent_layer_flag[CurrLayerId] is 1, the reference layer of the current layer is an inter-frame reference layer.

[0289] If sps_independent_layer_flag[CurrLayerId] is 0, that is, the current layer has inter-layer dependency, when encoding frames of this layer, the frame with layer number RefLayerId[CurrLayerId] in the same AU can be added to the reference frame list as a reference frame for inter-frame prediction. In other words, when sps_independent_layer_flag[CurrLayerId] is 0, the reference layer of the current layer is an inter-layer reference layer.

[0290] The layer identifier of the current layer can be recorded as the aforementioned CurrLayerId or curLayerId. In the following embodiments, CurrLayerId and curLayerId both represent the layer identifier of the current layer, and there is no substantial difference between the two.

[0291] The method for determining the information of the reference picture queue i (i is equal to 0 or 1) of the LayerId layer (LayerId is used to indicate a layer. For the current layer, the layer identifier of the current layer is recorded as curLayerId) is as follows (it should be noted that i in the following steps indicates the reference picture queue):

[0292] Initialize LibraryPictureExistFlag (knowledge image permission flag) and j (j indicates reference image j in reference image queue i) to 0, doiBase is the DOI of the current image (also the DIO of the layered image indicated by LayerId), and the value of j is 0 to NumOfRefPic[LayerId][i][RplsIndex[i]]-1. NumOfRefPic[LayerId][i][RplsIndex[i]] represents the number of reference images in the reference image queue configuration set i corresponding to the reference image queue i of the layered image indicated by LayerId (that is, it can be simply understood as the number of reference images of the current layer), and [RplsIndex[i] represents the reference image configuration set index corresponding to the reference image queue i, which is used to indicate the reference image configuration set i.

[0293] Perform the following operation NumOfRefPic[LayerId][i][RplsIndex[i]] times:

[0294] 1) If the value of LibraryIndexFlag[LayerId][i][RplsIndex[i]][j] is equal to 1 (that is, the reference image of the layered image indicated by LayerId is a knowledge image), then let RefPicDoiList[LayerId][i][j] be equal to ReferencedLibraryPictureIndex[LayerId][i][RplsIndex[i]][j], LibraryPictureExistFlag be equal to 1, the distance index DistanceIndex of the reference image j be equal to the POI of the current image minus 1 and multiplied by 2, and j be equal to j+1.

[0295] referenced_library_picture_index[LayerId][list][rpls][i] represents the referenced knowledge image index, i.e., it indicates the knowledge image referenced by the layered image indicated by LayerId. RefPicDoiList represents a list of reference image DOIs. For a particular image, this list contains the DOIs of all reference images for that image. Correspondingly, RefPicDoiList[LayerId][i][j] represents the DOI of reference image j in reference image list i for the layered image indicated by LayerId.

[0296] Execute the above step 1), and when the reference image of the layered image indicated by LayerId is a knowledge image, add the index of the knowledge image to RefPicDoiList.

[0297] 2) If the value of LibraryIndexFlag[LayerId][i][RplsIndex[i]][j] is equal to 0 (i.e., the reference image of the layered image indicated by LayerId is not a knowledge image), then let RefPicDoiList[LayerId][i][j] be equal to doiBase-DeltaDoi[LayerId][i][RplsIndex[i]][j], and let doiBase be equal to RefPicDoiList[LayerId][i][j]. DeltaDoi[LayerId][i][RplsIndex[i]][j] represents the difference between the DOI of the layered image indicated by LayerId and the DOI of the reference image of the layered image, and doiBase-DeltaDoi[LayerId][i][RplsIndex[i]][j] represents the DOI of the reference image of the layered image indicated by LayerId.

[0298] Furthermore, if RefPicDoiList[LayerId][i][j] is equal to the DOI of the current image (that is, the DOI of the reference image of the current layered image is equal to the DOI of the current layered image), then the distance index DistanceIndex of the reference image is equal to the POI of the current image minus 1 and multiplied by 2; otherwise, the distance index DistanceIndex is equal to the POI of the image whose DOI is RefPicDoiList[LayerId][i][j] multiplied by 2, and j is equal to j+1.

[0299] It can be seen that at this time, the distance index of the inter-layer reference frame is DistanceIndex, which is equal to the POI of the current image minus 1 and multiplied by 2. That is, the distance index of the 0th (when j=0) image in the reference image queue of the current layered coded image satisfies:

[0300] DistanceIndex=2×(POI_Currlayer-1), where DistanceIndex represents the distance index of the reference image, and POI_Currlayer represents the present order index (POI) of the current layered coded image. The POI of the current layered coded image is also the DOI of the current image and the DOI of the current layered image.

[0301] In an embodiment of the present application, in a layered coding scenario, the inter-layer reference frame may not be specifically identified. Instead, the inter-layer reference frame is located by taking advantage of the fact that the DOI of the reference frame is the same as the DOI of the current image, plus the reference layer number transmitted in the SPS (i.e., the spatial domain i-th layer dependency layer number ref_layer_id[i]). This saves bits and improves the encoding and decoding performance in the layered coding scenario.

[0302] In the encoding method provided in the embodiment of the present application, the knowledge image can be used as a long-term reference image. At the same time, according to whether the knowledge image is displayed, it is divided into a long-term reference (LTR) mode and a cross-random access point (RAP) reference (CRR) slicing mode. For example, Figure 9 is a schematic diagram of the encoding structure of the LTR mode (upper figure) and the CRR slicing mode (lower figure), wherein, in the LTR mode, the knowledge image is a display knowledge image, and all coded slices are in the same AU, which are numbered DOI / POI together with the display image to determine the playback order. In the CRR slicing mode, the knowledge image is a non-display knowledge image, and the code stream of the non-RL leading knowledge image is interwoven in the main video code stream in units of patch (slice), and each patch is an AU in a single layer. A non-RL leading knowledge image (non leading library picture of an RL picture) refers to a non-display knowledge image whose code stream order precedes an associated RL image, and an RL image (refrence library picture) refers to a P image or B image that uses only the knowledge image as a reference image for inter-frame prediction decoding).

[0303] In some embodiments, if the current picture uses a knowledge picture, each of the multiple layered coded pictures of the current picture has a knowledge picture. That is, if library_picture_enable_flag is equal to 1 (inter-frame prediction pictures that use knowledge pictures as reference pictures may exist in the coded video sequence), each layered coded picture has a knowledge picture.

[0304] In one implementation, an embodiment of the present application provides a two-layer reference-free hierarchical architecture. Figure 10 is a schematic diagram of the coding structure of the double-layer LTR (upper figure) and CRR slicing mode (lower figure). Since the P frame has two reference frames in the CRR mode, considering that the memory overhead of three reference frames is too large when the inter-layer reference is turned on, a two-layer reference-free hierarchical architecture is adopted (that is, the current image includes two layers, and the two layers are independently encoded). Both layers should have knowledge images. Of the two layered images of the current image, layer0 (layer 0) is a low-resolution coded image, and layer1 (layer 1) is a high-resolution coded image. The width and height of the low-resolution coded image are both 1 / 2 of the high-resolution image.

[0305] In some embodiments, when the knowledge image is a non-display knowledge image, the indexes of the image blocks of multiple layered coded images of the current image included in each access unit in the code stream are the same (i.e., the patchIndex is the same), and the image blocks here are slices (i.e., patches).

[0306] In some embodiments, when the knowledge image is a non-display knowledge image, the number of image blocks of multiple layered coded images of the current image included in each access unit in the code stream is the same, that is, the number of patches of different layers of the knowledge image mode 0 (slice mode) needs to be the same, and the image block here refers to a slice (i.e., patch).

[0307] Specifically, if LibraryStreamFlag is equal to 1, and LibraryPictureModeIndex is equal to 0, and NumOfLayers is greater than 1, for any unsigned integers k and m less than NumOfLayers (k is not equal to m), numPatchColumns[k] multiplied by numPatchRows[k] should be equal to numPatchColumns[m] multiplied by numPatchRows[m], that is, the number of patches contained in layer k of the current image is the same as the number of patches contained in layer m.

[0308] Refer to Figure 11, which shows the codestream structure for LTR mode (top) and CRR slicing mode (bottom). In CRR slicing mode, each patch of the non-RL pre-image is an AU, interleaved with the codestream of the display image. To ensure that the coded image at the same time is in the same AU, the non-display image of layer 1 and layer 0 must be divided into the same number of patches, and the layer units with the same patchIndex constitute an access unit, which is interleaved with the codestream of other display images.

[0309] In the coding method provided in the embodiment of the present application, a privacy protection method is used to distinguish between privacy protection modes (i.e., permission protection modes) in the use permission protection scenario. According to the description of the above embodiment, privacy_mode_index is used to indicate the permission protection mode, wherein mode 0 and mode 1 are privacy protection modes based on a hierarchical structure. This privacy protection mode divides the original image into a privacy image and a non-privacy image through pre-processed privacy detection, adopts a layered coding architecture, takes the non-privacy image as the base layer image, and takes the full image (or privacy image) as the enhancement layer image, and performs layered coding on both. Depending on whether cross-layer reference is adopted, it is divided into two modes, mode 0 (privacy_mode_index=0) is an independent coding mode based on a layered architecture, and mode 1 (i.e., privacy_mode_index=1) is a reference coding mode based on a layered architecture.

[0310] For example, Figure 12 is an algorithm flow chart for the privacy protection mode 0, and Figure 13 is an algorithm flow chart for the privacy protection mode 1. Referring to Figure 12, when the privacy protection mode is mode 0, no processing is performed on the input image (i.e., input YUV), and the original image (i.e., full image or enhanced layer image) is used for encoding to obtain a complete bitstream (also referred to as an enhanced layer bitstream), which is then transmitted to the decoding end. The decoder decodes the complete bitstream to obtain a complete image. In addition, the privacy area of ​​the input image (i.e., input YUV) is mosaic-masked (i.e., mosaic-masked), and the mosaic-masked image is encoded to obtain a privacy-protected bitstream (also referred to as a base layer bitstream), which is then transmitted to the decoding end. The decoder decodes the privacy-protected bitstream to obtain a mosaic-masked image (i.e., a privacy image).

[0311] Referring to Figure 13 , when the privacy protection mode is Mode 1, front-end detection is performed on the original captured image, dividing it into a privacy image (enhancement layer image) and a non-privacy image (base layer image), and obtaining information indicating the privacy region. The base layer image is then encoded to obtain a base layer codestream. When encoding the enhancement layer image, the reconstructed image of the base layer image (i.e., the base layer reconstructed image) serves as a reference image for the enhancement layer image (i.e., the enhancement layer image references the base layer image) to obtain an enhancement layer codestream. Information indicating the privacy region is also encoded. The base layer codestream, the enhancement layer codestream, and the encoded privacy region information (the encoded privacy region information is optional, meaning it can be transmitted or not) are interleaved to output a single codestream, which is then transmitted to the decoder. At the decoder, the enhancement layer codestream and base layer codestream are decoded to obtain the enhancement layer reconstructed image and base layer reconstructed image, respectively. The decoding process is the inverse of the encoding process and will not be further described here.

[0312] Correspondingly, the code stream obtained in S802 in the above embodiment also includes second indication information (PrivacyEnableFlag) and third indication information (PrivacyModeIndex). The second indication information is used to indicate whether the current image uses permission protection, and the third indication information is used to indicate the permission protection mode of the current image.

[0313] In some embodiments, when the second indication information indicates that the current image is protected by rights, and the third indication information indicates that the rights protection mode is layered independent coding mode or layered reference coding mode, the number of multiple layered coded pictures of the current image is 2 or 4. That is, if the value of PrivacyEnableFlag is equal to 1 and the value of PrivacyModeIndex is 0 or 1, the value of NumOfLayers should be equal to 2 or 4.

[0314] In the above embodiment, the code stream obtained in S802 also includes fourth indication information (SpsAllIndependentLayersFlag) and fifth indication information (SpsIndependentLayerFlag). The fourth indication information is used to indicate whether multiple layered coded images of the current image have inter-layer dependencies; the fifth indication information is used to indicate whether a layered coded image is independently coded.

[0315] In some embodiments, when the second indication information indicates that the current image is protected by rights and the third indication information indicates that the rights protection mode is layered independent coding mode, the fourth indication information indicates that there is no inter-layer dependency between the multiple layered coded images. That is, if the value of PrivacyEnableFlag is equal to 1 and the value of PrivacyModeIndex is equal to 0, the value of SpsAllIndependentLayersFlag should be '1'.

[0316] In some embodiments, when the second indication information indicates that the current image is protected by rights, the third indication information indicates that the rights protection mode is a layered reference coding mode, and the number of the plurality of layered coded images is 2, the fifth indication information indicates that the layered coded images with even layer identifiers in the plurality of layered coded images are independently coded, so that the layered coded images with odd layer identifiers are dependent on the layered coded images with even layer identifiers. That is, if the value of PrivacyEnableFlag is equal to 1, the value of PrivacyModeIndex is equal to 1, and the value of NumOfLayers is equal to 2, then the value of SpsIndependentLayerFlag[1] shall be '0', and the value of SpsIndependentLayerFlag[0] shall be '1'.

[0317] In some embodiments, when the second indication information indicates that the current image is protected by rights, the third indication information indicates that the rights protection mode is a layered reference coding mode, and the number of the plurality of layered coded images is 4, the layered coded images with even-numbered layer identifiers in the plurality of layered coded images are independently encoded; the layered coded images with odd-numbered layer identifiers in the plurality of layered coded images are not independently encoded, and the layered coded images with odd-numbered layer identifiers are dependent on the layered coded images with even-numbered layer identifiers. That is, if the value of PrivacyEnableFlag is equal to 1, the value of PrivacyModeIndex is equal to 1, and NumOfLayers is equal to 4, then the value of SpsIndependentLayerFlag[0] should be '1', the value of SpsIndependentLayerFlag[1] should be '0', the value of SpsIndependentLayerFlag[2] should be '1', and the value of SpsIndependentLayerFlag[3] should be '0'.

[0318] In some embodiments, when the second indication information indicates that the current image is protected by usage rights and the third indication information indicates that the rights protection mode is a layered reference coding mode, the resolution of the current layered coded image among multiple layered coded images is the same as the resolution of the reference image of the current layered coded image.

[0319] The resolution of a current layered coded image among the multiple layered coded images is the same as the resolution of a reference image of the current layered coded image, specifically including:

[0320] If the value of PrivacyEnableFlag is equal to 1 and the value of PrivacyModeIndex is equal to 1, a bitstream conforming to this application shall satisfy the value of HorizontalSize[RefLayerId[i]] (the horizontal size of the reference image) is equal to the value of HorizontalSize[i] (the horizontal size of the current layered coded image) and the value of VerticalSize[RefLayerId[i]] (the vertical size of the reference image) is equal to the value of VerticalSize[i] (the vertical size of the current layered coded image).

[0321] If the value of PrivacyModeFlag is 1, the value of PrivacyModeIndex is 0 or 1, and NumOfLayers is equal to 2 or 4, HorizontalSize[0] should be equal to HorizontalSize[1], that is, the horizontal size of layer 0 is equal to the horizontal size of layer 1.

[0322] If the value of PrivacyModeFlag is 1, the value of PrivacyModeIndex is 0 or 1, and NumOfLayers is equal to 4, HorizontalSize[2] should be equal to HorizontalSize[3], that is, the horizontal size of layer 2 is equal to the horizontal size of layer 3.

[0323] If the value of PrivacyModeFlag is 1, the value of PrivacyModeIndex is 0 or 1, and NumOfLayers is equal to 2 or 4, VerticalSize[0] should be equal to VerticalSize[1], that is, the vertical size of layer 0 is equal to the vertical size of layer 1.

[0324] If the value of PrivacyModeFlag is 1, the value of PrivacyModeIndex is 0 or 1, and NumOfLayers is equal to 4, VerticalSize[2] should be equal to VerticalSize[3], that is, the vertical size of layer 2 is equal to the vertical size of layer 3.

[0325] If the value of PrivacyModeFlag is 1, the value of PrivacyModeIndex is 0 or 1, and NumOfLayers is equal to 2 or 4, PicWidthInLuma[0] should be equal to PicWidthInLuma[1], that is, the horizontal coding size of the luminance component of layer 0 is equal to the horizontal coding size of the luminance component of layer 1.

[0326] If the value of PrivacyModeFlag is 1, the value of PrivacyModeIndex is 0 or 1, and NumOfLayers is equal to 4, PicWidthInLuma[2] should be equal to PicWidthInLuma[3], that is, the horizontal coding size of the luminance component of the second layer is equal to the horizontal coding size of the luminance component of the third layer.

[0327] If the value of PrivacyModeFlag is 1, the value of PrivacyModeIndex is 0 or 1, and NumOfLayers is equal to 2 or 4, PicHeightInLuma[0] should be equal to PicHeightInLuma[1], that is, the vertical encoding size of the luminance component of layer 0 is equal to the vertical encoding size of the luminance component of layer 1.

[0328] If the value of PrivacyModeFlag is 1, the value of PrivacyModeIndex is 0 or 1, and NumOfLayers is equal to 4, PicHeightInLuma[2] should be equal to PicHeightInLuma[3], that is, the vertical encoding size of the luminance component of the second layer is equal to the vertical encoding size of the luminance component of the third layer.

[0329] In the embodiment of the present application, the number of layers can be expanded to 4 layers, where the coded image resolutions of layer0 and layer1 are the same, i.e., low-resolution coded images, and the coded image resolutions of layer2 and layer3 are the same, i.e., high-resolution coded images. As can be seen, each resolution has both private and non-private code streams, meeting the needs of different users. For example, a user with high privileges can choose to decode either a low-resolution code stream or a high-resolution code stream. A user with low privileges can choose to decode either a low-resolution non-private code stream or a high-resolution non-private code stream.

[0330] Furthermore, the privacy area can be set. Layer0 and layer2 mosaic the privacy area through pre-processed privacy detection and encode and transmit it as a non-privacy image, while layer1 and layer3 do not perform any processing and use the original image for encoding and transmission.

[0331] Furthermore, an inter-layer reference mode can be used. For mode 0, all layers are independently encoded without inter-layer reference. For mode 1, the principle is that layers with high privacy permissions can refer to layers with low privacy permissions, but layers with low privacy permissions cannot refer to layers with high privacy permissions. Frames of layer1 can refer to frames of layer0 in the same AU when performing inter-frame prediction, i.e., RefLayerId[1]=0, layer0 is a layer with low privacy permissions (also called a low permission layer), and layer2 is a layer with high privacy permissions (also called a high permission layer). Frames of layer3 can refer to frames of layer2 in the same AU when performing inter-frame prediction, and frames of layer2 and layer0 are independently encoded, i.e., RefLayerId[3]=2.

[0332] As shown in Figure 14, an embodiment of the present application provides a decoding method for decoding multiple layered coded images of a current image. Based on the process of the encoding method described from the perspective of the encoding end, the decoding method provided by the embodiment of the present application includes S1401-S1403.

[0333] S1401. Parse the code stream to obtain the current layered coded image and reference image queue information of the current layered coded image, wherein the reference image queue information includes first indication information, and the first indication information is used to indicate the decoding order index (DOI) of the reference image of the current layered coded image.

[0334] It can be understood that parsing the code stream can obtain multiple layered coded images of the current image. This application takes a layered coded image (referred to as the current layered coded image) as an example to introduce the process of decoding the layered coded image.

[0335] In the embodiment of the present application, the image header is first decoded to obtain the display order index POI of the current layered coded image. Regarding the calculation of POI: if CurrLayerId is greater than 0 and sps_independent_layer_flag[CurrLayerId] is equal to 0 (inter-layer dependency is possible), and there is a picture picA in the current AU whose layerId is equal to RefLayerId, currPOI is equal to the POI of picA. Otherwise, it is derived according to the derivation method of the base layer: POI is equal to DOI+PictureOutputDelay-OutputReorderDelay+256×DOICycleCnt. Among them, PictureOutputDelay represents the image output delay (picture_output_delay), that is, the time required to wait from the completion of image decoding to output; OutputReorderDelay represents the image reordering delay (output_reorder_delay), which is the reordering delay caused by the inconsistency between the image encoding and decoding order and the display order; DOICycleCn represents the number of DOI flips.

[0336] Next, the reference picture queue information is exported (ie, the RefPicDoiList is exported). For details about the introduction of the RefPicDoiList, please refer to the relevant description in the above S801.

[0337] S1402. When the decoding order index of the reference image is equal to the decoding order index of the current layered coded image, the image in the decoded image buffer having the same reference layer identifier (RefLayerId[curLayerId]) as the current layered coded image and the same first indication information as the current layered coded image is added to the reference image queue of the current layered coded image.

[0338] In the embodiment of the present application, after the reference image queue information is derived, the decoded image buffer should be updated before decoding the current image. If the current image is a reference image, the decoded image buffer should not be updated.

[0339] For all displayed pictures in the decoded picture buffer that have the same LayerId as the current picture, if the picture is neither in reference picture queue 0 nor in reference picture queue 1, mark the picture as "not referenced"; otherwise, if the picture is in reference picture queue 0 or reference picture queue 1, mark the picture as "referenced". If a picture is marked as "not referenced", it should not be marked as "referenced" again in the future.

[0340] For a display image in the decoded image buffer that has the same DOI as the current image, if the image is in reference image queue 0 or reference image queue 1, the image is marked as "referenced"; otherwise, the original mark remains unchanged.

[0341] Perform the following operations on the decoded image buffer in sequence:

[0342] Remove the image:

[0343] 1) Remove all display images marked as "not referenced" and "output" from the decoded image buffer;

[0344] 2) If there is a knowledge image in the decoded image buffer, and the LibraryPictureExistFlag of the current image is equal to 1, and the knowledge image index of the knowledge image referenced by the current image is different from the knowledge image index of the knowledge image in the decoded image buffer, or the current image is a knowledge image, then the knowledge image in the decoded image buffer is removed and LibraryBufferEmpty is set to 1 (LibraryBufferEmpty is 1 indicating that there is no knowledge image in the decoded buffer).

[0345] Move in the image:

[0346] If LibraryPictureModeIndex is 2, LibraryBufferEmpty is 1, and the LibraryPictureExistFlag of the current image is equal to 1, the corresponding knowledge image is moved from the outside into the decoded image buffer and LibraryBufferEmpty is set to 0.

[0347] A bitstream conforming to this document shall meet the following requirements:

[0348] The temporal_id of the first NumRefActive[i] reference pictures in reference picture queue i (i is equal to 0 or 1) should be less than or equal to the temporal_id of the current picture. NumRefActive[i] represents the number of active reference pictures, that is, the maximum reference index of reference picture queue i when decoding the current picture. temporal_id represents the temporal layer identifier.

[0349] The LayerId of the first NumRefActive[i] reference pictures in reference picture queue i (i is equal to 0 or 1) should be less than or equal to the LayerId of the current picture.

[0350] During decoding, the total number of the current decoded image, "not yet output" display images, and knowledge images should not exceed the value of MaxDpbSize (buffer image number constraint). MaxDpbSize represents the maximum decoded image buffer size, which indicates the maximum decoded image buffer size required to decode the current bitstream.

[0351] After completing the update of the decoded image buffer, the decoding end constructs a reference image queue. The method for constructing the reference image queue i (i is equal to 0 or 1) of the LayerId layer (the LayerId layer is the current layer) is as follows:

[0352] Initialize j to 0.

[0353] Perform the following operation NumOfRefPic[LayerId][i][RplsIndex[i]] times:

[0354] 1) If the value of LibraryIndexFlag[LayerId][i][RplsIndex[i]][j] is equal to 1 (that is, the reference image of the current layered coded image is a knowledge image), move the knowledge image with the knowledge image index equal to RefPicDoiList[LayerId][i][j] from the decoded image buffer to the j-th position in the reference image queue i.

[0355] In one implementation, the knowledge image whose index is equal to RefPicDoiList[LayerId][i][j] and whose layer identifier is the same as the layer identifier of the current layered coded image is moved from the decoded image buffer to the j-th position in the reference image queue i.

[0356] 2) Otherwise (if the value of LibraryIndexFlag[LayerId][i][RplsIndex[i]][j] is equal to 0, that is, the reference picture of the current layered coded picture is not a knowledge picture), the reference picture queue is constructed according to the following operations:

[0357] a) If RefPicDoiList[LayerId][i][j] is equal to the DOI of the current picture, move the display picture in the decoded picture buffer whose decoding order index is equal to RefPicDoiList[LayerId][i][j] and whose LayerId is equal to RefLayerId[LayerId] to position j in reference picture queue i. A conforming bitstream shall satisfy that the display picture shall be present in the decoded picture buffer.

[0358] b) Move the display picture in the decoded picture buffer whose decoding order index is equal to RefPicDoiList[LayerId][i][j] and whose LayerId is the same as the current picture (the layer identifier is the same as the layer identifier of the current layered coded picture) into position j in reference picture queue i. Bitstreams conforming to this document shall satisfy that the display picture shall be present in the decoded picture buffer.

[0359] 3)j is equal to j+1.

[0360] S1403: Decode the current layered coded image based on the reference image queue.

[0361] In an embodiment of the present application, in conjunction with the encoding process, the decoding end can also parse the bitstream to obtain more indication information (syntax elements used for decoding), for example, to obtain second indication information, third indication information, fourth indication information, and fifth indication information from the bitstream, and decode the current layered coded image using a processing method that is reverse to that of the encoding end according to the instructions of each indication information.

[0362] As can be understood, after constructing the reference image queue for the current layered coded image, the distance index of the reference image is calculated, and based on the distance index of the current layered coded image and the distance index of the reference image, the current coded image is decoded. The image decoding process can be referenced to the decoding process in existing literature and will not be described in detail in this embodiment of the application.

[0363] For the relevant contents of S1401 to S1403, please refer to the above description of S801-S802, which will not be repeated here.

[0364] It is understandable that, in order to implement the above functions, the encoding device or decoding device includes hardware and / or software modules that perform the corresponding functions. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0365] In an embodiment of the present application, the encoding device or decoding device can divide the functional modules according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. In actual implementation, there may be other division methods.

[0366] In the case of dividing the functional modules according to the corresponding functions, Figure 15 shows a possible composition diagram of the encoding device involved in the above embodiment. As shown in Figure 15, the encoding device 1500 may include: an encoding unit 1501 and a processing unit 1502.

[0367] The encoding unit 1501 and the processing unit 1502 cooperate to perform S801 - S802 and more steps in the above method embodiment.

[0368] In the case of dividing the functional modules according to their respective functions, FIG16 shows a possible composition diagram of the decoding device involved in the above embodiment. As shown in FIG16 , the decoding device 1600 may include: a processing unit 1601 and a decoding unit 1602.

[0369] The processing unit 1601 and the decoding unit 1602 cooperate to execute S1401 - S1403 and more steps in the above method embodiment.

[0370] The present application also provides a chip. FIG17 shows a schematic diagram of the structure of a chip 1700. The chip 1700 includes one or more processors 1701 and an interface circuit 1702. Optionally, the chip 1700 may also include a bus 1703.

[0371] The processor 1701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above encoding and decoding method can be completed by a hardware integrated logic circuit in the processor 1701 or by software instructions.

[0372] Optionally, the processor 1701 may be a general-purpose processor, a digital signal processing (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods and steps disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor.

[0373] The interface circuit 1702 can be used to send or receive data, instructions or information. The processor 1701 can use the data, instructions or other information received by the interface circuit 1702 to process it, and can send the processing completion information through the interface circuit 1702.

[0374] Optionally, the chip also includes a memory, which may include a read-only memory and a random access memory, and provides operating instructions and data to the processor. Part of the memory may also include a non-volatile random access memory (NVRAM).

[0375] Optionally, the memory stores an executable software module or a data structure, and the processor can perform corresponding operations by calling an operation instruction stored in the memory (the operation instruction may be stored in an operating system).

[0376] Optionally, the chip can be used in an image processing device according to an embodiment of the present application. Optionally, the interface circuit 1702 can be used to output the execution result of the processor 1701. For the encoding and decoding methods provided in one or more embodiments of the present application, reference can be made to the aforementioned embodiments and will not be repeated here.

[0377] It should be noted that the corresponding functions of the processor 1701 and the interface circuit 1702 can be implemented through hardware design, software design, or a combination of hardware and software, and there is no limitation here.

[0378] FIG18 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1800 may be a processor or a chip or functional module in a processor. As shown in FIG18 , the electronic device 1800 includes a processor 1801 , a transceiver 1802 , and a communication line 1803 .

[0379] Among them, the processor 1801 is used to execute any step in the encoding and decoding method provided in the embodiment of the present application, and in the process of executing any step in the encoding and decoding provided in the embodiment of the present application, the transceiver 1802 and the communication line 1803 can be optionally called to complete the corresponding operation.

[0380] Furthermore, the electronic device 1800 may further include a memory 1804 . The processor 1801 , the memory 1804 and the transceiver 1802 may be connected via a communication line 1803 .

[0381] The processor 1801 is a processor, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 1801 may also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.

[0382] Transceiver 1802 is used to communicate with other devices or other communication networks, such as Ethernet, radio access networks (RAN), wireless local area networks (WLAN), etc. Transceiver 1802 can be a module, circuit, transceiver, or any device capable of implementing communication.

[0383] The transceiver 1802 is mainly used for sending and receiving commands and information, etc., and may include a transmitter and a receiver for sending and receiving commands and information, etc. respectively; operations other than sending and receiving commands and information, etc. are implemented by the processor.

[0384] The communication line 1803 is used to transmit information between the various components included in the electronic device 1800.

[0385] In one design, the processor can be considered as the logic circuit and the transceiver as the interface circuit.

[0386] The memory 1804 is used to store instructions, where the instructions may be computer programs.

[0387] Memory 1804 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM may be used, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). Memory 1804 may also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage media or other magnetic storage devices, etc. It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0388] It should be noted that memory 1804 can exist independently of processor 1801 or can be integrated with processor 1801. Memory 1804 can be used to store instructions, program code, or data. Memory 1804 can be located within or outside electronic device 1800, without limitation. Processor 1801 is configured to execute instructions stored in memory 1804 to implement the methods provided in the above embodiments of this application.

[0389] In one example, processor 1801 may include one or more processors, such as processor 0 (CPU0) and processor 1 (CPU1) in FIG. 18 .

[0390] As an optional implementation, the electronic device 1800 includes multiple processors. For example, in addition to the processor 1801 in FIG. 18 , it may also include a processor 1807 .

[0391] As an optional implementation, the electronic device 1800 further includes an output device 1805 and an input device 1806. For example, the input device 1806 is a keyboard, a mouse, a microphone, or a joystick, and the output device 1805 is a display screen, a speaker, or the like.

[0392] It should be pointed out that the electronic device 1800 can be a chip system or a device with a similar structure as shown in Figure 18. Among them, the chip system can be composed of chips, or it can include chips and other discrete devices. The actions, terms, etc. involved in the various embodiments of this application can refer to each other without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are only an example. Other names can also be used in the specific implementation without limitation. In addition, the component structure shown in Figure 18 does not constitute a limitation on the electronic device 1800. In addition to the components shown in Figure 18, the electronic device 1800 may include more or fewer components than those shown in Figure 18, or combine certain components, or arrange the components differently.

[0393] The processor and transceiver described in this application can be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit, a mixed-signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (NMOS), P-type metal oxide semiconductor (positive channel metal oxide semiconductor, PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0394] Figure 19 is a schematic diagram of the structure of an image processing device provided in an embodiment of the present application. The image processing device can be applied to the scenario shown in the above method embodiment. For ease of explanation, Figure 19 only shows the main components of the image processing device, including a processor 1901, a memory 1902, a control circuit 1903, and an input-output device 1904. The processor 1901 is mainly used to process communication protocols and communication data, execute software programs, and process data of software programs. The memory 1902 is mainly used to store software programs and data. The control circuit 1903 is mainly used for power supply and transmission of various electrical signals. The input-output device 1904 is mainly used to receive data input by the user and output data to the user.

[0395] When the image processing device is a processor 1901, the control circuit 1903 may be a motherboard, the memory 1902 includes a hard disk, RAM, ROM and other media with storage functions, the processor 1901 may include a baseband processor 1901 and a central processing unit, the baseband processor is mainly used to process the communication protocol and communication data, the central processing unit is mainly used to control the entire encoding device, execute software programs, and process software program data, the input and output devices 1904 include a display screen, a keyboard, and a mouse, etc.; the control circuit 1903 may further include or be connected to a transceiver circuit or transceiver, such as a network cable interface, for sending or receiving data or signals, such as for data transmission and communication with other devices. Furthermore, it may also include an antenna for sending and receiving wireless signals for data / signal transmission with other devices.

[0396] The present application also provides an encoding device, comprising: at least one processor, wherein when the at least one processor executes program code or instructions, the at least one processor implements the aforementioned related method steps to implement the encoding method of the aforementioned embodiment. Optionally, the encoding device may also include at least one memory for storing the program code or instructions.

[0397] The present application also provides a decoding device, comprising: at least one processor, which, when executing program code or instructions, implements the above-mentioned related method steps to implement the decoding method of the above-mentioned embodiment. Optionally, the decoding device may also include at least one memory for storing the program code or instructions.

[0398] An embodiment of the present application also provides a computer storage medium, which stores computer instructions. When the computer instructions are executed on an encoding device or a decoding device, the encoding device or the decoding device executes the above-mentioned related method steps to implement the encoding and decoding method in the above-mentioned embodiment.

[0399] The embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement the encoding and decoding method in the above-mentioned embodiment.

[0400] The present application also provides an encoding device, which may be a chip, integrated circuit, component, or module. Specifically, the device may include a connected processor and a memory for storing instructions, or the encoding device may include at least one processor for retrieving instructions from an external memory. When the device is running, the processor may execute the instructions, causing the chip to perform the encoding method described in each of the above method embodiments.

[0401] The present application also provides a decoding device, which can be a chip, integrated circuit, component, or module. Specifically, the device can include a connected processor and a memory for storing instructions, or the decoding device can include at least one processor for retrieving instructions from an external memory. When the device is running, the processor can execute the instructions, causing the chip to perform the decoding method described in each of the above method embodiments.

[0402] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions in accordance with the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a magnetic disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state drive (SSD)).

[0403] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0404] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0405] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0406] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0407] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.

[0408] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A decoding method, characterized in that, It includes: Parse the bitstream to obtain the current layer-coded image and the reference image queue information of the current layer-coded image; The reference image queue information includes first indication information; The first indication information is used to indicate the decoding order index DOI of the reference image of the current layer-coded image; When the decoding order index DOI of the reference image is equal to the decoding order index DOI of the current layer-coded image, add the image in the decoding image buffer whose layer identifier is the same as the reference layer identifier of the current layer-coded image and is the same as the first indication information of the current layer-coded image to the reference image queue of the current layer-coded image; Decode the current layer-coded image based on the reference image queue.

2. The method according to claim 1, wherein The bitstream includes multiple layer-coded images of the current image.

3. The method according to claim 1 or 2, wherein When the decoding order index DOI of the reference image is equal to the decoding order index DOI of the current layer-coded image, the distance index of the reference image of the current layer-coded image satisfies: DistanceIndex = 2×(POI_Currlayer - 1), where DistanceIndex represents the distance index and POI_Currlayer represents the display order index of the current layer-coded image.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Parse and obtain second indication information and third indication information from the bitstream; Wherein, the second indication information is used to indicate whether the current image uses rights protection; the third indication information is used to indicate the rights protection mode of the current image, and the rights protection mode includes any one of the following: layer-independent coding mode, layer-reference coding mode or single-layer coding mode.

5. The method according to claim 4, wherein When the second indication information indicates that the current image uses rights protection and the third indication information indicates that the rights protection mode is the layer-independent coding mode or the layer-reference coding mode, the number of the multiple layer-coded images of the current image is 2 or 4.

6. The method according to claim 4, wherein When the second indication information indicates that the current image uses rights protection and the third indication information indicates that the rights protection mode is the layer-independent coding mode, there is no inter-layer dependency among the multiple layer-coded images.

7. The method according to claim 4, wherein The value of the fourth indication information in the bitstream is a first value, and the first value is used to indicate that there is no inter-layer dependency among the multiple layer-coded images.

8. The method according to claim 4, wherein When the second indication information indicates that the current image uses rights protection and the third indication information indicates that the rights protection mode is the layer-independent coding mode or the layer-reference coding mode, the layer-coded image with an even layer identifier is a low-rights layer, and the layer-coded image with an odd layer identifier is a high-rights layer.

9. The method according to claim 4, wherein when the second indication information indicates that the current image uses permission protection, the third indication information indicates that the permission protection mode is a hierarchical reference coding mode, and the number of the multiple hierarchical coding images is 2, the hierarchical coding images with even hierarchical identifiers among the multiple hierarchical coding images are independently coded, and the hierarchical coding images with odd hierarchical identifiers are coded with reference to the hierarchical coding images with even hierarchical identifiers.

10. The method according to claim 4, wherein when the second indication information indicates that the current image uses permission protection, the third indication information indicates that the permission protection mode is a hierarchical reference coding mode, and the number of the multiple hierarchical coding images is 4, the hierarchical coding images with even hierarchical identifiers among the multiple hierarchical coding images are independently coded; the hierarchical coding images with odd hierarchical identifiers among the multiple hierarchical coding images are not independently coded, and the hierarchical coding images with odd hierarchical identifiers are dependent on the hierarchical coding images with even hierarchical identifiers.

11. The method according to claim 4, wherein when the second indication information indicates that the current image uses permission protection, the third indication information indicates that the permission protection mode is a hierarchical reference coding mode, the resolution of the current hierarchical coding image in the multiple hierarchical coding images is the same as that of the reference image of the current hierarchical coding image.

12. The method according to any one of claims 1 to 11, wherein if the current image uses a knowledge image, each of the multiple hierarchical coding images has a knowledge image.

13. The method according to claim 12, wherein when the knowledge image is a non-display knowledge image, the indexes of the image blocks of the multiple hierarchical coding images of the current image included in each access unit in the bitstream are the same.

14. The method according to claim 13, wherein when the knowledge image is a non-display knowledge image, the number of the image blocks of the multiple hierarchical coding images of the current image included in each access unit in the bitstream is the same.

15. A coding method, characterized in that, comprising: encoding multiple hierarchical images of a current image to obtain multiple hierarchical coding images; writing the multiple hierarchical coding images and reference image queue information into a bitstream; the reference image queue information includes first indication information; the first indication information is used to indicate the decoding order index DOI of the reference image of a hierarchical coding image, and when the reference image is an inter-layer reference image of the current hierarchical coding image, the decoding order index DOI of the reference image is equal to the decoding order index DOI of the hierarchical coding image.

16. The method according to claim 15, wherein when the decoding order index DOI of the reference image is equal to the decoding order index DOI of the current hierarchical coding image, the distance index of the reference image of the hierarchical coding image satisfies: DistanceIndex = 2×(POI_Currlayer - 1), where DistanceIndex represents the distance index and POI_Currlayer represents the display order index of the hierarchical encoded image.

17. The method according to claim 15 or 16, characterized in that the bitstream further includes: second indication information and / or third indication information; wherein, the second indication information is used to indicate whether the current image uses rights protection; the third indication information is used to indicate the rights protection mode of the current image, and the rights protection mode includes any one of the following: hierarchical independent coding mode, hierarchical reference coding mode or single-layer coding mode.

18. The method according to claim 17, characterized in that the bitstream further includes fourth indication information and / or fifth indication information; the fourth indication information is used to indicate whether there is an inter-layer dependency among multiple hierarchical encoded images of the current image; the fifth indication information is used to indicate whether a hierarchical encoded image is independently encoded; when the second indication information indicates that the current image uses rights protection and the third indication information indicates that the rights protection mode is the hierarchical independent coding mode or the hierarchical reference coding mode, the number of multiple hierarchical encoded images of the current image is 2 or 4.

19. The method according to claim 18, characterized in that when the second indication information indicates that the current image uses rights protection and the third indication information indicates that the rights protection mode is the hierarchical independent coding mode, the value of the fourth indication information is a first value, and the first value is used to indicate that there is no inter-layer dependency among the multiple hierarchical encoded images.

20. The method according to claim 18, characterized in that when the second indication information indicates that the current image uses rights protection and the third indication information indicates that the rights protection mode is the hierarchical independent coding mode or the hierarchical reference coding mode, the hierarchical encoded image with an even hierarchical identifier is a low-rights layer, and the hierarchical encoded image with an odd hierarchical identifier is a high-rights layer.

21. The method according to claim 18, characterized in that when the second indication information indicates that the current image uses rights protection, the third indication information indicates that the rights protection mode is the hierarchical reference coding mode, and the number of the multiple hierarchical encoded images is 2, the fifth indication information indicates that the hierarchical encoded image with an even hierarchical identifier among the multiple hierarchical encoded images is independently encoded, and the hierarchical encoded image with an odd hierarchical identifier depends on the hierarchical encoded image with an even hierarchical identifier.

22. The method according to claim 18, characterized in that When the second indication information indicates that the current image uses permission protection, the third indication information indicates that the permission protection mode is the hierarchical reference coding mode, and the number of the multiple hierarchical coded images is 4, the fifth indication information indicates that the hierarchical coded images with even hierarchical identifiers among the multiple hierarchical coded images are independently coded, and the hierarchical coded images with odd hierarchical identifiers among the multiple hierarchical coded images are not independently coded, and the hierarchical coded images with odd hierarchical identifiers are dependent on the hierarchical coded images with even hierarchical identifiers.

23. The method according to claim 18, wherein When the second indication information indicates that the current image uses permission protection, and the third indication information indicates that the permission protection mode is the hierarchical reference coding mode, the resolution of one hierarchical coded image among the multiple hierarchical coded images is the same as that of the reference image of the hierarchical coded image.

24. The method according to any one of claims 15 to 23, wherein If the current image uses a knowledge image, each of the multiple hierarchical coded images has a knowledge image.

25. The method according to claim 24, wherein When the knowledge image is a non-display knowledge image, the indexes of the image blocks of the multiple hierarchical coded images of the current image included in each access unit in the bitstream are the same.

26. The method according to claim 24, wherein When the knowledge image is a non-display knowledge image, the number of the image blocks of the multiple hierarchical coded images of the current image included in each access unit in the bitstream is the same.

27. A decoding device, comprising at least one processor and a memory, characterized in that, The at least one processor executes the program or instruction stored in the memory, so that the decoding device implements the method according to any one of claims 1 to 14.

28. An encoding device, comprising at least one processor and a memory, characterized in that, The at least one processor executes the program or instruction stored in the memory, so that the encoding device implements the method according to any one of claims 15 to 26.

29. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program runs on a computer or a processor, the computer or the processor implements the method according to any one of claims 1 to 26.

30. A computer program product comprising instructions, characterized in that, When the instruction runs on a computer or a processor, the computer or the processor implements the method according to any one of claims 1 to 26.

31. A chip, comprising at least one processor and a memory, characterized in that, The at least one processor executes the program or instruction stored in the memory, so that the chip implements the method according to any one of claims 1 to 26.

Citation Information

Patent Citations

  • Method and device for specifying reference image and method and device for processing reference image request

    CN110876083A

  • Reference picture management in layered video coding

    CN113892266A

  • Image encoding and decoding method and device

    CN116033147A

  • Identification of inter-layer reference pictures in coded video

    US20230111484A1

  • Image encoding / decoding method and apparatus for determining sublayer on basis of whether or not to make reference between layers, and method for transmitting bitstream

    US20230319302A1