Coding and decoding method and device

By performing image segmentation and mapping processing on the image to be encoded, the coded stream is compiled to avoid duplication of image segmentation between devices, which solves the problem of wasted computing power and realizes efficient image segmentation and recovery.

CN119946291APending Publication Date: 2025-05-06HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311468726.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When passing code streams between devices, multiple devices will repeatedly segment the same image, resulting in wasting of the device's computing power and increasing the computing power cost.

Method used

By performing image segmentation on the image to be encoded, a first segmented image is obtained and mapped to obtain a second segmented image. The segmented encoded image and mapping information are compiled into the code stream, so that the decoding end can restore the segmented image through the mapping information.

Benefits of technology

This avoids repeated segmentation operations of the same image by the device, reduces the waste of equipment computing power and costs, and simplifies the recovery process of segmented images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946291A_ABST
    Figure CN119946291A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a coding and decoding method and device, relates to the technical field of media, and can prevent equipment from repeatedly segmenting the same image. The method comprises the following steps: encoding a to-be-encoded image to obtain a code stream; performing image segmentation on the coded image to obtain a first segmented image; and performing mapping processing on the first segmented image to obtain a second segmented image. And determining mapping information according to the first segmented image. And encoding the second segmented image to obtain a segmented encoded image. And coding the segmented coded image, the mapping information and the image into a code stream. Wherein the mapping information comprises first information, second information and third information, the first information is used for indicating the number of pixel point types in the first segmented image, the second information is used for indicating pixel values corresponding to the pixel point types in the first segmented image, and the third information is used for indicating mapping pixel values corresponding to the pixel point types in the first segmented image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of media technology, and in particular to a coding and decoding method and device. Background Art

[0002] Image segmentation is an important part of image processing and machine vision technology in image understanding, and is also an important branch in the field of artificial intelligence (AI). Image segmentation is to classify each pixel in the image, determine the category of each point (such as background, person or car, etc.), and then divide the area. At present, image segmentation has been widely used in scenes such as autonomous driving and drone landing point determination.

[0003] In related technologies, when a device has image segmentation requirements (such as when the device is running heavy compression coding, license plate enhancement, and other services), the device needs to perform image segmentation operations based on the code stream. If the code stream is transmitted between multiple devices, multiple devices with image segmentation will repeatedly perform image segmentation operations on the code stream, which will cause a waste of device computing power and increase the device computing power cost.

[0004] Therefore, how to prevent a device from repeatedly performing image segmentation on the same image is one of the problems that those skilled in the art need to solve urgently. Summary of the invention

[0005] The embodiment of the present application provides a coding method and a device, which can avoid the device from repeatedly segmenting the same image. To achieve the above purpose, the embodiment of the present application adopts the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a coding method, the method comprising: encoding an image to be coded to obtain a code stream. Performing image segmentation on the coded image to obtain a first segmented image. Performing mapping processing on the first segmented image to obtain a second segmented image. Determining mapping information based on the first segmented image. Encoding the second segmented image to obtain a segmented coded image. Encoding the segmented coded image, the mapping information and the image into the code stream. The mapping information comprises first information, second information and third information, the first information being used to indicate the number of pixel point types in the first segmented image, the second information being used to indicate the pixel value corresponding to the pixel point type in the first segmented image, and the third information being used to indicate the mapped pixel value corresponding to the pixel point type in the first segmented image.

[0007] The method provided in the embodiment of the present application, by encoding the segmented image into the bitstream of the image to be encoded, enables the decoding end to obtain the segmented image through the decoded bitstream, thereby avoiding the device from repeatedly segmenting the same image. In addition, by encoding the mapping information of the segmented image into the bitstream, the decoding end can restore the segmented image through the mapping information, thereby reducing the difficulty of restoring the segmented image.

[0008] In a possible implementation manner, the first information and the second information may be determined according to the first segmented image; and the third information may be determined according to the second information.

[0009] It can be seen that the method provided in the embodiment of the present application can obtain the number of pixel point types in the first segmented image and the pixel values ​​corresponding to the pixel point types in the first segmented image through the first segmented image. Then, the mapping pixel value corresponding to the pixel point type in the first segmented image is determined through the pixel value corresponding to the pixel point type in the first segmented image.

[0010] In a possible implementation manner, the pixel values ​​in the second information may be shifted to obtain the mapped pixel values ​​in the third information.

[0011] It can be seen that the method provided in the embodiment of the present application can obtain the mapping pixel value in the third information through shift processing.

[0012] For example, the pixel value of a certain pixel in the second information is 4, and 4 is converted from decimal to binary to 100. 100 is shifted left by 5 bits to 10000000. 10000000 is converted from binary to decimal to 128. Then the mapped pixel value corresponding to the pixel with the pixel value of 4 in the second information is 128.

[0013] For another example, the pixel value of a certain pixel in the second information is 5, and 5 is converted from decimal to binary to 101. 101 is shifted left by 5 bits, and is 10100000. 10100000 is converted from binary to decimal to 160. Then the mapped pixel value corresponding to the pixel value 5 in the second information is 160.

[0014] In a possible implementation, the image to be encoded may be input into a target model to obtain the first segmented image, and a training set of the target model includes an image and a segmented image of the image.

[0015] It can be seen that the method provided in the embodiment of the present application can obtain a segmented image of the image to be encoded by inputting the image to be encoded into the network model.

[0016] In a possible implementation manner, the mapping information and the segmented coded image may be encoded into the extended information or supplementary enhancement information of the code stream.

[0017] It can be seen that the method provided in the embodiment of the present application can encode the segmented coded image and the mapping information into the code stream of the image to be encoded through extended information or supplementary enhancement information.

[0018] In a possible implementation manner, the mapping information further includes fourth information and fifth information, the fourth information is used to indicate the width of the first segmented image, and the fifth information is used to indicate the height of the first segmented image.

[0019] In a second aspect, an embodiment of the present application provides a decoding method, the method comprising: decoding a code stream to obtain a segmented coded image and mapping information, the segmented coded image being obtained based on a first segmented decoded image, the mapping information comprising first information, second information and third information, the first information being used to indicate the number of pixel point types in the first segmented image, the second information being used to indicate pixel values ​​corresponding to pixel point types in the first segmented image, and the third information being used to indicate mapped pixel values ​​corresponding to pixel point types in the first segmented image; decoding the segmented coded image to obtain the first segmented decoded image; and performing inverse mapping processing on the first segmented decoded image according to the mapping information to obtain a second segmented decoded image.

[0020] In a possible implementation manner, the pixel value of the pixel point in the second segmented decoded image may be determined according to the pixel value of the pixel point in the first segmented decoded image and the mapping information.

[0021] In a possible implementation, the mapped pixel value of the pixel in the first segmented decoded image can be determined according to the pixel value of the pixel in the first segmented decoded image and the mapping information. The pixel value of the pixel in the second segmented decoded image can be determined according to the mapped pixel value and the mapping information.

[0022] In a possible implementation manner, the mapping information further includes fourth information and fifth information, the fourth information is used to indicate the width of the first segmented image, and the fifth information is used to indicate the height of the first segmented image.

[0023] In a third aspect, an embodiment of the present application provides a coding device, which includes: a coding unit, a segmentation unit and a processing unit. The coding unit is used to encode the image to be encoded to obtain a code stream. The segmentation unit is used to perform image segmentation on the above-mentioned image to be encoded to obtain a first segmented image. The processing unit is used to perform mapping processing on the above-mentioned first segmented image to obtain a second segmented image. The processing unit is also used to determine mapping information based on the above-mentioned first segmented image, and the above-mentioned mapping information includes first information, second information and third information. The above-mentioned first information is used to indicate the number of pixel point types in the above-mentioned first segmented image, the above-mentioned second information is used to indicate the pixel value corresponding to the pixel point type in the above-mentioned first segmented image, and the above-mentioned third information is used to indicate the mapped pixel value corresponding to the pixel point type in the above-mentioned first segmented image. The coding unit is also used to encode the above-mentioned second segmented image to obtain a segmented coded image. The coding unit is also used to encode the above-mentioned segmented coded image and the above-mentioned mapping information and image into the above-mentioned code stream.

[0024] In a possible implementation manner, the processing unit is specifically configured to: determine the first information and the second information according to the first segmented image; and determine the third information according to the second information.

[0025] In a possible implementation manner, the processing unit is specifically configured to: perform shift processing on the pixel values ​​in the second information to obtain the mapped pixel values ​​in the third information.

[0026] In a possible implementation, the segmentation unit is specifically used to: input the image to be encoded into a target model to obtain the first segmented image, and the training set of the target model includes images and segmented images of the images.

[0027] In a possible implementation manner, the encoding unit is specifically used to: encode the mapping information and the segmented encoded image into the extended information or supplementary enhancement information of the code stream.

[0028] In a possible implementation manner, the mapping information further includes fourth information and fifth information, the fourth information is used to indicate the width of the first segmented image, and the fifth information is used to indicate the height of the first segmented image.

[0029] In a fourth aspect, an embodiment of the present application provides a decoding device, which includes: a decoding unit and a processing unit. The decoding unit is used to decode the code stream to obtain a segmented coded image and mapping information, the segmented coded image is obtained based on the first segmented decoded image, the mapping information includes first information, second information and third information, the first information is used to indicate the number of pixel types in the first segmented image, the second information is used to indicate the pixel value corresponding to the pixel type in the first segmented image, and the third information is used to indicate the mapped pixel value corresponding to the pixel type in the first segmented image. The decoding unit is used to decode the segmented coded image to obtain the first segmented decoded image. The processing unit is used to perform inverse mapping processing on the first segmented decoded image according to the mapping information to obtain the second segmented decoded image.

[0030] In a possible implementation manner, the processing unit is specifically configured to determine the pixel value of the pixel point in the second segmented decoded image according to the pixel value of the pixel point in the first segmented decoded image and the mapping information.

[0031] In a possible implementation, the processing unit is specifically used to: determine the mapping pixel value of the pixel point in the first segmented decoded image based on the pixel value of the pixel point in the first segmented decoded image and the mapping information; determine the pixel value of the pixel point in the second segmented decoded image based on the mapping pixel value and the mapping information.

[0032] In a possible implementation manner, the mapping information further includes fourth information and fifth information, the fourth information is used to indicate the width of the first segmented image, and the fifth information is used to indicate the height of the first segmented image.

[0033] In a fifth aspect, an embodiment of the present application further provides a coding device, comprising: at least one processor, which, when the at least one processor executes program code or instructions, implements the method described in the above first aspect or any possible implementation method thereof.

[0034] Optionally, the device may further include at least one memory, and the at least one memory is used to store the program code or instruction.

[0035] In a sixth aspect, an embodiment of the present application further provides a decoding device, comprising: at least one processor, when the at least one processor executes program code or instructions, it implements the method described in the above second aspect or any possible implementation method thereof.

[0036] Optionally, the device may further include at least one memory, and the at least one memory is used to store the program code or instruction.

[0037] In a seventh aspect, an embodiment of the present application further provides a chip, comprising: an input interface, an output interface, and at least one processor. Optionally, the chip further comprises a memory. The at least one processor is used to execute the code in the memory, and when the at least one processor executes the code, the chip implements the method described in the first aspect or any possible implementation thereof.

[0038] Optionally, the above chip may also be an integrated circuit.

[0039] In an eighth aspect, an embodiment of the present application further provides a computer-readable storage medium for storing a computer program, wherein the computer program includes methods for implementing the method described in the above-mentioned first aspect or any possible implementation thereof.

[0040] In a ninth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to implement the method described in the first aspect or any possible implementation thereof.

[0041] The encoding and decoding device, computer storage medium, computer program product and chip provided in this embodiment are all used to execute the encoding and decoding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the encoding and decoding method provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0043] Figure 1a An exemplary block diagram of a decoding system provided in an embodiment of the present application;

[0044] Figure 1b An exemplary block diagram of a video decoding system provided in an embodiment of the present application;

[0045] Figure 2 An exemplary block diagram of a video encoder provided in an embodiment of the present application;

[0046] Figure 3 An exemplary block diagram of a video decoder provided in an embodiment of the present application;

[0047] Figure 4 An exemplary schematic diagram of a candidate image block provided in an embodiment of the present application;

[0048] Figure 5An exemplary block diagram of a video decoding device provided in an embodiment of the present application;

[0049] Figure 6 An exemplary block diagram of a device provided in an embodiment of the present application;

[0050] Figure 7 A schematic diagram of an encoding method provided in an embodiment of the present application;

[0051] Figure 8 A schematic diagram of a decoding method provided in an embodiment of the present application;

[0052] Fig. 9 A schematic diagram of an encoding device provided in an embodiment of the present application;

[0053] Fig.10 A schematic diagram of a decoding device provided in an embodiment of the present application;

[0054] Fig.11 A schematic diagram of the structure of a chip provided in an embodiment of the present application;

[0055] Fig.12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0056] Fig.13 A schematic diagram of the structure of an encoding device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of the present application.

[0058] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0059] The terms "first" and "second" and the like in the description and drawings of the embodiments of the present application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.

[0060] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the embodiments of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices.

[0061] It should be noted that in the description of the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as having priority or advantage over other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific way.

[0062] Data encoding and decoding includes two parts: data encoding and data decoding. Data encoding is performed on the source side (or commonly referred to as the encoder side), and generally includes processing (e.g., compressing) the original data to reduce the amount of data required to represent the original data (thereby more efficiently storing and / or transmitting). Data decoding is performed on the destination side (or commonly referred to as the decoder side), and generally includes inverse processing relative to the encoder side to reconstruct the original data. The "encoding and decoding" of the data involved in the embodiments of the present application should be understood as the "encoding" or "decoding" of the data. The encoding part and the decoding part are also collectively referred to as encoding and decoding (encoding and decoding, CODEC).

[0063] In the case of lossless data encoding, the original data can be reconstructed, that is, the reconstructed original data has the same quality as the original data (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy data encoding, further compression is performed by quantization, etc. to reduce the amount of data required to represent the original data, but the decoder side cannot completely reconstruct the original data, that is, the quality of the reconstructed original data is lower or worse than the quality of the original data.

[0064] The embodiments of the present application can be applied to video data and other data with compression / decompression requirements. The following takes the encoding of video data (referred to as video encoding) as an example to illustrate the embodiments of the present application. Other types of data (such as image data, audio data, integer data and other data with compression / decompression requirements) can refer to the following description, and the embodiments of the present application will not be repeated. It should be noted that, compared with video encoding, the encoding process of audio data and integer data does not need to divide the data into blocks, but the data can be directly encoded.

[0065] Video coding generally refers to processing a sequence of images to form a video or video sequence. In the field of video coding, the terms "picture", "frame" or "image" can be used as synonyms.

[0066] Several video coding standards belong to the category of "lossy hybrid video codecs" (i.e., combining spatial and temporal prediction in the pixel domain with 2D transform coding in the transform domain for applying quantization). Each image in a video sequence is usually divided into a set of non-overlapping blocks, which are usually encoded at the block level. In other words, the encoder usually processes, i.e., encodes the video at the block (video block) level, for example, by generating a prediction block through spatial (intra-frame) prediction and temporal (inter-frame) prediction; subtracting the prediction block from the current block (currently processed / to be processed block) to obtain a residual block; transforming the residual block in the transform domain and quantizing the residual block to reduce the amount of data to be transmitted (compressed), while the decoder side applies the inverse processing part relative to the encoder to the encoded or compressed block to reconstruct the current block for representation. In addition, the encoder needs to repeat the processing steps of the decoder so that the encoder and the decoder generate the same predictions (e.g., intra-frame predictions and inter-frame predictions) and / or reconstructed pixels for processing, i.e., encoding subsequent blocks.

[0067] In the following embodiment of the decoding system 10, the encoder 20 and the decoder 30 are based on Figures 1a to 3 Give a description.

[0068] Figure 1a An exemplary block diagram of a decoding system 10 provided in an embodiment of the present application, for example, a video decoding system 10 (or simply a decoding system 10) that can utilize the techniques of the embodiments of the present application. The video encoder 20 (or simply an encoder 20) and the video decoder 30 (or simply a decoder 30) in the video decoding system 10 represent devices that can be used to perform various techniques according to various examples described in the embodiments of the present application.

[0069] like Figure 1a As shown, the decoding system 10 includes a source device 12, which is used to provide encoded image data 21 such as encoded images to a destination device 14 for decoding the encoded image data 21.

[0070] The source device 12 includes an encoder 20 , and may additionally or optionally include an image source 16 , a preprocessor (or a preprocessing unit) 18 such as an image preprocessor, and a communication interface (or a communication unit) 22 .

[0071] The image source 16 may include or may be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer-animated images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The above-mentioned image source may be any type of memory or storage for storing any of the above-mentioned images.

[0072] In order to distinguish the processing performed by the pre-processor (or pre-processing unit) 18 , the image (or image data) 17 may also be referred to as a raw image (or raw image data) 17 .

[0073] The preprocessor 18 is used to receive the original image data 17 and preprocess the original image data 17 to obtain a preprocessed image (or preprocessed image data) 19. For example, the preprocessing performed by the preprocessor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color adjustment, or denoising. It is understood that the preprocessing unit 18 may be an optional component.

[0074] The video encoder (or encoder) 20 is used to receive the pre-processed image data 19 and provide the encoded image data 21 (hereinafter referred to as Figure 2 etc. for further description).

[0075] The communication interface 22 in the source device 12 can be used to receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device through the communication channel 13 for storage or direct reconstruction.

[0076] The destination device 14 includes a decoder 30 and, in addition or alternatively, may include a communication interface (or communication unit) 28 , a post-processor (or post-processing unit) 32 , and a display device 34 .

[0077] The communication interface 28 in the destination device 14 is used to receive the encoded image data 21 (or any other processed version) directly from the source device 12 or from any other source device such as a storage device, for example, the storage device is a encoded image data storage device, and provide the encoded image data 21 to the decoder 30.

[0078] The communication interface 22 and the communication interface 28 can be used to send or receive encoded image data (or encoded data) 21 through a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any type of combination thereof.

[0079] For example, the communication interface 22 may be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or use any type of transmission coding or processing to process the encoded image data for transmission over a communication link or network.

[0080] The communication interface 28 corresponds to the communication interface 22 , for example, and can be used to receive transmission data and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain the encoded image data 21 .

[0081] Both the communication interface 22 and the communication interface 28 can be configured as follows Figure 1a A unidirectional communication interface, or a bidirectional communication interface, indicated by an arrow pointing from the source device 12 to the corresponding communication channel 13 of the destination device 14, and can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded image data transmission, etc.

[0082] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31 (hereinafter referred to as Figure 3 etc. for further description).

[0083] The post-processor 32 is used to post-process the decoded image data 31 (also called reconstructed image data) such as the decoded image to obtain the post-processed image data 33 such as the post-processed image. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color adjustment, cropping or resampling, or any other processing for generating the decoded image data 31 for display by the display device 34 or the like.

[0084] The display device 34 is used to receive the post-processed image data 33 to display the image to a user or viewer, etc. The display device 34 can be or include any type of display for displaying the reconstructed image, such as an integrated or external display screen or display. For example, the display screen can include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.

[0085] The decoding system 10 also includes a training engine 25, which is used to train the encoder 20 (especially the entropy coding unit 270 in the encoder 20) or the decoder 30 (especially the entropy decoding unit 304 in the decoder 30) to perform entropy coding on the image block to be encoded according to the estimated probability distribution. For a detailed description of the training engine 25, please refer to the following method test example.

[0086] although Figure 1a The source device 12 and the destination device 14 are shown as independent devices, but the device embodiment may also include the source device 12 and the destination device 14 or the functions of the source device 12 and the destination device 14 at the same time, that is, the source device 12 or the corresponding function and the destination device 14 or the corresponding function at the same time. In these embodiments, the source device 12 or the corresponding function and the destination device 14 or the corresponding function can be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.

[0087] According to the description, Figure 1a The presence and (exact) division of different units or functions in the source device 12 and / or the destination device 14 shown may vary depending on the actual device and application, which will be obvious to the skilled person.

[0088] Please refer to Figure 1b , Figure 1b An exemplary block diagram of a video decoding system 40 provided in an embodiment of the present application, wherein the encoder 20 (eg, the video encoder 20) or the decoder 30 (eg, the video decoder 30) or both may be configured as follows: Figure 1bThe processing circuits in the video decoding system 40 shown are implemented, for example, by one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding dedicated processors, or any combination thereof. Figure 2 and Figure 3 , Figure 2 An exemplary block diagram of a video encoder provided in an embodiment of the present application is shown in FIG. Figure 3 An exemplary block diagram of a video decoder provided in an embodiment of the present application. The encoder 20 may be implemented by a processing circuit 46 to include a reference Figure 2 The various modules discussed in encoder 20 and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented by processing circuit 46 to include reference Figure 3 The various modules discussed in the decoder 30 and / or any other decoder systems or subsystems described herein. The processing circuitry 46 described above may be used to perform the various operations discussed below. Figure 5 As shown, if part of the technology is implemented in software, the device can store the software instructions in a suitable non-transitory computer-readable storage medium, and use one or more processors to execute the instructions in hardware, thereby performing the technology of the embodiment of the present application. One of the video encoder 20 and the video decoder 30 can be integrated into a single device as part of a combined codec (encoder / decoder, CODEC), such as Figure 1b shown.

[0089] The source device 12 and the destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smart phone, a tablet or a tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, and a monitoring device, etc., and may not use or use any type of operating system. The source device 12 and the destination device 14 may also be devices in a cloud computing scenario, such as a virtual machine in a cloud computing scenario. In some cases, the source device 12 and the destination device 14 may be equipped with components for wireless communication. Therefore, the source device 12 and the destination device 14 may be wireless communication devices.

[0090] The source device 12 and the destination device 14 may install virtual scene applications (applications, APPs) such as virtual reality (VR) applications, augmented reality (AR) applications, or mixed reality (MR) applications, and may run VR applications, AR applications, or MR applications based on user operations (such as clicks, touches, slides, shakes, voice control, etc.). The source device 12 and the destination device 14 may collect images / videos of any object in the environment through cameras and / or sensors, and then display virtual objects on the display device based on the collected images / videos. The virtual objects may be virtual objects in VR scenes, AR scenes, or MR scenes (i.e., objects in the virtual environment).

[0091] It should be noted that in the embodiment of the present application, the virtual scene application in the source device 12 and the destination device 14 can be a built-in application of the source device 12 and the destination device 14 themselves, or it can be an application provided by a third-party service provider and installed by the user. There is no specific limitation on this.

[0092] In addition, the source device 12 and the destination device 14 may be installed with a real-time video transmission application, such as a live broadcast application. The source device 12 and the destination device 14 may collect images / videos through cameras and then display the collected images / videos on a display device.

[0093] In some cases, Figure 1a The video decoding system 10 shown is merely exemplary, and the techniques provided in the embodiments of the present application may be applicable to video encoding settings (e.g., video encoding or video decoding), which do not necessarily include any data communication between the encoding device and the decoding device. In other examples, the data is retrieved from a local memory, sent over a network, and so on. The video encoding device can encode the data and store the data in a memory, and / or the video decoding device can retrieve the data from the memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to a memory and / or retrieve and decode data from a memory.

[0094] Please refer to Figure 1b , Figure 1b An exemplary block diagram of a video decoding system 40 provided in an embodiment of the present application is shown in FIG. Figure 1b As shown, the video decoding system 40 may include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video encoder / decoder implemented by a processing circuit 46), an antenna 42, one or more processors 43, one or more memory storages 44 and / or a display device 45.

[0095] like Figure 1b As shown, imaging device 41, antenna 42, processing circuit 46, video encoder 20, video decoder 30, processor 43, memory storage 44 and / or display device 45 can communicate with each other. In different examples, video decoding system 40 can include only video encoder 20 or only video decoder 30.

[0096] In some instances, antenna 42 may be used to transmit or receive a coded bit stream of video data. In addition, in some instances, display device 45 may be used to present video data. Processing circuit 46 may include application-specific integrated circuit (ASIC) logic, graphics processor, general purpose processor, etc. Video decoding system 40 may also include an optional processor 43, which may similarly include application-specific integrated circuit (ASIC) logic, graphics processor, general purpose processor, etc. In addition, memory storage 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, memory storage 44 may be implemented by cache memory. In other instances, processing circuit 46 may include memory (e.g., cache, etc.) for implementing image buffers, etc.

[0097] In some examples, video encoder 20 implemented by logic circuits may include an image buffer (e.g., implemented by processing circuit 46 or memory storage 44) and a graphics processing unit (e.g., implemented by processing circuit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video encoder 20 implemented by processing circuit 46 to implement reference Figure 2 and / or the various modules discussed in connection with any other encoder system or subsystem described herein. Logic circuits may be used to perform the various operations discussed herein.

[0098] In some examples, video decoder 30 may be implemented in a similar manner by processing circuitry 46 to implement reference Figure 3The various modules discussed in connection with video decoder 30 and / or any other decoder system or subsystem described herein may be included in the video decoder 30 implemented by logic circuitry. In some examples, video decoder 30 implemented by logic circuitry may include an image buffer (implemented by processing circuitry 46 or memory storage 44) and a graphics processing unit (implemented by processing circuitry 46, for example). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video decoder 30 implemented by processing circuitry 46 to implement the video decoder 30 implemented by processing circuitry 46. Figure 3 and / or the various modules discussed above for any other decoder system or subsystem described herein.

[0099] In some examples, antenna 42 may be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data related to the encoded video frames discussed herein, indicators, index values, mode selection data, etc., such as data related to the encoded partitions (e.g., transform coefficients or quantized transform coefficients, (as discussed) optional indicators, and / or data defining the encoded partitions). Video decoding system 40 may also include video decoder 30 coupled to antenna 42 and used to decode the encoded bitstream. Display device 45 is used to present the video frames.

[0100] It should be understood that for the examples described with reference to the video encoder 20 in the embodiments of the present application, the video decoder 30 can be used to perform the reverse process. With respect to the signaling syntax elements, the video decoder 30 can be used to receive and parse such syntax elements and decode the related video data accordingly. In some examples, the video encoder 20 can entropy encode the syntax elements into an encoded video bitstream. In such examples, the video decoder 30 can parse such syntax elements and decode the related video data accordingly.

[0101] For ease of description, the embodiments of the present application are described with reference to the universal video coding (VVC) reference software or the high-efficiency video coding (HEVC) developed by the video coding joint working group (joint collaboration team on video coding, JCT-VC) of the ITU-T video coding experts group (VCEG) and the ISO / IEC motion picture experts group (MPEG). Those of ordinary skill in the art understand that the embodiments of the present application are not limited to HEVC or VVC.

[0102] Encoders and encoding methods

[0103] like Figure 2As shown, the video encoder 20 includes an input terminal (or input interface) 201, a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270 and an output terminal (or output interface) 272. The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254 and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The illustrated video encoder 20 may also be referred to as a hybrid video encoder or a hybrid video codec-based video encoder.

[0104] See also Figure 2 The inter-frame prediction unit is a trained target model (also called a neural network) that processes an input image or an image region or an image block to generate a prediction value of the input image block. For example, the neural network for inter-frame prediction is used to receive an input image or an image region or an image block, and generate a prediction value of the input image or an image region or an image block.

[0105] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208 and the mode selection unit 260 constitute the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-frame prediction unit 244 and the intra-frame prediction unit 254 constitute the backward signal path of the encoder, wherein the backward signal path of the encoder 20 corresponds to the signal path of the decoder (see Figure 3 The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded image buffer 230, the inter-frame prediction unit 244 and the intra-frame prediction unit 254 also constitute a “built-in decoder” of the video encoder 20.

[0106] Images and image segmentation (images and blocks)

[0107] The encoder 20 may be configured to receive, via an input 201 or the like, an image (or image data) 17, e.g., an image in a sequence of images forming a video or a video sequence. The received image or image data may also be a pre-processed image (or pre-processed image data) 19. For simplicity, the following description uses image 17. Image 17 may also be referred to as a current image or an image to be encoded (particularly when the current image is to be distinguished from other images in video encoding, such as previously encoded images and / or decoded images in the same video sequence, i.e., a video sequence also including the current image).

[0108] A (digital) image is or can be considered as a two-dimensional array or matrix of pixels with intensity values. The pixels in the array can also be called pixels (pixel or pel) (short for picture elements). The number of pixels in the array or image in the horizontal and vertical directions (or axes) determines the size and / or resolution of the image. In order to represent color, three color components are usually used, that is, the image can be represented as or include three pixel arrays. In the RBG format or color space, the image includes corresponding red, green and blue pixel arrays. However, in video coding, each pixel is usually represented in a brightness / chrominance format or color space, such as YCbCr, including a brightness component indicated by Y (sometimes also represented by L) and two chrominance components represented by Cb and Cr. The brightness (luma) component Y represents the brightness or grayscale level intensity (for example, the two are the same in grayscale images), while the two chrominance (abbreviated as chroma) components Cb and Cr represent the chrominance or color information components. Accordingly, an image in YCbCr format includes a luminance pixel array of luminance pixel values ​​(Y) and two chrominance pixel arrays of chrominance values ​​(Cb and Cr). An image in RGB format can be converted or transformed into YCbCr format and vice versa, a process also referred to as color conversion or transformation. If the image is black and white, the image may include only a luminance pixel array. Accordingly, an image may be, for example, a luminance pixel array in a monochrome format or a luminance pixel array and two corresponding chrominance pixel arrays in 4:2:0, 4:2:2 and 4:4:4 color formats.

[0109] In one embodiment, an embodiment of the video encoder 20 may include an image segmentation unit ( Figure 2), for dividing the image 17 into a plurality of (usually non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC) or coding tree blocks (CTB), or coding tree units (CTU) in the H.265 / HEVC and VVC standards. The partitioning unit may be used to use the same block size for all images in a video sequence and a corresponding grid of defined block sizes, or to change the block size between images or subsets of images or groups of images, and to partition each image into corresponding blocks.

[0110] In other embodiments, the video encoder may be configured to directly receive a block 203 of the image 17, for example, one, several or all blocks constituting the image 17. The image block 203 may also be referred to as a current image block or an image block to be encoded.

[0111] Like image 17, image block 203 is also or can be considered as a two-dimensional array or matrix composed of pixels with intensity values ​​(pixel values), but image block 203 is smaller than image 17. In other words, block 203 may include one pixel array (e.g., a brightness array in the case of monochrome image 17 or a brightness array or chrominance array in the case of a color image) or three pixel arrays (e.g., one brightness array and two chrominance arrays in the case of a color image 17) or any other number and / or type of arrays according to the color format adopted. The number of pixels in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Accordingly, the block can be an M×N (M columns×N rows) pixel array, or an M×N transform coefficient array, etc.

[0112] In one embodiment, Figure 2 The video encoder 20 shown is used to encode the image 17 block by block, for example, encoding and prediction are performed on each block 203.

[0113] In one embodiment, Figure 2 The illustrated video encoder 20 may also be used to partition and / or encode an image using slices (also referred to as video slices), wherein an image may be partitioned or encoded using one or more slices (usually non-overlapping). Each slice may include one or more blocks (e.g., coding tree units CTUs) or one or more groups of blocks (e.g., coding tiles in the H.265 / HEVC / VVC standard and bricks in the VVC standard).

[0114] In one embodiment, Figure 2The video encoder 20 shown can also be used to partition and / or encode an image using slices / coding block groups (also called video coding block groups) and / or coding blocks (also called video coding blocks), wherein an image can be partitioned or encoded using one or more slices / coding block groups (usually non-overlapping), each slice / coding block group may include one or more blocks (e.g., CTU) or one or more coding blocks, etc., wherein each coding block may be in a rectangular shape, etc., and may include one or more complete or partial blocks (e.g., CTU).

[0115] Residual calculation

[0116] The residual calculation unit 204 is used to calculate the residual block 205 (the prediction block 265 is described in detail later) based on the image block (or original block) 203 and the prediction block 265 in the following manner: for example, the pixel value of the prediction block 265 is subtracted from the pixel value of the image block 203 pixel by pixel (pixel by pixel) to obtain the residual block 205 in the pixel domain.

[0117] Transform

[0118] The transform processing unit 206 is used to perform discrete cosine transform (DCT) or discrete sine transform (DST) on the pixel values ​​of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207 may also be called transform residual coefficients, representing the residual block 205 in the transform domain.

[0119] The transform processing unit 206 may be used to apply an integerized approximation of the DCT / DST, such as the transform specified for H.265 / HEVC. This integerized approximation is typically scaled by a factor compared to an orthogonal DCT transform. In order to maintain the norm of the residual block processed by the forward transform and the inverse transform, other scaling factors are used as part of the transform process. The scaling factor is typically selected based on certain constraints, such as whether the scaling factor is a power of 2 for the shift operation, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. For example, a specific scaling factor is specified for the inverse transform on the encoder 20 side by the inverse transform processing unit 212 (and for the corresponding inverse transform on the decoder 30 side by, for example, the inverse transform processing unit 312), and correspondingly, a corresponding scaling factor may be specified for the forward transform on the encoder 20 side by the transform processing unit 206.

[0120] In one embodiment, the video encoder 20 (correspondingly, the transform processing unit 206) may be used to output transform parameters such as one or more transform types, for example, directly output or output after encoding or compression by the entropy coding unit 270, for example, so that the video decoder 30 may receive and use the transform parameters for decoding.

[0121] Quantification

[0122] The quantization unit 208 is used to quantize the transform coefficient 207 by, for example, scalar quantization or vector quantization to obtain a quantized transform coefficient 209 . The quantized transform coefficient 209 may also be referred to as a quantized residual coefficient 209 .

[0123] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded down to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, for scalar quantization, different degrees of scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to a finer quantization, while a larger quantization step size corresponds to a coarser quantization. A suitable quantization step size may be indicated by a quantization parameter (QP). For example, a quantization parameter may be an index of a predefined set of suitable quantization step sizes. For example, a smaller quantization parameter may correspond to fine quantization (a smaller quantization step size), a larger quantization parameter may correspond to coarse quantization (a larger quantization step size), and vice versa. Quantization may include dividing by a quantization step size, while a corresponding or inverse dequantization performed by an inverse quantization unit 210 or the like may include multiplying by a quantization step size. Embodiments according to some standards, such as HEVC, may be used to determine a quantization step size using a quantization parameter. In general, the quantization step size may be calculated using a fixed-point approximation of an equation containing division according to the quantization parameter. Other scaling factors may be introduced for quantization and dequantization to recover the norm of the residual block that may be modified due to the scaling used in the fixed-point approximation of the equations for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, a custom quantization table may be used and indicated from the encoder to the decoder in the bitstream, etc. Quantization is a lossy operation, where the larger the quantization step size, the greater the loss.

[0124] In one embodiment, the video encoder 20 (correspondingly, the quantization unit 208) may be configured to output a quantization parameter (QP), for example, directly output or output after being encoded or compressed by the entropy coding unit 270, for example, so that the video decoder 30 may receive and use the quantization parameter for decoding.

[0125] Dequantization

[0126] The inverse quantization unit 210 is used to perform inverse quantization of the quantization unit 208 on the quantized coefficients to obtain dequantized coefficients 211, for example, performing an inverse quantization scheme of the quantization scheme performed by the quantization unit 208 according to or using the same quantization step size as the quantization unit 208. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211, corresponding to the transform coefficients 207, but due to the loss caused by quantization, the dequantized coefficients 211 are usually not completely the same as the transform coefficients.

[0127] Inverse Transform

[0128] The inverse transform processing unit 212 is used to perform an inverse transform of the transform performed by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain a reconstructed residual block 213 (or a corresponding dequantized coefficient 213) in the pixel domain. The reconstructed residual block 213 may also be referred to as a transform block 213.

[0129] reconstruction

[0130] The reconstruction unit 214 (e.g., the summer 214) is used to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the pixel domain, for example, by adding the pixel point values ​​of the reconstructed residual block 213 and the pixel point values ​​of the prediction block 265.

[0131] Filtering

[0132] The loop filter unit 220 (or simply "loop filter" 220) is used to filter the reconstruction block 215 to obtain the filter block 221, or is generally used to filter the reconstructed pixel points to obtain the filtered pixel point value. For example, the loop filter unit is used to smoothly perform pixel conversion or improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chromascaling (LMCS) (i.e., an adaptive in-loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Figure 2 2 is shown as a loop filter, but in other configurations, the loop filter unit 220 can be implemented as a post-loop filter. The filter block 221 can also be called a filter reconstruction block 221.

[0133] In one embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) may be used to output loop filter parameters (e.g., SAO filter parameters, ALF filter parameters, or LMCS parameters), for example, directly or after being entropy encoded by the entropy encoding unit 270, such that the decoder 30 may receive and use the same or different loop filter parameters for decoding.

[0134] Decoded Image Buffer

[0135] The decoded picture buffer (DPB) 230 may be a reference picture memory for storing reference picture data for use by the video encoder 20 when encoding video data. The DPB 230 may be formed by any of a variety of memory devices, such as a dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The decoded picture buffer 230 may be used to store one or more filter blocks 221. The decoded picture buffer 230 may also be used to store other previous filter blocks of the same current picture or a different picture such as a previous reconstructed picture, such as a previously reconstructed and filtered block 221, and may provide a complete previously reconstructed, i.e., decoded picture (and corresponding reference blocks and pixels) and / or a partially reconstructed current picture (and corresponding reference blocks and pixels), such as for inter-frame prediction. The decoded image buffer 230 may also be used to store one or more unfiltered reconstructed blocks 215, or generally to store unfiltered reconstructed pixels, for example, reconstructed blocks 215 that have not been filtered by the loop filter unit 220, or reconstructed blocks or reconstructed pixels that have not undergone any other processing.

[0136] Mode selection (segmentation and prediction)

[0137] The mode selection unit 260 includes a segmentation unit 262, an inter-frame prediction unit 244, and an intra-frame prediction unit 254 for obtaining a decoded image from the decoded image buffer 230 or other buffers (eg, column buffers, Figure 2 The decoder 200 (not shown) receives or obtains original image data such as the original block 203 (current block 203 of the current image 17) and reconstructed image data, for example, filtered and / or unfiltered reconstructed pixels or reconstructed blocks of the same (current) image and / or one or more previously decoded images. The reconstructed image data is used as reference image data required for prediction such as inter-frame prediction or intra-frame prediction to obtain a prediction block 265 or a prediction value 265.

[0138] The mode selection unit 260 may be used to determine or select a partition for the current block (including no partition) and prediction mode (eg, intra-frame or inter-frame prediction mode), generate a corresponding prediction block 265 , calculate the residual block 205 , and reconstruct the reconstruction block 215 .

[0139] In one embodiment, the mode selection unit 260 may be used to select a segmentation and prediction mode (e.g., from prediction modes supported or available by the mode selection unit 260), the prediction mode providing the best match or minimum residual (minimum residual means better compression in transmission or storage), or providing minimum signaling overhead (minimum signaling overhead means better compression in transmission or storage), or considering or balancing both of the above. The mode selection unit 260 may be used to determine the segmentation and prediction mode based on rate distortion optimization (RDO), i.e., selecting the prediction mode that provides minimum rate distortion optimization. The terms "best", "lowest", "optimal", etc. herein do not necessarily mean "best", "lowest", "optimal" in general, but may also refer to situations where termination or selection criteria are met, for example, values ​​exceeding or falling below a threshold or other restrictions may result in a "suboptimal selection" but reduce complexity and processing time.

[0140] In other words, the partitioning unit 262 may be used to partition an image in a video sequence into a sequence of coding tree units (CTUs), and the CTU 203 may be further partitioned into smaller block portions or sub-blocks (again forming blocks), for example, by iteratively using quad-tree partitioning (QT), binary-tree partitioning (BT) or triple-tree partitioning (TT) or any combination thereof, and for performing prediction on each of the block portions or sub-blocks, for example, wherein the mode selection comprises selecting a tree structure of the partition block 203 and selecting a prediction mode to be applied to each of the block portions or sub-blocks.

[0141] The segmentation (eg, performed by segmentation unit 262) and prediction processes (eg, performed by inter-prediction unit 244 and intra-prediction unit 254) performed by video encoder 20 are described in detail below.

[0142] segmentation

[0143] The segmentation unit 262 may segment (or divide) an image block (or CTU) 203 into smaller parts, such as small blocks in a square or rectangular shape. For an image with three pixel arrays, a CTU consists of N×N luminance pixel blocks and two corresponding chrominance pixel blocks. The maximum allowed size of a luminance block in a CTU is specified as 128×128 in the developing universal video coding (VVC) standard, but may be specified as a value different from 128×128 in the future, such as 256×256. The CTUs of an image may be concentrated / grouped into slices / coding block groups, coding blocks, or bricks. A coding block covers a rectangular area of ​​an image, and a coding block may be divided into one or more bricks. A brick consists of multiple CTU rows within a coding block. A coding block that is not segmented into multiple bricks may be called a brick. However, a brick is a true subset of a coding block and is therefore not called a coding block. VVC supports two coding block group modes, namely raster scan slice / coding block group mode and rectangular slice mode. In raster scan CBG mode, a slice / CBG contains a sequence of CBs in a raster scan of CBs of an image. In rectangular slice mode, a slice contains multiple bricks of an image, which together form a rectangular region of the image. The bricks within a rectangular slice are arranged in the order of the bricks in the slice's raster scan. These smaller blocks (also called sub-blocks) can be further split into smaller parts. This is also called tree partitioning or hierarchical tree partitioning, where a root block at root tree level 0 (hierarchy level 0, depth 0), etc., can be recursively split into two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchy level 1, depth 1). These blocks can be split into two or more blocks at the next lower level, such as tree level 2 (hierarchy level 2, depth 2), etc., until the partitioning ends (because the end criteria are met, such as reaching the maximum tree depth or minimum block size). Blocks that are not further split are also called leaf blocks or leaf nodes of the tree. A tree divided into two parts is called a binary tree (BT), a tree divided into three parts is called a ternary tree (TT), and a tree divided into four parts is called a quadtree (QT).

[0144] For example, a coding tree unit (CTU) may be or include a CTB of luminance pixels, two corresponding CTBs of chrominance pixels of an image with three pixel arrays, or a CTB of pixels of a monochrome image, or a CTB of pixels of an image encoded using three independent color planes and syntax structures (for encoding pixels). Accordingly, a coding tree block (CTB) may be an N×N pixel block, where N may be set to a certain value so that the component is divided into CTBs, which is segmentation. A coding unit (coding unit, CU) may be or include a coding block of luminance pixels, two corresponding coding blocks of chrominance pixels of an image with three pixel arrays, or a coding block of pixels of a monochrome image, or a coding block of pixels of an image encoded using three independent color planes and syntax structures (for encoding pixels). Accordingly, a coding block (CB) may be an M×N pixel block, where M and N may be set to a certain value so that the CTB is divided into coding blocks, which is segmentation.

[0145] For example, in an embodiment, a coding tree unit (CTU) may be divided into multiple CUs according to HEVC by using a quadtree structure represented as a coding tree. A decision is made at the leaf-CU level whether to use inter-frame (temporal) prediction or intra-frame (spatial) prediction to encode an image area. Each leaf-CU may be further divided into one, two, or four PUs according to the PU partition type. The same prediction process is used within a PU, and relevant information is transmitted to the decoder in units of PUs. After the prediction process is applied according to the PU partition type to obtain the residual block, the leaf-CU may be partitioned into transform units (TUs) according to other quadtree structures similar to the coding tree for the CU.

[0146] For example, in an embodiment, according to the latest video coding standard currently under development (referred to as Versatile Video Coding (VVC), a combined quadtree of nested multi-type trees (e.g., binary trees and ternary trees) is used to partition a segment structure for partitioning a coding tree unit. In the coding tree structure within the coding tree unit, the CU may be a square or a rectangle. For example, a coding tree unit (CTU) is first partitioned by a quadtree structure. The quadtree leaf nodes are further partitioned by a multi-type tree structure. The multi-type tree structure has four partition types: vertical binary tree partition (SPLIT_BT_VER), horizontal binary tree partition (SPLIT_BT_HOR), vertical ternary tree partition ( SPLIT_TT_VER) and horizontal ternary tree splitting (SPLIT_TT_HOR). The leaf nodes of the multi-type tree are called coding units (CUs), unless the CU is too large for the maximum transform length, such segmentation is used for prediction and transform processing without any other splitting. In most cases, this means that the block sizes of CUs, PUs, and TUs in the coding block structure of the quadtree nested multi-type tree are the same. This exception occurs when the maximum supported transform length is less than the width or height of the color components of the CU. VVC has established a unique signaling mechanism for the split division information in the coding structure with quadtree nested multi-type trees. In the signaling mechanism, the coding tree unit ( CTU) is first split by the quadtree structure as the root of the quadtree. Then each quadtree leaf node (when large enough) is further split into a multi-type tree structure. In the multi-type tree structure, the first flag (mtt_split_cu_flag) is used to indicate whether the node is further split. When the node is further split, the second flag (mtt_split_cu_vertical_flag) is used to indicate the division direction, and then the third flag (mtt_split_cu_binary_flag) is used to indicate whether the division is a binary tree division or a ternary tree division. According to mtt_split_cu_vertical The decoder can derive the multi-type tree splitting mode (MttSplitMode) of the CU based on the values ​​of ical_flag and mtt_split_cu_binary_flag based on predefined rules or tables. It should be noted that for certain designs, such as the 64×64 luminance block and 32×32 chrominance pipeline design in the VVC hardware decoder, TT splitting is not allowed when the width or height of the luminance coding block is greater than 64. TT splitting is also not allowed when the width or height of the chrominance coding block is greater than 32. The pipeline design divides the image into multiple virtual pipeline data units (VPDUs), each VPDU is defined as a non-overlapping unit in the image. In the hardware decoder, consecutive VPDUs are processed simultaneously in multiple pipeline stages. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is necessary to keep the VPDU small.In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the ternary tree (TT) and binary tree (BT) partitioning may increase the VPDU size.

[0147] In addition, it should be noted that when a part of the tree node block exceeds the bottom or the right boundary of the image, the tree node block is forcibly divided until all the pixels of each coding CU are located within the image boundary.

[0148] For example, the intra sub-partitions (ISP) tool described above may vertically or horizontally divide the luma intra prediction block into two or four sub-partitions according to the block size.

[0149] In one example, mode select unit 260 of video encoder 20 may be used to perform any combination of the segmentation techniques described above.

[0150] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a (predetermined) prediction mode set. The prediction mode set may include, for example, an intra-frame prediction mode and / or an inter-frame prediction mode.

[0151] Intra-frame prediction

[0152] The intra-frame prediction mode set may include 35 different intra-frame prediction modes, for example, non-directional modes like DC (or mean) mode and plane mode, or directional modes as defined in HEVC, or may include 67 different intra-frame prediction modes, for example, non-directional modes like DC (or mean) mode and plane mode, or directional modes as defined in VVC. For example, several traditional angle intra-frame prediction modes are adaptively replaced with wide-angle intra-frame prediction modes for non-square blocks defined in VVC. For another example, in order to avoid the division operation of DC prediction, only the longer side is used to calculate the average value of non-square blocks. In addition, the intra-frame prediction result of the plane mode can also be modified using the position-dependent intraprediction combination (PDPC) method.

[0153] The intra prediction unit 254 is configured to generate an intra prediction block 265 by reconstructing pixels in adjacent blocks of the same current image according to an intra prediction mode in the intra prediction mode set.

[0154] The intra-frame prediction unit 254 (or generally the mode selection unit 260) is also used to output intra-frame prediction parameters (or generally information indicating the selected intra-frame prediction mode of the block) in the form of syntax elements 266 to the entropy coding unit 270 for inclusion in the encoded image data 21, so that the video decoder 30 can perform operations, such as receiving and using the prediction parameters for decoding.

[0155] The intra prediction modes in HEVC include DC prediction mode, plane prediction mode and 33 angle prediction modes, with a total of 35 candidate prediction modes. The current block can use the pixels of the reconstructed image blocks on the left and above as references for intra prediction. The image blocks in the surrounding area of ​​the current block used for intra prediction of the current block become reference blocks, and the pixels in the reference blocks are called reference pixels. Among the 35 candidate prediction modes, the DC prediction mode is applicable to the area with flat texture in the current block, and all pixels in this area use the average value of the reference pixels in the reference block as prediction; the plane prediction mode is applicable to image blocks with smoothly changing textures. The current block that meets this condition uses the reference pixels in the reference block for bilinear interpolation as the prediction of all pixels in the current block; the angle prediction mode uses the characteristics that the texture of the current block is highly correlated with the texture of the adjacent reconstructed image blocks, and copies the values ​​of the reference pixels in the corresponding reference block along a certain angle as the prediction of all pixels in the current block.

[0156] The HEVC encoder selects an optimal intra-frame prediction mode for the current block from 35 candidate prediction modes and writes the optimal intra-frame prediction mode into the video bitstream. To improve the coding efficiency of intra-frame prediction, the encoder / decoder derives three most likely modes from the optimal intra-frame prediction modes of the reconstructed image blocks in the surrounding area using intra-frame prediction. If the optimal intra-frame prediction mode selected for the current block is one of the three most likely modes, a first index is encoded to indicate that the selected optimal intra-frame prediction mode is one of the three most likely modes; if the selected optimal intra-frame prediction mode is not one of the three most likely modes, a second index is encoded to indicate that the selected optimal intra-frame prediction mode is one of the other 32 modes (other modes among the 35 candidate prediction modes except the aforementioned three most likely modes). The HEVC standard uses a 5-bit fixed-length code as the aforementioned second index.

[0157] The method by which the HEVC encoder derives the three most likely modes includes: selecting the optimal intra-frame prediction modes of the left adjacent image block and the upper adjacent image block of the current block and putting them into a set. If the two optimal intra-frame prediction modes are the same, only one is retained in the set. If the two optimal intra-frame prediction modes are the same and both are angle prediction modes, two angle prediction modes adjacent to the angle direction are selected to be added to the set; otherwise, the plane prediction mode, the DC mode, and the vertical prediction mode are selected in turn to be added to the set until the number of modes in the set reaches 3.

[0158] After the HEVC decoder performs entropy decoding on the bitstream, it obtains the mode information of the current block, which includes an indication flag indicating whether the optimal intra-frame prediction mode of the current block is among the three most likely modes, and the index of the optimal intra-frame prediction mode of the current block among the three most likely modes or the index of the optimal intra-frame prediction mode of the current block among the other 32 modes.

[0159] Inter prediction

[0160] In a possible implementation, the set of inter-prediction modes depends on the available reference picture (i.e., at least part of the previously decoded picture stored in the DBP 230 as mentioned above) and other inter-prediction parameters, for example, on whether to use the entire reference picture or only a part of the reference picture, such as a search window area around the area of ​​the current block, to search for the best matching reference block, and / or on whether to perform pixel interpolation such as half-pixel, quarter-pixel and / or 1 / 16 interpolation, for example.

[0161] In addition to the above prediction modes, skip mode and / or direct mode may also be employed.

[0162] For example, extended merge prediction, the merge candidate list of this mode consists of the following five candidate types in order: spatial MVP from spatially adjacent CUs, temporal MVP from collocated CUs, history-based MVP from FIFO tables, pairwise average MVP, and zero MV. Decoder side motion vector refinement (DMVR) based on bilateral matching can be used to increase the accuracy of the MV of the merge mode. Merge mode with MVD (MMVD) is derived from the merge mode with motion vector difference. The MMVD flag is sent immediately after the skip flag and merge flag are sent to specify whether the CU uses the MMVD mode. The CU-level adaptive motion vector resolution (AMVR) scheme can be used. AMVR supports encoding the MVD of the CU with different precisions. The MVD of the current CU is adaptively selected according to the prediction mode of the current CU. When the CU is encoded in merge mode, the combined inter / intra prediction (CIIP) mode can be applied to the current CU. The CIIP prediction is obtained by weighted averaging the inter and intra prediction signals. For affine motion compensated prediction, the affine motion field of the block is described by the motion information of 2 control points (4 parameters) or 3 control points (6 parameters) motion vectors. Subblock-based temporal motion vector prediction (SbTMVP) is similar to the temporal motion vector prediction (TMVP) in HEVC, but the motion vector of the sub-CU within the current CU is predicted. Bidirectional optical flow (BDOF), formerly known as BIO, is a simplified version that reduces calculations, especially in terms of the number of multiplications and the size of the multipliers. In the triangle partitioning mode, the CU is evenly divided into two triangular parts in two ways: diagonal partitioning and anti-diagonal partitioning. In addition, the bidirectional prediction mode is extended on the basis of simple averaging to support weighted averaging of two prediction signals.

[0163] The inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both in Figure 2). The motion estimation unit may be configured to receive or acquire the image block 203 (the current image block 203 of the current image 17) and the decoded image 231, or at least one or more previously reconstructed blocks, for example, reconstructed blocks of one or more other / different previously decoded images 231, for motion estimation. For example, a video sequence may include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 may be part of or form a sequence of images forming the video sequence.

[0164] For example, the encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different images in a plurality of other images, and provide the reference image (or reference image index) and / or the offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. The offset is also referred to as a motion vector (MV).

[0165] The motion compensation unit is used to obtain, for example, receive, inter-frame prediction parameters, and perform inter-frame prediction based on or using the inter-frame prediction parameters to obtain an inter-frame prediction block 246. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction block based on a motion / block vector determined by motion estimation, and may also include performing interpolation with sub-pixel accuracy. Interpolation filtering can generate pixel points of other pixels from pixel points of known pixels, thereby potentially increasing the number of candidate prediction blocks that can be used to encode the image block. Once a motion vector corresponding to a PU of a current image block is received, the motion compensation unit can locate the prediction block pointed to by the motion vector in one of the reference image lists.

[0166] The motion compensation unit may also generate syntax elements associated with blocks and video slices for use by the video decoder 30 when decoding image blocks of the video slice. In addition, or as an alternative to slices and corresponding syntax elements, coding block groups and / or coding blocks and corresponding syntax elements may be generated or used.

[0167] In the process of obtaining a candidate motion vector list in an advanced motion vector prediction (AMVP) mode, the motion vectors (MVs) that can be added to the candidate motion vector list as alternatives include MVs of image blocks that are spatially adjacent and temporally adjacent to the current block, wherein the MVs of image blocks that are spatially adjacent may include the MV of a left candidate image block located on the left side of the current block and the MV of an upper candidate image block located above the current block. For example, please refer to Figure 4 , Figure 4 An exemplary schematic diagram of a candidate image block provided in an embodiment of the present application is as follows: Figure 4As shown, the set of left candidate image blocks includes {A0, A1}, the set of upper candidate image blocks includes {B0, B1, B2}, and the set of temporally adjacent candidate image blocks includes {C, T}. All three sets can be added to the candidate motion vector list as alternatives. However, according to the existing coding standard, the maximum length of the candidate motion vector list of AMVP is 2, so it is necessary to determine the MV of adding up to two image blocks to the candidate motion vector list from the three sets according to the prescribed order. The order can be to give priority to the set of left candidate image blocks {A0, A1} of the current block (consider A0 first, and then consider A1 if A0 is not available), then consider the set of upper candidate image blocks {B0, B1, B2} of the current block (consider B0 first, and then consider B1 if B0 is not available, and then consider B2 if B1 is not available), and finally consider the set of temporally adjacent candidate image blocks {C, T} of the current block (consider T first, and then consider C if T is not available).

[0168] After obtaining the above candidate motion vector list, the optimal MV is determined from the candidate motion vector list by the rate distortion cost (RDcost), and the candidate motion vector with the smallest RD cost is used as the motion vector predictor (MVP) of the current block. The rate distortion cost is calculated by the following formula:

[0169] J=SAD+λR

[0170] Wherein, J represents RD cost, SAD is the sum of absolute differences (SAD) between the pixel value of the predicted block obtained after motion estimation using the candidate motion vector and the pixel value of the current block, R represents the bit rate, and λ represents the Lagrange multiplier.

[0171] The encoder passes the index of the determined MVP in the candidate motion vector list to the decoder. Furthermore, a motion search can be performed in the neighborhood centered on the MVP to obtain the actual motion vector of the current block. The encoder calculates the motion vector difference (MVD) between the MVP and the actual motion vector, and also passes the MVD to the decoder. The decoder parses the index, finds the corresponding MVP in the candidate motion vector list based on the index, parses the MVD, and adds the MVD to the MVP to obtain the actual motion vector of the current block.

[0172] In the process of obtaining the candidate motion information list in the merging mode, the motion information that can be added to the candidate motion information list as an alternative includes the motion information of the image blocks that are spatially adjacent or temporally adjacent to the current block, wherein the image blocks that are spatially adjacent and the image blocks that are temporally adjacent can refer to Figure 4 , the candidate motion information corresponding to the spatial domain in the candidate motion information list comes from the five spatially adjacent blocks (A0, A1, B0, B1 and B2). If the spatially adjacent block is not available or is intra-predicted, its motion information will not be added to the candidate motion information list. The candidate motion information of the current block in the time domain is obtained by scaling the MV of the block at the corresponding position in the reference frame according to the picture order count (POC) of the reference frame and the current frame. First, determine whether the block at position T in the reference frame is available. If not, select the block at position C. After obtaining the above candidate motion information list, the optimal motion information is determined from the candidate motion information list through the RD cost as the motion information of the current block. The encoder passes the index value (denoted as mergeindex) of the position of the optimal motion information in the candidate motion information list to the decoder.

[0173] Entropy Coding

[0174] The entropy coding unit 270 is used to apply an entropy coding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC scheme (CALVC), an arithmetic coding scheme, a binarization algorithm, a context adaptive binary arithmetic coding (CABAC), a syntax-based context-adaptive binary arithmetic coding (SBAC), a probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements to obtain coded image data 21 that can be output in the form of a coded bit stream 21 through an output terminal 272, so that the video decoder 30, etc. can receive and use the parameters for decoding. The coded bit stream 21 can be transmitted to the video decoder 30, or it can be stored in a memory for later transmission or retrieval by the video decoder 30.

[0175] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform based encoder 20 may directly quantize the residual signal without a transform processing unit 206 for certain blocks or frames. In another implementation, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.

[0176] Decoder and decoding method

[0177] like Figure 3 As shown, the video decoder 30 is used to receive the coded image data 21 (e.g., the coded bitstream 21) encoded by the encoder 20, and obtain a decoded image 331. The coded image data or bitstream includes information for decoding the coded image data, such as data representing image blocks of a coded video slice (and / or a coding block group or coding block) and related syntax elements.

[0178] exist Figure 3 In the example of FIG. 3 , decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter-frame prediction unit 344, and an intra-frame prediction unit 354. Inter-frame prediction unit 344 may be or include a motion compensation unit. In some examples, video decoder 30 may perform substantially the same as reference Figure 2 The decoding process of the video encoder 100 is the reverse of the encoding process described above.

[0179] As described above for encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer DPB 230, inter prediction unit 344, and intra prediction unit 354 also constitute a "built-in decoder" of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 122, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Therefore, the explanation of the corresponding units and functions of video encoder 20 applies accordingly to the corresponding units and functions of video decoder 30.

[0180] Entropy decoding

[0181] The entropy decoding unit 304 is used to parse the bit stream 21 (or generally the encoded image data 21) and perform entropy decoding on the encoded image data 21 to obtain the quantization coefficients 309 and / or the decoded encoding parameters ( Figure 3The entropy decoding unit 304 may be used to provide the inter-frame prediction parameters, intra-frame prediction parameters and / or other syntax elements to the mode application unit 360, and to provide other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice and / or video block level. In addition, or as an alternative to slices and corresponding syntax elements, coding block groups and / or coding blocks and corresponding syntax elements may be received or used.

[0182] Dequantization

[0183] The inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally information related to inverse quantization) and a quantization coefficient from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse quantize the decoded quantization coefficient 309 based on the quantization parameter to obtain an inverse quantization coefficient 311, which may also be referred to as a transform coefficient 311. The inverse quantization process may include using the quantization parameter calculated by the video encoder 20 for each video block in the video slice to determine a degree of quantization, and also determine a degree of inverse quantization that needs to be performed.

[0184] Inverse Transform

[0185] The inverse transform processing unit 312 may be configured to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain a reconstructed residual block 213 in the pixel domain. The reconstructed residual block 213 may also be referred to as a transform block 313. The transform may be an inverse transform, such as an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may also be configured to receive transform parameters or corresponding information from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304) to determine the transform applied to the dequantized coefficients 311.

[0186] reconstruction

[0187] The reconstruction unit 314 (eg, the summer 314 ) is used to add the reconstructed residual block 313 to the prediction block 365 to obtain the reconstructed block 315 in the pixel domain, for example, by adding the pixel point values ​​of the reconstructed residual block 313 and the pixel point values ​​of the prediction block 365 .

[0188] Filtering

[0189] The loop filter unit 320 (in or after the encoding loop) is used to filter the reconstruction block 315 to obtain a filter block 321, so as to smoothly perform pixel conversion or improve video quality, etc. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chromascaling (LMCS) (i.e., an adaptive in-loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Figure 3 3. Although shown as a loop filter in FIG. 3, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.

[0190] Decoded Image Buffer

[0191] The decoded video blocks 321 in one picture are then stored in a decoded picture buffer 330 which stores the decoded pictures 331 as reference pictures for subsequent motion compensation of other pictures and / or respective output displays.

[0192] The decoder 30 is used to output the decoded image 311 through the output terminal 312, etc., to display it to the user or for the user to view it.

[0193] predict

[0194] The inter-frame prediction unit 344 may be functionally the same as the inter-frame prediction unit 244 (particularly the motion compensation unit), and the intra-frame prediction unit 354 may be functionally the same as the inter-frame prediction unit 254, and may determine the division or segmentation and perform prediction based on the segmentation and / or prediction parameters or corresponding information received from the coded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304). The mode application unit 360 may be used to perform prediction (intra-frame or inter-frame prediction) of each block according to the reconstructed image, block or corresponding pixel point (filtered or unfiltered), and obtain a prediction block 365.

[0195] When the video slice is encoded as an intra coded (I) slice, the intra prediction unit 354 in the mode application unit 360 is used to generate a prediction block 365 for the image block of the current video slice based on the indicated intra prediction mode and data from the previously decoded block of the current image. When the video image is encoded as an inter coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) in the mode application unit 360 is used to generate a prediction block 365 for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. For inter prediction, these prediction blocks can be generated from one of the reference images in one of the reference image lists. The video decoder 30 can construct reference frame list 0 and list 1 using a default construction technique based on the reference images stored in the DPB 330. The same or similar processes may be applied to embodiments of coding block groups (e.g., video coding block groups) and / or coding blocks (e.g., video coding blocks) in addition to or as an alternative to slices (e.g., video slices), for example, a video may be encoded using I, P or B coding block groups and / or coding blocks.

[0196] The mode application unit 360 is used to determine prediction information for a video block of a current video slice by parsing motion vectors and other syntax elements, and use the prediction information to generate a prediction block for the current video block being decoded. For example, the mode application unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) for encoding the video block of the video slice, the inter-frame prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information of one or more reference picture lists for the slice, the motion vector for each inter-frame coded video block of the slice, the inter-frame prediction state for each inter-frame coded video block of the slice, and other information to decode the video block in the current video slice. In addition to slices (e.g., video slices) or as an alternative to slices, the same or similar process can be applied to embodiments of coding block groups (e.g., video coding block groups) and / or coding blocks (e.g., video coding blocks), for example, the video can be encoded using I, P, or B coding block groups and / or coding blocks.

[0197] In one embodiment, Figure 3The video encoder 30 may also be configured to partition and / or decode an image using slices (also referred to as video slices), wherein an image may be partitioned or decoded using one or more slices (usually non-overlapping). Each slice may include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., coding blocks in the H.265 / HEVC / VVC standard and bricks in the VVC standard).

[0198] In one embodiment, Figure 3 The video decoder 30 shown can also be used to segment and / or decode an image using slices / coding block groups (also called video coding block groups) and / or coding blocks (also called video coding blocks), wherein an image can be segmented or decoded using one or more slices / coding block groups (usually non-overlapping), each slice / coding block group may include one or more blocks (e.g., CTU) or one or more coding blocks, etc., wherein each coding block may be in a rectangular shape, etc., and may include one or more complete or partial blocks (e.g., CTU).

[0199] Other variations of the video decoder 30 may be used to decode the encoded image data 21. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, a non-transform based decoder 30 may directly dequantize the residual signal without the inverse transform processing unit 312 for certain blocks or frames. In another implementation, the video decoder 30 may have the dequantization unit 310 and the inverse transform processing unit 312 combined into a single unit.

[0200] It should be understood that the processing result of the current step can be further processed in the encoder 20 and the decoder 30 and then output to the next step. For example, after interpolation filtering, motion vector derivation or loop filtering, the processing result of interpolation filtering, motion vector derivation or loop filtering can be further operated, such as clipping or shifting operation.

[0201] It should be noted that the derived motion vector of the current block (including but not limited to the control point motion vector of the affine mode, the sub-block motion vector of the affine, plane, ATMVP mode, the temporal motion vector, etc.) can be further operated. For example, the value of the motion vector is limited to a predefined range according to the representation bit of the motion vector. If the representation bit of the motion vector is bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" represents the power. For example, if bitDepth is set to 16, the range is -32768 to 32767; if bitDepth is set to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of 4 4×4 sub-blocks in an 8×8 block) is limited so that the maximum difference between the integer parts of the above 4 4×4 sub-block MVs does not exceed N pixels, for example, not more than 1 pixel. Two methods of limiting motion vectors according to bitDepth are provided here.

[0202] Although the above embodiments mainly describe video coding and decoding, it should be noted that the embodiments of the decoding system 10, the encoder 20 and the decoder 30 and other embodiments described herein can also be used for still image processing or coding and decoding, that is, the processing or coding and decoding of a single image in video coding and decoding that is independent of any previous or consecutive images. In general, if the image processing is limited to a single image 17, the inter-frame prediction unit 244 (encoder) and the inter-frame prediction unit 344 (decoder) may not be available. All other functions (also called tools or techniques) of the video encoder 20 and the video decoder 30 can also be used for still image processing, such as residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354 and / or loop filtering 220 / 320, entropy encoding 270 and entropy decoding 304.

[0203] Please refer to Figure 5 , Figure 5 An exemplary block diagram of a video decoding device 500 provided in an embodiment of the present application. The video decoding device 500 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 500 may be a decoder, such as Figure 1a The video decoder 30 in the embodiment may also be an encoder, for example Figure 1a The video encoder 20 in.

[0204] The video decoding device 500 includes: an input port 510 (or input port 510) and a receiving unit (Rx) 520 for receiving data; a processor, a logic unit or a central processing unit (CPU) 530 for processing data; for example, the processor 530 here can be a neural network processor 530; a transmitter unit (Tx) 540 and an output port 550 (or output port 550) for transmitting data; and a memory 560 for storing data. The video decoding device 500 may also include an optical-to-electrical (OE) component and an electrical-to-optical (EO) component coupled to the input port 510, the receiving unit 520, the transmitting unit 540 and the output port 550 for the output or output of optical or electrical signals.

[0205] The processor 530 is implemented by hardware and software. The processor 530 can be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 530 communicates with the input port 510, the receiving unit 520, the sending unit 540, the output port 550, and the memory 560. The processor 530 includes a decoding module 570 (e.g., a decoding module 570 based on a neural network). The decoding module 570 implements the embodiments disclosed above. For example, the decoding module 570 performs, processes, prepares, or provides various encoding operations. Therefore, the decoding module 570 provides substantial improvements to the functions of the video decoding device 500 and affects the switching of the video decoding device 500 to different states. Alternatively, the decoding module 570 is implemented with instructions stored in the memory 560 and executed by the processor 530.

[0206] The memory 560 includes one or more disks, tape drives, and solid-state hard disks, and can be used as an overflow data storage device for storing such programs when such programs are selected for execution, and for storing instructions and data read during program execution. The memory 560 can be volatile and / or non-volatile, and can be a read-only memory (ROM), a random access memory (RAM), a ternary content-addressable memory (TCAM), and / or a static random-access memory (SRAM).

[0207] Please refer to Figure 6 , Figure 6An exemplary block diagram of a device 600 provided in an embodiment of the present application, the device 600 can be used as Figure 1a Either or both of the source device 12 and the destination device 14 in .

[0208] The processor 602 in the device 600 may be a central processing unit. Alternatively, the processor 602 may be any other type of device or devices that are currently available or will be developed in the future and are capable of manipulating or processing information. Although a single processor such as the processor 602 shown in the figure may be used to implement the disclosed implementation, using more than one processor is faster and more efficient.

[0209] In one implementation, the memory 604 in the apparatus 600 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 604. The memory 604 may include code and data 606 accessed by the processor 602 via a bus 612. The memory 604 may also include an operating system 608 and an application 610, wherein the application 610 includes at least one program that allows the processor 602 to perform the method described above. For example, the application 610 may include applications 1 to N, and also include a video decoding application that performs the method described above.

[0210] The apparatus 600 may also include one or more output devices, such as a display 618. In one example, the display 618 may be a touch-sensitive display that combines a display with a touch-sensitive element that may be used to sense touch input. The display 618 may be coupled to the processor 602 via the bus 612.

[0211] Although the bus 612 in the device 600 is described herein as a single bus, the bus 612 may include multiple buses. In addition, the auxiliary storage may be directly coupled to other components of the device 600 or accessed through a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Therefore, the device 600 may have a variety of configurations.

[0212] Please refer to Figure 7 , Figure 7 A coding method is provided in the embodiment of the present application. Figure 7 As shown, the encoding method may include:

[0213] S701: Encode the image to be encoded to obtain a code stream.

[0214] Exemplarily, the image to be encoded may be encoded by an encoder to obtain a code stream.

[0215] For example, the image to be encoded can be encoded by an H.264 encoder, an H.265 encoder, an H.266 encoder, a second-generation audio video coding standard (AVS2) encoder, a third-generation audio video coding standard (AVS3) encoder, a second-generation surveillance video and audio coding (SVAC2) encoder, a third-generation surveillance video and audio coding (SVAC3) encoder, or other encoders to obtain a code stream.

[0216] S702: Perform image segmentation on the image to be encoded to obtain a first segmented image.

[0217] In a possible implementation, the image to be encoded may be input into a target model to obtain the first segmented image, and a training set of the target model includes an image and a segmented image of the image.

[0218] Exemplarily, the target model can be an AI model

[0219] Exemplarily, the first segmented image may be a downsampled segmented image.

[0220] S703: Perform mapping processing on the first segmented image to obtain a second segmented image.

[0221] In a possible implementation, mapping processing may be performed on the first segmented image, and the first segmented image may be mapped into a pixel space of an encoder to obtain a second segmented image.

[0222] In a possible implementation, pixel values ​​of the pixels in the first segmented image may be mapped to ascending sorting numbers, and then the mapped pixel values ​​may be uniformly scaled to a preset value range.

[0223] For example, the total types of pixels in the first segmented image are 8, and the corresponding pixel values ​​are {0, 1, 2, 3, 4, 5, 6, 7}. The pixel values ​​of the pixels in the first segmented image are sorted from small to large {0, 1, 2, 3, 4, 5, 6, 7}, so the mapping rule is {0->0, 1->1, 2->2, 3->3, 4->4, 5->5, 6->6, 7->7}. Then, the mapped pixel values ​​are uniformly scaled to the numerical interval [0, 255] to obtain the second segmented image, and the pixel values ​​of the pixels in the second segmented image are {16, 48, 80, 112, 144, 176, 208, 240}.

[0224] Exemplarily, the pixel value L of the pixel point of the second segmented image may satisfy:

[0225]

[0226] Wherein, L is the pixel value of the pixel point of the second segmented image, N is the number of pixel point types in the first segmented image, and i is the sorting sequence number after the pixel value of the pixel point in the first segmented image is mapped. Indicates rounding down.

[0227] Among them, the pixel value of the pixel point of the first segmented image is 0, which indicates that the pixel point is a background pixel point; the pixel value of the pixel point of the first segmented image is 1, which indicates that the pixel point is a person pixel point; the pixel value of the pixel point of the first segmented image is 2, which indicates that the pixel point is a face pixel point; the pixel value of the pixel point of the first segmented image is 3, which indicates that the pixel point is a motor vehicle pixel point; the pixel value of the pixel point of the first segmented image is 4, which indicates that the pixel point is a non-motor vehicle pixel point; the pixel value of the pixel point of the first segmented image is 5, which indicates that the pixel point is an object pixel point; the pixel value of the pixel point of the first segmented image is 6, which indicates that the pixel point is a scene pixel point; the pixel value of the pixel point of the first segmented image is 7, which indicates that the pixel point is a license plate pixel point.

[0228] As another example, the total types of pixels in the first segmented image are 4, and the corresponding pixel values ​​are {0, 4, 6, 7}. The pixel values ​​of the pixels in the first segmented image are sorted from small to large {0, 1, 2, 3}, so the mapping rule is {0->0, 4->1, 6->2, 7->3}. Then, the mapped pixel values ​​are uniformly scaled to the numerical range [0, 255] to obtain the second segmented image, and the pixel values ​​of the pixels in the second segmented image are {32, 96, 160, 244}.

[0229] It should be noted that the above method of obtaining the second segmented image according to the first segmented image is only an exemplary description. The method of obtaining the second segmented image according to the first segmented image can also be implemented by other methods that can be thought of by those skilled in the art, and the embodiments of the present application are not limited to this.

[0230] S704: Determine mapping information according to the first segmented image.

[0231] Among them, the above-mentioned mapping information includes first information, second information and third information. The above-mentioned first information is used to indicate the number of pixel point types in the above-mentioned first segmented image, the above-mentioned second information is used to indicate the pixel values ​​corresponding to the pixel point types in the above-mentioned first segmented image, and the above-mentioned third information is used to indicate the mapping pixel values ​​corresponding to the pixel point types in the above-mentioned first segmented image.

[0232] In a possible implementation manner, the first information and the second information may be determined according to the first segmented image, and the third information may be determined according to the second information.

[0233] In a possible implementation manner, the pixel values ​​in the second information may be shifted to obtain the mapped pixel values ​​in the third information.

[0234] For example, as shown in Table 1, the total types of pixels in the first segmented image are 8, and the corresponding pixel values ​​are {0, 1, 2, 3, 4, 5, 6, 7}, that is, the pixel values ​​in the second information are {0, 1, 2, 3, 4, 5, 6, 7}. Converting {0, 1, 2, 3, 4, 5, 6, 7} from decimal to binary can obtain {000, 001, 010, 011, 100, 101, 110, 111}, and shifting {000, 001, 010, 011, 100, 101, 110, 111} left by 5 bits can obtain {000000000, 00100000, 01000000, 01100000, 10000000, 10100000, 1100000 00, 11100000}, converting {00000000, 00100000, 01000000, 01100000, 10000000, 10100000, 11000000, 11100000} from binary to decimal can obtain {0, 32, 64, 96, 128, 160, 192, 224}, that is, the pixel values ​​in the third information are {0, 32, 64, 96, 128, 160, 192, 224}.

[0235] Table 1

[0236]

[0237] As another example, as shown in Table 2, the total types of pixel points in the first segmented image are 4 types, and the corresponding pixel values ​​are {0, 4, 6, 7}, that is, the pixel values ​​in the second information are {0, 4, 6, 7}. Converting {0, 4, 6, 7} from decimal to binary yields {000, 100, 110, 111}, shifting {000, 001, 010, 011, 100, 101, 110, 111} left by 5 bits yields {000000000, 10000000, 11000000, 11100000}, and converting {000000000, 10000000, 11000000, 11100000} from binary to decimal yields {0, 128, 192, 224}, i.e., the pixel value in the third information is {0, 128, 192, 224}.

[0238] Table 2

[0239]

[0240] It should be noted that the above method of obtaining the third information according to the second information is only an exemplary description. The method of obtaining the third information according to the second information can also be implemented by other methods that can be thought of by those skilled in the art, and the embodiments of the present application are not limited to this.

[0241] In a possible implementation manner, the mapping information further includes fourth information and fifth information, the fourth information is used to indicate the width of the first segmented image, and the fifth information is used to indicate the height of the first segmented image.

[0242] S705: Encode the second segmented image to obtain a segmented coded image.

[0243] Exemplarily, the second segmented image may be encoded by an encoder to obtain a segmented encoded image.

[0244] For example, the second segmented image may be encoded by an H.264 encoder, an H.265 encoder, an H.266 encoder, an AVS2 encoder, an AVS3 encoder, an SVAC2 encoder, an SVAC3 encoder, or other encoders to obtain segmented encoded images.

[0245] In a possible implementation, when the encoder supports the YUV400 (a color encoding method) format, the second segmented image may be input into the encoder for encoding to obtain segmented encoded images.

[0246] In another possible implementation, when the encoder only supports the YUV420 format, the second segmented image can be converted into the YUV420 format (such as setting the UV component of the second segmented image to the median value of the pixel), and the second segmented image in the YUV420 format is input into the encoder for encoding to obtain a segmented encoded image.

[0247] In a possible implementation, the first parameter of the segmented coded image and the first parameter of the image to be coded (the coded image of the image to be coded) may be the same, wherein the first parameter includes a picture order count (POC) and / or a decoding order index (DOI).

[0248] It can be understood that if the image to be encoded is a random access point image, the segmented encoded images of the image to be encoded should also be random access point images.

[0249] S706: Encode the segmented coded image, the mapping information and the image into a bitstream.

[0250] In a possible implementation manner, the mapping information and the segmented coded image may be encoded into the extended information or supplementary enhancement information (SEI) of the code stream.

[0251] In a possible implementation, the segmented coded images, the mapping information, and the images may be encoded into a bitstream according to a preset syntax table.

[0252] Exemplarily, the segmented coded images, mapping information, and images may be encoded into a bitstream according to the syntax table shown in Table 3.

[0253] The map_class_num in the syntax table shown in Table 3, i.e., the number of pixel categories of the segmentation map, is used to indicate the number of pixel types in the first segmented image. The map_class_num may be an 8-bit unsigned integer.

[0254] The map_label[i] in the syntax table shown in Table 3, i.e., the segmentation map pixel label, is used to indicate the pixel value corresponding to the pixel point type in the first segmentation image. The map_label[i] may be an 8-bit unsigned integer.

[0255] The map_value[i] in the syntax table shown in Table 3, i.e., the segmentation image pixel label mapping value, is used to indicate the mapping pixel value corresponding to the pixel point type in the first segmentation image. Map_value[i] can be an 8-bit unsigned integer.

[0256] The map_width in the syntax table shown in Table 3, i.e. the segmentation map width, is used to indicate the width of the first segmented image and can be a 16-bit unsigned integer.

[0257] The map_height in the syntax table shown in Table 3, ie, the segmentation map height, is used to indicate the height of the first segmented image, and may be a 16-bit unsigned integer.

[0258] The map_Data in the syntax table shown in Table 3, namely the segmentation map data map_Data, is used to represent the segmentation of the coded image. The map_Data can be extracted and decoded by a decoder.

[0259] Table 3

[0260] Parameter Definition Descriptors SvacEx_aiseg_info(){ map_class_num u(8) for(i=0;i<map_class_num;i++){ map_label[i] u(8) map_value[i] u(8) } map_width u(16) map_height u(16) map_Data u(n) }

[0261] As another example, the segmented coded images, mapping information, and images may be encoded into the bitstream according to the syntax table shown in Table 4.

[0262] The map_class_num_minus1 in the syntax table shown in Table 4, namely the number of pixel categories in the segmentation map, is used to indicate the number of pixel types in the first segmentation image. The map_class_num_minus1 may be an 8-bit unsigned integer.

[0263] It should be noted that the number of pixel types in the first segmented image MapClassNum is map_class_num_minus1 + 1. For example, if map_class_num_minus1 is 3, the number of pixel types in the first segmented image is 4. For another example, if map_class_num_minus1 is 0, the number of pixel types in the first segmented image is 1.

[0264] The map_label[i] in the syntax table shown in Table 4, namely the segmentation map pixel label, is used to indicate the pixel value corresponding to the pixel point type in the first segmentation image. The map_label[i] may be an 8-bit unsigned integer.

[0265] It should be noted that map_label[0]=0 in the syntax table shown in Table 4.

[0266] The map_value[i] in the syntax table shown in Table 4, i.e., the segmentation image pixel label mapping value, is used to indicate the mapping pixel value corresponding to the pixel point type in the first segmentation image. Map_value[i] can be an 8-bit unsigned integer.

[0267] It should be noted that map_value[0] in the syntax table shown in Table 4=0.

[0268] It can be understood that, compared with the syntax table shown in Table 3, the syntax table shown in Table 4 can reduce the pixel categories that transmit the value 0 and their corresponding pixel mapping values.

[0269] Table 4

[0270]

[0271] The method provided in the embodiment of the present application, by encoding the segmented image into the bitstream of the image to be encoded, enables the decoding end to obtain the segmented image through the decoded bitstream, thereby avoiding the device from repeatedly segmenting the same image. In addition, by encoding the mapping information of the segmented image into the bitstream, the decoding end can restore the segmented image through the mapping information, thereby reducing the difficulty of restoring the segmented image.

[0272] Please refer to Figure 8 , Figure 8 A decoding method is provided in an embodiment of the present application. Figure 8 As shown, the decoding method may include:

[0273] S801. Decode the code stream to obtain segmented coded images and mapping information.

[0274] Among them, the above-mentioned segmented coded image is obtained based on the first segmented decoded image, and the above-mentioned mapping information includes first information, second information and third information. The above-mentioned first information is used to indicate the number of pixel point types in the first segmented image, the above-mentioned second information is used to indicate the pixel values ​​corresponding to the pixel point types in the above-mentioned first segmented image, and the above-mentioned third information is used to indicate the mapping pixel values ​​corresponding to the pixel point types in the above-mentioned first segmented image.

[0275] Exemplarily, the code stream may be decoded according to the syntax table shown in Table 3 to obtain mapping information (map_class_num, map_label[i], and map_value[i]) and a segmented coded image (map_Data).

[0276] As another example, the code stream may be decoded according to the syntax table shown in Table 4 to obtain mapping information (map_class_num_minus1, map_label[i], and map_value[i]) and a segmented coded image (map_Data).

[0277] In a possible implementation manner, the mapping information further includes fourth information and fifth information, the fourth information is used to indicate the width of the first segmented image, and the fifth information is used to indicate the height of the first segmented image.

[0278] S802: Decode the segmented coded image to obtain a first segmented decoded image.

[0279] In a possible implementation manner, the decoded segmented coded image may be decoded by a decoder to obtain a first segmented decoded image.

[0280] For example, the decoded segmented encoded image can be decoded by an H.264 decoder, an H.265 decoder, an H.266 decoder, an AVS2 decoder, an AVS3 decoder, an SVAC2 decoder, an SVAC3 decoder or other decoders to obtain a first segmented decoded image.

[0281] The decoder used for decoding the segmented coded image may correspond to the encoder used for encoding the second segmented image.

[0282] For example, if an H.265 encoder is used to encode the second segmented image to obtain a segmented coded image, then the decoder used to decode the segmented coded image is an H.265 decoder.

[0283] For another example, if an AVS3 encoder is used to encode the second segmented image to obtain a segmented coded image, then the decoder used to decode the segmented coded image is the AVS3 decoder.

[0284] S803 . Perform inverse mapping processing on the first segmented decoded image according to the mapping information to obtain a second segmented decoded image.

[0285] In a possible implementation manner, the pixel value of the pixel point in the second segmented decoded image may be determined according to the pixel value of the pixel point in the first segmented decoded image and the mapping information.

[0286] In a possible implementation, the mapped pixel value of the pixel in the first segmented decoded image can be determined according to the pixel value of the pixel in the first segmented decoded image and the mapping information. The pixel value of the pixel in the second segmented decoded image can be determined according to the mapped pixel value and the mapping information.

[0287] In a possible implementation, a mapping pixel value interval of a pixel point in the first segmented decoded image can be determined according to the mapping information, and a mapping pixel value of a pixel point in the first segmented decoded image can be determined according to the pixel value of the pixel point in the first segmented decoded image and the mapping pixel value interval. A pixel value of a pixel point in the second segmented decoded image can be determined according to the mapping pixel value and the mapping information.

[0288] Exemplarily, as shown in Tables 1 and 5, the total types of pixel points of the first segmented decoded image are 8 types, the corresponding pixel values ​​are {16, 48, 80, 112, 144, 176, 208, 240}, and the pixel values ​​in the third information in the mapping information are {0, 32, 64, 96, 128, 160, 192, 224}. According to the third information in the mapping information, it can be determined that the mapping pixel value intervals of the pixel points in the first segmented decoded image are {0-31, 32-63, 64-95, 96-127, 128-159, 160-191, 192-223, 223-255} respectively. According to the mapping pixel value interval, the mapping pixel value of the pixel point in the first segmented decoded image can be determined to be {0, 32, 64, 96, 128, 160, 192, 224}. According to Table 1 and the mapping pixel value of the pixel point in the first segmented decoded image, the pixel value of the pixel point in the second segmented decoded image can be determined to be {0, 1, 2, 3, 4, 5, 6, 7}.

[0289] The mapping pixel value interval of the pixel points in the first segmented decoded image is used to indicate the mapping relationship between the pixel values ​​of the pixel points in the first segmented decoded image and the pixel values ​​in the third information (the mapping pixel values ​​of the pixel points in the first segmented decoded image).

[0290] For example, the pixel value 0 in the third information corresponds to a mapping pixel value range of 0 to 31, which means that the pixel value in the third information corresponding to the pixel point with a pixel value of 0 to 31 in the first segmented decoded image (the mapping pixel value of the pixel point in the first segmented decoded image) is 0.

[0291] Table 5

[0292]

[0293] As another example, as shown in Table 2 and Table 6, the total types of pixels in the first segmented decoded image are 4, the corresponding pixel values ​​are {16, 144, 208, 240}, and the pixel values ​​in the third information in the mapping information are {0, 128, 192, 224}. According to the third information in the mapping information, the mapping pixel value intervals of the pixels in the first segmented decoded image can be determined to be {0-31, 128-159, 192-223, 223-255}. According to the mapping pixel value intervals, the mapping pixel values ​​of the pixels in the first segmented decoded image can be determined to be {0, 128, 192, 224}. According to Table 1 and the mapping pixel values ​​of the pixels in the first segmented decoded image, the pixel values ​​of the pixels in the second segmented decoded image can be determined to be {0, 4, 6, 7}.

[0294] Table 6

[0295]

[0296] In a possible implementation, the target pixel value of the pixel in the first segmented decoded image can be determined according to the mapping information. The mapping pixel value of the pixel in the first segmented decoded image is determined according to the target pixel value of the pixel in the first segmented decoded image. The pixel value of the pixel in the second segmented decoded image is determined according to the mapping pixel value and the mapping information. The target pixel of the pixel in the first segmented decoded image is the pixel value whose pixel value in the third information of the mapping information is closest to the pixel value of the pixel in the first segmented decoded image.

[0297] Exemplarily, as shown in Table 1 and Table 7, the total types of pixel points of the first segmented decoded image are 8 types, and the corresponding pixel values ​​are {16, 48, 80, 112, 144, 176, 208, 240}, and the pixel values ​​in the third information in the mapping information are {0, 32, 64, 96, 128, 160, 192, 224}, and the pixel value of the pixel point in the third information that is closest to the pixel value 16 is 0, the pixel value of the pixel point in the third information that is closest to the pixel value 48 is 32, the pixel value of the pixel point in the third information that is closest to the pixel value 80 is 64, the pixel value of the pixel point in the third information that is closest to the pixel value 112 is 96, and the pixel value of the pixel point in the third information that is closest to the pixel value 16 is 0. The pixel value closest to the pixel value 144 in the third information is 128, the pixel value closest to the pixel value 176 in the third information is 160, the pixel value closest to the pixel value 208 in the third information is 192, and the pixel value closest to the pixel value 240 in the third information is 224. It can be determined that the mapped pixel values ​​of the pixels in the first segmented decoded image are {0, 32, 64, 96, 128, 160, 192, 224}. According to Table 1 and the mapped pixel values ​​of the pixels in the first segmented decoded image, it can be determined that the pixel values ​​of the pixels in the second segmented decoded image are {0, 1, 2, 3, 4, 5, 6, 7}.

[0298] Table 7

[0299]

[0300] Exemplarily, as shown in Table 2 and Table 8, the total types of pixels in the first segmented decoded image are 4 types, and the corresponding pixel values ​​are {16, 144, 208, 240}, the pixel values ​​in the third information in the mapping information are {0, 128, 192, 224}, the pixel value closest to the pixel value 16 in the pixel values ​​of the pixels in the third information is 0, the pixel value closest to the pixel value 144 in the pixel values ​​of the pixels in the third information is 128, the pixel value closest to the pixel value 208 in the pixel values ​​of the pixels in the third information is 192, and the pixel value closest to the pixel value 240 in the pixel values ​​of the pixels in the third information is 224. It can be determined that the mapped pixel values ​​of the pixels in the first segmented decoded image are {0, 128, 192, 224}. According to Table 2 and the mapped pixel values ​​of the pixels in the first segmented decoded image, the pixel values ​​of the pixels in the second segmented decoded image can be determined to be {0, 4, 6, 7}.

[0301] Table 8

[0302]

[0303] In a possible implementation, the second segmented decoded image can be used for downstream services, for example, the second segmented decoded image can be used in applications such as secondary encoding compression, video colorization, live album, and video condensation.

[0304] The following will be combined Fig. 9 A coding device for executing the above coding method is introduced.

[0305] It is understandable that, in order to realize the above functions, the encoding device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0306] The embodiment of the present application can divide the functional modules of the encoding device according to the above method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0307] In the case of dividing each functional module into corresponding functional modules, Fig. 9 A possible schematic diagram of the composition of the encoding device involved in the above embodiment is shown. Fig. 9 As shown, the encoding device 900 may include: an encoding unit 901, a segmentation unit 902 and a processing unit 903.

[0308] The encoding unit 901 is used to encode the image to be encoded to obtain a code stream.

[0309] The segmentation unit 902 is configured to segment the image to be encoded to obtain a first segmented image.

[0310] The processing unit 903 is used to perform mapping processing on the first segmented image to obtain a second segmented image.

[0311] The processing unit 903 is also used to determine mapping information based on the above-mentioned first segmented image, and the above-mentioned mapping information includes first information, second information and third information. The above-mentioned first information is used to indicate the number of pixel point types in the above-mentioned first segmented image, the above-mentioned second information is used to indicate the pixel value corresponding to the pixel point type in the above-mentioned first segmented image, and the above-mentioned third information is used to indicate the mapping pixel value corresponding to the pixel point type in the above-mentioned first segmented image.

[0312] The encoding unit 901 is further configured to encode the second segmented image to obtain a segmented encoded image.

[0313] The encoding unit 901 is further configured to encode the segmented encoded image and the mapping information and image into the bitstream.

[0314] In a possible implementation manner, the processing unit 903 is specifically configured to: determine the first information and the second information according to the first segmented image; and determine the third information according to the second information.

[0315] In a possible implementation manner, the processing unit 903 is specifically configured to: perform shift processing on the pixel values ​​in the second information to obtain the mapped pixel values ​​in the third information.

[0316] In a possible implementation, the segmentation unit 902 is specifically configured to: input the image to be encoded into a target model to obtain the first segmented image, and a training set of the target model includes an image and a segmented image of the image.

[0317] In a possible implementation manner, the encoding unit 901 is specifically configured to: encode the mapping information and the segmented encoded image into the extended information or supplementary enhancement information of the code stream.

[0318] In a possible implementation manner, the mapping information further includes fourth information and fifth information, the fourth information is used to indicate the width of the first segmented image, and the fifth information is used to indicate the height of the first segmented image.

[0319] The following will be combined Fig.10 A decoding device for executing the above decoding method is introduced.

[0320] It is understandable that, in order to implement the above functions, the decoding device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0321] The embodiment of the present application can divide the functional modules of the decoding device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0322] In the case of dividing each functional module into corresponding functional modules, Fig.10 A possible schematic diagram of the composition of the decoding device involved in the above embodiment is shown. Fig.10 As shown, the decoding device 1000 may include: a decoding unit 1001 and a processing unit 1002 .

[0323] The decoding unit 1001 is used to decode the code stream to obtain a segmented coded image and mapping information, wherein the segmented coded image is obtained based on the first segmented decoded image, and the mapping information includes first information, second information and third information, wherein the first information is used to indicate the number of pixel point types in the first segmented image, the second information is used to indicate the pixel value corresponding to the pixel point type in the first segmented image, and the third information is used to indicate the mapped pixel value corresponding to the pixel point type in the first segmented image.

[0324] The decoding unit 1001 is further configured to decode the segmented coded image to obtain the first segmented decoded image.

[0325] The processing unit 1002 is configured to perform inverse mapping processing on the first segmented decoded image according to the mapping information to obtain a second segmented decoded image.

[0326] In a possible implementation manner, the processing unit 1002 is specifically configured to determine the pixel value of the pixel in the second segmented decoded image according to the pixel value of the pixel in the first segmented decoded image and the mapping information.

[0327] In a possible implementation, the processing unit 1002 is specifically used to: determine the mapping pixel value of the pixel point in the first segmented decoded image based on the pixel value of the pixel point in the first segmented decoded image and the mapping information; determine the pixel value of the pixel point in the second segmented decoded image based on the mapping pixel value and the mapping information.

[0328] In a possible implementation manner, the mapping information further includes fourth information and fifth information, the fourth information is used to indicate the width of the first segmented image, and the fifth information is used to indicate the height of the first segmented image.

[0329] The embodiment of the present application also provides a chip. Fig.11 1 shows a schematic diagram of the structure of a chip 1100. The chip 1100 includes one or more processors 1101 and an interface circuit 1102. Optionally, the chip 1100 may also include a bus 1103.

[0330] The processor 1101 may be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above encoding method or decoding method may be completed by an integrated logic circuit of hardware in the processor 1101 or by instructions in the form of software.

[0331] Optionally, the processor 1101 may be a general purpose processor, a digital signal processing (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods and steps disclosed in the embodiments of the present application may be implemented or executed. The general purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0332] The interface circuit 1102 can be used to send or receive data, instructions or information. The processor 1101 can use the data, instructions or other information received by the interface circuit 1102 to process, and can send the processing completion information through the interface circuit 1102.

[0333] Optionally, the chip also includes a memory, which may include a read-only memory and a random access memory, and provides operation instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory (NVRAM).

[0334] Optionally, the memory stores executable software modules or data structures, and the processor can perform corresponding operations by calling operation instructions stored in the memory (the operation instructions can be stored in the operating system).

[0335] Optionally, the chip can be used in the encoding device involved in the embodiment of the present application. Optionally, the interface circuit 1102 can be used to output the execution result of the processor 1101. The encoding method or decoding method provided in one or more embodiments of the embodiment of the present application can refer to the aforementioned embodiments, which will not be repeated here.

[0336] It should be noted that the corresponding functions of the processor 1101 and the interface circuit 1102 can be implemented through hardware design, software design, or a combination of hardware and software, and there is no limitation here.

[0337] Fig.12 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1200 may be a processor or a chip or a functional module in a processor. Fig.12 As shown, the electronic device 1200 includes a processor 1201 , a transceiver 1202 and a communication line 1203 .

[0338] Among them, the processor 1201 is used to execute any step in the encoding method or decoding method provided in the embodiment of the present application, and in the process of executing any step in the encoding method or decoding method provided in the embodiment of the present application, the transceiver 1202 and the communication line 1203 can be selectively called to complete the corresponding operation.

[0339] Furthermore, the electronic device 1200 may also include a memory 1204 . The processor 1201 , the memory 1204 and the transceiver 1202 may be connected via a communication line 1203 .

[0340] The processor 1201 is a processor, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 1201 may also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.

[0341] The transceiver 1202 is used to communicate with other devices or other communication networks, and other communication networks may be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. The transceiver 1202 may be a module, a circuit, a transceiver or any device capable of achieving communication.

[0342] The transceiver 1202 is mainly used for sending and receiving commands and information, etc., and may include a transmitter and a receiver, which respectively send and receive commands and information, etc.; operations other than sending and receiving commands and information, etc. are implemented by the processor.

[0343] The communication line 1203 is used to transmit information between the components included in the electronic device 1200.

[0344] In one design, the processor may be considered as the logic circuit and the transceiver may be considered as the interface circuit.

[0345] The memory 1204 is used to store instructions, where the instructions may be computer programs.

[0346] The memory 1204 may be a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memories. The nonvolatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM). The memory 1204 may also be a compact disc read-only memory (CD-ROM) or other optical disk storage, optical disk storage (including compact disc, laser disc, optical disk, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage media or other magnetic storage devices, etc. It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0347] It should be noted that the memory 1204 can exist independently of the processor 1201, or can be integrated with the processor 1201. The memory 1204 can be used to store instructions or program codes or some data, etc. The memory 1204 can be located inside the electronic device 1200, or outside the electronic device 1200, without limitation. The processor 1201 is used to execute the instructions stored in the memory 1204 to implement the method provided in the above embodiment of the present application.

[0348] In one example, the processor 1201 may include one or more processors, such as Fig.12 Processor 0 and processor 1 in.

[0349] As an optional implementation, the electronic device 1200 includes multiple processors, for example, Fig.12 In addition to the processor 1201, a processor 1207 may also be included.

[0350] As an optional implementation, the electronic device 1200 further includes an output device 1205 and an input device 1206. Exemplarily, the input device 1206 is a keyboard, a mouse, a microphone, a joystick, and the like, and the output device 1205 is a display screen, a speaker, and the like.

[0351] It should be noted that the electronic device 1200 may be a chip system or a Fig.12 Devices with similar structures in the chip system. Among them, the chip system can be composed of chips, or it can include chips and other discrete devices. The actions, terms, etc. involved in the various embodiments of this application can refer to each other without limitation. The message name or parameter name in the message exchanged between the various devices in the embodiment of this application is only an example, and other names can also be used in the specific implementation without limitation. In addition, Fig.12 The components shown in the figure do not constitute a limitation on the electronic device 1200. Fig.12 In addition to the components shown, the electronic device 1200 may include Fig.12 More or fewer components may be shown, or certain components may be combined, or the components may be arranged differently.

[0352] The processor and transceiver described in the present application can be implemented in an integrated circuit (IC), an analog IC, a radio frequency integrated circuit, a mixed signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (NMOS), P-type metal oxide semiconductor (positive channel metal oxide semiconductor, PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0353] Fig.13This is a schematic diagram of the structure of a coding device provided in an embodiment of the present application. The coding device can be applied to the scenario shown in the above method embodiment. For the convenience of explanation, Fig.13 Only the main components of the encoding device are shown, including a processor 1301, a memory 1302, a control circuit 1303, and an input-output device 1304. The processor 1301 is mainly used to process the communication protocol and communication data, execute the software program, and process the data of the software program. The memory 1302 is mainly used to store the software program and data. The control circuit 1303 is mainly used for power supply and transmission of various electrical signals. The input-output device 1304 is mainly used to receive data input by the user and output data to the user.

[0354] When the encoding device is a processor 1301, the control circuit 1303 can be a mainboard, the memory 1302 includes a hard disk, RAM, ROM and other media with storage functions, the processor 1301 can include a baseband processor 1301 and a central processing unit, the baseband processor is mainly used to process the communication protocol and communication data, the central processing unit is mainly used to control the entire encoding device, execute software programs, and process the data of the software programs, and the input and output device 1304 includes a display screen, a keyboard and a mouse, etc.; the control circuit 1303 can further include or connect a transceiver circuit or a transceiver, such as: a network cable interface, etc., for sending or receiving data or signals, such as data transmission and communication with other devices. Further, it can also include an antenna for sending and receiving wireless signals, and for data / signal transmission with other devices.

[0355] An embodiment of the present application also provides a coding device, which includes: at least one processor, when the at least one processor executes the program code or instruction, the above-mentioned related method steps are implemented to implement the encoding method or decoding method in the above-mentioned embodiment.

[0356] Optionally, the device may further include at least one memory, and the at least one memory is used to store the program code or instruction.

[0357] An embodiment of the present application also provides a computer storage medium, in which computer instructions are stored. When the computer instructions are executed on an encoding device, the encoding device executes the above-mentioned related method steps to implement the encoding method or decoding method in the above-mentioned embodiment.

[0358] The embodiments of the present application also provide a computer program product. When the computer program product is run on a computer, the computer is enabled to execute the above-mentioned related steps to implement the encoding method or decoding method in the above-mentioned embodiments.

[0359] The embodiment of the present application also provides a coding device, which can be a chip, an integrated circuit, a component or a module. Specifically, the device may include a connected processor and a memory for storing instructions, or the device includes at least one processor for obtaining instructions from an external memory. When the device is running, the processor can execute instructions so that the chip executes the encoding method or decoding method in the above-mentioned method embodiments.

[0360] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0361] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0362] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0363] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0364] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0365] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0366] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the above methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0367] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A coding method, characterized in that: The method comprises: Encoding the image to be encoded to obtain a code stream; Performing image segmentation on the image to be encoded to obtain a first segmented image; Performing mapping processing on the first segmented image to obtain a second segmented image; Determine mapping information according to the first segmented image, the mapping information comprising first information, second information and third information, the first information being used to indicate the number of pixel types in the first segmented image, the second information being used to indicate pixel values ​​corresponding to the pixel types in the first segmented image, and the third information being used to indicate mapped pixel values ​​corresponding to the pixel types in the first segmented image; encoding the second segmented image to obtain a segmented coded image; The divided coded images and the mapping information and images are encoded into the code stream.

2. The method according to claim 1, characterized in that The determining mapping information according to the first segmented image includes: determining the first information and the second information according to the first segmented image; The third information is determined according to the second information.

3. The method according to claim 2, characterized in that determining the third information according to the second information; The pixel values ​​in the second information are shifted to obtain the mapped pixel values ​​in the third information.

4. The method according to any one of claims 1 to 3, characterized in that The step of segmenting the image to be encoded to obtain a first segmented image includes: The image to be encoded is input into a target model to obtain the first segmented image, and a training set of the target model includes an image and a segmented image of the image.

5. The method according to any one of claims 1 to 4, characterized in that The step of encoding the mapping information and the segmented coded image into the bitstream comprises: The mapping information and the segmented coded images are encoded into the extended information or supplementary enhancement information of the code stream.

6. The method according to any one of claims 1 to 5, characterized in that The mapping information further includes fourth information and fifth information, wherein the fourth information is used to indicate a width of the first segmented image, and the fifth information is used to indicate a height of the first segmented image.

7. A decoding method, characterized in that: The method comprises: Decoding a bitstream to obtain segmented coded images and mapping information, wherein the segmented coded images are obtained based on the first segmented decoded image, and the mapping information includes first information, second information, and third information, wherein the first information is used to indicate the number of pixel point types in the first segmented image, the second information is used to indicate pixel values ​​corresponding to the pixel point types in the first segmented image, and the third information is used to indicate mapped pixel values ​​corresponding to the pixel point types in the first segmented image; Decoding the segmented coded image to obtain the first segmented decoded image; The first segmented decoded image is de-mapped according to the mapping information to obtain a second segmented decoded image.

8. The method according to claim 7, characterized in that The performing inverse mapping processing on the first segmented decoded image according to the mapping information to obtain a second segmented decoded image includes: The pixel values ​​of the pixels in the second segmented decoded image are determined according to the pixel values ​​of the pixels in the first segmented decoded image and the mapping information.

9. The method according to claim 8, characterized in that The determining the pixel value of the pixel point in the second segmented decoded image according to the pixel value of the pixel point in the first segmented decoded image and the mapping information includes: Determine a mapping pixel value of a pixel point in the first segmented decoded image according to the pixel value of the pixel point in the first segmented decoded image and the mapping information; The pixel value of a pixel point in the second segmented decoded image is determined according to the mapped pixel value and the mapping information.

10. The method according to any one of claims 7 to 9, characterized in that The mapping information further includes fourth information and fifth information, wherein the fourth information is used to indicate a width of the first segmented image, and the fifth information is used to indicate a height of the first segmented image.

11. A coding device, characterized in that: include: Coding unit, segmentation unit and processing unit; A coding unit, used for coding the image to be coded to obtain a code stream; A segmentation unit, configured to perform image segmentation on the image to be encoded to obtain a first segmented image; a processing unit, configured to perform mapping processing on the first segmented image to obtain a second segmented image; The processing unit is further used to determine mapping information according to the first segmented image, the mapping information including first information, second information and third information, the first information is used to indicate the number of pixel types in the first segmented image, the second information is used to indicate pixel values ​​corresponding to the pixel types in the first segmented image, and the third information is used to indicate mapped pixel values ​​corresponding to the pixel types in the first segmented image; The encoding unit is further used to encode the second segmented image to obtain a segmented encoded image; The encoding unit is further used to encode the segmented encoded image and the mapping information and image into the code stream.

12. The device according to claim 11, characterized in that The processing unit is specifically used for: determining the first information and the second information according to the first segmented image; The third information is determined according to the second information.

13. The device according to claim 12, characterized in that The processing unit is specifically used for: The pixel values ​​in the second information are shifted to obtain the mapped pixel values ​​in the third information.

14. The device according to any one of claims 11 to 13, characterized in that The segmentation unit is specifically used for: The image to be encoded is input into a target model to obtain the first segmented image, and a training set of the target model includes an image and a segmented image of the image.

15. The device according to any one of claims 11 to 14, characterized in that The encoding unit is specifically used for: The mapping information and the segmented coded images are encoded into the extended information or supplementary enhancement information of the code stream.

16. The device according to any one of claims 11 to 15, characterized in that The mapping information further includes fourth information and fifth information, wherein the fourth information is used to indicate a width of the first segmented image, and the fifth information is used to indicate a height of the first segmented image.

17. A decoding device, characterized in that: include: Decoding unit and processing unit; a decoding unit, configured to decode a bitstream to obtain segmented coded images and mapping information, wherein the segmented coded images are obtained based on the first segmented decoded images, and the mapping information includes first information, second information, and third information, wherein the first information is used to indicate the number of pixel point types in the first segmented image, the second information is used to indicate pixel values ​​corresponding to the pixel point types in the first segmented image, and the third information is used to indicate mapped pixel values ​​corresponding to the pixel point types in the first segmented image; A decoding unit, further configured to decode the segmented coded image to obtain the first segmented decoded image; A processing unit is used to perform inverse mapping processing on the first segmented decoded image according to the mapping information to obtain a second segmented decoded image.

18. The device according to claim 17, characterized in that The processing unit is specifically used for: The pixel values ​​of the pixels in the second segmented decoded image are determined according to the pixel values ​​of the pixels in the first segmented decoded image and the mapping information.

19. The device according to claim 18, characterized in that The processing unit is specifically used for: Determine a mapping pixel value of a pixel point in the first segmented decoded image according to the pixel value of the pixel point in the first segmented decoded image and the mapping information; The pixel value of a pixel point in the second segmented decoded image is determined according to the mapped pixel value and the mapping information.

20. The device according to any one of claims 17 to 19, characterized in that The mapping information further includes fourth information and fifth information, wherein the fourth information is used to indicate a width of the first segmented image, and the fifth information is used to indicate a height of the first segmented image.

21. A coding device, comprising at least one processor and a memory, characterized in that: The at least one processor executes a program or instruction stored in the memory so that the encoding device implements the method according to any one of claims 1 to 6.

22. A decoding device, comprising at least one processor and a memory, characterized in that: The at least one processor executes a program or instruction stored in the memory so that the encoding device implements the method according to any one of claims 7 to 10.

23. A computer-readable storage medium for storing a computer program, characterized in that: When the computer program is executed on a computer or a processor, the computer or the processor is enabled to implement the method according to any one of claims 1 to 10.

24. A computer program product, comprising instructions, characterized in that: When the instructions are executed on a computer or a processor, the computer or the processor is enabled to implement the method according to any one of claims 1 to 10.

25. A chip comprising at least one processor and a memory, characterized in that: The at least one processor executes a program or instruction stored in the memory, so that the chip implements the method according to any one of claims 1 to 10.