Image encoding and decoding methods and related devices

KR1020260139710APending Publication Date: 2026-09-22HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020267023927
Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2026-09-22

Smart Images

  • Figure PCT00043_ABST
    Figure PCT00043_ABST
Patent Text Reader

Abstract

The present application discloses an image encoding method applicable to the field of data processing technology. The image encoding method comprises the steps of: encoding a target image into a bitstream; and encoding a first identifier into the bitstream, wherein the first identifier represents an optical-electrical transfer function that matches the target image, and the target image includes a High Dynamic Range (HDR) image and / or a Standard Dynamic Range (SDR) image. That is, when the target image is encoded in the present application, the first identifier representing an optical-electrical transfer function that matches the target image may also be encoded into the bitstream to satisfy the display requirements of both SDR and HDR images at the decoder side, and as a result, after obtaining the first identifier through decoding, the decoder may render and display an image reconstructed according to the corresponding optical-electrical transfer function.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This application relates to the field of data processing technology, and in particular to image encoding and decoding methods and related devices. Background Technology

[0002] JPEG-AI is a machine learning-based image encoding and decoding standard. JPEG-AI-based encoding and decoding solutions are based on a self-encoder structure with color component separation, and the primary and secondary components of an image are encoded and decoded separately to implement image encoding and decoding.

[0003] The JPEG-AI-based image encoding process is as follows: an image in RGB format (a color space where R represents red, G represents green, and B represents blue) is converted into an image in YUV format (a color space where Y represents luminance, and U and V represent chroma) through color space conversion; then, downsampling is performed separately on the image with the Y component and the image with the UV component to reduce data throughput; then, the downsampled image with the Y component and the downsampled image with the UV component are encoded separately to obtain the bitstream of the Y component and the bitstream of the UV component. The JPEG-AI-based image decoding process is as follows: the bitstream of the Y component and the bitstream of the UV component are parsed separately to obtain the reconstructed image with the Y component and the reconstructed image with the UV component; then, upsampling is performed separately on the reconstructed image with the Y component and the reconstructed image with the UV component to obtain the reconstructed image in YUV format; Finally, the reconstructed image in YUV format is converted into a reconstructed image in RGB format through inverse color space conversion.

[0004] However, when an RGB format image is decoded and displayed, JPEG-AI performs processing by default according to an optical-electric transfer function based on a gamma function. The optical-electric transfer function [treats] standard dynamic range (SDR) images on conventional display devices (illumination is approximately 100 cd / m²). 2 It exhibits good performance in displaying on the image. However, in the case of high dynamic range (HDR) images, the photoelectric transfer function cannot be used to accurately display HDR images on the display device.

[0005] Embodiments of the present application provide an image encoding and decoding method and a related apparatus for accurately displaying HDR images.

[0006] According to a first embodiment, an image encoding method is provided comprising the steps of: encoding a target image into a bitstream; and encoding a first identifier into the bitstream, wherein the first identifier represents an optical-electric transfer function that matches the target image, and the target image includes a High Dynamic Range (HDR) image and / or a Standard Dynamic Range (SDR) image.

[0007] Optionally, the bitstream includes a picture header, and the first identifier is encoded in the picture header.

[0008] Optionally, the method further comprises the steps of: encoding a second identifier representing the chroma coordinates of each primary color of the target image into a bitstream; encoding a third identifier representing the pixel mapping range of the target image into a bitstream; and encoding a fourth identifier representing the coefficients of a transformation matrix used to perform a color space transformation on the target image into a bitstream.

[0009] Optionally, before encoding the target image into a bitstream, the method further includes the step of performing a color space conversion on the target image based on a second identifier, a third identifier, and a fourth identifier.

[0010] Optionally, the bitstream includes a picture header, and the second identifier, the third identifier, and the fourth identifier are encoded in the picture header.

[0011] Optionally, the method further includes the step of encoding a fifth identifier into a bitstream, wherein the fifth identifier represents the offset state of the sampling positions of the primary component of the target image and the secondary component of the target image.

[0012] Optionally, after color space conversion is performed on the target image, the target image comprises a primary component image and a secondary component image; before encoding the target image into a bitstream, the method further comprises the step of performing downsampling on the primary component image and the secondary component image based on a fifth identifier; and the step of encoding the target image into a bitstream comprises the step of encoding the downsampled image of the primary component and the downsampled image of the secondary component into the bitstream.

[0013] Optionally, the bitstream includes a picture header, and the fifth identifier is encoded in the picture header.

[0014] Optionally, the method further includes the step of encoding static metadata of a target image into a bitstream, wherein the static metadata includes source device information and image display information, the source device information represents color information of a mastering monitor generating the target image, and the image display information represents the display luma of the target image.

[0015] Optionally, source device information includes at least one of the chroma coordinates of the three primary colors of the mastering monitor, the chroma coordinates of the standard white light, the maximum display luma, and the minimum display luma.

[0016] Optionally, the image display information includes at least one of the maximum display luma or the maximum average display luma of the target image.

[0017] Optionally, the bitstream includes a picture header, and the picture header carries static metadata; or the bitstream includes a metadata header, and the metadata header carries static metadata; or the bitstream includes a metadata substream corresponding to the metadata header, and the metadata substream carries static metadata.

[0018] Optionally, the method further includes a step of encoding dynamic metadata of the target image into a bitstream, wherein the dynamic metadata includes color gamut volume change data of the target image.

[0019] Optionally, the bitstream includes a picture header, and the picture header includes dynamic metadata; or the bitstream includes a metadata header, and the metadata header includes dynamic metadata; or the bitstream includes a metadata substream corresponding to the metadata header, and the metadata substream includes dynamic metadata.

[0020] In conclusion, when a target image is encoded, a first identifier representing an optical-electrical transfer function matching the target image may also be encoded in the bitstream at the decoder side to satisfy the display requirements of both SDR images and HDR images, and as a result, the decoder can obtain the first identifier through decoding and then render and display an image reconstructed according to the corresponding optical-electrical transfer function.

[0021] According to a second aspect, an image decoding method is provided, comprising the steps of: parsing a bitstream to obtain a first identifier, wherein the first identifier represents an optical-electric transfer function that matches a target image, and the target image includes a High Dynamic Range (HDR) image and / or a Standard Dynamic Range (SDR) image; decoding the bitstream to obtain a reconstructed image of the target image; and performing an electro-optical conversion on the reconstructed image based on the first identifier.

[0022] Optionally, the bitstream includes a picture header, and the first identifier is included in the picture header.

[0023] Optionally, the method further comprises the steps of parsing a bitstream to obtain a second identifier representing the chroma coordinates of each primary color of a target image, parsing a bitstream to obtain a third identifier representing the mapping range of a target image, and parsing a bitstream to obtain a fourth identifier representing the coefficients of a transformation matrix used to perform color space transformation on a target image.

[0024] Optionally, before performing electro-optical conversion on an image reconstructed based on a first identifier, the method further includes the step of performing inverse color space conversion on an image reconstructed based on a second identifier, a third identifier, and a fourth identifier.

[0025] Optionally, the bitstream includes a picture header, and the second identifier, the third identifier, and the fourth identifier are included in the picture header.

[0026] Optionally, the method further includes the step of parsing a bitstream to obtain a fifth identifier, the fifth identifier representing the offset state of the sampling positions of the primary component of the target image and the secondary component of the target image.

[0027] Optionally, the step of parsing a bitstream to obtain a reconstructed image of a target image includes: a step of parsing a bitstream to obtain a reconstructed image of the primary component of the target image and a reconstructed image of the secondary component of the target image; and a step of performing upsampling on the reconstructed image of the primary component and the reconstructed image of the secondary component based on a fifth identifier to obtain a reconstructed image.

[0028] Optionally, the bitstream includes a picture header, and the fifth identifier is included in the picture header.

[0029] Optionally, the method further includes the step of parsing a bitstream to obtain static metadata of a target image, wherein the static metadata includes source device information and image display information, the source device information represents color information of a mastering monitor generating the target image, and the image display information represents the display luma of the target image.

[0030] Optionally, the method further includes the step of performing tone mapping on the reconstructed image based on static metadata.

[0031] Optionally, source device information includes at least one of the chroma coordinates of the three primary colors of the mastering monitor, the chroma coordinates of the standard white light, the maximum display luma, and the minimum display luma.

[0032] Optionally, the image display information includes at least one of the maximum display luma or the maximum average display luma of the target image.

[0033] Optionally, the bitstream includes a picture header, the picture header includes static metadata; or the bitstream includes a metadata header, the metadata header includes static metadata; or the bitstream includes a metadata substream corresponding to the metadata header, the metadata substream includes static metadata.

[0034] Optionally, the method further includes the step of parsing a bitstream to obtain dynamic metadata of a target image, wherein the dynamic metadata includes color gamut volume change data of the target image.

[0035] Optionally, the method further includes the step of performing tone mapping on the reconstructed image based on static metadata and dynamic metadata.

[0036] Optionally, the bitstream includes a picture header, and the picture header includes dynamic metadata; or the bitstream includes a metadata header, and the metadata header includes dynamic metadata; or the bitstream includes a metadata substream corresponding to the metadata header, and the metadata substream includes dynamic metadata.

[0037] In conclusion, when the bitstream is decoded, an optical-electric transfer function matching the target image can be determined based on a first identifier obtained through decoding, and accordingly, the reconstructed image of the target image is accurately rendered and displayed according to the optical-electric transfer function, thereby satisfying the display requirements for both SDR and HDR images on the decoder side and ensuring the accuracy of the target image display.

[0038] According to a third embodiment, an encoding device is provided. The encoding device has the function of implementing the operation of an image encoding method according to a first embodiment. The encoding device includes at least one module, and at least one module is configured to implement the image encoding method according to a first embodiment.

[0039] According to a fourth embodiment, a decoding device is provided. The decoding device has the function of implementing the operation of an image decoding method according to a second embodiment. The decoding device includes at least one module, and at least one module is configured to implement the image decoding method according to a second embodiment.

[0040] According to a fifth embodiment, an encoding device is provided. The encoding device includes a processor, the processor is coupled to memory, the memory is configured to store a program or instruction, and when the program or instruction is executed by the processor, the encoding device is able to perform an image encoding method according to a first embodiment.

[0041] According to a sixth embodiment, a decoding device is provided. The decoding device includes a processor, the processor is coupled to memory, the memory is configured to store a program or instruction, and when the program or instruction is executed by the processor, the decoding device is able to perform an image decoding method according to a second embodiment.

[0042] According to a seventh embodiment, an encoding and decoding system is provided. The encoding and decoding system includes an encoding device according to a fifth embodiment and / or a decoding device according to a sixth embodiment.

[0043] According to the eighth embodiment, a computer-readable storage medium containing program code is provided. When the program code is executed on a computer, the computer is able to perform the method according to the first embodiment.

[0044] According to the ninth embodiment, a computer-readable storage medium containing program code is provided. When the program code is executed on a computer, the computer is able to perform the method according to the second embodiment.

[0045] According to the tenth embodiment, a computer program product comprising instructions is provided. When the instructions are executed on a computer, the computer is able to perform the method according to the first embodiment.

[0046] According to the eleventh embodiment, a computer program product comprising instructions is provided. When the instructions are executed on a computer, the computer is able to perform the method according to the second embodiment.

[0047] According to the 12th embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a bitstream obtained according to the method according to the first embodiment.

[0048] According to the 13th embodiment, an apparatus for storing a bitstream is provided, comprising at least one storage medium and a communication interface. The communication interface is configured to receive or transmit the bitstream. At least one storage medium is configured to store the bitstream. The bitstream is obtained through encoding by an encoder according to the image encoding method according to the first embodiment.

[0049] According to the 14th embodiment, a bitstream storage method is provided, comprising the steps of: receiving a bitstream through a communication interface; and storing the bitstream in one or more storage media. The bitstream is obtained by an encoder through encoding according to the image encoding method according to the first embodiment.

[0050] According to the 15th embodiment, a bitstream distribution system comprising at least one storage medium and a video stream device is provided. At least one storage medium is configured to store a bitstream, and the bitstream is obtained by encoding by an encoder according to an image encoding method according to the first embodiment. The video stream device is configured to transmit a bitstream within the at least one storage medium to a decoder in response to a request from a decoder.

[0051] According to the 16th embodiment, a bitstream distribution method is provided, comprising the steps of: receiving a first request; selecting a bitstream from at least one storage medium in response to the first request; and transmitting the bitstream to a destination device. At least one storage medium is configured to store a bitstream, and the bitstream is obtained by an encoder through encoding according to the image encoding method described above according to the first embodiment.

[0052] According to the 17th embodiment, a bitstream processing system is provided comprising an image source device, an encoder, one or more storage media, and a destination device. The image source device is configured to provide image data. The encoder is configured to acquire image data from the image source device through an interface and to encode the image data to acquire one or more bitstreams, and the bitstreams are acquired by the encoder through encoding according to the image encoding method according to the first embodiment. The encoder is configured to store one or more bitstreams in one or more storage media. Alternatively, the encoder is configured to encapsulate one or more bitstreams to acquire a transmission bitstream. The encoder is configured to transmit the transmission bitstream to a destination device through a communication link or a communication network. The destination device is configured to decapsulate the transmission bitstream to acquire one or more bitstreams. The destination device is configured to decode one or more bitstreams to acquire decoded data.

[0053] The technical effects obtained in the second through seventeenth embodiments are similar to the technical effects obtained through the corresponding technical means in the first embodiment. Further details are not explained here. Brief explanation of the drawing

[0054] FIG. 1 is a drawing of an implementation environment according to one embodiment of the present application. FIG. 2 is a drawing of an AI encoder model according to one embodiment of the present application. FIG. 3 is a drawing of an AI decoder model according to one embodiment of the present application. FIG. 4 is a drawing of a JPEG-AI encoder model according to one embodiment of the present application. FIG. 5 is a drawing of a JPEG-AI decoder model according to one embodiment of the present application. FIG. 6 is a schematic flowchart of an image encoding method according to one embodiment of the present application. FIG. 7 is a diagram of a bitstream structure according to one embodiment of the present application. FIG. 8 is a diagram of a chroma sampling location according to one embodiment of the present application. FIG. 9 is a diagram of another bitstream structure according to one embodiment of the present application. FIG. 10 is a schematic flowchart of an image decoding method according to one embodiment of the present application. FIG. 11 is a diagram of the structure of an encoding device according to one embodiment of the present application. FIG. 12 is a diagram of the structure of a decoding device according to one embodiment of the present application. Specific details for implementing the invention

[0055] To further clarify the purpose, technical solution, and advantages of the present application, an implementation of the present application will be described in detail below with reference to the accompanying drawings.

[0056] To aid understanding, before describing in detail the image encoding method and image decoding method provided in the embodiments of the present application, the terms, implementation environment, and application scenarios in the embodiments of the present application are first described.

[0057] First, nouns in the embodiments of the present application are explained.

[0058] 1. High Dynamic Range (HDR)

[0059] Dynamic range represents the ratio of the maximum value to the minimum value of a variable in many fields. In the case of digital images, dynamic range represents the ratio of the maximum grayscale value to the minimum grayscale value within the range in which the image can be displayed. Currently, in most color digital images, each of the R, G, and B channels uses 1 byte, or 8 bits, for storage. Specifically, when the representation range of each channel is grayscale levels from 0 to 255, the dynamic range of the image is 0 to 255. In the real world, the dynamic range of a scene is 10 -3 to 10 6 When it falls within the range, that dynamic range is called High Dynamic Range (HDR). In contrast, the dynamic range of a typical image is Low Dynamic Range (LDR). The imaging process of a digital camera is actually a mapping from the High Dynamic Range of the real world to the Low Dynamic Range of the photograph.

[0060] 2. Standard dynamic range opto-electric transfer function

[0061] Standard dynamic range images correspond to high dynamic range images. Conventional 8-bit images in formats such as JPEG can be considered standard dynamic range images. Before the advent of cameras capable of capturing HDR images, conventional cameras could record light information captured only within a specific range by controlling exposure values. The maximum illuminance information of a display device cannot reach the luminance information of the real world. Furthermore, display devices are used for browsing images. Therefore, an optical-electrical transfer function is required. In early displays, the optical-electrical transfer function of a CRT display is the gamma function. The BT.1886 standard published by the International Telecommunication Union-Radiocommunication Sector (ITU-R) defines the optical-electrical transfer function based on the gamma function. An image obtained after tone mapping is performed according to the transfer function, and subsequently quantized into an 8-bit image, is a conventional SDR image. In other words, the SDR image and the aforementioned optical-electrical transfer function correspond to conventional display devices (illuminance is approximately 100 cd / m²). 2 It is performed well in the case.

[0062] 3. HDR Photoelectric Transfer Function

[0063] However, with the upgrade of display devices, the illuminance range of the display devices continuously increases. The illuminance information for existing consumer-grade HDR displays is 600 cd / m². 2 And, the illumination information for high-end HDR displays is 2000 cd / m² 2It can reach [value], which is much larger than the illuminance information of an SDR display device. The optical-electronic transfer function of the BT.1886 standard cannot adequately represent the display performance of an HDR display device. Therefore, an improved optical-electronic transfer function is required to adapt to upgrades of display devices, and the idea for the optical-electronic transfer function is derived from the mapping function of a tone mapping (TM) algorithm. The mapping function is adjusted as the optical-electronic transfer function (OETF).

[0064] Currently, there are three common optoelectric transfer functions: the perception quantization (PQ) optoelectric transfer function, the hybrid log-gamma (HLG) optoelectric transfer function, and the scene luminance fidelity (SLF) optoelectric transfer function. These three optoelectric transfer functions are the transfer functions specified in the audio video coding standard (AVS).

[0065] Unlike conventional gamma transfer functions, the PQ optical-electrical transfer function is a cognitive quantization transfer function proposed based on the human eye's luminance perception model. The PQ optical-electrical transfer function represents the transformation relationship between the linear signal value of an image pixel and the non-linear signal value in the PQ domain. The HLG optical-electrical transfer function is obtained by improving the conventional gamma curve. The HLG optical-electrical transfer function uses the conventional gamma curve in the low segment and supplements the logarithmic curve in the high segment. The HLG optical-electrical transfer function represents the transformation relationship between the linear signal value of an image pixel and the non-linear signal value in the HLG domain. Assuming that the SLF optical-electrical transfer function satisfies the optical characteristics of the human eye, the SLF optical-electrical transfer curve represents the transformation relationship between the linear signal value of an image pixel and the non-linear signal value in the SLF domain.

[0066] 4. Dynamic range mapping when displaying

[0067] Dynamic range mapping methods are primarily applied to the adaptation between a front-end HDR signal and a back-end HDR terminal display device. For example, the front-end collects a 4,000-nit optical signal, but the back-end HDR terminal display device (e.g., a TV set) only has an HDR display capability of 500 nits. Therefore, the method of mapping the 4,000-nit signal to a 500-nit device is a high-to-low tone mapping process. In another example, the front-end captures a 100-nit SDR signal, and the display end can display a 2,000-nit television signal. Properly displaying the 100-nit signal on a 2,000-nit device is another low-to-high tone mapping process.

[0068] Dynamic range mapping methods can be classified into static mapping and dynamic mapping. In static mapping methods, a single data point is used to perform the entire tone mapping process based on the same video content or the same hard disk content; that is, there is usually a uniform processing curve. This method has the advantage of conveying less information and having a simpler processing procedure, but it has the disadvantage that information may be lost in some scenes because the same curve is used for tone mapping in every scene. For example, if the curve focuses on protecting bright areas, some details may be lost or even invisible in some extremely dark scenes. Consequently, the user experience may be affected. In dynamic mapping methods, dynamic adjustments are performed based on specific areas of the content, each scene, or each frame. This method has the advantage of better processing results because different curve processing is performed based on specific areas, each scene, or each frame. In this way, while the processing results are better, each frame or each scenario must convey relevant scenario information, and a large amount of information is conveyed.

[0069] 5. ITU-T T.35 Terminal Identifier Assignment

[0070] ITU-T T.35 (abbreviated as T.35 in this application) is the T.35 specification published by ITU-T, that is, a procedure for assigning non-standard extended codes defined by the ITU. ITU-T T.35 consists of three parts: Country code, Terminal provider code, and Terminal provider oriented code. Country codes are defined in Annex A and B of ITU-T T.35, and Terminal provider codes are assigned by the ITU governing body. Each terminal provider to which a provider code has been assigned by the governing body manages the Terminal provider oriented code.

[0071] T.35 can be used in various contexts, such as SEI messages in video codecs. Data transmitted via T.35 messages may include HDR dynamic metadata, HDR static metadata, HDR display mapping information, etc. However, the information transmitted via T.35 is not limited to video and may be information of an arbitrary nature.

[0072] 6. Metadata

[0073] Key information about the recorded video, scene, or image within a frame, such as the average, maximum, and minimum values ​​within the scene. HDR static metadata, also known as fixed metadata, is used to control the color and detail of each frame of an image within the entire video based on the same metadata. Conversely, HDR dynamic metadata, also known as variable metadata, is used to control the color and detail of each frame of an image within the video based on different metadata.

[0074] Next, the implementation environment in the embodiment of the present application is described.

[0075] FIG. 1 is a diagram of an implementation environment according to an embodiment of the present application. The implementation environment includes a source device (10), a destination device (20), a link (30), and a storage device (40). The source device (10) can generate an encoded image. Accordingly, the source device (10) may also be referred to as an image encoding device or an encoder side. The destination device (20) can decode the encoded image generated by the source device (10). Accordingly, the destination device (20) may also be referred to as an image decoding device or a decoder side. The link (30) can receive the encoded image generated by the source device (10) and transmit the encoded image to the destination device (20). The storage device (40) can receive the encoded image generated by the source device (10) and can store the encoded image. In this case, the destination device (20) can directly obtain the encoded image from the storage device (40). Alternatively, the storage device (40) may correspond to a file server or other intermediate storage device capable of storing the encoded image generated by the source device (10). In this case, the destination device (20) may transmit or download the encoded image stored in the storage device (40) via streaming.

[0076] The source device (10) and the destination device (20) may each include one or more processors and memory coupled to one or more processors. The memory may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, any other medium accessible to a computer that can be used to store program code required in the form of instructions or data structures. For example, the source device (10) and the destination device (20) may each include a mobile phone, a smartphone, a personal digital assistant (PDA), a wearable device, a palmtop computer PPC (Pocket PC), a tablet computer, a smart in-vehicle infotainment system, a smart television, a smart sound box, a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a handheld phone such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video game console, a vehicle-mounted computer, etc.

[0077] The link (30) may include one or more media or devices capable of transmitting an encoded image from a source device (10) to a destination device (20). In a possible implementation, the link (30) may include one or more communication media capable of enabling the source device (10) to directly transmit the encoded image to the destination device (20) in real time. In this embodiment of the application, the source device (10) may modulate the encoded image based on a communication standard. The communication standard may be a wireless communication protocol, etc., and may transmit the modulated image to the destination device (20). One or more communication media may include a wireless communication medium and / or a wired communication medium. For example, one or more communication media may include a radio frequency (RF) spectrum or one or more physical transmission lines. One or more communication media may be part of a packet-based network. The packet-based network may be a local area network, a wide area network, a global network (e.g., the Internet), etc. One or more communication media may include a router, a switch, a base station, other devices that facilitate communication from a source device (10) to a destination device (20), etc. This is not specifically limited to the embodiments of the present application.

[0078] In a possible implementation, the storage device (40) may store a received encoded image transmitted by the source device (10), and the destination device (20) may directly obtain the encoded image from the storage device (40). In this case, the storage device (40) may include any one of a plurality of distributed or local access data storage media. For example, any one of the plurality of distributed or local access data storage media may be a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded image.

[0079] In a possible implementation, the storage device (40) may correspond to a file server or other intermediate storage device capable of storing an encoded image generated by the source device (10), and the destination device (20) may transmit or download the image stored in the storage device (40) via streaming. The file server may be any type of server capable of storing the encoded image and transmitting the encoded image to the destination device (20). In a possible implementation, the file server may include a network server, a file transfer protocol (FTP) server, a network attached storage (NAS) device, a local disk drive, etc. The destination device (20) may obtain the encoded image through any standard data connection (including an internet connection). Any standard data connection may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL) or a cable modem), or a combination of a wireless channel and a wired connection suitable for obtaining the encoded image stored in the file server. The transmission of the encoded image from the storage device (40) may be a streaming transmission, a download transmission, or a combination thereof.

[0080] The implementation environment illustrated in FIG. 1 is merely a possible implementation. Furthermore, the technology of the embodiments of the present application is applicable not only to the source device (10) capable of encoding the image illustrated in FIG. 1 and the destination device (20) capable of decoding the encoded image, but also to other devices capable of encoding the image and other devices capable of decoding the encoded image. This is not specifically limited to the embodiments of the present application.

[0081] In the implementation environment illustrated in FIG. 1, the source device (10) includes a data source (120), an encoder (100), and an output interface (140). In some embodiments, the output interface (140) may include a modulator / demodulator (modem) and / or a transmitter. The transmitter may also be referred to as an emitter. The data source (120) may include an image capture device (e.g., a camera), an archive containing previously captured images, a feed-in interface for receiving images from an image content provider, and / or a computer graphics system for generating images, or a combination of these image sources.

[0082] A data source (120) can transmit an image to an encoder (100), and the encoder (100) can obtain an encoded image by encoding the received image transmitted by the data source (120). The encoder can transmit the encoded image to an output interface. In some embodiments, the source device (10) transmits the encoded image directly to the destination device (20) through the output interface (140). In other embodiments, the encoded image may alternatively be stored in a storage device (40), so that the destination device (20) subsequently obtains the encoded image for decoding and / or display.

[0083] In the implementation environment illustrated in FIG. 1, the destination device (20) includes an input interface (240), a decoder (200), and a display device (220). In some embodiments, the input interface (240) includes a receiver and / or a modem. The input interface (240) may receive an encoded image from a link (30) and / or a storage device (40) and then transmit the encoded image to the decoder (200). The decoder (200) may decode the received encoded image to obtain a decoded image. The decoder may transmit the decoded image to the display device (220). The display device (220) may be integrated with the destination device (20) or placed outside the destination device (20). Typically, the display device (220) displays the decoded image. The display device (220) may be any one of a plurality of types of display devices. For example, the display device (220) may be a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0084] Although not illustrated in FIG. 1, in some embodiments, the encoder (100) and the decoder (200) may be integrated with the encoder and the decoder, respectively, and may include a suitable multiplexer-demultiplexer (MUX-DEMUX) unit or other hardware and software to encode both audio and video in a common data stream or separate data streams. In some embodiments, where applicable, the MUX-DEMUX unit may comply with other protocols such as the ITU H.223 multiplexer protocol or the user datagram protocol (UDP).

[0085] Each of the encoder (100) and the decoder (200) may be any one of the following circuits: one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. Where the technology in the embodiments of the present application is partially implemented in software, the device may store instructions for the software in a suitable non-volatile computer-readable storage medium and execute instructions in hardware through one or more processors to implement the technology in the embodiments of the present application. Any of the aforementioned contents (including hardware, software, a combination of hardware and software, etc.) may be considered as one or more processors. Each of the encoder (100) and the decoder (200) may be included in one or more encoders or decoders. An encoder or decoder may be integrated as part of a combined encoder / decoder (codec) within a corresponding device.

[0086] In an embodiment of the present application, the encoder (100) may generally be referred to as “signaling” or “transmitting” some information to another device, e.g., a decoder (200). The terms “signaling” or “transmitting” may generally refer to the transmission of syntax elements and / or other data used to decode a compressed image. Such transmission may occur in real time or near real time. Alternatively, such communication may occur after a period of time, for example, when the syntax elements of the bitstream encoded during encoding are stored in a computer-readable storage medium. The decoding device may then retrieve the syntax elements at any time after the syntax elements have been stored in the medium.

[0087] It should be noted that the application scenarios and implementation environments described in the embodiments of this application are intended to more clearly explain the technical solutions in the embodiments of this application, but do not constitute limitations on the technical solutions provided in the embodiments of this application. A person skilled in the art will understand that, as application scenarios and implementation environments evolve, the technical solutions provided in the embodiments of this application may also be applicable to similar technical problems.

[0088] Finally, an application scenario in an embodiment of the present application is described.

[0089] Encoding and decoding technologies are techniques that represent original signal data in a lossy or lossless manner using a small number of bits, based on data characteristics such as spatial, visual, and statistical redundancy, and enable the effective transmission and storage of information such as images, videos, and audio. This technology plays a crucial role in the current media era, where the types of information and the volume of data being transmitted or stored are increasing.

[0090] Image compression is used as an example. Image compression includes lossy and lossless compression. Lossy compression achieves a high compression ratio at the cost of some degradation in image quality, while lossless compression does not result in the loss of image detail. Conventional image / video compression algorithms have been developed over decades, leading to the formation of mature compression standards such as High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC). Conventional image / video compression algorithms are widely used due to their advantages, such as comprehensively supported features, excellent versatility, and good hardware support. However, there is a bottleneck in improving encoding efficiency.

[0091] As deep learning outperforms conventional algorithms in multiple computer vision tasks, such as image recognition and target detection, more researchers are beginning to explore deep learning-based image / video compression methods. Unlike conventional image algorithms, where image encoding, quantization, and entropy encoding are individually optimized through manual design, all modules of AI image compression algorithms (encoder networks, entropy estimation networks, decoder networks, etc.) are optimized collectively. Consequently, AI image compression solutions offer superior compression performance. However, deep learning is a resource-intensive algorithm that not only requires high computing costs but also consumes a significant amount of memory. Despite the increasing availability of computing resources, optimizing the training and inference of deep learning networks for model implementation remains critical. In particular, an increasing number of models are shifting from the server side to the edge, to resource-constrained devices such as smartphones and embedded devices. Deploying complex models to resource-constrained devices is a challenge that must be addressed using current deep learning technologies.

[0092] Deep learning-based image compression methods are typically associated with AI compression models and JPEG-AI compression models. For better understanding, the encoding and decoding processes of the two compression models are briefly described here.

[0093] (1) AI compression model

[0094] Figure 2 is a diagram of an AI encoder model. The AI ​​encoder model is also called an encoder and is used on the encoder side. During encoding, an image (x) to be encoded is input into an encoder network to obtain a feature map (y), and the feature map (y) is input into a hyper-encoder network to obtain hyper-previous features (z). Entropy encoding is performed based on a specified probability distribution for the hyper-previous features (z) to encode the hyper-previous features (z) into a bitstream. Additionally, the hyper-previous features (z) are input into a hyper-decoder network to obtain a previous feature (p) and a variance (σ). Based on the previous feature (p) and the feature map (y), an mean value (μ) is obtained through a context network, and then entropy encoding is performed based on the probability distribution N(μ, σ) for the feature map (y) to encode the feature map (y) into a bitstream.

[0095] The bit sequence obtained by performing entropy encoding on the hyper-predicted feature (z) is a partial bit sequence included in the bitstream. The partial bit sequence can be referred to as the hyper-predicted bitstream, and bs z It is indicated as. The bit sequence obtained by performing entropy encoding on the feature map (y) is a partial bit sequence included in the bitstream. The partial bit sequence can be referred to as an image bitstream, and bs y It is displayed as.

[0096] In some embodiments, the context network includes a context model and a probability distribution estimation model. In this case, a feature map (y) is input into the context model to form context features ( ) can be obtained. Prior features (p) and context features ( Based on ), the mean value (μ) is estimated based on a probability distribution estimation model. Of course, this is merely an example. In practice, implementation may be performed in alternative ways. This is not limited to the embodiments of this application. Furthermore, the foregoing provides an explanation using entropy encoding as an example. In practice, encoding may be performed in alternatively different encoding methods. This is also not limited to the embodiments of this application.

[0097] Figure 3 is a diagram of an AI decoder model. The AI ​​decoder model is also called a decoder and is used on the decoder side. During decoding, the hyper-predicted bitstream (bs) contained in the bitstream z Entropy decoding is performed based on a specified probability distribution for ) to hyper-prior features ( Acquire ) and hyper pre-features( ) is input into a hyperdecoder network and pre-features( ) and variance (σ) are obtained. Additional information is determined from the decoded partial image features, and prior features ( The mean value (μ) is obtained through a context network based on ) and additional information. Based on the probability distribution N(μ, σ), the image bitstream (bs) included in the bitstream y feature map( ) is obtained. Feature map( ) is input into the decoder network and reconstructed image( Acquires ).

[0098] In some embodiments, the context network includes a context model and a probability distribution estimation model. In this case, prior features ( Based on ), the mean value (μ) corresponding to the partial image feature can be estimated based on a probability distribution estimation model, and the partial image feature is an image bitstream (bs) included in the bitstream based on the probability distribution N(μ, σ) of the partial image feature.y It is obtained through entropy decoding from ). For other image features, additional information is determined from the decoded image features, and this additional information is input into the context model to form context features ( Acquires ), and prior features( ) and context features( Based on a probability distribution estimation model, the mean value (μ) corresponding to other image features is estimated by referring to ). Of course, this is merely an example. In practice, implementation may be performed in alternative ways. This is not limited to the embodiments of this application. Furthermore, the foregoing is explained using entropy decoding as an example. In practice, decoding may be performed in other decoding methods corresponding to the encoding method. This is also not limited to the embodiments of this application.

[0099] It should be noted that if a probability distribution estimation network is modeled using a Gaussian model (e.g., a single Gaussian model or a hybrid Gaussian model), the estimated probability distribution includes the mean and variance. If a probability distribution estimation network is modeled using a Laplacian distribution model, the estimated probability distribution includes location parameters and scale parameters. If a probability distribution estimation network is modeled using a logistic distribution model, the estimated probability distribution includes the mean and scale parameters. The foregoing is explained using a Gaussian model as an example.

[0100] Additionally, the hyperdecoder network includes a hyperdecoder prediction network and a hyperscale decoder network. The hyperdecoder prediction network is configured to determine prior features, and the hyperscale decoder network is configured to determine variance. Specifically, hyper-prior features are input into the hyperdecoder prediction network to obtain prior features, and hyper-prior features are input into the hyperscale decoder network to obtain variance.

[0101] (2) JPEG-AI compression model

[0102] The JPEG-AI compression model implements image encoding and decoding by encoding and decoding each separated component of the image to be encoded based on an autoencoder structure in which color components are separated. Additionally, the basic structure of the JPEG-AI compression model is identical to the structure of the AI ​​compression model described above. However, the image input components of the JPEG-AI compression model are YUV components, and encoding and decoding operations are performed individually for the Y component and the UV component. During encoding, the Y component is individually encoded using the architecture shown in FIG. 2, and the UV component is encoded based on the data of the Y component. During decoding, the Y component is individually decoded using the architecture shown in FIG. 3, and the UV component is decoded based on the data of the Y component.

[0103] Figure 4 is a diagram of the JPEG-AI encoder model. The JPEG-AI encoder model is also referred to as an encoder and is used on the encoder side. During encoding, the image (x) to be encoded in RGB (color mode, where R represents red, G represents green, and B represents blue) format is converted into an image in YUV (color encoding method, where Y represents luminance, i.e., grayscale value, and U and V represent chroma) format through color space conversion, and the image of the Y component (x Y ) and UV component image(x UV Obtain ), and the image of the Y component (x Y ) is input into the encoder network and the feature map of the Y component (y Y Obtain ), and the feature map of the Y component (y Y ) is input into the hyperencoder network and the hyper-pre-features of the Y component (z Y Obtaining ), and based on a specified probability distribution, hyper-prior features of the Y component (z YEntropy encoding is performed on ) to obtain the hyper-pre-features of the Y component (z Y Encodes ) into a bitstream. Hyper-predicted features of the Y component (z Y ) is input into the hyperdecoder network and the prior features of the Y component (p Y ) and variance (σ Y ) obtain, and prior features (p Y ) and variance (σ Y Based on ), the mean value (μ) through the context network Y After obtaining ), the probability distribution N(μ of the Y component Y , σ Y Feature map of the Y component (y based on ) Y Entropy encoding is performed on ) to obtain the feature map of the Y component (y Y Encodes ) into a bitstream.

[0104] Also, the image of the Y component (x Y ) is an auxiliary image of the next downsampled UV component ( Converted to ), and UV component image (x UV ) and auxiliary image of UV component( ) is input into the encoder network and the feature map of the UV components (y UV ) obtain and UV component feature map(y UV ) is input into the hyperencoder network and the hyperpre-features of the UV components (z UV Acquires ), and based on a specified probability distribution, hyper-pre-features of UV components (z UV Entropy encoding is performed on ) to obtain hyper-pre-features of UV components (z UV Encodes ) into a bitstream. Hyper-predicted features of UV components (z UV ) is input into a hyperdecoder network and the prior features of the UV component (p UV ) and variance (σ UV ) obtained, and prior features of UV components (p UV ) and feature map(y UVBased on ), the average value of the UV component (μ) through the context network UV After ) is obtained, the probability distribution N(μ of the UV component UV , σ UV Feature map of UV components (y based on ) UV Entropy encoding is performed on ) to obtain the feature map of the UV components (y UV Encodes ) into a bitstream.

[0105] Hyper-pre-features of the Y component (z Y The bit sequence obtained by performing entropy encoding on ) is part of the bit sequence included in the bitstream, and the part of the bit sequence is the hyper-predicted bitstream of the Y component (bs zY It can be referred to as ). Feature map of the Y component (y Y The bit sequence obtained by performing entropy encoding on ) is part of the bit sequence included in the bitstream, and part of the bit sequence is the image bitstream of the Y component (bs yY It can be referred to as ). Hyper-predicate characteristics of UV components (z UV The bit sequence obtained by performing entropy encoding on ) is part of the bit sequence included in the bitstream, and part of the bit sequence is a hyper-predicted bitstream of UV components (bs zUV It can be referred to as a feature map of the UV component (y UV The bit sequence obtained by performing entropy encoding on ) is part of the bit sequence included in the bitstream, and part of the bit sequence is the image bitstream of the UV component (bs yUV It can be referred to as ).

[0106] Figure 5 is a diagram of the JPEG-AI decoder model. The JPEG-AI decoder model is also called a decoder and is used on the decoder side. During decoding, the hyper-pre-bitstream (bs) of the Y component contained in the bitstream. zY Entropy decoding is performed based on a specified probability distribution for ) to the hyper-prior features of the Y component ( ) obtain and hyper-pre-features of the Y component ( ) is input into the hyperdecoder network and the prior features of the Y component ( ) and variance (σ Y ) is obtained. Additional information is determined from the decoded partial image features, and the prior features of the Y component ( Based on ) and additional information, the average value of the Y component (μ) through the context network Y ) is obtained. The probability distribution N(μ of the Y component Y , σ Y Based on ), a feature map of the Y component ( ) is obtained. Feature map of the Y component( ) is input into the decoder network and the reconstructed image of the Y component ( Acquires ).

[0107] Also, the feature map of the Y component (y Y ) is an auxiliary feature of the next downsampled UV component ( It is converted into ). Then, based on a specified probability distribution, the hyper-pre-bitstream of the UV components included in the bitstream (bs zUV Hyper-pre-features of UV components by performing entropy decoding on ) ) obtained, and hyper pre-characteristics of UV components ( This is input into a hyperdecoder network, and the pre-features of the UV component ( ) and variance (σ UV ) is obtained. Additional information is determined from the decoded partial image features, and the prior features of the UV components ( The average value of UV components through a context network based on ) and additional information ( ) is obtained. Probability distribution N( of the UV component , Based on ), a feature map of UV components ( ) is obtained. UV component feature map( ) and auxiliary features of UV components( ) is input into the decoder network and the reconstructed image of the UV components ( Acquires ).

[0108] Then, the reconstructed image of the Y component ( ) and reconstructed image of UV components( ) is processed through an Inter Channel Correlation Information (ICCI) filtering network to obtain a reconstructed image in YUV format. Finally, the format of the reconstructed image is converted from YUV format to RGB format through inverse color space conversion to obtain the reconstructed image ( Acquires ).

[0109] The foregoing describes the general encoding and decoding processes of the JPEG-AI compression model. For some details, please refer to the relevant descriptions of the AI ​​compression model mentioned above. Detailed information is not provided here.

[0110] In other words, before encoding an image, the JPEG-AI compression model must perform color space conversion on the target image to be encoded and perform downsampling preprocessing on the component images. After the reconstructed images of the target image components are obtained through decoding and reconstruction, upsampling is performed on the reconstructed images, and inverse color space conversion is performed on the reconstructed images. However, when the JPEG-AI decoder model performs inverse color space conversion, the coefficients of the transformation matrix used are applicable to SDR images. Therefore, an accurate HDR image cannot be reconstructed. Furthermore, when upsampling is performed on the reconstructed images of the components, the offset state of the sampling position of chroma information relative to luminance information differs from the offset state of downsampling performed on the decoder side; in this case, an accurate HDR image cannot be reconstructed. Additionally, regarding the reconstructed images, the display device on the decoder side typically performs tone mapping according to the opto-electric transfer function corresponding to SDR, and thus cannot accurately map and display the reconstructed HDR image.

[0111] Based on this, in order to implement encoding, decoding, transmission, and display of HDR images, the present application provides an image encoding method and an image decoding method that add image information and HDR metadata information that can be used to accurately reconstruct an HDR image and display an HDR image during encoding, thereby enabling the decoder side to reconstruct an accurate image based on the image information. Additionally, the HDR image can be effectively displayed based on the HDR metadata information.

[0112] As illustrated in FIG. 6, the present application provides an image encoding method. The method is applied to an encoder, and the method comprises the following steps.

[0113] Step 601: Encode the target image into a bitstream.

[0114] For a specific implementation of encoding a target image into a bitstream, refer to the relevant description of the aforementioned process for encoding and decoding an image using the JPEG-AI compression model. Details are not described again in this specification.

[0115] Step 602: Encode a first identifier into a bitstream, wherein the first identifier represents an optical-electric transfer function that matches a target image, and the target image includes an HDR image and / or an SDR image.

[0116] The optoelectric transfer function is used to accurately display an image on a display by converting a digital signal into an analog signal capable of driving the physical output of the display. When the target image is an HDR image, the first identifier represents an optoelectric transfer function corresponding to the HDR image, for example, a PQ optoelectric transfer function or an HLG optoelectric transfer function. When the target image is an SDR image, the first identifier represents an optoelectric transfer function corresponding to the SDR image, for example, an optoelectric transfer function based on a gamma function.

[0117] In some embodiments, when the first identifier is encoded into a bitstream, the first identifier may be encoded into the bitstream in the form of an 8-bit unsigned integer, that is, the first identifier corresponds to a value in the range of 0 to 255. Of course, the first identifier may be encoded into the bitstream in other ways. This is not limited to the embodiments of the present application.

[0118] When the first identifier is encoded into a bitstream in the form of an 8-bit unsigned integer, Table 1 provides an example of a correspondence, and based on this correspondence, the photoelectric transfer function represented by the first identifier is determined.

[0119] First identifier Photoelectric transfer function Reference standard file 0 Reserved For future use by ITU-T | ISO / IEC 1 V = α * L c 0.45 - ( α- 1 ), 1 >= L c If >= β, V = 4.500 * L c β > L c If >= 0 Rec. ITU-R BT.709-6 Rec. ITU-R BT.1361-0 Historical color gamut system (historical) (functionally equivalent to values ​​6, 14, and 15) 2 Unspecified Image characteristics are unknown or determined by the application. 3 Reserved For future use by ITU-T | ISO / IEC 4 Assumed display gamma 2.2 Rec. ITU-R BT.470-6 System M (Historical) National Television System Committee 1953 Color Television Transmission Standard Recommendations U.S. Federal Communications Commission (2003) Federal Regulations Part 47 73.682 (a) (20)Rec. ITU-R BT.1700-0 625 PAL and 625 SECAM 5 Assumed display gamma 2.8 Rec. ITU-R BT.470-6 System B, G (Historical) 6 V = α * L c 0.45 - ( α - 1 )1 >= L c If >= β, V = 4.500 * L c β > L c If >= 0 Rec. ITU-R BT.601-7 525 or 625 Rec. ITU-R BT.1358-1 525 or 625 (Historical) Rec. ITU-R BT.1700-0 NTSCSMPTE ST 170 (2004) (Functionally equivalent to values ​​1, 14, and 15) 7 V = α * L c 0.45 - ( α - 1 )1 >= L c If >= β, V = 4.0 * L c β > L c If >= 0 SMPTE ST 240 (1999) 8 V = L c, 1 > L c If >= 0 Linear transfer characteristics 9 V = 1.0 + Log10( L c ) ÷21 >= L c If >= 0.01, V = 0.00.01 > L c If >= 0 Log forwarding characteristics (100:1 range) 10 V = 1.0 + Log10( L c ) ÷2.5,1 >= L c For the case where >= Sqrt( 10 ) ÷ 1000; V = 0.0, Sqrt( 10 ) ÷ 1000 > L c If >= 0 Log forwarding properties (range of 100 * Sqrt( 10 ) : 1) 11 V = α * L c 0.45 - ( α - 1 )L c If >= β, V = 4.500 * L c β > L c In the case of -β, V = -α * ( -L c ) 0.45 + ( α - 1 )-β >= L c In the case of IEC 61966-2-4 12 V = α * L c 0.45 - ( α - 1 )1.33 > L c If >= β, V = 4.500 * L c β > L c If >= -γ, V = -( α * ( -4 * L c ) 0.45 - ( α - 1 ) ) ÷ 4 - γ >= L c If >= -0.25 Rec. ITU-R BT.1361-0 Extended Color Gamut System (Historical) 13 - If the Matrix Coefficients are equal to 0, then V = α * L c ( 1÷2.4 ) - ( α - 1 )1 > L c If >= β, V = 12.92 * L c β > L c If >= 0 Otherwise, V = α * L c ( 1÷2.4 ) - ( α - 1 )L c If >= β, V = 12.92 * L c β > L c In the case of -β, V = -α * ( -L c )( 1÷2.4) + ( α - 1 )-β >= L c In the case of IEC 61966-2-1 sRGB (Matrix Coefficients equal to 0) IEC 61966-2-1 sYCC (Matrix Coefficients equal to 5) 14 V = α * L c 0.45 - ( α - 1 )1 >= L c If >= β, V = 4.500 * L c β > L c If >= 0 Rec. ITU-R BT.2020-2 (10-bit system) (functionally equivalent to values ​​1, 6, and 15) 15 V = α * L c 0.45 - ( α - 1 )1 >= L c If >= β, V = 4.500 * L c β > L c If >= 0 Rec. ITU-R BT.2020-2 (12-bit system) (functionally equivalent to values ​​1, 6, and 14) 16 V = ( ( c1 + c2 * L o n ) ÷( 1 + c3* L o n ) ) m c1 = c3 - c2 + 1 = 107 ÷ 128 = 0.8359375 c2 = 2413 ÷ 128 = 18.8515625 c3 = 2392 ÷ 128 = 18.6875 m = 2523 ÷ 32 = 78.84375 n = 653 ÷ 4096 = 0.1593017578125 all L o Regarding the value SMPTE ST 2084 (2014)Rec. ITU-R BT.2100-2 Cognitive Quantization (PQ) Systems for 10, 12, 14, and 16-bit Systems 17 V = ( 48 * L o ÷5 2.37 )( 1 ÷2.6 all L o Regarding the value, for peak white, L o Case 1 is generally intended to correspond to a reference output luminance level of 48 candela per square meter. SMPTE ST 428-1 (2019) 18 V = a * Ln( 12 * L c - b ) + c1 >= L c For the case of 1 ÷ 12, V = Sqrt(3) * L c 0.5 1 ÷ 12 >= L c If >= 0, a = 0.17883277, b = 0.28466892, c = 0.55991073 ARIB STD-B67 (2015)Rec. ITU-R BT.2100-2 Hybrid Log-Gamma (HLG) System 19-255 Reserved For future use by ITU-T | ISO / IEC

[0120] From Table 1, it can be seen that when the target image is an SDR image, the first identifier of the target image may be 13, and the opto-electric transfer function corresponding to the first identifier is the opto-electric transfer function defined in the IEC 61966-2-1 standard. The opto-electric transfer function supports opto-electric conversion for the SDR image. When the target image is an HDR image, the first identifier encoded in the bitstream is 16 or 18, where 16 represents the PQ opto-electric transfer function and 18 represents the SLF opto-electric transfer function. Refer to the middle column of Table 1 for the specific function representation of the first identifier. L o is the linear signal value before transformation, and L o The value of is normalized to [0, 1]. L c is the linear signal value before transformation, and L c The value range of is [0, 12]. V is the nonlinear signal value after conversion, and the value range of V is [0, 1]. In addition, the corresponding opto-electric transfer function is also recorded and described in the relevant standard, and you may also refer to the standard file provided in the right column of Table 1.

[0121] It should be noted that Table 1 is merely an example. In specific applications, the opto-electric transfer function corresponding to the identifier value in the table may be adjusted based on relevant standards for image / video encoding and decoding. This is not limited to the embodiments of this application.

[0122] FIG. 7 is a diagram of a bitstream structure defined by JPEG-AI. The bitstream structure includes a start of codestream marker (SOC), a picture header marker (PIH) and a corresponding picture header, a tools header marker (TOH) and a corresponding (encoding and decoding) tool header, a start of quality map marker (SOQ) and a corresponding bitstream (Q-stream), a plurality of picture bitstreams, and an end of codestream marker (EOC).

[0123] Multiple image bitstreams include a codestream of hyper-priority features (Z-stream), a codestream of primary component residuals (RY-stream), and a codestream of secondary component residuals (RUV-stream). Each image bitstream corresponds to a start marker. The codestream of hyper-priority features follows the start of z-stream marker (SOZ), the codestream of primary component residuals follows the start of residual stream for primary component marker (SORP), and the codestream of secondary component residuals follows the start of residual stream for secondary component marker (SORS).

[0124] Based on this bitstream structure, the implementation process of step 602 may encode the first identifier into the picture header.

[0125] Regardless of whether the encoded target image is an SDR image or an HDR image, an optical-electric transfer function matching the target image can be determined, and a first identifier representing the optical-electric transfer function is also encoded in the bitstream during encoding, so that the decoder side can display a reconstructed image corresponding to the target image based on the first identifier and according to the matched optical-electric transfer function.

[0126] In some embodiments, before the encoder encodes the target image, if the target image is in RGB format, a color space conversion must be performed on the target image to convert the RGB format image into a YUV format image. However, JPEG-AI supports only two color space conversion modes: a preset BT.709 conversion matrix and a user-defined conversion matrix. BT.709 is the default color space for SDR images. However, in other ITU-R RGB-YUV conversion standards, such as ITU-R BT.2020 and ITU-R BT.2100, HDR television standards are specified, and the color space specified in the BT.2020 standard is inherited and used. Therefore, when encoding and decoding HDR images, the encoder and decoder must also support BT.2020 color space conversion (which corresponds to the BT.2100 color space).

[0127] Based on this, when encoding the target image, the encoder side may additionally encode information related to the color space conversion matrix used by the encoder side to perform color space conversion on the target image into the bitstream, so that the decoder side converts the reconstructed image in YUV format into a reconstructed image in RGB format based on the information related to the color space conversion matrix.

[0128] In a possible implementation, the image encoding method provided in this embodiment of the application further comprises: encoding a second identifier representing the chroma coordinates of each primary color of a target image into a bitstream; encoding a third identifier representing the pixel mapping range of a target image into a bitstream; and encoding a fourth identifier representing the coefficients of a transformation matrix used to perform a color space transformation on a target image into a bitstream.

[0129] In this case, before encoding the target image into a bitstream, the encoder performs a color space conversion on the target image based on the second identifier, the third identifier, and the fourth identifier.

[0130] In other words, the encoder performs a color space conversion on the target image using a corresponding color space conversion matrix based on the type of the target image (i.e., an HDR image or an SDR image), and then encodes the converted target image into a bitstream. Additionally, the encoder encodes the identifiers of the information associated with the color space conversion matrix—namely, the second identifier, the third identifier, and the fourth identifier—into the bitstream.

[0131] Similarly, if the bitstream includes a picture header, the second identifier, the third identifier, and the fourth identifier may be encoded in the picture header.

[0132] In some embodiments, a second identifier representing the chroma coordinates of each primary color of the target image may be encoded in the bitstream in the form of an 8-bit unsigned integer and corresponds to values ​​from 0 to 255. For example, the chroma of CIE 1931 specified in ISO / CIE 11664-1 includes x and y coordinates. Table 2 illustrates a method for determining corresponding chroma coordinates based on the second identifier. That is, Table 2 illustrates exemplary correspondences for finding chroma coordinates corresponding to red, green, blue, and standard white light based on the second identifier.

[0133] Second identifier Primary color chroma coordinates Reference standard file 0 Reserved For future use by ITU-T | ISO / IEC 1 primary color x y green 0.300 0.600 blue 0.150 0.060 red 0.640 0.330 White D65 0.3127 0.3290 Rec. ITU-R BT.709-6 Rec. ITU-R BT.1361-0 Historical Color Gamut Systems and Extended Color Gamut Systems (Historical) IEC 61966-2-1 sRGB or sYCC IEC 61966-2-4 SMPTE RP 177 (1993) Appendix B 2 Unspecified Image characteristics are unknown or determined by the application. 3 Reserved For future use by ITU-T | ISO / IEC 4 primary color x y green 0.21 0.71 blue 0.14 0.08 red 0.67 0.33 White C 0.310 0.316 Rec. ITU-R BT.470-6 System M (Historical) National Television System Committee 1953 Color Television Transmission Standard Recommendations U.S. Federal Communications Commission (2003) Federal Regulations Part 47 73.682 (a) (20) 5 primary color x y green 0.29 0.60 blue 0.15 0.06 red 0.64 0.33 White D65 0.3127 0.3290 Rec. ITU-R BT.470-6 System B, G (Historical) Rec. ITU-R BT.601-7 625 Rec. ITU-R BT.1358-0 625 (Historical) Rec. ITU-R BT.1700-0 625 PAL and 625 SECAM 6 primary color x y green 0.310 0.595 blue 0.155 0.070 red 0.630 0.340 White D65 0.3127 0.3290 Rec. ITU-R BT.601-7 525 Rec. ITU-R BT.1358-1 525 or 625 (historical) Rec. ITU-R BT.1700-0 NTSCSMPTE ST 170 (2004) (functionally equivalent to value 7) 7 primary color x y green 0.310 0.595 blue 0.155 0.070 red 0.630 0.340 White D65 0.3127 0.3290 SMPTE ST 240 (1999) (functionally equivalent to value 6) 8 primary color x y green 0.243 0.692 (Wratten 58) Blue 0.145 0.049 (Wratten 47) Red 0.681 0.319 (Wratten 25) White C 0.310 0.316 Standard film (color filter using light source C) 9 primary color x y green 0.170 0.797 blue 0.131 0.046 red 0.708 0.292 White D65 0.3127 0.3290 Rec. ITU-R BT.2020-2Rec. ITU-R BT.2100-2 10 primary color x y green (Y) 0.0 1.0 Blue (Z) 0.0 0.0 Red (X) 1.0 0.0 center white 1 ÷ 3 1 ÷ 3 SMPTE ST 428-1 (2019)(CIE 1931 XYZ as in ISO / CIE 11664-1) 11 primary color x y green 0.265 0.690 blue 0.150 0.060 red 0.680 0.320 white 0.314 0.351 SMPTE RP 431-2 (2011) 12 primary color x y green 0.265 0.690 blue 0.150 0.060 red 0.680 0.320 White D65 0.3127 0.3290 SMPTE EG 432-1 (2010) 13-21 Reserved For future use by ITU-T | ISO / IEC 22 primary color x y green 0.295 0.605 blue 0.155 0.077 red 0.630 0.340 White D65 0.3127 0.3290 Corresponding industrial standards not identified 23-255 Reserved For future use by ITU-T | ISO / IEC

[0134] From Table 2, it can be seen that when the target image is an HDR image, the second identifier encoded in the bitstream is 9, the green coordinate corresponding to chroma x is 0.170, and the green coordinate corresponding to chroma y is 0.797; the blue coordinate corresponding to chroma x is 0.131, and the blue coordinate corresponding to chroma y is 0.046; the red coordinate corresponding to chroma x is 0.708, and the red coordinate corresponding to chroma y is 0.292; the standard white light coordinate corresponding to chroma x is 0.3127, and the standard white light coordinate corresponding to chroma y is 0.3290. The aforementioned related information regarding HDR images and chroma primary colors is included in both the ITU-R BT.2020-2 standard and the ITU-R BT.2100-2 standard, and can also be found in the ITU-R BT.2020-2 standard and ITU-R BT.2100-2 standard files.

[0135] In some embodiments, a third identifier indicating a pixel mapping range of a target image may be 0 or 1. If the input signal is an RGB image, the third identifier is 1, indicating that full-range mapping is performed for all pixels of the image. If the input signal is a luminance and chroma image (e.g., a YUV format image), and the range of black levels and luminance and chroma signals is not specified, the third identifier may be 0, indicating that mapping is performed for all pixels of the image within a limited range.

[0136] Optionally, for images in YUV format, the identifier value of the third identifier may be marked as 1 to perform full-range mapping for all pixels of the image. This is not limited to the embodiments of the present application.

[0137] In some embodiments, when a fourth identifier representing the coefficients of a transformation matrix used to perform color space transformation on a target image is encoded in a bitstream, the fourth identifier may be encoded in the bitstream in the form of an 8-bit unsigned integer and corresponds to a value in the range of 0 to 255. For example, the chroma of CIE 1931 specified in ISO / CIE 11664-1 includes x and y coordinates. Table 3 below provides an example of a correspondence, and the fourth identifier and the coefficients of the transformation matrix are determined based on this correspondence.

[0138] Fourth identifier coefficients of the transformation matrix Reference standard file 0 unit Identity matrix. Typically used for GBR (commonly referred to as RGB), but may also be used for YZX (commonly referred to as XYZ); see IEC 61966-2-1 sRGBSMPTE ST 428-1 (2019) Equations 48 to 50. 1 K R = 0.2126; K B = 0.0722 Rec. ITU-R BT.709-6 Rec. ITU-R BT.1361-0 Historical Color Gamut Systems and Extended Color Gamut Systems (Historical) IEC 61966-2-4 xvYCC 709 Refer to SMPTE RP 177 (1993) Appendix B, Equations 45 through 47. 2 Unspecified Image characteristics are unknown or determined by the application 3 Reserved For future use by ITU-T | ISO / IEC 4 K R = 0.30; K B = 0.11 U.S. Federal Communications Commission (2003) Federal Regulations Part 47 73.682 (a) (20) See mathematical formulas 45 to 47 5 K R = 0.299; K B = 0.114 Rec. ITU-R BT.470-6 System B, G (Historical) Rec. ITU-R BT.601-7 625 Rec. ITU-R BT.1358-0 625 (Historical) Rec. ITU-R BT.1700-0 625 PAL and 625 SECAMIEC 61966-2-1 sYCCIEC 61966-2-4 xvYCC 601 (Functionally identical to value 6) Refer to mathematical formulas 45 through 47. 6 K R = 0.299; K B = 0.114 Rec. ITU-R BT.601-7 525 Rec. ITU-R BT.1358-1 525 or 625 (historical) Rec. ITU-R BT.1700-0 NTS CSMPT ST 170 (2004) (functionally equivalent to value 5) See Equations 45 through 47 7 K R = 0.212; K B = 0.087 Refer to SMPTE ST 240 (1999) Equations 45 through 47. 8 YCgCo or YCgCo-R For YCgCo, refer to Equations 51 through 57 (BitDepth C BitDepth Y (In the case identical to) For YCgCo-R, refer to Equations 58 to 65 (BitDepth C BitDepth Y (If identical to + 1) 9 K R = 0.2627; K B = 0.0593 Rec. ITU-R BT.2020-2 (Non-constant Luminance) Rec. ITU-R BT.2100-2 Y′ Refer to Equations 45 to 47 10 K R = 0.2627; K B = 0.0593 Rec. ITU-R BT.2020-2 (Constant Luminance) Refer to Equations 66 to 75 11 Y′′ Z D′ X Refer to SMPTE ST 2085 (2015) Equations 76 through 78. 12 Refer to mathematical formulas 39 through 44. Refer to mathematical formulas 45 to 47 for the chromaticity derivation and non-luminance system. 13 Refer to mathematical formulas 39 through 44. Refer to chromaticity derivation constant and luminance system mathematical formulas 66 to 75. 14 IC T C P Rec. ITU-R BT.2100-2 IC T C P For TransferCharacteristics value 16 (PQ), refer to Equations 79 through 81. For TransferCharacteristics value 18 (HLG), refer to Equations 82 through 84. 15 IPT-C2 Refer to SMPTE ST 2128 (202x) Formulas 85 through 87 16 YCgCo-Re Refer to mathematical formulas 58 through 65. 17 YCgCo-Ro Refer to mathematical formulas 58 through 65. 18-255 Reserved For future use by ITU-T | ISO / IEC

[0139] From Table 3, when the target image is an HDR image, the fourth identifier encoded in the bitstream is 9, and the corresponding matrix coefficient (K R ) is 0.2627, and K B It can be seen that is 0.0593. In addition, matrix coefficients are included in both the ITU-R BT.2020-2 standard and the ITU-R BT.2100-2 standard, and can also be seen in the ITU-R BT.2020-2 standard and ITU-R BT.2100-2 standard files.

[0140] In this embodiment of the present application, considering that the chroma coordinates of each primary color of the target image, the pixel mapping range of the target image, and the coefficients of the transformation matrix used to perform color space transformation for the target image jointly determine the transformation matrix used to perform color space transformation for the target image, when the encoder performs color space transformation for the target image, relevant information representing the transformation matrix, namely, the second identifier, the third identifier, and the fourth identifier, may also be encoded in the bitstream so that the decoder can perform color space transformation processing for a reconstructed HDR image in YUV format based on the second identifier, the third identifier, and the fourth identifier.

[0141] In some embodiments, after converting a target image in RGB format to a target image in YUV format through color space conversion, the encoder may reduce data throughput by separately performing downsampling processing on the Y component image and the UV component image. However, in different sampling modes, the sampling positions of the Y component image and the UV component image are offset. After acquiring the reconstructed Y component image and the reconstructed UV component image through decoding, the decoder side performs upsampling on the reconstructed image. If the offset state of the sampling position differs from the offset state of the encoder side, the accuracy of the reconstructed image is also affected.

[0142] Based on this, the encoder may additionally encode a fifth identifier in the bitstream that indicates the offset state of the sampling positions of the first component and the second component of the target image. That is, after color space conversion is performed on the target image, the target image includes an image of the first component and an image of the second component. Before encoding the target image into the bitstream, the encoder may perform downsampling on the image of the first component and the image of the second component based on the fifth identifier, and then encode the downsampled image of the first component and the downsampled image of the second component into the bitstream.

[0143] In the case of an image in YUV format, the primary component may be the Y component and the secondary component may be the UV component.

[0144] A 420 sampling mode is used as an example. Refer to the drawings of the chroma sampling locations shown in FIGS. 8a through 8f. The lumina (Y) sample area in the top-left corner of the image is indicated using a black solid line box, and the chroma (UV) sample area in the top-left corner of the image is indicated using a black solid line box. The center point of the lumina sample area is the lumina sampling point, and the center point of the chroma sample area is the chroma sampling point. FIGS. 8a through 8f show a total of six offset states (a) through (f).

[0145] For the six types of offset states listed in FIGS. 8a to 8f, Table 4 indicates the offset state between the chroma sampling position and the lumina sampling position in 420 sampling mode. That is, Table 4 lists the correspondence between the fifth identifier and the offset state of the chroma sampling position, corresponds to the figures in FIGS. 8a to 8f, and lists the horizontal offset value and the vertical offset value of the chroma sampling position relative to the lumina sampling position in each offset state.

[0146] Identifier value Horizontal offset value Vertical offset value 0 0 0.5 1 0.5 0.5 2 0 0 3 0.5 0 4 0 1 5 0.5 1

[0147] When the identifier value of the fifth identifier is 0, corresponding to FIG. 8a, the sampling position of the chroma sampling point has a horizontal offset value of 0 and a vertical offset value of 0.5 relative to the lumina sampling point; when the identifier value of the fifth identifier is 1, corresponding to FIG. 8b, the sampling position of the chroma sampling point has a horizontal offset value of 0.5 and a vertical offset value of 0.5 relative to the lumina sampling point; when the identifier value of the fifth identifier is 2, corresponding to FIG. 8c, the sampling position of the chroma sampling point has a horizontal offset value of 0 and a vertical offset value of 0 relative to the lumina sampling point; when the identifier value of the fifth identifier is 3, corresponding to FIG. 8d, the sampling position of the chroma sampling point has a horizontal offset value of 0.5 and a vertical offset value of 0 relative to the lumina sampling point; when the identifier value of the fifth identifier is 4, corresponding to FIG. 8e, the sampling position of the chroma sampling point has a horizontal offset value of 0 and a vertical offset value of 1 relative to the lumina sampling point; When the identifier value of the fifth identifier is 5, corresponding to FIG. 8f, it should be understood that the sampling position of the chroma sampling point has a horizontal offset value of 0.5 and a vertical offset value of 1 compared to the luminance sampling point.

[0148] From FIGS. 8a to 8f and Table 4, it can be seen that when the fifth identifier is 1, the chroma sampling point is located in the middle row and middle column of the luminance sampling point. In this case, if the pixel coordinates of the luminance sampling point Y are (x, y), the pixel coordinates of the corresponding chroma sampling points U and V are (x / 2, y / 2). When the fifth identifier is 2, the chroma sampling point is located in the top-left corner of the luminance sampling point. In this case, if the pixel coordinates of the luminance sampling point Y are (x, y), the pixel coordinates of the corresponding chroma sampling points of samples U and V are (floor(x / 2), floor(y / 2)).

[0149] It should be noted that in this 420 sampling mode, the sampling density of luminance samples is lower than that of chroma to reduce the need for data transmission and storage. Additionally, in 420 sampling mode, in the horizontal direction, the sampling density of luminance is full resolution, but the sampling density of chroma is only half that of luminance; in the vertical direction, both luminance and chroma are sampled at half the full resolution.

[0150] It should be understood that the foregoing example is described using only the 420 sampling mode. For other sampling modes, such as the 422 sampling mode, identifiers indicating the offset status of the sampling positions of the first and second components in the sampling mode may be encoded in the bitstream in the same way. This is not limited to the embodiments of the present application.

[0151] Likewise, if the bitstream includes a picture header, the fifth identifier may be encoded in the picture header.

[0152] In this embodiment of the present application, when downsampling is performed on the image of the first component and the image of the second component after color space conversion is performed on the target image, a fifth identifier representing the offset state of the sampling position of the first component of the target image and the second component of the target image may be encoded in the bitstream so that the decoder can perform corresponding upsampling processing on the reconstructed image of the Y component and the reconstructed image of the UV component based on the fifth identifier.

[0153] In some embodiments, to accurately render and display a target image, the image encoding method provided in this embodiment of the application further comprises the step of encoding static metadata of the target image into a bitstream, wherein the static metadata includes source device information and image display information. The source device information represents color information of a mastering monitor generating the target image, and the image display information represents the display luma of the target image.

[0154] In one example, source device information includes at least one of the chroma coordinates of the three primary colors of the mastering monitor, the chroma coordinates of the standard white light, the maximum display luma, and the minimum display luma. The chroma coordinates of the three primary colors include the x coordinates of the three primary colors and the y coordinates of the three primary colors, and the chroma coordinates of the standard white light include the x coordinates of the standard white light and the y coordinates of the standard white light.

[0155] Optionally, each item of source device information may be represented in the form of a 16-bit unsigned integer. This is not limited to the embodiments of the present application.

[0156] In one example, the image display information includes at least one of the maximum display luma or the maximum average display luma of the target image.

[0157] Maximum display luma is the maximum luma value that can be reached by a single pixel of the target image. Maximum display luma is 1 cd / m² 2 It can be represented in units using a 16-bit unsigned integer, and the value range is 1 cd / m 2 Up to 65535 cd / m 2is. If the target image is an image sequence or any frame of a video, the average luminance value of all pixels in each frame of the image within the image sequence or video is calculated first, and then the maximum value of the average luminance value corresponding to each frame of the image is the aforementioned maximum average display luminance.

[0158] Optionally, if there is only one frame of an independent image, the maximum average display luma may be a default value. This is not limited to the embodiments of the present application.

[0159] As illustrated in FIG. 7, the bitstream includes a picture header, and the picture header includes static metadata. That is, the first to fifth identifiers and the static metadata can all be encoded in the picture header.

[0160] In some embodiments, where the target image is an image sequence or a frame of a video, the image encoding method provided in this embodiment of the application further comprises the step of encoding dynamic metadata of the target image into a bitstream to accurately render and display the target image, wherein the dynamic metadata includes color gamut volume change data of the target image.

[0161] Color gamut volume change data is used to describe data related to changes in content and scenarios within an image as the amount of time or frames changes. Color gamut volume change data includes color change information, luminance change information, etc., of the target image.

[0162] As illustrated in FIG. 7, the bitstream includes a picture header, and the picture header includes static metadata. That is, the first to fifth identifiers, and static metadata and / or dynamic metadata can all be encoded in the picture header.

[0163] In some embodiments, in this embodiment of the application, when static metadata and / or dynamic metadata of a target image are encoded into a bitstream, a metadata header may also be added to the bitstream as shown in FIG. 9 to transmit the static metadata and / or dynamic metadata of the target image based on the metadata header.

[0164] In some embodiments, where the bitstream includes a metadata substream corresponding to a metadata header, the metadata substream may also be used to convey static metadata and dynamic metadata of the target image.

[0165] In the present embodiment of the application, it should be understood that only the data information of the target image to be encoded into the bitstream, e.g., the first to fifth identifiers, and static metadata and / or dynamic metadata, is limited, but the specific location of the data information when the data information is encoded into the bitstream is not limited. In other words, the picture header, metadata header, and metadata substream of the above example may all contain the relevant data information. Of course, alternatively, all the above data information may be contained only in the picture header, metadata header, or metadata substream. This is not limited to the embodiments of the application.

[0166] Based on the above-described embodiment, in order to include the first to fifth identifiers and static metadata and dynamic metadata of the target image in the picture header of the bitstream, in this embodiment of the present application, four syntax elements including image information and metadata information related to the target image are added to the picture header of the bitstream. Table 5 shows examples of syntax elements of the picture header.

[0167] picture_header( ) { descriptor PIH u(16) picture_header_size u(16) img_width u(16) img_height u(16) picture_format u(2) bit_depth u(1) bit_c_ver u(1) bit_c_hor u(1) bit_s_ver u(1) bit_s_hor u(1) coding_independent_code_points() mastering_display_color_volume() content_light_level_information () dynamic_metadata() color_transform_header() model_header() }

[0168] picture_header_size is used to define the spatial size of the picture header, and the picture header defines picture-related information, for example, img_width is the width of the encoded picture; img_height is the height of the encoded image; picture_format is the picture format and is used to define whether the encoded picture is a YUV picture in 4:4:4 sampling format, a YUV picture in 4:2:2 sampling format, a YUV picture in 4:2:0 sampling format, or an RGB picture; bit_depth is quantization information; and model_header() is a syntax related to the encoder model. For the specific meanings of these parameters, refer to the relevant definitions of the high-level syntax in the encoding file header of the JPEG-AI standard. Details are not described in this embodiment of the application.

[0169] In the aforementioned picture header, coding_independent_code_points() is a newly added phrase to convey image features in the present application, mastering_display_color_volume() is a newly added phrase to convey source device information in the present application, content_light_level_information() is a newly added phrase to convey image display information in the present application, and dynamic_metadata() is a newly added phrase to convey dynamic metadata of the target image in this embodiment of the present application.

[0170] It should be noted that in the syntax element table illustrated in this embodiment of the application, the right descriptor column is the data type and bit information of the variable defined in the corresponding syntax. For example, U(16) is a 16-bit unsigned integer.

[0171] In one example, the syntax elements of the image feature syntax coding_independent_code_points() are shown in Table 6.

[0172] coding_independent_code_points() { descriptor colour_primaries u(8) matrix_coefficients u(8) image_full_range_flag u(1) transfer_characteristics u(16) chroma420_sample_loc_type u(8) }

[0173] In Table 6, colour_primaries is a byte corresponding to a second identifier representing the chroma coordinates of each primary color of the target image, matrix_coefficients is a byte corresponding to a fourth identifier representing the coefficients of the transformation matrix used to perform color space transformation on the target image, image_full_range_flag is a byte corresponding to a third identifier representing the pixel mapping range of the target image, transfer_characteristics is a byte corresponding to a first identifier representing the photoelectric transfer function matching the target image, and chroma420_sample_loc_type is a byte corresponding to a fifth identifier representing the offset state of the sampling location of the primary component of the target image and the secondary component of the target image.

[0174] In one example, the syntax elements of source device information mastering_display_color_volume() are shown in Table 7.

[0175] mastering_display_color_volume() { descriptor for ( i=0;i<3;i++ ) { mastering_display_colour_primaries_x [i] u(16) mastering_display_colour_primaries_y [i] u(16) } mastering_display_white_point_chromaticity_x u(16) mastering_display_white_point_chromaticity_y u(16) mastering_display_maximum_luminance u(16) mastering_display_minimum_luminance u(16) }

[0176] In Table 7, mastering_display_colour_primaries_x [i] and mastering_display_colour_primaries_y [i] are bytes corresponding to the chroma coordinates of the three primary colors of the mastering monitor, mastering_display_white_point_chromaticity_x and mastering_display_white_point_chromaticity_y are bytes corresponding to the chroma coordinates of the standard white light of the mastering monitor, mastering_display_maximum_luminance is a byte corresponding to the maximum display luminance of the mastering monitor, and mastering_display_minimum_luminance is the minimum display luminance of the mastering monitor.

[0177] For example, the syntax elements of image display information content_light_level_information() are shown in Table 8.

[0178] content_light_level_information() { 디스크립터 maximum_content_light_level u(16) maximum_frame_average_light_level u(16) }

[0179] In Table 8, maximum_content_light_level is the byte corresponding to the maximum display luma of the target image, and maximum_frame_average_light_level is the byte corresponding to the maximum average display luma of the target image.

[0180] For example, the syntax elements of dynamic metadata dynamic_metadata() are shown in Table 9.

[0181] dynamic_metadata{ 디스크립터 payloadSize = 0 while( next_bits( 8 ) = = 0xFF ) { payloadSize += 255 } last_payload_size_byte u(8) payloadSize += last_payload_size_byte itu_t_t35_information(payloadSize) { itu_t_t35_country_code b(8) if( itu_t_t35_country_code != 0xFF ) i = 1 else { itu_t_t35_country_code_extension_byte b(8) i = 2 } do { itu_t_t35_payload_byte b(8) i++ } while( i < payloadSize ) }

[0182] In Table 9 above, itu_t_t35_country_code is a byte having a value designated as a country code by ITU-T T.35 Annex A, itu_t_t35_country_code_extension_byte must be a byte having a value designated as a country code by ITU-T T.35 Annex B, itu_t_t35_payload_byte must be a byte containing data registered in accordance with ITU-T T.35, and the ITU-T T.35 terminal provider code and terminal provider oriented code must be included in the first or more bytes of itu_t_t35_payload_byte. The format is specified by the organization publishing the terminal provider code. Any remaining itu_t_t35_payload_byte data must be data having the syntax and semantics specified by the entity identified by the ITU-T T.35 country code and terminal provider code.

[0183] Table 5 is used as an example, and it should be understood that all data information added in this embodiment of the application is included in the picture header. In certain applications, Table 5 may be modified to include some data in the picture header. This is not limited to the embodiments of the application.

[0184] Based on the aforementioned embodiments, in this embodiment of the present application, the metadata header of the bitstream may further include static metadata and dynamic metadata of the target image. In one example, Table 10 below shows the syntax elements of the metadata header.

[0185] metadata_header( ) { 디스크립터 MDH u(16) metadata_header_size u(16) static_metadata() dynamic_metadata() }

[0186] MDH is the metadata header marker, the metadata header follows the metadata header marker, and metadata_header_size is the size of the metadata header.

[0187] For example, the syntax element table of static metadata is shown in Table 11 below.

[0188] static_metadata{ 디스크립터 mastering_display_color_volume () content_light_level_information () }

[0189] Refer to Table 7 for specific details regarding the syntax elements of mastering_display_color_volume(). Refer to Table 8 for the syntax elements of content_light_level_information(). Details are not explained again here.

[0190] In conclusion, when encoding a target image, the encoder may encode a first identifier representing an optical-electrical transfer function matching the target image into the bitstream to satisfy the display requirements of both SDR images and HDR images on the decoder side, and as a result, after obtaining the first identifier through decoding, the decoder can render and display an image reconstructed according to the corresponding optical-electrical transfer function.

[0191] The following describes an image decoding method provided in an embodiment of the present application.

[0192] As illustrated in FIG. 10, the present application provides an image decoding method. The method is applied to a decoder, and the method comprises the following steps.

[0193] Step 1001: Parse the bitstream to obtain a first identifier, the first identifier represents an optical-electric transfer function that matches a target image, and the target image includes an HDR image and / or an SDR image.

[0194] The bitstream includes a picture header, and the first identifier is included in the picture header.

[0195] Optionally, the bitstream includes a picture header, and the second identifier, the third identifier, and the fourth identifier are included in the picture header. In this case, the second identifier, the third identifier, and the fourth identifier may be further parsed from the bitstream.

[0196] The second identifier represents the chroma coordinates of each primary color of the target image, the third identifier represents the mapping range of the target image, and the fourth identifier represents the coefficients of the transformation matrix used to perform color space transformation on the target image.

[0197] Optionally, the bitstream includes a picture header, and the fifth identifier is included in the picture header. In this case, the fifth identifier can be further parsed from the bitstream.

[0198] Optionally, the bitstream includes a picture header, and static metadata and / or dynamic metadata of the target image are also included in the picture header. In this case, the static metadata and / or dynamic metadata of the target image can be parsed from the bitstream.

[0199] Step 1002: Decode the bitstream to obtain a reconstructed image of the target image.

[0200] When the fifth identifier is parsed from the bitstream, the implementation process of step 1002 may include parsing the bitstream to obtain a reconstructed image of the primary component of the target image and a reconstructed image of the secondary component of the target image; and performing upsampling on the reconstructed image of the primary component and the reconstructed image of the secondary component based on the fifth identifier to obtain a reconstructed image.

[0201] Optionally, when the second identifier, the third identifier, and the fourth identifier are parsed from the bitstream, before performing step 1003, the decoder determines the transformation matrix used when the encoder side performs color space transformation on the target image by searching Tables 1 through 3 based on the second identifier, the third identifier, and the fourth identifier, and performs inverse color space transformation on the reconstructed image based on the transformation matrix.

[0202] Step 1003: Perform electro-optical conversion on the reconstructed image based on the first identifier.

[0203] In some embodiments, when static metadata of a target image is parsed from a bitstream, tone mapping is performed on the reconstructed image based on the static metadata to display the reconstructed image.

[0204] In some embodiments, when dynamic metadata of a target image is parsed from a bitstream, tone mapping is performed on a reconstructed image based on the dynamic metadata to display the reconstructed image.

[0205] It should be noted that when the decoder performs an image decoding method, the specific meaning and execution order of the relevant technical features refer to the aforementioned embodiment in which the encoder performs an image encoding method. Further details are not described in this specification.

[0206] In conclusion, when decoding the bitstream, the decoder can determine an optical-electrical transfer function that matches the target image based on a first identifier obtained through decoding, and accordingly, the reconstructed image of the target image is accurately rendered and displayed according to the optical-electrical transfer function, thereby satisfying the display requirements of both SDR and HDR images on the decoder side and guaranteeing the accuracy of the target image display.

[0207] FIG. 11 is a diagram of the structure of an encoding device according to one embodiment of the present application. The encoding device may be an encoder as shown in FIG. 1. As shown in FIG. 11, the encoding device (1100) includes an image encoding module (1101) and an information encoding module (1102).

[0208] The image encoding module (1101) is configured to encode the target image into a bitstream.

[0209] The information encoding module (1102) is configured to encode a first identifier into a bitstream, the first identifier represents an optical-electric transfer function that matches a target image, and the target image includes a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image.

[0210] Optionally, the bitstream includes a picture header, and the first identifier is encoded in the picture header.

[0211] Optionally, the information encoding module (1102) also,

[0212] Encode a second identifier representing the chroma coordinates of each primary color of the target image into a bitstream;

[0213] Encoding a third identifier representing the pixel mapping range of the target image into a bitstream;

[0214] It is configured to encode a fourth identifier representing the coefficients of the transformation matrix used to perform color space transformation on the target image into the bitstream.

[0215] Optionally, the device further includes an image conversion module configured to perform color space conversion on a target image based on a second identifier, a third identifier, and a fourth identifier.

[0216] Optionally, the bitstream includes a picture header, and the second identifier, third identifier, and fourth identifier are encoded in the picture header.

[0217] Optionally, the information encoding module (1102) also,

[0218] It is configured to encode a fifth identifier into the bitstream, and the fifth identifier represents the offset state of the sampling positions of the first component of the target image and the second component of the target image.

[0219] Optionally, after color space conversion is performed on the target image, the target image includes an image of the first component and an image of the second component; the device further includes a downsampling module configured to perform downsampling on the image of the first component and the image of the second component based on a fifth identifier.

[0220] Specifically, the image encoding module (1101) is,

[0221] It is configured to encode the downsampled image of the first component and the downsampled image of the second component into a bitstream.

[0222] Optionally, the bitstream includes a picture header, and the fifth identifier is encoded in the picture header.

[0223] Optionally, the information encoding module (1102) also,

[0224] It is configured to encode static metadata of the target image into a bitstream, the static metadata including source device information and image display information, the source device information represents the color information of the mastering monitor generating the target image, and the image display information represents the display luma of the target image.

[0225] Optionally, source device information includes at least one of the chroma coordinates of the three primary colors of the mastering monitor, the chroma coordinates of the standard white light, the maximum display luma, and the minimum display luma.

[0226] Optionally, the image display information includes at least one of the maximum display luma or the maximum average display luma of the target image.

[0227] Optionally, the bitstream includes a picture header, and the picture header includes static metadata; or the bitstream includes a metadata header, and the metadata header includes static metadata; or the bitstream includes a metadata substream corresponding to the metadata header, and the metadata substream includes static metadata.

[0228] Optionally, the information encoding module (1102) also,

[0229] It is configured to encode the dynamic metadata of the target image into a bitstream, and the dynamic metadata includes color gamut volume change data of the target image.

[0230] Optionally, the bitstream includes a picture header, and the picture header includes dynamic metadata; or the bitstream includes a metadata header, and the metadata header includes dynamic metadata; or the bitstream includes a metadata substream corresponding to the metadata header, and the metadata substream includes dynamic metadata.

[0231] In the present embodiment of the application, when encoding a target image, the encoding device may also encode a first identifier representing an optical-electrical transfer function matching the target image into the bitstream so as to satisfy the display requirements of both SDR images and HDR images at the decoder side, and as a result, after obtaining the first identifier through decoding, the decoder may render and display an image reconstructed according to the corresponding optical-electrical transfer function.

[0232] It should be noted that during image encoding by the encoding device provided in the above-described embodiment, the division of the aforementioned functional modules is used only as an example for illustrative purposes. In practice, the aforementioned functions may be assigned to and implemented in different functional modules based on requirements. Specifically, the internal structure of the device is divided into different functional modules to implement all or part of the aforementioned functions. Furthermore, the encoding device and the image encoding method embodiment provided in the above-described embodiment belong to the same concept. For details regarding the specific implementation process of the encoding device, refer to the method embodiment. Details are not described again here.

[0233] FIG. 12 is a diagram of the structure of a decoding device according to one embodiment of the present application. The decoding device may be the decoder shown in FIG. 1. As shown in FIG. 12, the device includes an information decoding module (1201), an image decoding module (1202), and an image conversion module (1203).

[0234] The information decoding module (1201) is configured to parse a bitstream to obtain a first identifier, the first identifier representing an optical-electric transfer function that matches a target image, and the target image includes a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image.

[0235] The image decoding module (1202) is configured to decode the bitstream to obtain a reconstructed image of the target image.

[0236] The image conversion module (1203) is configured to perform electro-optical conversion on the reconstructed image based on the first identifier.

[0237] Optionally, the bitstream includes a picture header, and the first identifier is included in the picture header.

[0238] Optionally, the information decoding module (1201) also,

[0239] Parse the bitstream to obtain a second identifier - the second identifier represents the chroma coordinates of each primary color of the target image -;

[0240] Parsing the bitstream to obtain a third identifier - the third identifier represents the mapping range of the target image -;

[0241] It is configured to parse the bitstream to obtain a fourth identifier—the fourth identifier represents the coefficients of the transformation matrix used to perform color space conversion on the target image.

[0242] Optionally, the device further includes an image inverse conversion module configured to perform an inverse color space conversion on an image reconstructed based on a second identifier, a third identifier, and a fourth identifier.

[0243] Optionally, the bitstream includes a picture header, and the second identifier, the third identifier, and the fourth identifier are included in the picture header.

[0244] Optionally, the information decoding module (1201) also,

[0245] It is configured to parse the bitstream to obtain a fifth identifier, and the fifth identifier represents the offset state of the sampling positions of the first component of the target image and the second component of the target image.

[0246] Optionally, the image decoding module (1202) specifically,

[0247] Parsing the bitstream to obtain a reconstructed image of the primary components of the target image and a reconstructed image of the secondary components of the target image;

[0248] It is configured to obtain a reconstructed image by performing upsampling on the reconstructed image of the primary component and the reconstructed image of the secondary component based on the fifth identifier.

[0249] Optionally, the bitstream includes a picture header, and the fifth identifier is included in the picture header.

[0250] Optionally, the information decoding module (1201) also,

[0251] It is configured to parse a bitstream to obtain static metadata of a target image, the static metadata including source device information and image display information, the source device information represents the color information of the mastering monitor generating the target image, and the image display information represents the display luma of the target image.

[0252] Optionally, the device further includes a tone mapping module configured to perform tone mapping on a reconstructed image based on static metadata.

[0253] Optionally, source device information includes at least one of the chroma coordinates of the three primary colors of the mastering monitor, the chroma coordinates of the standard white light, the maximum display luma, and the minimum display luma.

[0254] Optionally, the image display information includes at least one of the maximum display luma or the maximum average display luma of the target image.

[0255] Optionally, the bitstream includes a picture header, the picture header includes static metadata; or the bitstream includes a metadata header, the metadata header includes static metadata; or the bitstream includes a metadata substream corresponding to the metadata header, the metadata substream includes static metadata.

[0256] Optionally, the information decoding module (1201) also,

[0257] It is configured to parse the bitstream to obtain dynamic metadata of the target image, and the dynamic metadata includes color gamut volume change data of the target image.

[0258] Optionally, the tone mapping module also,

[0259] It is configured to perform tone mapping on reconstructed images based on static metadata and dynamic metadata.

[0260] Optionally, the bitstream includes a picture header, and the picture header includes dynamic metadata; or the bitstream includes a metadata header, and the metadata header includes dynamic metadata; or the bitstream includes a metadata substream corresponding to the metadata header, and the metadata substream includes dynamic metadata.

[0261] In the present embodiment of the application, when decoding a bitstream, the decoding device can determine an optical-electrical transfer function that matches a target image based on a first identifier obtained through decoding, and accordingly, the reconstructed image of the target image is accurately rendered and displayed according to the optical-electrical transfer function, thereby satisfying the display requirements for both SDR images and HDR images on the decoder side and ensuring the accuracy of the target image display.

[0262] It should be noted that during image decoding by the decoding device provided in the above-described embodiment, the division of the aforementioned functional modules is used only as an example for illustrative purposes. In practice, the aforementioned functions may be assigned to different functional modules and implemented based on requirements. Specifically, the internal structure of the device is divided into different functional modules to implement all or part of the aforementioned functions. Furthermore, the decoding device and the image decoding method embodiments provided in the above-described embodiment belong to the same concept. For details regarding the specific implementation process of the decoding device, refer to the method embodiments. Details are not described again herein.

[0263] One embodiment of the present application further provides an encoding device. The encoding device includes a processor, the processor is coupled to a memory, the memory is configured to store a program or instruction, and when the program or instruction is executed by the processor, the encoding device is able to perform the aforementioned image encoding method.

[0264] One embodiment of the present application further provides a decoding device. The decoding device includes a processor, the processor is coupled to a memory, the memory is configured to store a program or instruction, and when the program or instruction is executed by the processor, the decoding device is able to perform the aforementioned image decoding method.

[0265] One embodiment of the present application further provides an encoding and decoding system. The encoding and decoding system includes an encoding device and / or a decoding device.

[0266] One embodiment of the present application further provides a computer-readable storage medium comprising program code. When the program code is executed on a computer, the computer becomes able to perform the image encoding method described above.

[0267] One embodiment of the present application further provides a computer-readable storage medium comprising program code. When the program code is executed on a computer, the computer becomes able to perform the image decoding method described above.

[0268] One embodiment of the present application further provides a computer program product comprising instructions. When the instructions are executed on a computer, the computer becomes able to perform the aforementioned image encoding method.

[0269] One embodiment of the present application further provides a computer program product comprising instructions. When the instructions are executed on a computer, the computer becomes able to perform the image decoding method described above.

[0270] One embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a bitstream obtained according to the image encoding method described above.

[0271] One embodiment of the present application further provides a bitstream storage device comprising at least one storage medium and a communication interface. The communication interface is configured to receive or transmit a bitstream. At least one storage medium is configured to store a bitstream. The bitstream is obtained by an encoder through encoding according to the image encoding method described above.

[0272] One embodiment of the present application further provides a bitstream storage method comprising the steps of receiving a bitstream through a communication interface and storing the bitstream in one or more storage media. The bitstream is obtained by an encoder through encoding according to the image encoding method described above.

[0273] One embodiment of the present application further provides a bitstream distribution system comprising at least one storage medium and a video stream device. At least one storage medium is configured to store a bitstream, and the bitstream is obtained by an encoder through encoding according to the image encoding method described above. The video stream device is configured to transmit the bitstream in the at least one storage medium to the decoder in response to a request from the decoder.

[0274] One embodiment of the present application further provides a bitstream distribution method comprising the steps of receiving a first request, selecting a bitstream from at least one storage medium in response to the first request, and transmitting the bitstream to a destination device. At least one storage medium is configured to store a bitstream, and the bitstream is obtained through encoding by an encoder according to the image encoding method described above.

[0275] One embodiment of the present application further provides a bitstream processing system comprising an image source device, an encoder, one or more storage media, and a destination device. The image source device is configured to provide image data. The encoder is configured to acquire image data from the image source device through an interface and to encode the image data to acquire one or more bitstreams, and the bitstreams are acquired by the encoder through encoding according to the image encoding method described above. The encoder is configured to store one or more bitstreams in one or more storage media. Alternatively, the encoder is configured to encapsulate one or more bitstreams to acquire a transmission bitstream. The encoder is configured to transmit the transmission bitstream to a destination device through a communication link or a communication network. The destination device is configured to decapsulate the transmission bitstream to acquire one or more bitstreams. The destination device is configured to decode one or more bitstreams to acquire decoded data.

[0276] All or part of the foregoing embodiments may be implemented by software, hardware, firmware, or any combination thereof. Where software is used to implement the embodiments, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product comprises one or more computer instructions. When computer instructions are loaded into and executed by a computer, a procedure or function according to an embodiment of the present application is created in whole or in part. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device. Computer instructions may be stored on a computer-readable storage medium or transferred from a computer-readable storage medium to another computer-readable storage medium. For example, computer instructions may be transferred from a website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, or digital subscriber line (DSL)) or wireless (e.g., infrared, radio, or microwave). A computer-readable storage medium may be any available medium accessible by a computer, or a data storage device such as a server or data center that incorporates one or more available media. Available media may be magnetic media (e.g., floppy disks, hard disks, or magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), semiconductor media (e.g., solid-state disks (SSDs)), etc. It should be noted that the computer-readable storage media mentioned in the embodiments of this application may be non-volatile storage media, that is, non-transient storage media.

[0277] In this specification, "plural" should be understood to mean two or more. In the description of the embodiments of this application, " / " means "or" unless otherwise specified. For example, A / B may represent A or B. In this specification, "and / or" describes only the association between related objects and indicates that three relationships may exist. For example, A and / or B may represent the following three cases: A alone exists, both A and B exist, and B alone exists. Additionally, to clearly describe the technical solution in the embodiments of this application, terms such as "first" and "second" are used in the embodiments of this application to distinguish identical or similar items that essentially provide the same function or purpose. Those skilled in the art will understand that terms such as "first" and "second" do not limit quantity or order of execution, and that terms such as "first" and "second" do not indicate a clear difference.

[0278] It should be noted that in the embodiments of this application, information (including, but not limited to, user equipment information, user's personal information, etc.), data (including, but not limited to, data used for analysis, stored data, displayed data, etc.), and signals are used with permission by the user or full permission by all parties, and that the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0279] The foregoing description is merely an example of the present application, but is not intended to limit the present application. Any modification, equivalent substitution, or improvement made without departing from the principles of the present application shall be included within the scope of protection of the present application.

Claims

Claim 1 An image encoding method comprising the steps of encoding a target image into a bitstream and encoding a first identifier into the bitstream, wherein the first identifier represents an optical-electric transfer function that matches the target image, and the target image includes a High Dynamic Range (HDR) image and / or a Standard Dynamic Range (SDR) image. Claim 2 An image encoding method according to claim 1, wherein the bitstream includes a picture header and the first identifier is encoded in the picture header. Claim 3 An image encoding method according to claim 1 or 2, further comprising the steps of: encoding a second identifier representing the chroma coordinates of each primary color of the target image into the bitstream; encoding a third identifier representing the pixel mapping range of the target image into the bitstream; and encoding a fourth identifier representing the coefficients of a transformation matrix used to perform color space transformation on the target image into the bitstream. Claim 4 An image encoding method according to claim 3, wherein, before encoding the target image into the bitstream, the method further comprises the step of performing a color space conversion on the target image based on the second identifier, the third identifier, and the fourth identifier. Claim 5 An image encoding method according to claim 3 or 4, wherein the bitstream includes a picture header, and the second identifier, the third identifier, and the fourth identifier are encoded in the picture header. Claim 6 An image encoding method according to any one of claims 1 to 5, wherein the method further comprises the step of encoding a fifth identifier representing the offset state of the sampling position of the primary component of the target image and the secondary component of the target image into the bitstream. Claim 7 In claim 6, after a color space conversion is performed on the target image, the target image includes the image of the first component and the image of the second component, and before encoding the target image into the bitstream, the method, The method further includes the step of performing downsampling on the image of the first component and the image of the second component based on the fifth identifier, and The step of encoding the above target image into the above bitstream is, A method comprising the step of encoding the downsampled image of the first component and the downsampled image of the second component into the bitstream. Image encoding method. Claim 8 An image encoding method according to claim 6 or 7, wherein the bitstream includes a picture header and the fifth identifier is encoded in the picture header. Claim 9 An image encoding method according to any one of claims 1 to 8, wherein the method further comprises the step of encoding static metadata of the target image into the bitstream, wherein the static metadata includes source device information and image display information, wherein the source device information represents color information of a mastering monitor generating the target image, and the image display information represents the display luma of the target image. Claim 10 An image encoding method according to claim 9, wherein the source device information includes at least one of the chroma coordinates of the three primary colors of the mastering monitor, the chroma coordinates of standard white light, the maximum display luma, and the minimum display luma. Claim 11 An image encoding method according to claim 9 or 10, wherein the image display information comprises at least one of the maximum display luma or the maximum average display luma of the target image. Claim 12 An image encoding method according to any one of claims 9 to 11, wherein the bitstream comprises a picture header and the picture header comprises static metadata; or the bitstream comprises a metadata header and the metadata header comprises static metadata; or the bitstream comprises a metadata substream corresponding to the metadata header and the metadata substream comprises static metadata. Claim 13 An image encoding method according to any one of claims 9 to 12, wherein the method further comprises the step of encoding dynamic metadata of the target image into the bitstream, and the dynamic metadata includes color gamut volume change data of the target image. Claim 14 An image encoding method according to claim 13, wherein the bitstream includes a picture header and the picture header includes the dynamic metadata; or the bitstream includes a metadata header and the metadata header includes the dynamic metadata; or the bitstream includes a metadata substream corresponding to the metadata header and the metadata substream includes the dynamic metadata. Claim 15 An image decoding method comprising: a step of parsing a bitstream to obtain a first identifier, wherein the first identifier represents an optical-electric transfer function that matches a target image, and the target image includes a High Dynamic Range (HDR) image and / or a Standard Dynamic Range (SDR) image; a step of decoding the bitstream to obtain a reconstructed image of the target image; and a step of performing an electric-optical conversion on the reconstructed image based on the first identifier. Claim 16 An image decoding method according to claim 15, wherein the bitstream includes a picture header, and the first identifier is included in the picture header. Claim 17 An image decoding method according to claim 15 or 16, wherein the method further comprises the steps of parsing the bitstream to obtain a second identifier representing the chroma coordinates of each primary color of the target image, parsing the bitstream to obtain a third identifier representing the mapping range of the target image, and parsing the bitstream to obtain a fourth identifier representing the coefficients of a transformation matrix used to perform color space transformation on the target image. Claim 18 An image decoding method according to claim 17, wherein, before performing electro-optical conversion on the reconstructed image based on the first identifier, the method further comprises the step of performing inverse color space conversion on the reconstructed image based on the second identifier, the third identifier, and the fourth identifier. Claim 19 An image decoding method according to claim 17 or 18, wherein the bitstream includes a picture header, and the second identifier, the third identifier, and the fourth identifier are included in the picture header. Claim 20 An image decoding method according to any one of claims 15 to 19, wherein the method further comprises the step of parsing the bitstream to obtain a fifth identifier representing the offset state of the sampling positions of the primary component of the target image and the secondary component of the target image. Claim 21 An image decoding method according to claim 20, wherein the step of parsing the bitstream to obtain a reconstructed image of the target image comprises: a step of parsing the bitstream to obtain a reconstructed image of the primary component of the target image and a reconstructed image of the secondary component of the target image; and a step of performing upsampling on the reconstructed image of the primary component and the reconstructed image of the secondary component based on the fifth identifier to obtain the reconstructed image. Claim 22 An image decoding method according to claim 20 or 21, wherein the bitstream includes a picture header and the fifth identifier is included in the picture header. Claim 23 An image decoding method according to any one of claims 15 to 22, wherein the method further comprises the step of parsing the bitstream to obtain static metadata of the target image, wherein the static metadata includes source device information and image display information, wherein the source device information represents color information of a mastering monitor generating the target image, and the image display information represents the display luma of the target image. Claim 24 An image decoding method according to claim 23, wherein the method further comprises the step of performing tone mapping on the reconstructed image based on the static metadata. Claim 25 An image decoding method according to claim 23 or 24, wherein the source device information comprises at least one of the chroma coordinates of the three primary colors of the mastering monitor, the chroma coordinates of the standard white light, the maximum display luma, and the minimum display luma. Claim 26 An image decoding method according to any one of claims 23 to 25, wherein the image display information comprises at least one of the maximum display luma or the maximum average display luma of the target image. Claim 27 An image decoding method according to any one of claims 23 to 26, wherein the bitstream comprises a picture header and the picture header comprises static metadata; or the bitstream comprises a metadata header and the metadata header comprises static metadata; or the bitstream comprises a metadata substream corresponding to the metadata header and the metadata substream comprises static metadata. Claim 28 An image decoding method according to any one of claims 23 to 27, wherein the method further comprises the step of parsing the bitstream to obtain dynamic metadata of the target image, and the dynamic metadata includes color gamut volume change data of the target image. Claim 29 An image decoding method according to claim 28, wherein the method further comprises the step of performing tone mapping on the reconstructed image based on the static metadata and the dynamic metadata. Claim 30 An image decoding method according to claim 29, wherein the bitstream comprises a picture header and the picture header comprises the dynamic metadata; or the bitstream comprises a metadata header and the metadata header comprises the dynamic metadata; or the bitstream comprises a metadata substream corresponding to the metadata header and the metadata substream comprises the dynamic metadata. Claim 31 An encoding device comprising an image encoding module configured to encode a target image into a bitstream and an information encoding module configured to encode a first identifier into the bitstream, wherein the first identifier represents an optical-electrical transfer function that matches the target image, and the target image includes a High Dynamic Range (HDR) image and / or a Standard Dynamic Range (SDR) image. Claim 32 A decoding device comprising: an information decoding module configured to parse a bitstream to obtain a first identifier, wherein the first identifier represents an optical-electric transfer function that matches a target image, and the target image includes a High Dynamic Range (HDR) image and / or a Standard Dynamic Range (SDR) image; an image decoding module configured to decode the bitstream to obtain a reconstructed image of the target image; and an image conversion module configured to perform an electric-optical conversion on the reconstructed image based on the first identifier. Claim 33 An encoding device comprising a processor, wherein the processor is coupled to a memory, wherein the memory is configured to store a program or instruction, and when the program or instruction is executed by the processor, the encoding device is capable of performing a method according to any one of claims 1 to 14. Claim 34 A decoding device comprising a processor, wherein the processor is coupled to a memory, wherein the memory is configured to store a program or instruction, and when the program or instruction is executed by the processor, the decoding device is capable of performing a method according to any one of claims 15 to 30. Claim 35 An encoding and decoding system, wherein the encoding and decoding system comprises an encoding device according to claim 33 and / or a decoding device according to claim 34. Claim 36 A computer-readable storage medium comprising program code, wherein when the program code is executed on a computer, the computer is able to perform an image encoding method according to any one of claims 1 to 14. Claim 37 A computer-readable storage medium comprising program code, wherein when the program code is executed on a computer, the computer is able to perform an image decoding method according to any one of claims 15 to 30. Claim 38 A computer program product comprising instructions, wherein when the instructions are executed on a computer, the computer is enabled to perform a method according to any one of claims 1 to 14. Claim 39 A computer program product comprising instructions, wherein when the instructions are executed on a computer, the computer is able to perform a method according to any one of claims 15 to 30. Claim 40 A computer-readable storage medium, wherein the computer-readable storage medium stores a bitstream obtained by one or more processors by performing a method according to any one of claims 1 to 14. Claim 41 A bitstream storage device comprising at least one storage medium and a communication interface, wherein the communication interface is configured to receive or transmit a bitstream, the at least one storage medium is configured to store the bitstream, and the bitstream is obtained through encoding by an encoder according to an image encoding method according to any one of claims 1 to 14. Claim 42 A bitstream storage method comprising the steps of receiving a bitstream through a communication interface and storing the bitstream in one or more storage media, wherein the bitstream is obtained by encoding by an encoder according to an image encoding method according to any one of claims 1 to 14. Claim 43 A bitstream distribution system comprising at least one storage medium and a video stream device, wherein the at least one storage medium is configured to store a bitstream, the bitstream is obtained through encoding by an encoder according to an image encoding method according to any one of claims 1 to 14, and the video stream device is configured to transmit the bitstream within the at least one storage medium to the decoder in response to a request from a decoder. Claim 44 A bitstream distribution method comprising the steps of receiving a first request, selecting a bitstream from at least one storage medium in response to the first request, and transmitting the target bitstream to a destination device, wherein the at least one storage medium is configured to store a bitstream, and the bitstream is obtained through encoding by an encoder according to an image encoding method according to any one of claims 1 to 14. Claim 45 A bitstream processing system comprising an image source device, an encoding device, one or more storage media, and a destination device, wherein the image source device is configured to provide image data, the encoding device is configured to acquire the image data of the image source device through an interface and to encode the image data to acquire one or more bitstreams, wherein the bitstreams are acquired through encoding by an encoder according to an image encoding method according to any one of claims 1 to 14, the encoding device is configured to store the one or more bitstreams in the one or more storage media, or the encoding device is configured to encapsulate the one or more bitstreams to acquire a transmission bitstream, the encoding device is configured to transmit the transmission bitstream to the destination device through a communication link or a communication network, the destination device is configured to decapsulate the transmission bitstream to acquire the one or more bitstreams, and the destination device is configured to decode the one or more bitstreams to acquire decoded data.