Image encoding method, image decoding method and related device
By adding photoelectric conversion function identifiers during the encoding process, the problem that the JPEG AI decoding model cannot display HDR images is solved, and the accurate reconstruction and display of HDR images is realized.
Patent Information
- Application Number
- PCT/CN2024/073143
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-24
AI Technical Summary
The JPEG AI decoding model cannot accurately display high dynamic range (HDR) images because the photoelectric conversion function it uses is suitable for standard dynamic range (SDR) images, resulting in inaccurate display of HDR images after decoding.
The encoding process is added to indicate the photoelectric conversion function matching the target image, and the photoelectric conversion is used during decoding to adapt to the display needs of HDR and SDR images.
The accurate display of HDR images is achieved, ensuring the reconstruction and display quality of HDR images on the decoding end.
Smart Images

Figure CN2024073143_24072025_PF_FP_ABST
Abstract
Description
Image encoding and decoding method and related device Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an image encoding and decoding method and related devices. Background Art
[0002] JPEG AI is an image coding and decoding standard based on machine learning. The coding and decoding scheme based on JPEG AI is based on the autoencoder structure with color component separation. It realizes image coding and decoding by encoding and decoding the main and secondary components in the image separately.
[0003] The image encoding process based on JPEG AI is as follows: through color space conversion, the image in RGB (a color space, R represents red, G represents green, and B represents blue) format is converted into an image in YUV (a color space, Y represents brightness, U and V represent chrominance) format; then, the Y component image and the UV component image are downsampled separately to reduce the amount of data processing; then, the downsampled Y component image and UV component image are encoded separately to obtain the Y component code stream and the UV component code stream. The image decoding process based on JPEG AI is as follows: the Y component code stream and the UV component code stream are parsed separately to obtain the reconstructed Y component image and the reconstructed UV component image; then, the reconstructed Y component image and the reconstructed UV component image are upsampled separately to obtain the reconstructed image in YUV format; finally, through inverse color space conversion, the reconstructed image in YUV format is converted into the reconstructed image in RGB format.
[0004] However, when displaying RGB format images after decoding, JPEG AI defaults to using a photoelectric transfer function based on the "Gamma" function for processing. This photoelectric conversion function works well on traditional display devices (with an illumination of around 100cd / m2) for displaying standard dynamic range (SDR) images, but for high dynamic range (HDR) images, the above photoelectric conversion function cannot accurately display HDR images on the display device.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide an image encoding and decoding method and related devices, which can achieve accurate display of HDR images.
[0007] In a first aspect, an image encoding method is provided, comprising: encoding a target image into a bitstream; encoding a first identifier into the bitstream, wherein the first identifier indicates a photoelectric conversion function that matches the target image, and the target image includes a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image.
[0008] Optionally, the code stream includes an image header, and the first identifier is encoded into the image header.
[0009] Optionally, the method further includes: encoding a second identifier into the code stream, the second identifier indicating the chromaticity coordinates of each primary color of the target image; encoding a third identifier into the code stream, the third identifier indicating the pixel mapping range of the target image; and encoding a fourth identifier into the code stream, the fourth identifier indicating the coefficients of the conversion matrix used for color space conversion of the target image.
[0010] Optionally, before encoding the target image into a code stream, the method further includes: performing color space conversion on the target image based on the second identifier, the third identifier, and the fourth identifier.
[0011] Optionally, the code stream includes an image header, and the second identifier, the third identifier, and the fourth identifier are encoded into the image header.
[0012] Optionally, the method further includes: encoding a fifth identifier into the code stream, wherein the fifth identifier indicates an offset of sampling positions of the primary component and the secondary component of the target image.
[0013] Optionally, the target image includes an image of a primary component and an image of a secondary component after color space conversion; before encoding the target image into the bitstream, the method further includes: downsampling the image of the primary component and the image of the secondary component based on the fifth identifier; encoding the target image into the bitstream includes: encoding the downsampled image of the primary component and the downsampled image of the secondary component into the bitstream.
[0014] Optionally, the code stream includes an image header, and the fifth identifier is encoded into the image header.
[0015] Optionally, the method further includes: encoding static metadata of the target image into the code stream, the static metadata including source device information and image display information, the source device information indicating the color information of the main monitor that generates the target image, and the image display information indicating the display brightness of the target image.
[0016] Optionally, the source device information includes at least one of chromaticity coordinates of three primary colors of the main monitor, chromaticity coordinates of standard white light, maximum display brightness, and minimum display brightness.
[0017] Optionally, the image display information includes at least one of a maximum display brightness and a maximum average display brightness of the target image.
[0018] Optionally, the codestream includes an image header, and the image header carries the static metadata; or, the codestream includes a metadata header, and the metadata header carries the static metadata; or, the codestream includes a metadata sub-codestream corresponding to the metadata header, and the metadata sub-codestream carries the static metadata.
[0019] Optionally, the method further includes: encoding dynamic metadata of the target image into the bitstream, the dynamic metadata including color gamut volume change data of the target image.
[0020] Optionally, the codestream includes an image header, and the image header carries the dynamic metadata; or, the codestream includes a metadata header, and the metadata header carries the dynamic metadata; or, the codestream includes a metadata sub-codestream corresponding to the metadata header, and the metadata sub-codestream carries the dynamic metadata.
[0021] To sum up, when encoding the target image, the first identifier indicating the photoelectric conversion function that matches the target image can also be encoded into the code stream to simultaneously meet the display requirements of SDR images and HDR images at the decoding end, so that the decoder can use the corresponding photoelectric conversion function to render and display the reconstructed image after decoding the first identifier.
[0022] In a second aspect, an image decoding method is provided, comprising: parsing a first identifier from a bitstream, the first identifier indicating a photoelectric conversion function matching a target image, the target image including a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image; decoding the bitstream to obtain a reconstructed image of the target image; and performing photoelectric conversion on the reconstructed image based on the first identifier.
[0023] Optionally, the code stream includes an image header, and the first identifier is carried in the image header.
[0024] Optionally, the method further includes: parsing a second identifier from the code stream, the second identifier indicating the chromaticity coordinates of each primary color of the target image; parsing a third identifier from the code stream, the third identifier indicating the mapping range of the target image; and parsing a fourth identifier from the code stream, the fourth identifier indicating the coefficients of the conversion matrix used for color space conversion of the target image.
[0025] Optionally, before performing photoelectric conversion on the reconstructed image based on the first identifier, the method further includes: performing color space inverse conversion on the reconstructed image based on the second identifier, the third identifier and the fourth identifier.
[0026] Optionally, the code stream includes an image header, and the second identifier, the third identifier, and the fourth identifier are carried in the image header.
[0027] Optionally, the method further includes: parsing a fifth identifier from the code stream, where the fifth identifier indicates an offset of sampling positions of a primary component and a secondary component of the target image.
[0028] Optionally, parsing the code stream to obtain the reconstructed image of the target image includes: parsing the code stream to obtain the reconstructed image of the primary component and the reconstructed image of the secondary component of the target image; and upsampling the reconstructed image of the primary component and the reconstructed image of the secondary component based on the fifth identifier to obtain the reconstructed image.
[0029] Optionally, the code stream includes an image header, and the fifth identifier is carried in the image header.
[0030] Optionally, the method further includes: parsing static metadata of the target image from the code stream, the static metadata including source device information and image display information, the source device information indicating color information of a primary monitor that generates the target image, and the image display information indicating display brightness of the target image.
[0031] Optionally, the method further comprises: performing tone mapping on the reconstructed image based on the static metadata.
[0032] Optionally, the source device information includes at least one of chromaticity coordinates of three primary colors of the main monitor, chromaticity coordinates of standard white light, maximum display brightness, and minimum display brightness.
[0033] Optionally, the image display information includes at least one of a maximum display brightness and a maximum average display brightness of the target image.
[0034] Optionally, the codestream includes an image header, and the image header carries the static metadata; or, the codestream includes a metadata header, and the metadata header carries the static metadata; or, the codestream includes a metadata sub-codestream corresponding to the metadata header, and the metadata sub-codestream carries the static metadata.
[0035] Optionally, the method further includes: parsing dynamic metadata of the target image from the code stream, where the dynamic metadata includes color gamut volume change data of the target image.
[0036] Optionally, the method further includes: performing tone mapping on the reconstructed image based on the static metadata and the dynamic metadata.
[0037] Optionally, the codestream includes an image header, and the image header carries the dynamic metadata; or, the codestream includes a metadata header, and the metadata header carries the dynamic metadata; or, the codestream includes a metadata sub-codestream corresponding to the metadata header, and the metadata sub-codestream carries the dynamic metadata.
[0038] To sum up, when decoding the code stream, the photoelectric conversion function that matches the target image can be determined based on the decoded first identifier, and then the photoelectric conversion function can be used to accurately render and display the reconstructed image of the target image, so as to simultaneously meet the display requirements of SDR images and HDR images at the decoding end, thereby ensuring the accuracy when displaying the target image.
[0039] In a third aspect, a coding device is provided, wherein the coding device has the function of implementing the image coding method of the first aspect. The coding device includes at least one module, wherein the at least one module is used to implement the image coding method of the first aspect.
[0040] In a fourth aspect, a decoding device is provided, wherein the decoding device has the function of implementing the image decoding method in the second aspect. The decoding device includes at least one module, wherein the at least one module is used to implement the image decoding method provided in the second aspect.
[0041] In a fifth aspect, a coding device is provided, comprising: a processor, the processor being coupled to a memory, the memory being used to store programs or instructions, and when the program or instructions are executed by the processor, the coding device executes the image coding method described in the first aspect above.
[0042] In a sixth aspect, a decoding device is provided, comprising: a processor, the processor being coupled to a memory, the memory being used to store programs or instructions, and when the program or instructions are executed by the processor, the decoding device executes the image decoding method described in the second aspect above.
[0043] In a seventh aspect, a coding and decoding system is provided, which includes the coding device described in the fifth aspect and / or the decoding device described in the sixth aspect.
[0044] In an eighth aspect, a computer-readable storage medium is provided, comprising a program code, wherein when the program code is run on a computer, the computer executes the method described in the first aspect.
[0045] In a ninth aspect, a computer-readable storage medium is provided, comprising a program code, wherein when the program code is run on a computer, the computer executes the method described in the second aspect.
[0046] In a tenth aspect, a computer program product is provided, comprising instructions, which, when executed on a computer, cause the computer to execute the method described in the first aspect.
[0047] In an eleventh aspect, a computer program product is provided, comprising instructions, which, when executed on a computer, cause the computer to execute the method described in the second aspect.
[0048] In a twelfth aspect, a computer-readable storage medium is provided, on which a code stream obtained according to the method described in the first aspect is stored.
[0049] In the thirteenth aspect, a device for storing a code stream is provided, comprising at least one storage medium and a communication interface; the communication interface is used to receive or send a code stream; the at least one storage medium is used to store the code stream; the code stream is encoded by an encoder according to the image encoding method described in the first aspect above.
[0050] In a fourteenth aspect, a method for storing a code stream is provided, comprising: receiving a code stream through a communication interface; and storing the code stream in one or more storage media, wherein the code stream is encoded by an encoder according to the image encoding method described in the first aspect.
[0051] In a fifteenth aspect, a system for distributing code streams is provided, comprising at least one storage medium and a video streaming device; the at least one storage medium is used to store the code stream, which is encoded by an encoder according to the image encoding method described in the first aspect above; the video streaming device is used to respond to a request from a decoder so that the code stream in the at least one storage medium is sent to the decoder.
[0052] In a sixteenth aspect, a method for distributing a code stream is provided, comprising: receiving a first request; selecting a code stream from at least one storage medium in response to the first request; and sending the code stream to a destination device; the at least one storage medium is used to store the code stream, and the code stream is encoded by an encoder according to the image encoding method described in the first aspect above.
[0053] In a seventeenth aspect, a system for processing a code stream is provided, comprising an image source device, an encoder, one or more storage media, and a destination device; the image source device is used to provide image data; the encoder is used to obtain the image data of the image source device through an interface, and encode the image data to obtain one or more code streams, wherein the code streams are encoded by the encoder according to the image encoding method described in the first aspect above; the encoder is used to store the one or more code streams in one or more storage media; or the encoder is used to encapsulate the one or more code streams to obtain a transmission code stream; the encoder is used to transmit the transmission code stream to the destination device via a communication link or a communication network; the destination device is used to decapsulate the transmission code stream to obtain the one or more code streams; and the destination device is used to decode the one or more code streams to obtain decoded data.
[0054] The technical effects obtained in the above-mentioned second to seventeenth aspects are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] FIG1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0056] FIG2 is a schematic diagram of an AI coding model provided in an embodiment of the present application;
[0057] FIG3 is a schematic diagram of an AI decoding model provided in an embodiment of the present application;
[0058] FIG4 is a schematic diagram of a JPEG-AI encoding model provided in an embodiment of the present application;
[0059] FIG5 is a schematic diagram of a JPEG-AI decoding model provided in an embodiment of the present application;
[0060] FIG6 is a schematic diagram of a flow chart of an image encoding method provided in an embodiment of the present application;
[0061] FIG7 is a schematic diagram of a code stream structure provided in an embodiment of the present application;
[0062] FIG8 is a schematic diagram of chroma sampling positions provided in an embodiment of the present application;
[0063] FIG9 is a schematic diagram of another code stream structure provided in an embodiment of the present application;
[0064] FIG10 is a schematic flow chart of an image decoding method provided in an embodiment of the present application;
[0065] FIG11 is a schematic structural diagram of an encoding device provided in an embodiment of the present application;
[0066] FIG12 is a schematic structural diagram of a decoding device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0068] To facilitate understanding, before explaining in detail the image encoding method and image decoding method provided in the embodiments of the present application, the nouns, implementation environment and application scenarios involved in the embodiments of the present application are first introduced.
[0069] First, the terms involved in the embodiments of the present application are explained.
[0070] 1. High dynamic range (HDR)
[0071] Dynamic range is used in many fields to express the ratio of a variable's maximum to minimum values. In digital images, dynamic range represents the ratio between the maximum and minimum grayscale values within the image's displayable range. In most current color digital images, each R, G, and B channel uses one byte, or 8 bits, for storage. This means each channel represents a grayscale range of 0 to 255, with 0 to 255 representing the image's dynamic range. In the real world, a dynamic range of 10⁻³ to 10⁶ in the same scene is called high dynamic range (HDR); in contrast, the dynamic range of an ordinary image is low dynamic range (LDR). The imaging process of a digital camera is essentially a mapping of the real-world high dynamic range to the low dynamic range of a photograph.
[0072] 2. Standard dynamic range photoelectric conversion function
[0073] Standard dynamic range images are the opposite of high dynamic range images. Traditionally used 8-bit images in formats like JPEG can be considered standard dynamic range images. Before the advent of cameras capable of capturing HDR images, traditional cameras could only record captured light information within a certain range by controlling the exposure value. Because the maximum illumination information of a display device cannot match the brightness information in the real world, and we view images through display devices, a photoelectric transfer function is required. Early display devices, such as CRT monitors, used a gamma function. This gamma-based photoelectric transfer function is defined in the BT.1886 standard published by the International Telecommunication Union-Radiocommunications Sector (ITU-R). Using this transfer function for tone mapping and then quantizing it to 8 bits creates a traditional SDR image. In other words, SDR images and the aforementioned photoelectric transfer function perform well on traditional display devices (with an illumination of approximately 100 cd / m²).
[0074] 3. HDR Photoelectric Transfer Function
[0075] However, with the upgrade of display devices, the illumination range of display devices continues to increase. The existing consumer-grade HDR display is 600cd / m2, and the illumination information of high-end HDR display can reach 2000cd / m2, which far exceeds the illumination information of SDR display devices. The photoelectric conversion function in the BT.1886 standard cannot well express the display performance of HDR display devices. Therefore, an improved photoelectric transfer function is needed to adapt to the upgrade of display devices. The idea of the photoelectric transfer function comes from the mapping function in the tone mapping (TM) algorithm. Appropriate adjustment of the mapping function is the photoelectric transfer function (OETF).
[0076] At present, there are three common photoelectric conversion functions: perception quantization (PQ) photoelectric conversion function, hybrid log-gamma (HLG) photoelectric conversion function and scene luminance fidelity (SLF) photoelectric conversion function. These three photoelectric conversion functions are the conversion functions specified by the audio video coding standard (AVS).
[0077] Unlike the traditional "Gamma" transfer function, the PQ photoelectric transfer function is a perceptual quantization transfer function proposed based on the brightness perception model of the human eye. The PQ photoelectric transfer function represents the conversion relationship between the linear signal value of an image pixel and the nonlinear signal value in the PQ domain. The HLG photoelectric transfer function is improved based on the traditional Gamma curve. The HLG photoelectric transfer function applies the traditional Gamma curve in the low segment and supplements the log curve in the high segment. The HLG photoelectric transfer function represents the conversion relationship between the linear signal value of an image pixel and the nonlinear signal value in the HLG domain. The SLF photoelectric transfer function is based on the premise of meeting the optical characteristics of the human eye. The SLF photoelectric transfer curve represents the conversion relationship between the linear signal value of an image pixel and the nonlinear signal value in the SLF domain.
[0078] 4. Dynamic Range Mapping in Display
[0079] The dynamic range mapping method is mainly used in the adaptation of the front-end HDR signal and the back-end HDR terminal display device. For example, the front-end collects a 4000nit light signal, while the back-end HDR terminal display device (such as a TV) has an HDR display capability of only 500nit. Therefore, how to map the 4000nit signal to the 500nit device is a high-to-low tone-mapping process. Another example is that the front-end collects a 100nit SDR signal, while the display end is a 2000nit TV signal. If the 100nit signal is better displayed on the 2000nit device, it is another low-to-high tone-mapping process.
[0080] Among them, this dynamic range mapping method can be divided into static and dynamic. The static mapping method is to use a single data to perform an overall tone mapping process based on the same video content or the same hard disk content, that is, the processing curve is usually the same. The advantage of this method is that it carries less information and has a relatively simple processing flow; the disadvantage of this method is that it uses the same curve for tone mapping for each scene, which can lead to information loss in some scenes. For example, if the curve focuses on protecting bright areas, some details will be lost in some extremely dark scenes, affecting the experience. The dynamic mapping method dynamically adjusts the content based on specific areas, each scene, or each frame. The advantage of this method is that different curves are processed according to specific areas, each scene, or each frame, which will produce better processing results. However, each frame or each scene must carry relevant scene information, and the amount of information carried is relatively large.
[0081] 5. ITU-T T.35 Terminal Identification Code Allocation
[0082] ITU-T T.35 (referred to as T.35 in this application) is the T.35 specification published by the ITU-T, "Procedure for the Allocation of ITU-Defined Non-Standard Extension Codes." It consists of three components: the country code, the terminal provider code, and the terminal provider-oriented code. The country code is defined in Appendices A and B of ITU-T T.35. The terminal provider code is allocated by the ITU administration. Each terminal provider that has been assigned a provider code by the administration administers its own terminal provider-oriented code.
[0083] T.35 can be used in various contexts, such as SEI messages of video codecs, where the data carried by T.35 messages may include HDR dynamic metadata, HDR static metadata, HDR display mapping information, etc. However, the information transmitted through T.35 is not limited to video and can be information of any nature.
[0084] 6. Metadata
[0085] Recording key information about the image in a video, scene, or frame, such as the average, maximum, and minimum values within a scene. HDR static metadata, also known as fixed metadata, uses the same metadata throughout the entire video to control the color and detail of each frame. In contrast, HDR dynamic metadata, also known as variable metadata, uses different metadata for each frame in the video to control color and detail.
[0086] Secondly, the implementation environment involved in the embodiments of this application is introduced.
[0087] Please refer to Figure 1, which is a schematic diagram of an implementation environment provided by an embodiment of the present application. The implementation environment includes a source device 10, a destination device 20, a link 30, and a storage device 40. The source device 10 can generate encoded images. Therefore, the source device 10 can also be referred to as an image encoding device or encoding end. The destination device 20 can decode the encoded images generated by the source device 10. Therefore, the destination device 20 can also be referred to as an image decoding device or decoding end. The link 30 can receive the encoded images generated by the source device 10 and transmit the encoded images to the destination device 20. The storage device 40 can receive the encoded images generated by the source device 10 and store them. Under such conditions, the destination device 20 can directly obtain the encoded images from the storage device 40. Alternatively, the storage device 40 can correspond to a file server or another intermediate storage device that can store the encoded images generated by the source device 10. Under such conditions, the destination device 20 can stream or download the encoded images stored by the storage device 40.
[0088] The source device 10 and the destination device 20 may each include one or more processors and a memory coupled to the one or more processors, wherein the memory may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium that can be used to store desired program code in the form of computer-accessible instructions or data structures. For example, the source device 10 and the destination device 20 may each include a mobile phone, a smartphone, a personal digital assistant (PDA), a wearable device, a pocket PC (PPC), a tablet computer, a smart car computer, a smart TV, a smart speaker, a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video game console, an in-vehicle computer, or the like.
[0089] Link 30 may include one or more media or devices capable of transmitting encoded images from source device 10 to destination device 20. In one possible implementation, link 30 may include one or more communication media that enable source device 10 to send encoded images directly to destination device 20 in real time. In an embodiment of the present application, source device 10 may modulate the encoded images based on a communication standard, such as a wireless communication protocol, and may transmit the modulated images to destination device 20. The one or more communication media may include wireless and / or wired communication media, such as radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, such as a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from source device 10 to destination device 20, although this embodiment of the present application does not specifically limit this.
[0090] In one possible implementation, the storage device 40 may store the received encoded image sent by the source device 10, and the destination device 20 may directly obtain the encoded image from the storage device 40. Under such conditions, the storage device 40 may include any of a variety of distributed or locally accessible data storage media, for example, any of the various distributed or locally accessible data storage media may be a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded images.
[0091] In one possible implementation, storage device 40 may correspond to a file server or another intermediate storage device that can store the encoded images generated by source device 10, and destination device 20 may stream or download the images stored on storage device 40. The file server may be any type of server capable of storing and transmitting the encoded images to destination device 20. In one possible implementation, the file server may include a network server, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Destination device 20 may obtain the encoded images via any standard data connection, including an internet connection. Any standard data connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of the two suitable for obtaining the encoded images stored on the file server. The transmission of the encoded images from storage device 40 may be streaming, downloading, or a combination of the two.
[0092] The implementation environment shown in FIG1 is only one possible implementation method, and the technology of the embodiment of the present application is applicable not only to the source device 10 that can encode images and the destination device 20 that can decode encoded images shown in FIG1 , but also to other devices that can encode images and decode encoded images, and the embodiment of the present application does not make specific limitations on this.
[0093] In the implementation shown in FIG1 , source device 10 includes a data source 120, an encoder 100, and an output interface 140. In some embodiments, output interface 140 may include a modem and / or a transmitter, where the transmitter may also be referred to as a transmitter. Data source 120 may include an image capture device (e.g., a camera), an archive containing previously captured images, a feed interface for receiving images from an image content provider, and / or a computer graphics system for generating images, or a combination of these sources of images.
[0094] The data source 120 may send an image to the encoder 100, and the encoder 100 may encode the image received from the data source 120 to generate an encoded image. The encoder may send the encoded image to an output interface. In some embodiments, the source device 10 directly sends the encoded image to the destination device 20 via the output interface 140. In other embodiments, the encoded image may also be stored on the storage device 40 for later retrieval by the destination device 20 for decoding and / or display.
[0095] In the implementation environment shown in FIG1 , the destination device 20 includes an input interface 240, a decoder 200, and a display device 220. In some embodiments, the input interface 240 includes a receiver and / or a modem. The input interface 240 may receive encoded images via the link 30 and / or from the storage device 40, and then transmit the encoded images to the decoder 200. The decoder 200 may decode the received encoded images to obtain decoded images. The decoder may transmit the decoded images to the display device 220. The display device 220 may be integrated with the destination device 20 or may be external to the destination device 20. Generally, the display device 220 displays the decoded images. The display device 220 may be any of a variety of types of display devices, for example, the display device 220 may be a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0096] Although not shown in FIG1 , in some aspects, encoder 100 and decoder 200 can be integrated with an encoder and decoder, respectively, and can include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software for encoding both audio and video in a common data stream or in separate data streams. In some embodiments, the MUX-DEMUX units can conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP), if applicable.
[0097] The encoder 100 and the decoder 200 can each be any of the following circuits: one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology of the embodiments of the present application is implemented in part by software, the device can store instructions for the software in a suitable non-volatile computer-readable storage medium, and can use one or more processors to execute the instructions in hardware to implement the technology of the embodiments of the present application. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) can be regarded as one or more processors. Each of the encoder 100 and the decoder 200 can be included in one or more encoders or decoders, and any of the encoders or decoders can be integrated as part of a combined encoder / decoder (encoder / decoder) in the corresponding device.
[0098] Embodiments of the present application may generally refer to encoder 100 as "signaling" or "sending" certain information to another device, such as decoder 200. The terms "signaling" or "sending" may generally refer to the transmission of syntax elements and / or other data used to decode a compressed image. This transmission may occur in real time or near real time. Alternatively, this communication may occur over time, such as when syntax elements are stored in the encoded bitstream to a computer-readable storage medium during encoding, and a decoding device may then retrieve the syntax elements at any time after they are stored to this medium.
[0099] It should be noted that the application scenarios and implementation environments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of application scenarios and implementation environments, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0100] Finally, the application scenarios involved in the embodiments of this application are introduced.
[0101] Coding and decoding technology refers to a technology that uses data characteristics such as spatial redundancy, visual redundancy and statistical redundancy to represent the original signal data with fewer bits, either losslessly or losslessly. It can achieve effective transmission and storage of information such as images, videos and audio, and plays an important role in the current media era where the types and amounts of information transmitted or stored are increasing.
[0102] Taking image compression as an example, there are two types of compression: lossy and lossless. Lossy compression achieves a higher compression ratio at the expense of a certain degree of image quality, while lossless compression does not cause loss of image detail. Traditional image / video compression algorithms have undergone decades of development, resulting in mature compression standards such as High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC). These traditional image / video compression algorithms have been widely used due to their comprehensive support, versatility, and robust hardware compatibility. However, they have also encountered bottlenecks in improving coding efficiency.
[0103] As deep learning surpasses traditional algorithms in many computer vision tasks, such as image recognition and object detection, a growing number of researchers are exploring deep learning-based image / video compression methods. Unlike traditional image algorithms, which manually optimize image encoding, quantization, and entropy coding separately, AI image compression algorithms optimize each module (encoding network, entropy estimation network, decoding network, etc.) as a whole. This results in superior compression performance. However, deep learning is a resource-intensive algorithm, requiring both significant computational cost and memory consumption. Despite increasing computing resources, optimizing the training and inference of deep learning networks remains crucial for model implementation. In particular, as more models are moving from servers to resource-constrained devices like smartphones and embedded systems, deploying complex models on resource-constrained devices remains a crucial challenge for current deep learning technologies.
[0104] Deep learning-based image compression methods typically involve AI compression models and JPEG-AI compression models. To facilitate understanding, this article briefly introduces the encoding and decoding processes of these two compression models.
[0105] (1) AI compression model
[0106] Please refer to Figure 2, which is a schematic diagram of an AI coding model. The AI coding model is also called an encoder and is applied to the encoding end. During encoding, the image x to be encoded is input into the encoding network to obtain a feature map y, and the feature map y is input into the super-coding network to obtain a super-prior feature z. The super-prior feature z is entropy encoded according to the specified probability distribution to encode the super-prior feature z into the bitstream. In addition, the super-prior feature z is input into the super-decoding network to obtain the prior feature p and the variance σ. Based on the prior feature p and the feature map y, the mean μ is obtained through the context network, and then based on the probability distribution N(μ, σ), the feature map y is entropy encoded to encode the feature map y into the bitstream.
[0107] Among them, the bit sequence obtained by entropy coding the super prior feature z is a partial bit sequence included in the code stream, which can be called a super prior bit stream, denoted as bs z The bit sequence obtained by entropy coding the feature map y is a partial bit sequence included in the code stream. This partial bit sequence can be called the image bit stream, denoted as bs y .
[0108] In some embodiments, the context network includes a context model and a probability distribution estimation model. In this case, the feature map y can be input into the context model to obtain the context feature φ. In combination with the prior feature p and the context feature φ, the mean μ is estimated by the probability distribution estimation model. Of course, this is just an example. In actual applications, it can also be implemented in other ways, and the embodiments of the present application do not limit this. Moreover, the above is introduced using entropy coding as an example. In actual applications, it can also be encoded by other encoding methods, and the embodiments of the present application do not limit this.
[0109] Please refer to Figure 3, which is a schematic diagram of an AI decoding model. The AI decoding model is also called a decoder. It is applied to the decoding end. When decoding, it calculates the super-prior bit stream bs included in the bit stream according to the specified probability distribution. z Perform entropy decoding to obtain super-prior features The super prior feature Input to the super decoding network to obtain prior features And variance σ. Determine the side information from the decoded partial image features and combine the prior features And side information, through the context network to get the mean μ, based on the probability distribution N(μ,σ), from the image bit stream bs included in the code stream y The feature map is decoded by medium entropy The feature map Input to the decoding network to obtain the reconstructed image
[0110] In some embodiments, the context network includes a context model and a probability distribution estimation model. In this case, based on the prior features The probability distribution estimation model can be used to estimate the mean μ corresponding to some image features. Based on the probability distribution N(μ,σ) of these image features, the image bit stream bs included in the code stream is obtained. y For other image features, the side information is determined from the decoded image features and input into the context model to obtain the context features. Combining prior features and contextual features The mean μ corresponding to other image features is estimated using a probability distribution estimation model. Of course, this is just an example, and in actual applications, it can be implemented in other ways, which are not limited in the present embodiment. Moreover, the above description uses entropy decoding as an example, and in actual applications, it can also be decoded using other decoding methods corresponding to the encoding method, which are also not limited in the present embodiment.
[0111] It should be noted that if the above-mentioned probability distribution estimation network uses a Gaussian model (such as a single Gaussian model or a mixed Gaussian model) for modeling, the estimated probability distribution includes a mean and a variance. If the above-mentioned probability distribution estimation network uses a Laplace distribution model for modeling, the estimated probability distribution includes a location parameter and a scale parameter. If the above-mentioned probability distribution estimation network uses a logistic distribution model for modeling, the estimated probability distribution includes a mean and a scale parameter. The above is explained using the Gaussian model as an example.
[0112] In addition, the superdecoding network includes a superdecoding prediction network and a superscale decoding network. The superdecoding prediction network is used to determine the prior features, and the superscale decoding network is used to determine the variance. That is, the super prior features are input into the superdecoding prediction network to obtain the prior features, and the super prior features are input into the superscale decoding network to obtain the variance.
[0113] (2) JPEG-AI compression model
[0114] The JPEG-AI compression model is based on an autoencoder structure with color component separation, and realizes the encoding and decoding of the image by encoding and decoding each separated component in the image to be encoded. Moreover, the basic structure of the JPEG-AI compression model is the same as that of the above-mentioned AI compression model, but the image input components of the JPEG-AI compression model are YUV components, and the Y component and UV component are encoded and decoded separately. Among them, when encoding, the Y component is encoded alone using the architecture shown in Figure 2 above, and the UV component is encoded by combining the data of the Y component. When decoding, the Y component is decoded alone using the architecture shown in Figure 3 above, and the UV component is decoded by combining the data of the Y component.
[0115] Please refer to Figure 4, which is a schematic diagram of a JPEG-AI encoding model. The JPEG-AI encoding model is also called an encoder. It is applied to the encoding end. During encoding, the image x in RGB format (a color mode where R represents red, G represents green, and B represents blue) is converted into an image in YUV format (a color encoding method where Y represents brightness, i.e., grayscale value, and U and V represent chrominance) through the color space to obtain the Y component image x Y and UV components of the image x UV , the Y component of the image x Y Input to the encoding network to obtain the feature map y of the Y component Y , the feature map y of the Y component Y Input to the super encoding network to obtain the super prior feature z of the Y component Y , according to the specified probability distribution of the Y component of the hyper-prior feature z Y Perform entropy coding to convert the super-prior feature z of the Y component into Y The super prior feature z of the Y component is encoded into the bitstream. Y Input to the super decoding network to obtain the prior feature p of the Y component Y and variance σ Y , based on the prior feature p Y and variance σ Y , get the mean μ through the context network Y , and then based on the probability distribution N(μ Y ,σ Y ), the feature map y of the Y component Y Perform entropy coding to transform the feature map y of the Y component Y Encode into the code stream.
[0116] In addition, the Y component of the image x Y Auxiliary image x' converted to UV component after downsampling UV , the UV component of the image x UVAnd the auxiliary image x' of the UV component UV Input to the encoding network to obtain the feature map y of the UV component UV , the UV component feature map y UV Input to the super encoding network to obtain the super prior feature z of the UV component UV , according to the specified probability distribution of the UV component hyper-prior feature z UV Entropy coding is performed to convert the ultra-prior feature z of the UV component into UV Encode into the bitstream. The ultra-prior feature z of the UV component UV Input to the super decoding network to obtain the prior feature p of the UV component UV and variance σ UV , based on the prior feature p of the UV component UV and feature map y UV , the mean value μ of the UV component is obtained through the context network UV , and then based on the probability distribution N(μ UV ,σ UV ), the characteristic map y of the UV component UV Perform entropy coding to convert the feature map y of the UV component UV Encode into the code stream.
[0117] Among them, the super prior feature z of the Y component Y The bit sequence obtained by entropy coding is a partial bit sequence included in the code stream, which can be called the super-prior bit stream bs of the Y component. zY . Feature map y of Y component Y The bit sequence obtained by entropy coding is a partial bit sequence included in the code stream, which can be called the image bit stream bs of the Y component. yY The hyper-prior feature z of the UV component UV The bit sequence obtained by entropy coding is a partial bit sequence included in the code stream, which can be called the super priori bit stream bs of the UV component. zUV . Characteristic map y of UV component UV The bit sequence obtained by entropy coding is a partial bit sequence included in the code stream, which can be called the image bit stream bs of the UV component. yUV .
[0118] Please refer to Figure 5, which is a schematic diagram of a JPEG-AI decoding model. The JPEG-AI decoding model is also called a decoder. It is applied to the decoding end. When decoding, according to the specified probability distribution, the super-prior bit stream bs of the Y component included in the code stream is calculated. zY Perform entropy decoding to obtain the super-prior features of the Y component The hyper-prior feature of the Y component Input to the super decoding network to obtain the prior features of the Y component and variance σ Y Determine the side information from the decoded partial image features and combine it with the prior features of the Y component And side information, the mean μ of the Y component is obtained through the context network Y , based on the probability distribution N(μ Y ,σ Y ), entropy decode the feature map of the Y component from the image bit stream of the Y component included in the code stream The feature map of the Y component Input to the decoding network to obtain the reconstructed Y component image
[0119] In addition, the feature map y of the Y component Y Auxiliary feature y' converted to UV component after downsampling UV Then, according to the specified probability distribution, the super-prior bit stream bs of the UV component included in the code stream is zUV Perform entropy decoding to obtain the ultra-prior features of the UV component The ultra-prior features of the UV component Input to the super decoding network to obtain the prior features of the UV component and variance σ UV Determine the side information from the decoded partial image features and combine them with the prior features of the UV component And side information, the mean μ of the UV component is obtained through the context network UV , based on the probability distribution of UV components N(μ UV ,σ UV ), entropy decode the feature map of UV component from the image bit stream of UV component included in the code stream The characteristic map of the UV component and the auxiliary feature y′ of the UV component UV Input to the decoding network to obtain the reconstructed UV component image
[0120] Afterwards, the reconstructed Y component image is filtered through the Inter Channel Correlation Information (ICCI) network. and UV component of the image Processing is performed to obtain a reconstructed image in YUV format, and finally the reconstructed image in YUV format is converted into RGB format through inverse color space conversion to obtain a reconstructed image
[0121] The above introduces the general process of encoding and decoding of the JPEG-AI compression model. For some details, please refer to the above explanation of the AI compression model, which will not be repeated here.
[0122] That is, before encoding the image, the JPEG-AI compression model needs to perform color space conversion on the target image to be encoded and pre-process the downsampling of the component images. After decoding and reconstructing the reconstructed component image of the target image, the reconstructed component image is upsampled and the reconstructed image is inversely converted in color space. However, when the JPEG AI decoding model performs inverse color space conversion, the coefficients of the conversion matrix used are applicable to SDR images, so it is impossible to reconstruct an accurate HDR image; moreover, when the reconstructed component image is upsampled, the offset of the sampling position of the chrominance information relative to the luminance information is different from the downsampling at the decoding end, and an accurate HDR image cannot be reconstructed. In addition, for the reconstructed image, the display device at the decoding end usually uses the photoelectric conversion function corresponding to SDR for tone mapping, and cannot accurately map and display the reconstructed HDR image.
[0123] Based on this, in order to realize the encoding, decoding, transmission and display of HDR images, the present application provides an image encoding method and an image decoding method, so as to add some image information and HDR metadata information that can guide the accurate reconstruction and display of HDR images during encoding, so that the decoding end can combine these image information to reconstruct an accurate image; moreover, combined with the HDR metadata information, the HDR image can be effectively displayed.
[0124] Referring to FIG6 , the present application provides an image encoding method, which is applied to an encoder and includes the following steps.
[0125] Step 601: Encode the target image into a bitstream.
[0126] Among them, the specific implementation method of encoding the target image into the bitstream can be found in the relevant description of the image encoding and decoding process of the above-mentioned JPEG AI compression model, which will not be repeated here.
[0127] Step 602: Encode a first identifier into the bitstream, wherein the first identifier indicates a photoelectric conversion function that matches the target image, and the target image includes an HDR image and / or an SDR image.
[0128] The photoelectric conversion function is used to convert a digital signal into an analog signal that can drive the physical output of the display so that the image can be correctly displayed on the display. When the target image is an HDR image, the first identifier indicates the photoelectric conversion function corresponding to the HDR image, such as the PQ photoelectric conversion function, the HLG photoelectric conversion function, etc.; when the target image is an SDR image, the first identifier indicates the photoelectric conversion function corresponding to the SDR image, such as the photoelectric transfer function based on the "Gamma" function.
[0129] In some embodiments, when the first identifier is encoded into the code stream, the first identifier can be encoded into the code stream in the form of an 8-bit unsigned integer, that is, the above-mentioned first identifier corresponds to a value of 0-255; of course, the first identifier can also be encoded into the code stream in other ways, and the embodiments of the present application do not limit this.
[0130] When the first identifier is encoded into the bitstream in the form of an 8-bit unsigned integer, the following Table 1 provides an exemplary corresponding relationship, so as to determine the photoelectric conversion function indicated by the first identifier based on the corresponding relationship.
[0131] Table 1
[0132] As can be seen from Table 1 above, when the target image is an SDR image, its first identifier can be 13, and its corresponding photoelectric conversion function is the photoelectric conversion function defined in the IEC IEC61966-2-1 standard, which supports photoelectric conversion of SDR images. When the target image is an HDR image, the first identifier encoded into the bitstream is 16 or 18, 16 indicates the PQ photoelectric conversion function, and 18 indicates the SLG photoelectric conversion function. The specific function expression can be found in the middle column of Table 1, L o is the linear signal value before conversion, and its value is normalized to [0, 1], L c is the linear signal value before conversion, with a value range of [0, 12], and V is the nonlinear signal value after conversion, with a value range of [0, 1]. The corresponding photoelectric conversion function is also documented and described in the relevant standards and can be found in the standard documents listed in the column to the right of Table 1.
[0133] It should be noted that the above Table 1 is only an example. In specific applications, the photoelectric conversion function corresponding to the identification value in the above table can also be adjusted according to the relevant standards of image / video encoding and decoding. The embodiment of the present application does not limit this.
[0134] See Figure 7, which shows a schematic diagram of the codestream structure defined by JPEG AI. The codestream structure includes a start of codestream marker (SOC), a picture header marker (PIH) and the corresponding picture header, a tools header marker (TOH) and the corresponding (codec) tool header, a start of quality map marker (SOQ) and the corresponding Q-stream, multiple picture codestreams, and an end of codestream marker (EOC).
[0135] Among them, multiple image code streams include a coding stream of super prior features (Z-stream), a coding stream of primary component residuals (RY-stream), and a coding stream of secondary component residuals (RUV-stream). Each image code stream corresponds to a start marker. The coding stream of super prior features follows the start of z-stream marker (SOZ), the coding stream of primary component residuals follows the start of residual stream for primary component marker (SORP), and the coding stream of secondary component residuals follows the start of residual stream for secondary component marker (SORS).
[0136] Based on this code stream structure, the implementation process of step 602 may be: encoding the first identifier into the above-mentioned picture header.
[0137] It can be seen from this that no matter whether the encoded target image is an SDR image or an HDR image, the photoelectric conversion function that matches the target image can be determined, and the first identifier indicating the photoelectric conversion function can also be encoded into the bitstream during encoding, so that the decoding end can use the matching photoelectric conversion function to display the reconstructed image corresponding to the target image based on the first identifier.
[0138] In some embodiments, before the encoder encodes the target image, if the target image is in RGB format, the target image needs to be converted into a color space to convert the RGB format image into a YUV format image. However, the color space conversion of JPEG AI only supports two modes: the preset BT.709 conversion matrix and the user-defined conversion matrix. Among them, BT.709 is the default color space for SDR images. However, in other RGB and YUV conversion standards of ITU-R, such as ITU-R BT.2020, ITU-R BT.2100, etc., HDR TV standards are specified and the color space specified by the BT.2020 standard is inherited and used. Therefore, when encoding and decoding HDR images, the encoder and decoder are required to support BT.2020 color space (consistent with the BT.2100 color space) conversion.
[0139] Based on this, when the decoding end encodes the target image, it can also encode the information related to the color space conversion matrix used by the encoding end to perform color space conversion on the target image into the bitstream to guide the decoding end to convert the reconstructed image in YUV format into a reconstructed image in RGB format based on the information related to the color space conversion matrix.
[0140] In one possible implementation, the image encoding method provided by an embodiment of the present application further includes: encoding a second identifier into a code stream, the second identifier indicating the chromaticity coordinates of each primary color of the target image; encoding a third identifier into the code stream, the third identifier indicating the pixel mapping range of the target image; and encoding a fourth identifier into the code stream, the fourth identifier indicating the coefficients of the conversion matrix used for color space conversion of the target image.
[0141] In this case, the encoder performs color space conversion on the target image based on the second identifier, the third identifier, and the fourth identifier before encoding the target image into the code stream.
[0142] In other words, the encoder uses the corresponding color space conversion matrix according to the type of the target image (i.e., HDR image or SDR image), performs color space conversion on the target image, and then encodes the converted target image into the bitstream; at the same time, an identifier related to the color space conversion matrix, i.e., a second identifier, a third identifier, and a fourth identifier, is encoded into the bitstream.
[0143] Similarly, in the case where the code stream includes a picture header, the second identifier, the third identifier, and the fourth identifier may also be encoded into the picture header.
[0144] In some embodiments, the second identifier used to indicate the chromaticity coordinates of each primary color of the target image can also be encoded into the bitstream in the form of an 8-bit unsigned integer, corresponding to values 0 to 255. Taking the CIE 1931 chromaticity specified in ISO / CIE 11664-1, which includes x and y coordinates, as an example, Table 2 below provides a method for determining the corresponding chromaticity coordinates based on the second identifier. That is, Table 2 provides an exemplary correspondence for finding the chromaticity coordinates corresponding to red, green, blue, and standard white light based on the second identifier.
[0145] Table 2
[0146] As shown in Table 2 above, when the target image is an HDR image, the second identifier encoded into the bitstream is 9, the green coordinate corresponding to chromaticity x is 0.170, and the green coordinate corresponding to chromaticity y is 0.797; the blue coordinate corresponding to chromaticity x is 0.131, and the blue coordinate corresponding to chromaticity y is 0.046; the red coordinate corresponding to chromaticity x is 0.708, and the red coordinate corresponding to chromaticity y is 0.292; the standard white light coordinate corresponding to chromaticity x is 0.3127, and the standard white light coordinate corresponding to chromaticity y is 0.3290. The relevant information on the above HDR images and chromaticity primaries is covered in the ITU-R BT.2020-2 and ITU-R BT.2100-2 standards and can also be found in the ITU-R BT.2020-2 and ITU-R BT.2100-2 standard documents.
[0147] In some embodiments, the third flag for indicating the pixel mapping range of the target image can be 0 or 1. When the input signal is an RGB image, the third flag is 1 to indicate that all pixels in the image are mapped in the full range; when the input signal is a luminance and chrominance image (such as a YUV format image), if the black level and range of the luminance and chrominance signals are not specified, the third flag can be 0 to indicate that all pixels in the image can be mapped within a limited range.
[0148] Optionally, for an image in YUV format, the identification value of the third identification may be recorded as 1 to perform full-range mapping on all pixels in the image, and this embodiment of the present application does not impose any restrictions on this.
[0149] In some embodiments, when a fourth identifier indicating the coefficients of the conversion matrix used to perform color space conversion on a target image is encoded into the codestream, the fourth identifier may also be encoded into the codestream in the form of a 1-bit unsigned integer, corresponding to values 0 to 255. Taking the CIE 1931 chromaticity specified in ISO / CIE 11664-1, including x and y coordinates, as an example, Table 3 below provides an exemplary correspondence relationship for determining the fourth identifier and the coefficients of the conversion matrix based on this correspondence relationship.
[0150] Table 3
[0151] As can be seen from Table 3 above, when the target image is an HDR image, the fourth identifier encoded into the bitstream is 9, and the corresponding matrix coefficient is K R is 0.2627, K B The matrix coefficients are also mentioned in the ITU-R BT.2020-2 standard and the ITU-R BT.2100-2 standard, and can also be found in the ITU-R BT.2020-2 standard and the ITU-R BT.2100-2 standard documents.
[0152] In an embodiment of the present application, the chromaticity coordinates of each primary color of the target image, the pixel mapping range of the target image, and the coefficients of the conversion matrix used for color space conversion of the target image are jointly determined to determine the conversion matrix used when performing color space conversion on the target image. Therefore, when the encoder performs color space conversion on the target image, the relevant information indicating the above-mentioned conversion matrix, that is, the second identifier, the third identifier and the fourth identifier, can also be encoded into the bitstream, so that the decoder can perform color space conversion processing on the reconstructed YUV format HDR image based on the second identifier, the third identifier and the fourth identifier.
[0153] In some embodiments, after the encoder converts the target image in RGB format into the target image in YUV format through color space conversion, it can downsample the Y component image and the UV component image separately to reduce the amount of data processing. However, under different sampling modes, the sampling positions of the Y component image and the UV component image will be offset. After the decoder decodes the reconstructed image of the Y component and the reconstructed image of the UV component, when it upsamples the reconstructed image, if the sampling position offset is different from that of the encoder, it will also affect the accuracy of the reconstructed image.
[0154] Based on this, the encoder can also encode a fifth flag into the bitstream, indicating the offset between the sampling positions of the primary and secondary components of the target image. In other words, after color space conversion, the target image includes a primary and secondary component image. Before encoding the target image into the bitstream, the encoder can downsample the primary and secondary components based on the fifth flag and then encode the downsampled primary and secondary components into the bitstream.
[0155] For an image in YUV format, the primary component may be a Y component, and the secondary component may be a UV component.
[0156] Taking the 420 sampling mode as an example, refer to the chroma sampling position diagram shown in Figure 8. The luminance (Y) sample area in the upper left corner of the image is represented by a black solid line frame, and the chroma (UV) sample area in the upper left corner of the image is represented by a black solid line frame; among them, the center point of the luminance sample area is the luminance sampling point, and the center point of the chroma sample area is the chroma sampling point. Figure 8 shows a total of six offset situations (a)-(f).
[0157] For the six offset conditions listed in Figure 8, the following Table 4 lists the offset conditions between the chroma sampling position and the luma sampling position in the 420 sampling mode. That is, Table 4 lists the correspondence between the fifth identifier and the chroma sampling position offset conditions, and corresponds to the schematic diagram of Figure 8, lists the horizontal and vertical offset values of the chroma sampling position relative to the luma sampling position in each offset condition.
[0158] Table 4
[0159] It should be understood that when the identification value of the fifth identifier is 0, corresponding to (a) shown in Figure 8, the sampling position of the chroma sampling point is offset by 0 in the horizontal direction and 0.5 in the vertical direction compared to the luminance sampling point; when the identification value of the fifth identifier is 1, corresponding to (b) shown in Figure 8, the sampling position of the chroma sampling point is offset by 0.5 in the horizontal direction and 0.5 in the vertical direction compared to the luminance sampling point; when the identification value of the fifth identifier is 2, corresponding to (c) shown in Figure 8, the sampling position of the chroma sampling point is offset by 0 in the horizontal direction and 0.5 in the vertical direction compared to the luminance sampling point. The shift value is 0; when the identification value of the fifth identifier is 3, corresponding to (d) shown in Figure 8, the sampling position of the chroma sampling point is offset by 0.5 in the horizontal direction and 0 in the vertical direction compared with the luminance sampling point; when the identification value of the fifth identifier is 4, corresponding to (e) shown in Figure 8, the sampling position of the chroma sampling point is offset by 0 in the horizontal direction and 1 in the vertical direction compared with the luminance sampling point; when the identification value of the fifth identifier is 5, corresponding to (f) shown in Figure 8, the sampling position of the chroma sampling point is offset by 0.5 in the horizontal direction and 1 in the vertical direction compared with the luminance sampling point.
[0160] As can be seen from Figure 8 and Table 4 above, when the fifth flag is 1, the chroma sampling point is located in the middle row and column of the luma sampling point. At this time, if the pixel coordinates of the luma sampling point Y are (x, y), then the pixel coordinates of the corresponding chroma sampling points U and V will be (x / 2, y / 2). When the fifth flag is 2, the chroma sampling point is located in the upper left corner of the luma sampling point. At this time, if the pixel coordinates of the luma sampling point Y are (x, y), then the pixel coordinates of the corresponding chroma sampling points U and V will be (floor(x / 2), floor(y / 2)).
[0161] It should be noted that in this 420 sampling mode, the sampling density of luma samples is lower than that of chroma to reduce data transmission and storage requirements. In addition, in the 420 sampling mode, the sampling density of luma samples in the horizontal direction is full resolution, while the sampling density of chroma samples is only half of that of luma. In the vertical direction, both luma and chroma samples are sampled at half of the full resolution.
[0162] It should be understood that the above example is only illustrated using the 420 sampling mode. For other sampling modes such as the 422 sampling mode, the same method can be used to encode the identifier indicating the offset of the sampling positions of the primary component and the secondary component under the sampling mode into the bitstream. The embodiments of the present application do not limit this.
[0163] Similarly, when the code stream includes a picture header, the fifth identifier can also be encoded into the picture header.
[0164] In an embodiment of the present application, after the target image is converted into a color space, if the image of the primary component and the image of the secondary component are downsampled, a fifth identifier indicating the offset of the sampling positions of the primary component and the secondary component of the target image can also be encoded into the bitstream, so that the decoder can perform upsampling processing on the reconstructed image of the Y component and the reconstructed image of the UV component based on the fifth identifier.
[0165] In some embodiments, to accurately render and display a target image, the image encoding method provided in embodiments of the present application further includes encoding static metadata of the target image into a bitstream, where the static metadata includes source device information and image display information. The source device information indicates the color information of the primary monitor that generated the target image, and the image display information indicates the display brightness of the target image.
[0166] As an example, the source device information includes at least one of the chromaticity coordinates of the three primary colors of the master monitor, the chromaticity coordinates of standard white light, the maximum display brightness, and the minimum display brightness. The chromaticity coordinates of the three primary colors include the three primary color x-coordinates and the three primary color y-coordinates, and the chromaticity coordinates of standard white light include the standard white light x-coordinate and the standard white light y-coordinate.
[0167] Optionally, each item in the source device information may be represented in the form of a 16-bit unsigned integer, which is not limited in this embodiment of the present application.
[0168] As an example, the image display information includes at least one of a maximum display brightness and a maximum average display brightness of the target image.
[0169] The maximum display brightness is the maximum brightness that a single pixel in the target image can achieve. This maximum display brightness can be expressed as a 6-bit unsigned integer, with 1 cd / m² as the unit, and the value range is 1 cd / m² to 65535 cd / m². When the target image is any frame in an image sequence or video, the average brightness of all pixels in each frame of the image sequence or video is first calculated. The maximum value of the average brightness values corresponding to each frame is then used as the maximum average display brightness.
[0170] Optionally, when there is only one independent image frame, the above-mentioned maximum average display brightness can be a default value, and this embodiment of the present application does not limit this.
[0171] As shown in Figure 7, the code stream includes an image header, which carries static metadata. That is, the first identifier to the fifth identifier and the static metadata can all be encoded into the image header.
[0172] In some embodiments, when the target image is a frame in a certain image sequence or a video, in order to accurately render and display the target image, the image encoding method provided in the embodiment of the present application also includes: encoding the dynamic metadata of the target image into the bitstream, and the dynamic metadata includes the color gamut volume change data of the target image.
[0173] The color gamut volume change data is used to describe data related to changes in content and scenes in an image over time or in the number of frames. The color gamut volume change data includes color change information and brightness change information of the target image.
[0174] As shown in Figure 7, the code stream includes an image header, which carries static metadata. That is, the first identifier to the fifth identifier, as well as the static metadata and / or dynamic metadata can all be encoded into the image header.
[0175] In some embodiments, when the embodiments of the present application encode static metadata and / or dynamic metadata of the target image into the bitstream, as shown in Figure 9, a metadata header can also be added to the bitstream to carry the static metadata and / or dynamic metadata of the target image through the metadata header.
[0176] In some embodiments, when the codestream includes a metadata sub-codestream corresponding to the metadata header, the metadata sub-codestream may also be used to carry static metadata and dynamic metadata of the target image.
[0177] It should be understood that the embodiments of the present application only limit the data information of the target image that needs to be included in the codestream, such as the first identifier to the fifth identifier, and the static metadata and / or dynamic metadata, and do not limit the specific location of this data information when it is included in the codestream. That is, the image header, metadata header, and metadata sub-codestream in the above examples can all carry relevant data information. Of course, all of the above data information can also be carried only in the image header, metadata header, or metadata sub-codestream, and the embodiments of the present application do not impose any restrictions on this.
[0178] Based on the above embodiment, to implement a method for carrying the first through fifth identifiers, as well as static and dynamic metadata of the target image, in the picture header of a bitstream, this embodiment of the present application adds four syntax elements to the picture header of the bitstream to carry image information and metadata related to the target image. Table 5 provides an example of syntax elements for the picture header.
[0179] Table 5
[0180] Among them, picture_header_size is used to define the spatial size of the image header, which defines the relevant information of the image, such as img_width is the width of the encoded image; img_height is the height of the encoded image; picture_format is the image format, which is used to define the encoded image as a YUV image with a 4:4:4 sampling format, a YUV image with a 4:2:2 sampling format, a YUV image with a 4:2:0 sampling format, or an RGB image; bit_depth is the quantization information; model_header() is the syntax related to the encoding model. The specific meanings of these parameters can be found in the relevant definitions of high-level syntax in the encoding file header in the JPEG AI standard, and the embodiments of this application will not be repeated here.
[0181] In the above-mentioned image header, coding_independent_code_points() is a syntax newly added in this application for carrying image characteristics, mastering_display_color_volume() is a syntax newly added in this application for carrying source device information, content_light_level_information() is a syntax newly added in this application for carrying image display information, and dynamic_metadata() is a syntax newly added in this application embodiment for carrying dynamic metadata of the target image.
[0182] It should be noted that in the syntax element table shown in the embodiment of the present application, the Descriptor column on the right corresponds to the data type and bit information of the variable defined by the syntax. For example, U(16) refers to a variable that is a 16-bit unsigned integer.
[0183] As an example, the syntax elements of the syntax coding_independent_code_points() of the image characteristics are shown in Table 6.
[0184] Table 6
[0185] In Table 6, colour_primaries is the byte corresponding to the second identifier indicating the chromaticity coordinates of each primary color of the target image, matrix_coefficients is the byte corresponding to the fourth identifier indicating the coefficients of the conversion matrix used for color space conversion of the target image, image_full_range_flag is the byte corresponding to the third identifier indicating the pixel mapping range of the target image, transfer_characteristics is the byte corresponding to the first identifier indicating the photoelectric conversion function matching the target image, and chroma420_sample_loc_type is the byte corresponding to the fifth identifier indicating the offset of the sampling positions of the primary component and the secondary component of the target image.
[0186] As an example, the syntax elements of source device information mastering_display_color_volume() are shown in Table 7.
[0187] Table 7
[0188] In Table 7, mastering_display_colour_primaries_x[i] and mastering_display_colour_primaries_y[i] are the bytes corresponding to the chromaticity coordinates of the three primary colors of the main monitor, mastering_display_white_point_chromaticity_x and mastering_display_white_point_chromaticity_y are the bytes corresponding to the chromaticity coordinates of the standard white light of the main monitor, mastering_display_maximum_luminance is the byte corresponding to the maximum display brightness of the main monitor, and mastering_display_minimum_luminance is the minimum display brightness of the main monitor.
[0189] As an example, the syntax elements of the image display information content_light_level_information() are shown in Table 8.
[0190] Table 8
[0191] In Table 8, maximum_content_light_level is the byte corresponding to the maximum display brightness of the target image, and maximum_frame_average_light_level is the byte corresponding to the maximum average display brightness of the target image.
[0192] As an example, the syntax elements of dynamic metadata dynamic_metadata() are shown in Table 9.
[0193] Table 9
[0194] In Table 9 above, itu_t_t35_country_code is a byte with the value specified as the country code by ITU-T T.35 Annex A, itu_t_t35_country_code_extension_byte shall be a byte with the value specified as the country code by ITU-T T.35 Annex B, itu_t_t35_payload_byte shall be a byte containing data registered in accordance with ITU-T T.35, the ITU-T T.35 terminal provider code and the code for the terminal provider (terminal provider targeting code) shall be contained in the first byte or bytes of itu_t_t35_payload_byte, and the format shall be specified by the administration issuing the terminal provider code. Any remaining itu_t_t35_payload_byte data shall be data with the syntax and semantics specified by the entity identified by the ITU-T T.35 country code and terminal provider code.
[0195] It should be understood that the above Table 5 is used as an example, and all the data information newly added in the embodiment of the present application is carried in the image header. In specific applications, the above Table 5 can also be modified to carry part of the data in the image header. The embodiment of the present application does not limit this.
[0196] Based on the above embodiment, the embodiment of the present application can also carry static metadata and dynamic metadata of the target image in the metadata header of the code stream. As an example, the following Table 10 provides a syntax element of the metadata header.
[0197] Table 10
[0198] Among them, MDH is the metadata header mark, the metadata header mark is the metadata header, and metadata_header_size is the size of the metadata header.
[0199] As an example, the syntax element table of static metadata is shown in Table 11 below.
[0200] Table 11
[0201] For specific contents of the syntax elements of mastering_display_color_volume(), please refer to Table 7 above, and for syntax elements of content_light_level_information(), please refer to Table 8 above, which will not be repeated here.
[0202] To sum up, when the encoder encodes the target image, it can also encode the first identifier indicating the photoelectric conversion function that matches the target image into the bitstream to simultaneously meet the display requirements of SDR images and HDR images at the decoding end, so that the decoder can use the corresponding photoelectric conversion function to render and display the reconstructed image after decoding the first identifier.
[0203] Next, the image decoding method provided in the embodiment of the present application is introduced.
[0204] Please refer to Figure 10. The present application provides an image decoding method, which is applied to a decoder and includes the following steps.
[0205] Step 1001: Parse a first identifier from a bitstream, where the first identifier indicates a photoelectric conversion function that matches a target image, where the target image includes an HDR image and / or an SDR image.
[0206] The code stream includes an image header, and the first identifier is carried in the image header.
[0207] Optionally, the code stream includes an image header, and the second identifier, the third identifier, and the fourth identifier are carried in the image header. In this case, the second identifier, the third identifier, and the fourth identifier may also be parsed from the code stream.
[0208] The second identifier indicates the chromaticity coordinates of each primary color of the target image; the third identifier indicates the mapping range of the target image; and the fourth identifier indicates the coefficients of the conversion matrix used for color space conversion of the target image.
[0209] Optionally, the code stream includes an image header, and the fifth identifier is carried in the image header. In this case, the fifth identifier can also be parsed from the code stream.
[0210] Optionally, the code stream includes an image header, and the static metadata and / or dynamic metadata of the target image are also included in the image header. In this case, the static metadata and / or dynamic metadata of the target image can be parsed from the code stream.
[0211] Step 1002: Decode the code stream to obtain a reconstructed image of the target image.
[0212] When the fifth identifier is parsed from the bitstream, the implementation process of step 1002 may be: parsing the bitstream to obtain a reconstructed image of the primary component and a reconstructed image of the secondary component of the target image; then, based on the fifth identifier, upsampling the reconstructed image of the primary component and the reconstructed image of the secondary component to obtain a reconstructed image.
[0213] Optionally, when the second identifier, the third identifier and the fourth identifier are parsed from the code stream, before executing the above step 1003, the decoder determines the conversion matrix used by the encoding end when performing color space conversion on the target image by looking up the above Tables 1-3 based on the second identifier, the third identifier and the fourth identifier, and thereby performs inverse color space conversion on the reconstructed image based on the conversion matrix.
[0214] Step 1003: Based on the first identifier, perform photoelectric conversion on the reconstructed image.
[0215] In some embodiments, when static metadata of a target image is parsed from a code stream, tone mapping may be performed on the reconstructed image based on the static metadata to display the reconstructed image.
[0216] In some embodiments, after parsing the dynamic metadata of the target image from the bitstream, tone mapping may be performed on the reconstructed image based on the dynamic metadata to display the reconstructed image.
[0217] It should be noted that when the above decoder executes the image decoding method, the specific meaning and execution order of the relevant technical features can be referred to the embodiment of the above encoder executing the image encoding method, and will not be repeated here.
[0218] To sum up, when decoding the code stream, the decoder can determine the photoelectric conversion function that matches the target image based on the decoded first identifier, and then use the photoelectric conversion function to accurately render and display the reconstructed image of the target image, so as to simultaneously meet the display requirements of SDR images and HDR images at the decoding end, and ensure the accuracy when displaying the target image.
[0219] FIG11 is a schematic diagram of the structure of an encoding device provided in an embodiment of the present application, and the encoding device may be the encoder shown in FIG1 . Referring to FIG11 , the encoding device 1100 includes: an image encoding module 1101 and an information encoding module 402 .
[0220] The image encoding module 1101 is used to encode the target image into a code stream;
[0221] The information encoding module 1102 is configured to encode a first identifier into the bitstream, wherein the first identifier indicates a photoelectric conversion function that matches the target image, and the target image includes a high dynamic range HDR image and / or a standard dynamic range SDR image.
[0222] Optionally, the code stream includes an image header, and the first identifier is encoded into the image header.
[0223] Optionally, the information encoding module 1102 is further configured to:
[0224] encoding a second identifier into the bitstream, wherein the second identifier indicates the chromaticity coordinates of each primary color of the target image;
[0225] encoding a third identifier into the bitstream, wherein the third identifier indicates a pixel mapping range of the target image;
[0226] A fourth identifier is encoded into the code stream, where the fourth identifier indicates coefficients of a conversion matrix used for performing color space conversion on the target image.
[0227] Optionally, the device further comprises:
[0228] An image conversion module is configured to perform color space conversion on the target image based on the second identifier, the third identifier, and the fourth identifier.
[0229] Optionally, the code stream includes an image header, and the second identifier, the third identifier, and the fourth identifier are encoded into the image header.
[0230] Optionally, the information encoding module 1102 is further configured to:
[0231] A fifth identifier is encoded into the code stream, where the fifth identifier indicates the offset of the sampling positions of the primary component and the secondary component of the target image.
[0232] Optionally, the target image includes a primary component image and a secondary component image after color space conversion; and the device further includes:
[0233] a downsampling module, configured to downsample the image of the primary component and the image of the secondary component based on the fifth identifier;
[0234] The image encoding module 1101 is specifically used for:
[0235] The downsampled image of the primary component and the downsampled image of the secondary component are encoded into the code stream.
[0236] Optionally, the code stream includes an image header, and the fifth identifier is encoded into the image header.
[0237] Optionally, the information encoding module 1102 is further configured to:
[0238] Static metadata of the target image is encoded into the bitstream, the static metadata including source device information and image display information, the source device information indicating color information of a main monitor generating the target image, and the image display information indicating display brightness of the target image.
[0239] Optionally, the source device information includes at least one of chromaticity coordinates of three primary colors of the main monitor, chromaticity coordinates of standard white light, maximum display brightness, and minimum display brightness.
[0240] Optionally, the image display information includes at least one of a maximum display brightness and a maximum average display brightness of the target image.
[0241] Optionally, the codestream includes an image header, and the image header carries the static metadata; or, the codestream includes a metadata header, and the metadata header carries the static metadata; or, the codestream includes a metadata sub-codestream corresponding to the metadata header, and the metadata sub-codestream carries the static metadata.
[0242] Optionally, the information encoding module 1102 is further configured to:
[0243] The dynamic metadata of the target image is encoded into the bitstream, where the dynamic metadata includes color gamut volume change data of the target image.
[0244] Optionally, the codestream includes an image header, and the image header carries the dynamic metadata; or, the codestream includes a metadata header, and the metadata header carries the dynamic metadata; or, the codestream includes a metadata sub-codestream corresponding to the metadata header, and the metadata sub-codestream carries the dynamic metadata.
[0245] In an embodiment of the present application, when encoding a target image, the encoding device may also encode a first identifier indicating a photoelectric conversion function that matches the target image into the bitstream to simultaneously meet the display requirements of SDR images and HDR images at the decoding end, so that the decoder can use the corresponding photoelectric conversion function to render and display the reconstructed image after decoding the first identifier.
[0246] It should be noted that the encoding device provided in the above embodiments, when performing image encoding, is illustrated only by the division of the above-described functional modules. In actual applications, the above-described functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the encoding device provided in the above embodiments and the image encoding method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0247] FIG12 is a schematic diagram of the structure of a decoding device provided in an embodiment of the present application, which may be the decoder shown in FIG1. Referring to FIG12, the device includes: an information decoding module 1201, an image decoding module 1202, and an image conversion module 1203.
[0248] An information decoding module 1201 is configured to parse a first identifier from a bitstream, where the first identifier indicates a photoelectric conversion function that matches a target image, where the target image includes a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image;
[0249] An image decoding module 1202 is configured to decode the code stream to obtain a reconstructed image of the target image;
[0250] The image conversion module 1203 is configured to perform photoelectric conversion on the reconstructed image based on the first identifier.
[0251] Optionally, the code stream includes an image header, and the first identifier is carried in the image header.
[0252] Optionally, the information decoding module 1201 is further configured to:
[0253] Parsing a second identifier from the code stream, where the second identifier indicates the chromaticity coordinates of each primary color of the target image;
[0254] Parsing a third identifier from the code stream, where the third identifier indicates a mapping range of the target image;
[0255] A fourth identifier is parsed from the code stream, where the fourth identifier indicates coefficients of a conversion matrix used for performing color space conversion on the target image.
[0256] Optionally, the device further comprises:
[0257] An image inverse conversion module is configured to perform color space inverse conversion on the reconstructed image based on the second identifier, the third identifier, and the fourth identifier.
[0258] Optionally, the code stream includes an image header, and the second identifier, the third identifier, and the fourth identifier are carried in the image header.
[0259] Optionally, the information decoding module 1201 is further configured to:
[0260] A fifth identifier is parsed from the code stream, where the fifth identifier indicates an offset of sampling positions of a primary component and a secondary component of the target image.
[0261] Optionally, the image decoding module 1202 is specifically configured to:
[0262] Parsing the code stream to obtain a reconstructed image of a primary component and a reconstructed image of a secondary component of the target image;
[0263] Based on the fifth identifier, up-sampling is performed on the reconstructed image of the main component and the reconstructed image of the secondary component to obtain the reconstructed image.
[0264] Optionally, the code stream includes an image header, and the fifth identifier is carried in the image header.
[0265] Optionally, the information decoding module 1201 is further configured to:
[0266] Static metadata of the target image is parsed from the code stream, the static metadata including source device information and image display information, the source device information indicating color information of a primary monitor generating the target image, and the image display information indicating display brightness of the target image.
[0267] Optionally, the device further comprises:
[0268] A tone mapping module is configured to perform tone mapping on the reconstructed image based on the static metadata.
[0269] Optionally, the source device information includes at least one of chromaticity coordinates of three primary colors of the main monitor, chromaticity coordinates of standard white light, maximum display brightness, and minimum display brightness.
[0270] Optionally, the image display information includes at least one of a maximum display brightness and a maximum average display brightness of the target image.
[0271] Optionally, the codestream includes an image header, and the image header carries the static metadata; or, the codestream includes a metadata header, and the metadata header carries the static metadata; or, the codestream includes a metadata sub-codestream corresponding to the metadata header, and the metadata sub-codestream carries the static metadata.
[0272] Optionally, the information decoding module 1201 is further configured to:
[0273] Dynamic metadata of the target image is parsed from the code stream, where the dynamic metadata includes color gamut volume change data of the target image.
[0274] Optionally, the tone mapping module is further configured to:
[0275] The reconstructed image is tone mapped based on the static metadata and the dynamic metadata.
[0276] Optionally, the codestream includes an image header, and the image header carries the dynamic metadata; or, the codestream includes a metadata header, and the metadata header carries the dynamic metadata; or, the codestream includes a metadata sub-codestream corresponding to the metadata header, and the metadata sub-codestream carries the dynamic metadata.
[0277] In an embodiment of the present application, when decoding the code stream, the decoding device can determine the photoelectric conversion function that matches the target image based on the decoded first identifier, and then use the photoelectric conversion function to accurately render and display the reconstructed image of the target image, so as to simultaneously meet the display requirements of SDR images and HDR images at the decoding end, thereby ensuring the accuracy when displaying the target image.
[0278] It should be noted that the decoding device provided in the above embodiments, when performing image decoding, is merely illustrated by the division of the aforementioned functional modules. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, i.e., the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the decoding device provided in the above embodiments and the image decoding method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be further described here.
[0279] An embodiment of the present application also provides an encoding device, which includes: a processor, the processor is coupled to a memory, the memory is used to store programs or instructions, and when the program or instructions are executed by the processor, the encoding device performs the above-mentioned image encoding method.
[0280] An embodiment of the present application also provides a decoding device, which includes: a processor, the processor is coupled to a memory, the memory is used to store programs or instructions, and when the program or instructions are executed by the processor, the decoding device performs the above-mentioned image decoding method.
[0281] An embodiment of the present application further provides a coding and decoding system, which includes the coding device and / or the decoding device.
[0282] An embodiment of the present application further provides a computer-readable storage medium, comprising a program code, wherein when the program code is executed on a computer, the computer is enabled to execute the above-mentioned image encoding method.
[0283] An embodiment of the present application further provides a computer-readable storage medium, comprising a program code, which, when executed on a computer, enables the computer to execute the above-mentioned image decoding method.
[0284] An embodiment of the present application further provides a computer program product, comprising instructions, which, when executed on a computer, enable the computer to execute the above-mentioned image encoding method.
[0285] An embodiment of the present application further provides a computer program product, comprising instructions, which, when executed on a computer, enable the computer to execute the above-mentioned image decoding method.
[0286] An embodiment of the present application further provides a computer-readable storage medium, on which a code stream obtained according to the above-mentioned image encoding method is stored.
[0287] An embodiment of the present application also provides a device for storing a code stream, comprising at least one storage medium and a communication interface; the communication interface is used to receive or send a code stream; the at least one storage medium is used to store the code stream; the code stream is encoded by an encoder according to the above-mentioned image encoding method.
[0288] An embodiment of the present application further provides a method for storing a code stream, comprising: receiving a code stream through a communication interface; and storing the code stream in one or more storage media, wherein the code stream is encoded by an encoder according to the above-mentioned image encoding method.
[0289] An embodiment of the present application also provides a system for distributing code streams, comprising at least one storage medium and a video streaming device; the at least one storage medium is used to store the code stream, which is encoded by an encoder according to the above-mentioned image encoding method; the video streaming device is used to respond to a request from a decoder so that the code stream in the at least one storage medium can be sent to the decoder.
[0290] An embodiment of the present application also provides a method for distributing a code stream, comprising: receiving a first request; selecting a code stream from at least one storage medium in response to the first request; and sending the code stream to a destination device; the at least one storage medium is used to store the code stream, wherein the code stream is encoded by an encoder according to the above-mentioned image encoding method.
[0291] An embodiment of the present application also provides a system for processing a code stream, including an image source device, an encoder, one or more storage media, and a destination device; the image source device is used to provide image data; the encoder is used to obtain the image data of the image source device through an interface, and encode the image data to obtain one or more code streams, wherein the code streams are encoded by the encoder according to the above-mentioned image encoding method; the encoder is used to store the one or more code streams in one or more storage media; or the encoder is used to encapsulate the one or more code streams to obtain a transmission code stream; the encoder is used to transmit the transmission code stream to the destination device via a communication link or a communication network; the destination device is used to decapsulate the transmission code stream to obtain the one or more code streams; and the destination device is used to decode the one or more code streams to obtain decoded data.
[0292] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.
[0293] It should be understood that the "plurality" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0294] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0295] The above description is an embodiment provided for this application and is not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. An image encoding method, characterized in that, The method includes: Encoding a target image into a bitstream; Encoding a first identifier into the bitstream, where the first identifier indicates a photoelectric conversion function matching the target image, and the target image includes a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image.
2. The method according to claim 1, wherein The bitstream includes an image header, and the first identifier is encoded into the image header.
3. The method according to claim 1 or 2, characterized in that The method further includes: Encoding a second identifier into the bitstream, where the second identifier indicates the chromaticity coordinates of each primary color of the target image; Encoding a third identifier into the bitstream, where the third identifier indicates the pixel mapping range of the target image; Encoding a fourth identifier into the bitstream, where the fourth identifier indicates the coefficients of a conversion matrix used for color space conversion of the target image.
4. The method according to claim 3, characterized in that Before encoding the target image into the bitstream, the method further includes: Performing color space conversion on the target image based on the second identifier, the third identifier, and the fourth identifier.
5. The method according to claim 3 or 4, characterized in that, The bitstream includes an image header, and the second identifier, the third identifier, and the fourth identifier are encoded into the image header.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Encoding a fifth identifier into the bitstream, where the fifth identifier indicates the offset of the sampling positions of the main component and the sub-component of the target image.
7. The method according to claim 6, characterized in that, After color space conversion, the target image includes an image of the main component and an image of the sub-component; before encoding the target image into the bitstream, the method further includes: Performing downsampling on the image of the main component and the image of the sub-component based on the fifth identifier. Encoding the target image into the bitstream includes: Encoding the downsampled image of the main component and the downsampled image of the sub-component into the bitstream.
8. The method according to claim 6 or 7, characterized in that, The bitstream includes an image header, and the fifth identifier is encoded into the image header.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: Encoding static metadata of the target image into the bitstream, where the static metadata includes source device information and image display information, the source device information indicates the color information of the main monitor that generates the target image, and the image display information indicates the display brightness of the target image.
10. The method according to claim 9, wherein The source device information includes at least one of the chromaticity coordinates of the three primary colors of the main monitor, the chromaticity coordinates of the standard white light, the maximum display brightness, and the minimum display brightness.
11. The method according to claim 9 or 10, characterized in that, The image display information includes at least one of the maximum display brightness and the maximum average display brightness of the target image.
12. The method according to any one of claims 9 to 11, characterized in that, The bitstream includes an image header that carries the static metadata; or, the bitstream includes a metadata header that carries the static metadata; or, the bitstream includes a metadata sub-bitstream corresponding to the metadata header, and the metadata sub-bitstream carries the static metadata.
13. The method according to any one of claims 9-12, characterized in that, The method further includes: Encoding dynamic metadata of the target image into the bitstream, where the dynamic metadata includes data on the change in the gamut volume of the target image.
14. The method according to claim 13, wherein The bitstream includes an image header that carries the dynamic metadata; alternatively, the bitstream includes a metadata header that carries the dynamic metadata; alternatively, the bitstream includes a metadata sub-bitstream corresponding to the metadata header, and the metadata sub-bitstream carries the dynamic metadata.
15. An image decoding method, characterized in that The method includes: Parsing a first identifier from the bitstream, where the first identifier indicates a photoelectric conversion function that matches a target image, and the target image includes a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image; Decoding the bitstream to obtain a reconstructed image of the target image; Performing photoelectric conversion on the reconstructed image based on the first identifier.
16. The method according to claim 15, characterized in that, The bitstream includes an image header, and the first identifier is carried in the image header.
17. The method according to claim 15 or 16, characterized in that, The method further includes: Parsing a second identifier from the bitstream, where the second identifier indicates the chromaticity coordinates of each primary color of the target image; Parsing a third identifier from the bitstream, where the third identifier indicates the mapping range of the target image; Parsing a fourth identifier from the bitstream, where the fourth identifier indicates the coefficients of a conversion matrix used for color space conversion of the target image.
18. The method according to claim 17, wherein Before performing photoelectric conversion on the reconstructed image based on the first identifier, the method further includes: Performing inverse color space conversion on the reconstructed image based on the second identifier, the third identifier, and the fourth identifier.
19. The method according to claim 17 or 18, characterized in that, The bitstream includes an image header, and the second identifier, the third identifier, and the fourth identifier are carried in the image header.
20. The method according to any one of claims 15-19, characterized in that, The method further includes: Parsing a fifth identifier from the bitstream, where the fifth identifier indicates the offset situation of the sampling positions of the main component and the sub-component of the target image.
21. The method according to claim 20, wherein The parsing the bitstream to obtain a reconstructed image of the target image includes: Parsing the bitstream to obtain a reconstructed image of the main component and a reconstructed image of the sub-component of the target image; Performing upsampling on the reconstructed image of the main component and the reconstructed image of the sub-component based on the fifth identifier to obtain the reconstructed image.
22. The method according to any one of claims 20 or 21, characterized in that The bitstream includes an image header, and the fifth identifier is carried in the image header.
23. The method according to any one of claims 15-22, characterized in that, The method further includes: Parsing the static metadata of the target image from the bitstream, where the static metadata includes source device information and image display information, the source device information indicates the color information of the main monitor that generates the target image, and the image display information indicates the display brightness of the target image.
24. The method according to claim 23, characterized in that, The method further includes: Performing tone mapping on the reconstructed image based on the static metadata.
25. The method according to claim 23 or 24, characterized in that, The source device information includes at least one of the chromaticity coordinates of the three primary colors of the main monitor, the chromaticity coordinates of the standard white light, the maximum display brightness, and the minimum display brightness.
26. The method according to any one of claims 23-25, characterized in that, The image display information includes at least one of the maximum display brightness and the maximum average display brightness of the target image.
27. The method according to any one of claims 23-26, characterized in that, The bitstream includes an image header that carries the static metadata; alternatively, the bitstream includes a metadata header that carries the static metadata; alternatively, the bitstream includes a metadata sub-bitstream corresponding to the metadata header, and the metadata sub-bitstream carries the static metadata.
28. The method according to any one of claims 23-27, characterized in that, The method further includes: Parse the dynamic metadata of the target image from the bitstream, where the dynamic metadata includes the gamut volume change data of the target image.
29. The method according to claim 28, wherein The method further includes: Perform tone mapping on the reconstructed image based on the static metadata and the dynamic metadata.
30. The method according to claim 29, wherein The bitstream includes an image header that carries the dynamic metadata; alternatively, the bitstream includes a metadata header that carries the dynamic metadata; alternatively, the bitstream includes a metadata sub-bitstream corresponding to the metadata header, and the metadata sub-bitstream carries the dynamic metadata.
31. An encoding device, characterized in that, The encoding device includes: An image encoding module for encoding a target image into a bitstream; An information encoding module for encoding a first identifier into the bitstream, where the first identifier indicates a photoelectric conversion function matching the target image, and the target image includes a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image.
32. A decoding device, characterized in that, The decoding device includes: An information decoding module for parsing a first identifier from the bitstream, where the first identifier indicates a photoelectric conversion function matching the target image, and the target image includes a high dynamic range (HDR) image and / or a standard dynamic range (SDR) image; An image decoding module for decoding the bitstream to obtain a reconstructed image of the target image; An image conversion module for performing photoelectric conversion on the reconstructed image based on the first identifier.
33. A coding device, characterized in that, The encoding device includes: a processor, where the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the encoding device executes the method according to any one of claims 1 to 14.
34. A decoding device, characterized in that, The decoding device includes: a processor, where the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the decoding device executes the method according to any one of claims 15 to 30.
35. A codec system, characterized in that, The encoding / decoding system includes the encoding device according to claim 33, and / or the decoding device according to claim 34.
36. A computer-readable storage medium, characterized in that, Includes program code that, when running on a computer, causes the computer to execute the image encoding method according to any one of claims 1 to 14.
37. A computer-readable storage medium, characterized in that, Includes program code that, when running on a computer, causes the computer to execute the image decoding method according to any one of claims 15 to 30.
38. A computer program product, characterized in that, Includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1 to 14.
39. A computer program product, characterized in that, Includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 15 to 30.
40. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a bitstream obtained by executing the method according to any one of claims 1 - 14 by one or more processors.
41. A device for storing a bitstream, characterized in that, Includes at least one storage medium and a communication interface; The communication interface is used to receive or send a bitstream; The at least one storage medium is used to store the bitstream; The bitstream is encoded by an encoder according to the image encoding method according to any one of claims 1 to 14.
42. A method for storing a bitstream, characterized in that, Includes: Receive a bitstream through a communication interface; Store the bitstream in one or more storage media, where the bitstream is Encoded by an encoder according to the image encoding method of any one of claims 1 to 14.
43. A system for distributing a bitstream, characterized in that, Comprising at least one storage medium and a video stream device; The at least one storage medium is used to store a bitstream, where the bitstream is encoded by an encoder according to the image encoding method of any one of claims 1 to 14; The video stream device is configured to, in response to a request from a decoder, cause the bitstream in the at least one storage medium to be sent to the decoder.
44. A method for distributing a bitstream, characterized in that, Comprising: Receive a first request; In response to the first request, select a bitstream from at least one storage medium; Send the target bitstream to the destination device; The at least one storage medium is used to store a bitstream, where the bitstream is encoded by an encoder according to the image encoding method of any one of claims 1 to 14.
45. A system for processing a bitstream, characterized in that, Comprising an image source device, an encoder device, one or more storage media, and a destination device; The image source device is configured to provide image data; The encoder device is configured to obtain the image data of the image source device through an interface and encode the image data to obtain one or more bitstreams, where the bitstreams are encoded by the encoder according to the image encoding method of any one of claims 1 to 14; The encoder device is configured to store the one or more bitstreams in one or more storage media; Or The encoder device is configured to encapsulate the one or more bitstreams to obtain a transport bitstream; The encoder device is configured to transmit the transport bitstream to the destination device through a communication link or a communication network; The destination device is configured to de-encapsulate the transport bitstream to obtain the one or more bitstreams; The destination device is configured to decode the one or more bitstreams to obtain decoded data.
Citation Information
Patent Citations
Method for improving viewing experience of high dynamic range video
CN106657714A
Signal reshaping and coding for HDR and wide color gamut signals
CN107852511A
High dynamic range video self-adaptive preprocessing method
CN110933416A
End-to-end implementation method for high dynamic range video
CN111669532A
HDR video processing method, encoding device and decoding device
CN113630563A