Tone mapping method and apparatus
Patent Information
- Application Number
- PCT/CN2025/147177
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2025-12-30
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025147177_01102026_PF_FP_ABST
Abstract
Description
A tone mapping method and apparatus
[0001] This application claims priority to Chinese Patent Application No. 202510359912.6, filed with the State Intellectual Property Office of China on March 24, 2025, entitled "A Tone Mapping Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of image processing, and more particularly to a tone mapping method and apparatus. Background Technology
[0003] A core function of an imaging system is image display. Good display quality can realistically reproduce the original scene, giving the viewer the same feeling as directly observing the original scene—a sense of immersion or realism. Tone mapping is an important component of an image reproduction system; it maps the lighting of the original scene to the luminous intensity of the display device.
[0004] In end-to-end imaging systems for high dynamic range (HDR) content, the conversion between electrical signals processed by the machine and light signals perceptible to the human eye is involved. Tone mapping is typically implemented after the electrical signals are converted into light signals to adapt to the dynamic range of the display device. For example, the dynamic range of the image signal acquired by the front end is mapped to the dynamic range of the display device through tone mapping, thus adapting the front end image signal to the display device.
[0005] However, current tone mapping schemes are only suitable for fixed display states of display devices (i.e., fixed dynamic range), and do not take into account dynamic changes in the dynamic range of display devices (such as image scaling). This results in images exceeding the human eye's perception capacity when display states change, leading to abnormal display effects (such as unclear display, banding, or other issues). Summary of the Invention
[0006] This application provides a tone mapping method and apparatus that enables images to be displayed normally under various display conditions, giving users a better and more realistic perceptual experience.
[0007] Firstly, a tone mapping method is provided, which can be applied to an encoding device. The method specifically includes: performing tone mapping on a first image according to mapping parameters to obtain a second image; if the second image does not meet human visual perception requirements, adjusting the mapping parameters and re-performing tone mapping on the first image according to the adjusted mapping parameters; if the second image meets human visual perception requirements, generating a bitstream of the first image, the bitstream including the first image with an initial bit width and metadata of the mapped second image. Here, human visual perception refers to the range of image quality perceived by the human eye; the mapping parameters include the bit width of the first image and / or the metadata of the first image.
[0008] The solution provided in this application adjusts the mapping parameters used for tone mapping to ensure that the tone-mapped image meets the perceptual capabilities of the human eye. This ensures that the mapped image can be perceived by the human eye under various display conditions, achieving the goal of normal image display in all situations. This provides users with a better and more realistic perceptual experience across different display environments, increasing the immersive feeling of the displayed image.
[0009] One possible implementation is that the mapping parameter is the bit width of the first image. If the second image obtained by mapping from the first image with the first bit width satisfies human visual perception, the method may further include: generating a bitstream of difference information. This difference information is used to indicate the difference between the first image with the initial bit width and the first image with the first bit width. In this way, the decoding device can adjust the first image to the first bit width based on the difference information, thereby obtaining a mapped image that satisfies human visual perception for display.
[0010] Another possible implementation is to use the metadata of the first image as the mapping parameter. If the second image obtained by performing tone mapping on the first image according to the first metadata satisfies human visual perception, then the metadata of the mapped second image is the first metadata.
[0011] Another possible implementation is that the aforementioned difference information includes: a first image with an initial bit width and a mask image of the first image with a first bit width.
[0012] Another possible implementation is that the aforementioned difference information includes metadata extracted from the mask image of the first image with the initial bit width and the first image with the first bit width.
[0013] Another possible implementation involves the second image not meeting human visual perception requirements, including: the first quality index of the second image relative to the first image with the initial bit width does not meet human visual perception requirements. The first quality index reflects the degree of detail retention of the second image relative to the first image with the initial bit width. Using the first quality index, it is possible to intuitively and quickly determine whether the second image meets human visual perception requirements from a fidelity perspective.
[0014] Another possible implementation is that the first quality metric is the peak signal-to-noise ratio (PSNR), which does not meet the human eye's perception ability, including: the PSNR ratio is less than or equal to a first threshold.
[0015] Another possible implementation is that the first quality indicator is the structural similarity index (SSIM), which does not meet the requirements of human visual perception, including: SSIM being less than or equal to the second threshold.
[0016] Another possible implementation is that the first quality metric is the feature similarity index (FSIM), and the first quality metric does not meet the human visual perception ability, including: FSIM is less than or equal to the third threshold.
[0017] Another possible implementation is that the first quality metric is visual information fidelity (VIF). The first quality metric does not meet the human eye's perception ability, including: VIF is less than or equal to the fourth threshold.
[0018] Another possible implementation is that the first quality metric is gradient magnitude similarity deviation (GMSD). The first quality metric does not meet the human eye's perception ability, including: GMSD is greater than or equal to the fifth threshold.
[0019] Another possible implementation is that the first quality metric is learned perceptual image patch similarity (LPIPS). The first quality metric does not meet the human eye's perception ability, including: LPIPS greater than or equal to the sixth threshold.
[0020] Another possible implementation involves a second image that does not meet human visual perception requirements, including: a second quality index of the second image that does not meet human visual perception requirements. The second quality index reflects the absolute quality of the second image; using the second quality index allows for a direct and quick assessment of whether the second image meets human visual perception requirements from the perspective of image display quality.
[0021] Another possible implementation is that the second quality metric is the first evaluation result obtained by the blind / referenceless image spatial quality evaluator (BRISQUE). The second quality metric does not meet the human eye's perception ability, including: the first evaluation result is greater than or equal to the seventh threshold.
[0022] Another possible implementation is that the second quality index is the second evaluation result obtained by the natural image quality evaluator (NIQE). The second quality index does not meet the human eye's perception ability, including: the second evaluation result is greater than or equal to the eighth threshold.
[0023] Another possible implementation is that the second quality indicator is information entropy, which does not meet the requirements of human visual perception, including: information entropy is less than or equal to the ninth threshold.
[0024] Another possible implementation is that the second quality index is the sharpness index, which does not meet the human eye's perception ability, including: the sharpness index is less than or equal to the tenth threshold.
[0025] Another possible implementation is tone mapping, which includes luminance mapping and / or chrominance mapping.
[0026] Another possible implementation is that the aforementioned metadata includes: luminance metadata for describing luminance characteristics, and / or chrominance metadata for describing color characteristics.
[0027] Another possible implementation involves metadata including: luminance metadata describing luminance characteristics, and / or chrominance metadata describing color characteristics, as well as environmental data. This environmental data describes the environmental characteristics at the time the first image was acquired.
[0028] Secondly, another tone mapping method is provided for use in a decoding device. This method may include: receiving a bitstream from an encoding device, the bitstream including a first image and second metadata; performing tone mapping on the first image based on the second metadata to obtain a second image; if the second image does not meet human visual perception requirements, adjusting the bit width of the first image, and re-performing tone mapping on the adjusted first image based on the second metadata; if the second image meets human visual perception requirements, displaying the second image.
[0029] The solution provided in this application adjusts the mapping parameters used for tone mapping to ensure that the tone-mapped image meets the perceptual capabilities of the human eye. This ensures that the mapped image can be perceived by the human eye under various display conditions, achieving the goal of normal image display in all situations. This provides users with a better and more realistic perceptual experience across different display environments, increasing the immersive feeling of the displayed image.
[0030] In one possible implementation, the method may further include: receiving a bitstream including difference information, which indicates the difference between a first image with an initial bit width and a first image with a first bit width. Correspondingly, adjusting the bit width of the first image includes: adjusting the first image to the first bit width based on the difference information. With the difference information transmitted from the encoding side, the decoding device can adjust the first image to the first bit width based on the difference information, thereby obtaining a mapped image that satisfies human visual perception for display. This improves the efficiency of the decoding device in adjusting the bit width of the first image.
[0031] The specific implementation of the tone mapping method provided in the second aspect of this application can be referred to the specific implementation of the first aspect above, and will not be repeated here.
[0032] Thirdly, another tone mapping method is provided for use in a decoding device. The method may include: receiving a bitstream from an encoding device, the bitstream including a first image and second metadata; performing tone mapping on the first image according to the second metadata to obtain a second image; if the second image does not meet the human eye's perception ability, adjusting the brightness range and / or chromaticity range of the second image until the second image meets the human eye's perception ability, and then displaying the second image.
[0033] The solution provided in this application adjusts the parameter range of the image obtained after tone mapping to ultimately display a mapped image that meets the perceptual capabilities of the human eye. This ensures that the mapped image can be perceived by the human eye under various display conditions, achieving the goal of normal image display in all situations. This provides users with a better and more realistic perceptual experience in various display states, increasing the immersive feeling of the image display.
[0034] The specific implementation of the tone mapping method provided in the third aspect of this application can be referred to the specific implementation of the first aspect above, and will not be repeated here.
[0035] Fourthly, a tone mapping apparatus is provided, which is applied to an encoding device. The apparatus includes: a mapping unit, a matching unit, and a processing unit. Wherein:
[0036] A mapping unit is used to perform tone mapping on a first image according to mapping parameters to obtain a second image. The mapping parameters include the bit width of the first image and / or the metadata of the first image.
[0037] The matching unit is used to determine whether the second image meets the requirements of human visual perception. Here, human visual perception refers to the range within which the human eye perceives image quality.
[0038] The processing unit is used to adjust the mapping parameters if the matching unit determines that the second image does not meet the human eye's perceptual capabilities. If the matching unit determines that the second image meets the human eye's perceptual capabilities, it generates a bitstream of the first image, which includes the first image with an initial bit width and the metadata of the mapped second image.
[0039] The mapping unit is also used to re-perform tone mapping on the first image based on the mapping parameters adjusted by the adjustment unit.
[0040] One possible implementation is that the mapping parameter is the bit width of the first image. Given a first image with a first bit width, and a second image obtained through mapping that satisfies human visual perception, the processing unit further generates a bitstream containing difference information. This difference information indicates the difference between the first image with the initial bit width and the first image with a first bit width of [missing information]. In this way, the decoding device, based on the difference information, can adjust the first image to the first bit width, thereby obtaining a mapped image that satisfies human visual perception for display.
[0041] Another possible implementation is to use the metadata of the first image as the mapping parameter. If the second image obtained by performing tone mapping on the first image according to the first metadata satisfies human visual perception, then the metadata of the mapped second image is the first metadata.
[0042] Another possible implementation is that the aforementioned difference information includes: a first image with an initial bit width and a mask image of the first image with a first bit width.
[0043] Another possible implementation is that the aforementioned difference information includes metadata extracted from the mask image of the first image with the initial bit width and the first image with the first bit width.
[0044] Another possible implementation involves the second image not meeting human visual perception requirements, including: the first quality index of the second image relative to the first image with the initial bit width does not meet human visual perception requirements. The first quality index reflects the degree of detail retention of the second image relative to the first image with the initial bit width. Using the first quality index, it is possible to intuitively and quickly determine whether the second image meets human visual perception requirements from a fidelity perspective.
[0045] Another possible implementation is that the first quality indicator is PSNR, and the first quality indicator does not meet the human eye's perception ability, including: the PSNR ratio is less than or equal to a first threshold.
[0046] Another possible implementation is that the first quality metric is SSIM, which does not meet the human eye's perception ability, including: SSIM being less than or equal to a second threshold.
[0047] Another possible implementation is that the first quality metric is FSIM, which does not meet the human eye's perception ability, including: FSIM being less than or equal to a third threshold.
[0048] Another possible implementation is that the first quality indicator is VIF, which does not meet the human eye's perception ability, including: VIF is less than or equal to the fourth threshold.
[0049] Another possible implementation is that the first quality metric is GMSD, and the first quality metric does not meet the human eye's perception ability, including: GMSD is greater than or equal to the fifth threshold.
[0050] Another possible implementation is that the first quality metric is LPIPS, which does not meet the requirements of human visual perception, including: LPIPS being greater than or equal to the sixth threshold.
[0051] Another possible implementation involves a second image that does not meet human visual perception requirements, including: a second quality index of the second image that does not meet human visual perception requirements. The second quality index reflects the absolute quality of the second image; using the second quality index allows for a direct and quick assessment of whether the second image meets human visual perception requirements from the perspective of image display quality.
[0052] Another possible implementation is that the second quality metric is the first evaluation result obtained by BRISQUE, and the second quality metric does not meet the human eye perception ability, including: the first evaluation result is greater than or equal to the seventh threshold.
[0053] Another possible implementation is that the second quality indicator is the second evaluation result obtained from NIQE. The second quality indicator does not meet the human eye perception ability, including: the second evaluation result is greater than or equal to the eighth threshold.
[0054] Another possible implementation is that the second quality indicator is information entropy, which does not meet the requirements of human visual perception, including: information entropy is less than or equal to the ninth threshold.
[0055] Another possible implementation is that the second quality index is the sharpness index, which does not meet the human eye's perception ability, including: the sharpness index is less than or equal to the tenth threshold.
[0056] Another possible implementation is tone mapping, which includes luminance mapping and / or chrominance mapping.
[0057] Another possible implementation is that the aforementioned metadata includes: luminance metadata for describing luminance characteristics, and / or chrominance metadata for describing color characteristics.
[0058] Another possible implementation involves metadata including: luminance metadata describing luminance characteristics, and / or chrominance metadata describing color characteristics, as well as environmental data. This environmental data describes the environmental characteristics at the time the first image was acquired.
[0059] The tone mapping apparatus provided in the fourth aspect of this application is used to implement the tone mapping method provided in the first aspect above. The specific implementation can be referred to the specific implementation of the first aspect above, and will not be repeated here.
[0060] Fifthly, another tone mapping device is provided for use in decoding equipment. This device includes: a receiving unit, a mapping unit, a matching unit, and a processing unit. Wherein:
[0061] A receiving unit is used to receive a bitstream from an encoding device, the bitstream including a first image and second metadata.
[0062] The mapping unit is used to perform tone mapping on the first image based on the second metadata to obtain the second image.
[0063] The matching unit is used to determine whether the second image meets the human eye's perceptual capabilities.
[0064] The processing unit is configured to adjust the bit width of the first image if the matching unit determines that the second image does not meet human visual perception requirements, and to display the second image if the matching unit determines that the second image meets human visual perception requirements.
[0065] The mapping unit is also used to: re-perform tone mapping on the first image after adjusting the bit width based on the second metadata.
[0066] In one possible implementation, the receiving unit is further configured to: receive a bitstream including difference information. The difference information is used to indicate the difference between the first image with an initial bit width and the first image with a first bit width. Specifically, the processing unit is configured to: adjust the first image to a first bit width based on the difference information.
[0067] The tone mapping apparatus provided in the fifth aspect of this application is used to implement the tone mapping method provided in the second aspect above. The specific implementation can be referred to the specific implementation of the second aspect above, and will not be repeated here.
[0068] Sixthly, another tone mapping device is provided for use in decoding equipment. This device includes: a receiving unit, a mapping unit, a matching unit, and a processing unit. Wherein:
[0069] A receiving unit is used to receive a bitstream from an encoding device, the bitstream including a first image and second metadata.
[0070] The mapping unit is used to perform tone mapping on the first image based on the second metadata to obtain the second image.
[0071] The matching unit is used to determine whether the second image meets the human eye's perceptual capabilities.
[0072] The processing unit is configured to adjust the tonal parameter range of the first image until the second image meets the human eye's perceptual requirements if the matching unit determines that the second image does not meet the human eye's perceptual requirements. If the matching unit determines that the second image meets the human eye's perceptual requirements, the second image is displayed.
[0073] The tone mapping apparatus provided in the sixth aspect of this application is used to implement the tone mapping method provided in the third aspect above. The specific implementation can be referred to the specific implementation of the third aspect above, and will not be repeated here.
[0074] A seventh aspect provides a computing device including a processor and a memory. The processor is configured to execute instructions stored in the memory to cause the computing device to perform operational steps of the methods described in the first, second, or third aspect or any possible implementation thereof.
[0075] Eighthly, a computer-readable storage medium is provided, comprising: computer software instructions; when the computer software instructions are executed in a processor, causing the processor to perform operational steps of the method as described in the first, second, or third aspect or any of the possible implementations above.
[0076] Ninthly, a computer program product is provided that, when run on a computer, causes the computer to perform the operational steps of the method as described in the first, second, or third aspect or any of the possible implementations above.
[0077] In a tenth aspect, a chip is provided, comprising: a processor and a power supply circuit; wherein the power supply circuit is used to supply power to the processor; and the processor is used to perform operational steps of the method in the first aspect, the second aspect, the third aspect, or any possible implementation thereof.
[0078] The technical effects of any of the design methods in aspects four through ten can be found in the technical effects of different design methods in aspects one, two, or three, and will not be repeated here.
[0079] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0080] Figure 1 is a schematic diagram of the operation of an HDR end-to-end system;
[0081] Figure 2 is a schematic diagram of the codeword allocation in the PQ encoding scheme;
[0082] Figure 3 is a schematic diagram of the architecture of an imaging system provided in an embodiment of this application;
[0083] Figure 4 is a schematic diagram of the internal structure of a terminal device provided in an embodiment of this application;
[0084] Figure 5 is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0085] Figure 6 is a flowchart illustrating a tone mapping method provided in an embodiment of this application;
[0086] Figure 7 is a flowchart illustrating another tone mapping method provided in an embodiment of this application;
[0087] Figure 8 is a schematic diagram of the tone mapping principle of an encoding device provided in an embodiment of this application;
[0088] Figure 9 is a schematic diagram of the tone mapping principle of a decoding device provided in an embodiment of this application;
[0089] Figure 10 is a schematic diagram of the tone mapping process provided in Embodiment 1 of this application;
[0090] Figure 11 is a schematic diagram of the tone mapping process provided in Embodiment 2 of this application;
[0091] Figure 12 is a schematic diagram of the tone mapping process provided in Embodiment 3 of this application;
[0092] Figure 13 is a schematic diagram of the tone mapping process provided in Embodiment 4 of this application;
[0093] Figure 14 is a schematic diagram of a tone mapping device provided in an embodiment of this application;
[0094] Figure 15 is a schematic diagram of another tone mapping device provided in an embodiment of this application. Detailed Implementation
[0095] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between the related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. A and B can be singular or plural.
[0096] In the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0097] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0098] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0099] It is understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, various embodiments throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in multiple embodiments. It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0100] It is understood that some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current solution on which they are based, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in certain scenarios. Correspondingly, the apparatus given in the embodiments of this application can also implement these features or functions, which will not be elaborated here.
[0101] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments are consistent and can be mutually referenced. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following descriptions of the embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0102] For ease of understanding, the main terms used in this application are explained below.
[0103] The dynamic range of an image refers to the range of brightness differences between the darkest and brightest parts of the image that it can display.
[0104] High dynamic range (HDR) images are images with a dynamic range between 0.001 and 10,000 nits. nit is a unit of illumination.
[0105] Standard dynamic range (SDR) images refer to images whose dynamic range is generally between 1 nit and 100 nits.
[0106] Tone mapping (TM) is a technique for adjusting the range of image parameters, specifically the adjustment methods between different parameter ranges. For example, tone mapping can convert an HDR image to a low dynamic range (LDR) image. Tone mapping aims to preserve as much detail and contrast as possible in an HDR image while adapting to the display limitations of LDR devices. For instance, tone mapping can adjust the tonal parameters of an image. SDR can also be referred to as LDR in contrast to HDR.
[0107] Metadata: Information used to record key features of images in a video or scene frame. For example, metadata can record information such as the average brightness of the scene, maximum and minimum values, etc.
[0108] Tone mapping is mainly used for adapting front-end signals to terminal display devices. For example, if the front-end acquires or produces a 4000 nit light signal, but the terminal display device only has a display capability of 500 nits, how to best map the 4000 nit signal to the 500 nit device is a tone mapping process from high dynamic range to low dynamic range, and vice versa.
[0109] Current industry-standard dynamic range mapping (DLL) methods involve generating a tone mapping curve (the horizontal axis represents brightness before processing, and the vertical axis represents brightness after processing) based on the image's metadata and the target image (an image matching the display capabilities). Tone mapping is then performed based on this curve. Therefore, during dynamic range adjustment, metadata containing key information or features from the source image acts as a "bridge" between the source and the display device, guiding the generation of the mapping curve. DLL methods include static mapping methods and dynamic mapping methods.
[0110] Static mapping methods are based on immutable static metadata. For example, a static mapping method can perform an overall tone mapping process from a single data set based on the same video or image content; that is, the tone mapping curve is usually the same throughout the process. The advantage of static mapping methods is that they carry less metadata and have a simpler processing flow. However, the disadvantage is that information loss can occur in some scenarios. For instance, if the curve focuses on protecting bright areas, details may be lost or even completely invisible in extremely dark scenes, affecting the user experience.
[0111] Dynamic mapping methods are based on dynamic metadata that changes with each frame or scene. By applying different mapping curves to each scene or frame according to the image and video characteristics, dynamic mapping can adapt to more diverse scenarios and achieve better display adaptation, thus becoming the mainstream choice.
[0112] The scenario described in this application is a set of image frames that express the same meaning, and a scenario may include one or more image frames.
[0113] Image frame transmission in an imaging system is end-to-end transmission of HDR content, involving an end-to-end system for HDR content. This system includes encoding and decoding devices. In an end-to-end HDR content system, signal transmission involves the mutual conversion between electrical and optical signals; this conversion process is called encoding. In the decoding device, the electrical signal is converted back to an optical signal, and after being restored to the original image's optical signal, tone mapping is used to adapt to the display device's dynamic range.
[0114] Taking a perceptual quantizer (PQ) encoding system as an example, the transmission process of an HDR end-to-end system is explained. The workflow of an HDR end-to-end system is shown in Figure 1: The camera acts as the encoding device, deploying an optical-electro transfer function (OETF), which specifically includes an optical-optical transfer function (OOTF) and an inverse electro-optical transfer function (EOTF). Scene light enters the camera, is processed and encoded to form a machine-perceptible electrical signal, and is transmitted. The display device, acting as the decoding device, deploys an EOTF. After the electrical signal enters the display device, it undergoes electro-optical conversion to restore linear display light, a signal perceptible to the human eye. In the process illustrated in Figure 1, the electro-optic / optical conversion is based on the PQ function. The principle is to allocate codewords based on human eye perception, assigning more codewords to areas more sensitive to human vision and fewer codewords to less sensitive areas. Figure 2 shows the codeword allocation of the PQ encoding scheme. As can be seen from Figure 2, the scheme always allocates more codewords to areas with lower brightness and only a small number of codewords to areas with higher brightness.
[0115] Although the PQ encoding scheme considers factors sensitive to human visual perception, it is a photoelectric conversion scheme and is unrelated to the process of adapting the light signal to the display device through tone mapping. Current tone mapping schemes typically only consider a fixed display state (i.e., a fixed dynamic range). However, in reality, the display side faces many more state changes, such as image scaling, causing the dynamic range of the display to change dynamically. Therefore, if a tone mapping scheme that only considers a fixed display state is implemented, the image will display normally in that fixed state. But once the display state changes, abnormal phenomena may occur that are visible to the human eye, such as unclear areas or color banding.
[0116] Based on this, this application provides a tone mapping method. By adjusting the mapping parameters used for tone mapping, it ensures that the tone-mapped image meets the perceptual capabilities of the human eye, so that the mapped image can be perceived by the human eye under various display states, achieving the goal of normal image display under various display states. This allows users to have a better and more realistic perceptual experience under various display states, increasing the immersive feeling of the image display.
[0117] The tone mapping method provided in this application can be applied to the imaging system illustrated in Figure 3. As shown in Figure 3, the imaging system includes an encoding device 301, a decoding device 302, and a display screen 303.
[0118] In the imaging system illustrated in Figure 3, encoding device 301 encodes the image and generates metadata for tone mapping. Then, encoding device 301 transmits a bitstream containing the metadata and the image to decoding device 302. Upon receiving the bitstream containing the metadata and the image, decoding device 302 first decodes and reconstructs the image, then performs tone mapping according to the received metadata, adapting the image to the display screen 303. Finally, the image is displayed on the display screen 303.
[0119] In one possible implementation, the encoding device 301 can capture images. For example, the encoding device 301 can have an image capturing function. For example, the encoding device 301 can be a camera.
[0120] In another possible implementation, the encoding device 301 can also acquire images transmitted by other devices. For example, the encoding device 301 can be a professional image editing tool (such as video editing software, image color correction software, etc.). Alternatively, the encoding device 301 can be a program or software with automatic metadata analysis and generation capabilities.
[0121] In one possible implementation, the decoding device 302 and the display screen 303 can be deployed in a single hardware device (as shown in Figure 3), or they can be decoupled and deployed in different hardware devices (not shown in Figure 3).
[0122] In another possible implementation, the imaging system shown in Figure 3 can be deployed in a hardware device, with the encoding device 301 and the decoding device 302 being different functional units in the imaging system.
[0123] For example, the imaging system or decoding device 302 shown in Figure 3 can be a user terminal, such as a mobile phone, tablet computer, personal computer or other forms, which is not limited in the embodiments of this application.
[0124] The internal structure of the imaging system or decoding device 302 shown in Figure 3 will be illustrated below with reference to Figure 4, taking the imaging system or decoding device 302 shown in Figure 3 as an example.
[0125] Please refer to Figure 4, which is a structural schematic diagram of a terminal device provided in this application. As shown in Figure 4, the terminal device includes: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a power management module 140, an antenna, a wireless communication module 160, an audio module 170, a speaker 170A, a speaker interface 170B, a microphone 170C, a sensor module 180, buttons 190, an indicator 191, a display screen 192, a camera 193, etc. The sensor module 180 may include sensors such as a distance sensor, a proximity sensor, a fingerprint sensor, a temperature sensor, a touch sensor, and an ambient light sensor.
[0126] The structure illustrated in this embodiment does not constitute a specific limitation on the terminal device. In other embodiments, the terminal device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0127] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0128] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0129] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, and / or a USB interface, etc.
[0130] The power management module 140 is used to connect to a power source. The power management module 140 can also be connected to the processor 110, internal memory 121, display screen 192, camera 193, and wireless communication module 160, etc. The power management module 140 receives power input and supplies power to the processor 110, internal memory 121, display screen 192, camera 193, and wireless communication module 160, etc. In some embodiments, the power management module 140 can also be located within the processor 110.
[0131] The wireless communication function of the terminal device can be implemented through an antenna and a wireless communication module 160. The wireless communication module 160 can provide solutions for wireless communication applications on the terminal device, including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.
[0132] The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via an antenna, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, perform frequency modulation and amplification on them, and then convert them into electromagnetic waves for radiation via the antenna. In some embodiments, the antenna of the terminal device and the wireless communication module 160 are coupled, enabling the terminal device to communicate with networks and other devices via wireless communication technology.
[0133] In this embodiment, the terminal device can communicate with the media server via the wireless communication module 160 and the antenna.
[0134] The terminal device implements display functions through a GPU, a display screen 192, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 192 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0135] The display screen 192 is used to display text, images, and videos, etc. The display screen 192 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.
[0136] In this embodiment, the display screen 192 is used to display the interface provided by the terminal device. These interfaces can be used to select audio or video resources to be played, adjust the target loudness, or perform other functions.
[0137] The terminal device can implement shooting functions through an ISP, camera 193, video codec, GPU, display 192, and application processor. The ISP is used to process the data fed back by the camera 193. In some embodiments, the ISP can be located in the camera 193.
[0138] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the terminal device may include one or N cameras 193, where N is a positive integer greater than 1. For example, this embodiment does not limit the position of the cameras 193 on the terminal device.
[0139] The terminal device may not include a camera, meaning the aforementioned camera 193 is not integrated into the terminal device (e.g., a television). The terminal device can connect an external camera 193 via an interface (e.g., USB interface 130). This external camera 193 can be fixed to the terminal device using an external fastener (e.g., a camera holder with a clip). For example, the external camera 193 can be fixed to the edge of the terminal device's display screen 192, such as the upper edge, using an external fastener.
[0140] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in terminal devices, such as image recognition, speech recognition, and text understanding.
[0141] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal device. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, text, images, and video files can be saved on the external storage card.
[0142] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of the terminal device by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of the terminal device (such as audio data, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0143] The terminal device can implement audio functions through an audio module 170, a speaker 170A, a microphone 170C, a speaker interface 170B, and an application processor. Examples include music playback and recording. In this application, the microphone 170C can be used to receive voice commands issued by the user to the terminal device. The speaker 170A can be used to provide feedback to the user on the terminal device's decision commands.
[0144] Audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. Audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, audio module 170 may be located in processor 110, or some functional modules of audio module 170 may be located in processor 110. Speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. Microphone 170C, also called a "microphone" or "microphone," is used to convert sound signals into electrical signals.
[0145] The speaker jack 170B is used to connect wired speakers. The speaker jack 170B can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0146] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. The terminal device can receive button input and generate key signal inputs related to user settings and function control of the terminal device.
[0147] Indicator 191 can be an indicator light, which can be used to indicate whether the terminal device is in a powered-on, standby, or powered-off state. For example, an indicator light that is off indicates that the terminal device is powered off; an indicator light that is green or blue indicates that the terminal device is powered on; and an indicator light that is red indicates that the terminal device is in a standby state.
[0148] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device. It may have more or fewer components than those shown in Figure 4, may combine two or more components, or may have different component configurations. For example, the terminal device may also include components such as speakers. The various components shown in Figure 4 can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing or application-specific integrated circuits.
[0149] The internal structure of the image system, decoding device 301, or decoding device 302 shown in Figure 3 will be illustrated below with reference to Figure 5.
[0150] Please refer to Figure 5, which is a schematic diagram of a computing device provided in this application. This computing device can be the imaging system shown in Figure 3, or decoding device 301, or decoding device 302.
[0151] As shown in Figure 5, the computing device may include a processor 5010, a bus 5020, a memory 5030, and a communication interface 5040. The processor 5010, the memory 5030, and the communication interface 5040 are connected via the bus 5020.
[0152] It should be understood that in this embodiment, the processor 5010 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0153] The processor 5010 may also be a GPU, NPU, microprocessor, ASIC, or one or more integrated circuits used to control the execution of the program in this application.
[0154] The communication interface 5040 is used to enable communication between the computing device and external devices or components.
[0155] Bus 5020 is used to transfer information between the aforementioned components (such as processor 5010 and memory 5030). In addition to a data bus, bus 5020 may also include a power bus, control bus, and status signal bus. However, for clarity, all buses are labeled as bus 5020 in the figure. Bus 5020 can be a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. It is worth noting that Figure 5 only uses a computing device including one processor 5010 and one memory 5030 as an example. Here, processor 5010 and memory 5030 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business requirements.
[0156] The memory 5030 can be a pool of volatile memory or a pool of non-volatile memory, or it can include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0157] For example, the processor 5010 can perform some or all of the functions of the solution provided in this application by running or executing software programs and / or modules stored in the memory 5030.
[0158] The tone mapping method provided in this application will now be described in detail with reference to the accompanying drawings.
[0159] It should be noted that in the solution provided in this application, the tone mapping device can process a single image frame or a sequence of image frames. The following embodiments of this application illustrate the process of the tone mapping device processing a single image frame (hereinafter referred to as an image). When the tone mapping device processes a sequence of image frames including multiple image frames, a single frame image in the image frame sequence, or a frame image in a scene in the image frame sequence, can be processed with reference to the process described in the following embodiments, and will not be repeated here.
[0160] Figure 6 is a schematic flowchart of a tone mapping method provided in this application, which is applied to a tone mapping apparatus. This tone mapping apparatus can be deployed in an encoding device within an image system. For example, the tone mapping apparatus can be deployed in the encoding device 301 of the image system illustrated in Figure 3.
[0161] As shown in Figure 6, the tone mapping method provided in this embodiment includes:
[0162] S601, The tone mapping device performs tone mapping on the first image according to the mapping parameters to obtain the second image.
[0163] The first image is the object to which the tone mapping device performs tone mapping in this cycle. In different rounds of tone mapping, the bit width of the first image can be the same or different.
[0164] In one possible implementation, the image captured by the encoding device, or the image captured by another device and transmitted to the encoding device, is the first image with an initial bit width. When performing the first tone mapping on the first image, the bit width of the first image is the initial bit width. When performing the Nth tone mapping on the first image, the bit width of the first image is either the initial bit width or the adjusted bit width.
[0165] The mapping parameters are the basis for tone mapping. The mapping parameters on the encoding device side may include the metadata of the first image and / or the bit width of the first image.
[0166] The image's metadata refers to the image's features represented in vector form. The image's bit width is the number of binary bits required to represent the color information of each pixel in the image; it is also known as color depth.
[0167] In one possible implementation, the metadata includes: luminance data for describing luminance characteristics, and / or chrominance data for describing chrominance characteristics.
[0168] In another possible implementation, in addition to the luminance data used to describe luminance characteristics and / or the chrominance data used to describe chrominance characteristics, the metadata may also include environmental data. The environmental data describes the environmental characteristics at the time the first image was acquired.
[0169] In one possible implementation, the tone mapping device can extract metadata from the image. For example, it can use algorithms from standards such as HDR10+ or HDR Vivid to extract metadata.
[0170] Specifically, performing tone mapping on the first image refers to generating a tone mapping curve based on metadata, and adjusting the tone parameters (brightness and / or chroma) of the first image with a certain bit width according to the tone mapping curve. The encoding end can generate the tone mapping curve based on the mastering display's tone parameters when generating the tone mapping curve from the metadata; these parameters can be configured according to actual needs.
[0171] Specifically, tone mapping includes luminance mapping and / or color mapping. Luminance mapping is used to adjust luminance parameters, while color mapping is used to adjust chrominance parameters.
[0172] A tone mapping curve indicates how tone parameter values in one tone parameter range are mapped to tone parameter values in another tone parameter range. This application does not limit the specific scheme for generating tone mapping curves based on metadata.
[0173] For example, taking the hue parameter as the luminance parameter, the horizontal axis of the corresponding tone mapping curve can be the luminance value before mapping, and the vertical axis can be the luminance value after mapping. Based on the tone mapping curve corresponding to this luminance, the luminance values of pixels in the HDR image can be adjusted to map the HDR image to an SDR image.
[0174] S602, The tone mapping device determines whether the second image meets the human eye's perception ability.
[0175] Specifically, an image meeting human visual perception requirements means that, under human visual perception, the image content is lossless (can be displayed normally). Conversely, an image failing to meet human visual perception requirements means that, under human visual perception, the image content is damaged or distorted (cannot be displayed normally). Human visual perception capability refers to the range within which the human eye perceives image quality. Indicators describing human visual perception capability, as well as the perceptual range for each indicator, can be configured according to actual needs.
[0176] In one possible implementation, the quality index of the image is used to determine whether the second image meets the perceptual requirements of the human eye.
[0177] The following example illustrates in detail the implementation of determining whether the second image satisfies the human eye's perception capability in S602.
[0178] 1. Based on the quality index of the second image relative to the first image with the initial bit width, determine whether the second image meets the human eye's perception requirements.
[0179] In implementation 1, the second image does not meet human visual perception requirements. Specifically, the first quality index of the second image relative to the first image with the initial bit width does not meet human visual perception requirements. The first quality index reflects the degree of detail retention of the second image compared to the first image with the initial bit width. The content of the first quality index can be configured according to actual needs.
[0180] When the primary quality indicator is different, the evaluation criteria for failing to meet human visual perception requirements also differ, as shown in the following examples:
[0181] Example 1: The first quality indicator is PSNR. The first quality indicator does not meet the human eye's perception ability, including: the PSNR ratio is less than or equal to the first threshold.
[0182] PSNR is the difference between two images calculated based on mean squared error; a higher PSNR value indicates less distortion. Therefore, the PSNR of the second image can be configured to be less than or equal to a first threshold to indicate that the second image does not meet the perceptual requirements of the human eye.
[0183] The first threshold is a threshold value for judging whether the human eye can recognize the distortion based on PSNR. The value of the first threshold can be configured according to actual needs.
[0184] Example 2: The first quality indicator is SSIM. The first quality indicator does not meet the human eye's perception ability, including: SSIM is less than or equal to the second threshold.
[0185] SSIM comprehensively evaluates image quality based on brightness, contrast, and structural similarity, and is more sensitive to detailed texture information. A higher SSIM indicates greater image similarity. Therefore, a second image with an SSIM less than or equal to a second threshold can be configured to indicate that the second image does not meet human visual perception requirements.
[0186] Alternatively, SSIM can be replaced by multi-scale structural similarity index (multi-scale SSIM, MS-SSIM), which improves the ability to evaluate details through multi-resolution. The first quality metric is MS-SSIM, and the first quality metric does not meet the requirements of human visual perception, including: MS-SSIM being less than or equal to the second threshold.
[0187] The second threshold is a threshold value for determining whether the human eye can recognize the distortion based on SSIM. The value of the second threshold can be configured according to actual needs.
[0188] Example 3: The first quality indicator is FSIM. The first quality indicator does not meet the human eye's perception ability, including: FSIM is less than or equal to the third threshold.
[0189] FSIM utilizes phase consistency and gradient magnitude features to emphasize the fidelity of edge and texture details. A higher FSIM value indicates better structure preservation. Therefore, the FSIM of the second image can be configured to be less than or equal to a third threshold to indicate that the second image does not meet the perceptual requirements of the human eye.
[0190] The third threshold is a threshold value used to determine whether the human eye can detect distortion based on FSIM. The value of the third threshold can be configured according to actual needs.
[0191] Example 4: The first quality indicator is VIF. The first quality indicator does not meet the human eye's perception ability, including: VIF is less than or equal to the fourth threshold.
[0192] Visual information retention (VIF) measures the degree to which visual information is preserved in an image. Combined with statistical information about the natural scene, it is sensitive to detail loss. A higher VIF value indicates better image quality. Therefore, a VIF value less than or equal to a fourth threshold can be configured to indicate that the second image does not meet human visual perception requirements.
[0193] The fourth threshold is a threshold value for judging whether the human eye can recognize distortion based on VIF. The value of the fourth threshold can be configured according to actual needs.
[0194] Example 5: The first quality indicator is GMSD. The first quality indicator does not meet the human eye's perception ability, including: GMSD is greater than or equal to the fifth threshold.
[0195] GMSD, based on the local similarity of gradient magnitude, directly reflects the degradation of edges and details. The smaller the GMSD value, the more stable the gradient similarity, the less image distortion, and the higher the quality. Therefore, a GMSD of the second image greater than or equal to a fifth threshold can be configured to indicate that the second image does not meet the perceptual requirements of the human eye.
[0196] The fifth threshold is a threshold value based on GMSD to determine whether the human eye can recognize the distortion. The value of the fifth threshold can be configured according to actual needs.
[0197] Example 6: The first quality indicator is LPIPS. The first quality indicator does not meet the human eye's perception ability, including: LPIPS is greater than or equal to the sixth threshold.
[0198] LPIPS uses pre-trained deep neural networks (such as VGG and AlexNet) to extract features and calculate the distance between spatial features, making it sensitive to semantic information. The smaller the LPIPS value, the more similar the image is to human perception. Therefore, a LPIPS value greater than or equal to a sixth threshold can be configured to indicate that the second image does not meet human perceptual requirements.
[0199] The sixth threshold is a threshold value for judging whether the human eye can recognize distortion based on LPIPS. The value of the sixth threshold can be configured according to actual needs.
[0200] It should be understood that Examples 1 to 6 above are merely illustrative examples and are not intended to limit the content of the second quality index or the specific implementation of judging whether the second image conforms to human visual perception based on the second quality index.
[0201] Implementation 2: Based on the quality index of the second image, determine whether the second image meets the human eye's perception requirements.
[0202] In implementation 2, the second image does not meet the requirements of human visual perception. Specifically, the second quality index of the second image does not meet the requirements of human visual perception. The second quality index reflects the quality of the second image. The content of the second quality index can be configured according to actual needs.
[0203] When the second quality indicator is different, the evaluation content that fails to meet the human eye's perception ability will also be different, as shown in the following examples:
[0204] Example a: The second quality indicator is the first evaluation result obtained by BRISQUE. The second quality indicator does not meet the human eye perception ability, including: the first evaluation result is greater than or equal to the seventh threshold.
[0205] BRISQUE evaluates image quality using a support vector machine based on statistical features of natural scenes, and is sensitive to blur and noise. The first evaluation result obtained by BRISQUE is negatively correlated with image quality. Therefore, it can be configured that if the first evaluation result of a second image is greater than or equal to a seventh threshold, the second image does not meet the perceptual requirements of the human eye.
[0206] The seventh threshold is a threshold value used to determine whether the human eye can detect distortion based on the evaluation results of BRISQUE. The value of the seventh threshold can be configured according to actual needs.
[0207] Example b: The second quality indicator is the second assessment result obtained from NIQE. The second quality indicator does not meet the human eye perception ability, including: the second assessment result is greater than or equal to the eighth threshold.
[0208] NIQE models the statistical features of natural images and calculates the degree of deviation from an ideal model, making it suitable for assessing detail degradation. A smaller value in the second evaluation result obtained from NIQE indicates that the tested image is closer to the statistical features of a natural image. Therefore, it can be configured that a second evaluation result greater than or equal to an eighth threshold indicates that the second image does not meet human visual perception requirements.
[0209] The eighth threshold is a threshold value used to determine whether the human eye can detect distortion based on the NIQE evaluation results. The value of the eighth threshold can be configured according to actual needs.
[0210] Example c: The second quality indicator is information entropy. The second quality indicator does not meet the human eye's perception ability, including: information entropy is less than or equal to the ninth threshold.
[0211] Information entropy calculates the entropy value of the grayscale distribution of an image. A decrease in entropy value may reflect the loss of detailed information. Therefore, it can be configured that if the information entropy of the second image is less than or equal to the ninth threshold, it indicates that the second image does not meet the perceptual requirements of the human eye.
[0212] The ninth threshold is a threshold value based on information entropy to determine whether the human eye can recognize the distortion. The value of the ninth threshold can be configured according to actual needs.
[0213] Example d: The second quality index is the sharpness index. The second quality index does not meet the human eye's perception ability, including: the sharpness index is less than or equal to the tenth threshold.
[0214] The sharpness index directly measures the clarity of details by quantizing image sharpness through gradient energy, Laplacian response, or frequency domain fractionation. A higher sharpness index value indicates a sharper image. Therefore, a sharpness index of the second image less than or equal to the tenth threshold can be configured to indicate that the second image does not meet the perceptual requirements of the human eye.
[0215] The tenth threshold is a threshold value based on the sharpness index to determine whether the human eye can recognize the distortion. The value of the tenth threshold can be configured according to actual needs.
[0216] It should be understood that the examples a to d above are merely illustrative examples and are not intended to limit the content of the second quality index or the specific implementation of judging whether the second image conforms to human visual perception based on the second quality index.
[0217] Furthermore, in S602, determining whether the second image satisfies the human eye's perception capability can be achieved using the methods described in implementation 1 and / or implementation 2.
[0218] In another possible implementation, a matching model is configured to determine whether an image meets the perceptual requirements of the human eye. The matching model is a model trained on a large amount of sample data. The input to the matching model can be a second image, or the second image and a first image with an initial bit width, and the output is whether the second image meets the perceptual requirements of the human eye. For example, this matching model can be an AI model, a deep learning model, or something else. This application does not limit the type or specific implementation of this matching model.
[0219] For example, the matching model can be designed based on one or more of the following principles: PQ / HLG (hybrid log-gamma) quantization perception, depth perception, brightness perception, detail perception, and resolution principle.
[0220] Of course, the specific implementation of determining whether the second image meets the human eye's perception capability can be configured according to actual needs. The specific implementation of the matching algorithm is not limited in the embodiments of this application.
[0221] Furthermore, if it is determined in S602 that the second image meets the requirements of human visual perception, then the tone mapping in S601 is valid, and the tone mapping process for the current image can be terminated, proceeding to S604. If it is determined in S602 that the second image does not meet the requirements of human visual perception, then S603 is executed.
[0222] S603, The tone mapping device adjusts the mapping parameters.
[0223] Specifically, in S603, some or all of the parameters in the mapping parameters can be adjusted. For example, in S603, the bit width of the first image can be adjusted, and / or the metadata of the first image can be adjusted.
[0224] In one possible implementation, the bit width of the first image is increased by a first preset step size to complete one adjustment of the mapping parameters. The size of the first preset step size can be configured according to actual needs.
[0225] In another possible implementation, the values of one or more features in the metadata of the first image are increased according to a second preset step size to complete one adjustment of the mapping parameters. The size of the second preset step size can be configured according to actual needs.
[0226] After adjusting the mapping parameters in S603, the process in S601 is re-executed to re-perform tone mapping and enter the next round of tone mapping. This process is repeated iteratively until the tone-mapped image is determined to meet human visual perception in S602. At this point, the tone mapping process for the current image ends, and S604 is executed. The difference is that when S601 is re-executed after S603, tone mapping is re-performed on the first image according to the adjusted mapping parameters.
[0227] S604. The tone mapping device generates a bitstream of a first image, which includes the first image with an initial bit width and metadata of the mapped second image.
[0228] The metadata of the second image obtained through mapping refers to the metadata of the second image obtained from the first image that satisfies the perceptual capabilities of the human eye. This allows the decoding device to perform tone mapping on the first image according to the metadata of the second image obtained through mapping, resulting in a mapped image that satisfies the perceptual capabilities of the human eye, and then displaying the mapped image. In this way, the image can be displayed normally in various display states, and the human eye can have a better and more realistic sensory experience.
[0229] Furthermore, when the mapping parameter is a bit width, and the second image obtained by mapping a first image with a first bit width satisfies human visual perception, the method provided in this application embodiment further includes: generating a bitstream including difference information. This difference information is used to indicate the difference between the first image with the initial bit width and the first image with a first bit width of 1.
[0230] In one possible implementation, the difference information can be: a first image with an initial bit width and a mask image of the first image with a first bit width.
[0231] In another possible implementation, the difference information can be: metadata extracted from the mask image of the first image with the initial bit width and the first image with the first bit width.
[0232] Of course, the content of this difference information can be configured according to actual needs. It is only necessary to show the difference between the first image with the initial bit width and the first image with the first bit width.
[0233] One possible implementation includes a first image with an initial bit width, a bitstream of metadata mapped to a second image, and a bitstream including difference information. This bitstream can be one or multiple bitstreams, and the embodiments of this application are not limited thereto.
[0234] Furthermore, during the execution of S601 to S604, if the mapping parameters are adjusted for a preset number of times, or if the mapping parameters have reached a limit value and a second image that meets the human eye's perception ability is still not obtained, the current process can be terminated, and a first image including the initial bit width and the initial metadata bitstream can be generated.
[0235] Furthermore, if the mapping parameter is metadata, the method in this embodiment may further include: smoothing the mapping curves of adjacent image frames and / or the metadata to ensure that the adjustment process is imperceptible to the user.
[0236] Figure 7 is a schematic flowchart of another tone mapping method provided in this application, which is applied to a tone mapping apparatus. This tone mapping apparatus can be deployed in a decoding device within an image system. For example, the tone mapping apparatus can be deployed in decoding device 302 of the image system illustrated in Figure 3. The tone mapping method illustrated in Figure 7 can be used alone or in combination with the tone mapping method used in Figure 6.
[0237] As shown in Figure 7, the tone mapping method provided in this application embodiment includes:
[0238] S701, The tone mapping device receives a bitstream from the encoding device, the bitstream including a first image and second metadata.
[0239] Specifically, the bitstream received in S701 is the bitstream generated by the encoding device. This bitstream can be the bitstream generated in S604, or the image and metadata included in the bitstream can be an image with the initial bit width and initial metadata.
[0240] S702, The tone mapping device performs tone mapping on the first image according to the second metadata to obtain the second image.
[0241] The first image is the object to which the tone mapping device performs tone mapping in this cycle. In different rounds of tone mapping, the bit width of the first image can be the same or different.
[0242] Specifically, S701 performs tone mapping on the first image. The tone mapping device adapts and adjusts the second metadata according to the parameters of the display device corresponding to the decoding device, and then generates a tone mapping curve according to the adapted and adjusted second metadata. According to the indication of the tone mapping curve, the tone parameters (brightness and / or chromaticity) of the first image are adjusted.
[0243] The details of tone mapping have already been described in S601 and will not be repeated here.
[0244] S703, The tone mapping device determines whether the second image meets the human eye's perception requirements.
[0245] The specific implementation of S703 is described in the aforementioned S602 and will not be repeated here.
[0246] S704, The tone mapping device adjusts the bit width of the first image.
[0247] In one possible implementation, the tone mapping device in S703 can increase the bit width of the first image by a first preset step size to complete one adjustment of the mapping parameters. The size of the first preset step size can be configured according to actual needs.
[0248] In another possible implementation, the tone mapping device also receives a bitstream from the encoding device including difference information, which is used to indicate the difference between the first image with an initial bit width and the first image with a first bit width. Specifically, in S704, the first image is adjusted to the first bit width according to the difference information.
[0249] For information on the differences, please refer to the description in S604 above, and it will not be repeated here.
[0250] After S703, the process of S702 is re-executed to re-perform tone mapping on the first image. This process is repeated iteratively until the tone-mapped second image is deemed satisfactory to human visual perception in S703. At this point, the tone mapping process ends, and S705 is executed. The difference is that when S702 is re-executed after S703, tone mapping on the first image is re-performed according to the adjusted mapping parameters.
[0251] S705, The tone mapping device displays a second image.
[0252] Specifically, the tone mapping device can display a second image through a corresponding display screen. This display screen can be deployed in the same decoding device as the tone mapping device, or it can be deployed in a display device connected to the decoding device.
[0253] Furthermore, during the execution of S701 to S705, if the bit width is adjusted a preset number of times, or if the bit width reaches a limit value and a second image that satisfies human visual perception is still not obtained, the current process can be terminated and the image mapped from the first image based on the initial metadata can be displayed.
[0254] The solution provided in this application adjusts the mapping parameters used for tone mapping to ensure that the tone-mapped image meets the perceptual capabilities of the human eye. This ensures that the mapped image can be perceived by the human eye under various display conditions, achieving the goal of normal image display in all situations. This provides users with a better and more realistic perceptual experience across different display environments, increasing the immersive feeling of the displayed image.
[0255] The principle of the solution provided in this application can be illustrated as shown in Figure 8 or Figure 9.
[0256] Figure 8 illustrates the tone mapping principle of the encoding device. As shown in Figure 8, the encoding device extracts mapping parameters from the image / video, performs tone mapping on the image according to the mapping parameters, and obtains a mapped image. Then, it determines whether the mapped image meets the requirements of human visual perception. If the mapped image meets the requirements of human visual perception, it outputs metadata and the image / video. If the mapped image does not meet the requirements of human visual perception, it updates the mapping parameters and re-performs tone mapping on the image until a mapped image that meets the requirements of human visual perception is obtained.
[0257] Figure 9 illustrates the tone mapping principle of the decoding device. As shown in Figure 9, the decoding device receives an image / video and metadata, performs tone mapping on the image, and obtains a mapped image. Then, it determines whether the mapped image meets the requirements of human visual perception. If the mapped image meets the requirements of human visual perception, it is displayed. If the mapped image does not meet the requirements of human visual perception, the mapping parameters are updated, and tone mapping is re-performed on the image until a mapped image that meets the requirements of human visual perception is obtained.
[0258] The solution provided in this application will be illustrated below through specific embodiments.
[0259] Example 1:
[0260] The scenario in Example 1 is an encoding scenario. The mapping parameter is the bit width of the input image. Figure 10 illustrates the tone mapping process in Example 1. As shown in Figure 10, the tone mapping process may include:
[0261] The input image has an initial bit width, which is initialized and obtained. The input image is quantized to a specific bit width. This specific bit width is the bit width required under current conditions (such as device capabilities or other requirements), and its value is not limited. Metadata is extracted from the input image to obtain initial luminance metadata. Based on the initial luminance metadata, dynamic range tone mapping is performed on the input image to obtain a mapped image. Then, the input image and the mapped image are matched against human visual perception capabilities to determine if the mapped image meets human visual perception requirements. If the current mapped image meets human visual perception requirements, a bitstream of the input image, including the luminance metadata used in this tone mapping and the initial bit width, is generated. If the current mapped image does not meet human visual perception requirements, the process enters an iterative loop, increasing the bit width of the input image and re-entering the mapping process.
[0262] Furthermore, when the mapped image meets the perceptual requirements of the human eye, a high-bit metadata or a high-bit mask image can be generated as output based on the final bit-width image and the initial bit-width image (the original input image). The high-bit metadata or high-bit mask image is used to indicate the difference between the final bit-width image and the initial bit-width image.
[0263] High-bit metadata or high-bit mask images can be encoded and transmitted together with the initial metadata, or they can be encoded and output separately.
[0264] Furthermore, the adjustment of luminance-related hue parameters described in Embodiment 1 can also be an adjustment of chroma-related hue parameters, or an adjustment of both luminance-related and chroma-related hue parameters. Correspondingly, the metadata in Embodiment 1 can be luminance metadata and / or chroma metadata. The aforementioned high-bit metadata can be luminance high-bit metadata and / or chroma high-bit metadata. The aforementioned high-bit mask image can be a luminance high-bit mask image and / or a chroma high-bit mask image.
[0265] As can be seen, the scheme in Implementation Example 1 adds a human eye perception ability matching step, which can combine the input image and the mapped image to match the perception ability, and is used to determine whether the processed image meets the human eye perception ability and whether it can achieve visual losslessness or greater realism.
[0266] Example 2:
[0267] The scenario in Example 2 is an encoding-side scenario. The mapping parameters are the metadata of the input image. Figure 11 illustrates the tone mapping process in Example 2. As shown in Figure 11(a), the brightness tone mapping process may include:
[0268] Metadata is extracted from the input image to obtain luminance metadata. Based on the luminance metadata, dynamic luminance range tone mapping is performed on the input image to obtain a mapped image. Then, the input image and the mapped image are matched against human visual perception capabilities to determine if they meet these requirements. If the mapped image meets human visual perception capabilities, a bitstream of the input image, including the luminance metadata used in this tone mapping and the initial bit width, is generated. If the mapped image does not meet human visual perception capabilities, the process enters an iterative loop, updates the luminance metadata, and re-enters the mapping process.
[0269] Furthermore, the adjustment of luminance-related hue parameters described above in Embodiment 2 can also be the adjustment of chroma-related hue parameters, or the adjustment of both luminance-related and chroma-related hue parameters. Accordingly, the metadata in Embodiment 2 can be luminance metadata and / or chroma metadata. For example, as shown in Figure 11(b), the chroma-hue mapping process is illustrated, and will not be described in detail here.
[0270] Example 3:
[0271] The scenario in Example 3 is a decoding scenario, and the mapping parameter is the image bit width. Figure 12 illustrates the tone mapping process in Example 3.
[0272] As shown in Figure 12(a), the tone mapping process for dynamic brightness range can include:
[0273] Based on the received luminance metadata, the received input image undergoes dynamic luminance range tone mapping to obtain a mapped image. Then, the input image and the mapped image are matched against human visual perception capabilities to determine if they meet these requirements. If the mapped image meets these requirements, it is displayed. If it does not meet these requirements, the process enters an iterative loop, updates the image bit width, and re-enters the mapping process.
[0274] In one possible implementation, in this embodiment, if the encoding end transmits high-bit brightness metadata or a high-bit mask image in addition to transmitting metadata and the image, the bit width of the input image is updated based on the high-bit brightness metadata or the high-bit mask image.
[0275] In another possible implementation, in this embodiment, if the encoding end only transmits metadata and images, the image bit width is increased according to a preset step size to update the input image bit width.
[0276] Furthermore, the adjustment of luminance-related hue parameters described above in Embodiment 3 can also be the adjustment of chroma-related hue parameters, or the adjustment of both luminance-related and chroma-related hue parameters. Accordingly, the metadata in Embodiment 3 can be luminance metadata and / or chroma metadata. For example, as shown in Figure 12(b), the chroma-hue mapping process is illustrated, and will not be described in detail here.
[0277] Example 4:
[0278] Example 4 is a decoding scenario. Figure 13 illustrates the tone mapping process in Example 4. As shown in Figure 13(a), the tone mapping process for dynamic brightness range can include:
[0279] Based on the received luminance metadata, the received image undergoes dynamic luminance range tone mapping to obtain a mapped image. Then, the input image and the mapped image are matched against human visual perception capabilities to determine if they meet those capabilities. If the mapped image meets human visual perception capabilities, it is displayed. If the mapped image does not meet human visual perception capabilities, an iterative loop is entered, updating the dynamic range of the mapped image until it meets human visual perception capabilities and is then displayed.
[0280] For example, the initial mapping requires a peak brightness of 4000 nits, but after matching with human eye perception capabilities, it is found that processing to 4000 nits based on existing materials and specifications does not meet human eye perception capabilities. The dynamic range can be updated and adjusted, for example, by adjusting the peak brightness to 3000 nits, until it meets human eye perception capabilities before displaying.
[0281] Furthermore, in the scenario of Embodiment 4, updates can also be made for the dynamic chroma range, or the dynamic range of luminance can be updated together with the chroma range. For example, as shown in Figure 13(b), the chroma-tone mapping process is illustrated, and will not be described in detail here.
[0282] The foregoing mainly describes the solution provided in this application. Accordingly, this application also provides a tone mapping apparatus for executing the solution in the above method embodiments.
[0283] In some embodiments, the tone mapping apparatus includes hardware structures and / or software modules corresponding to the execution of each function in order to achieve the above-described functions. Those skilled in the art will readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0284] This application embodiment can divide the tone mapping device into functional modules according to the above method embodiment. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0285] On one hand, this application provides a tone mapping device 140, which is used to implement the function of the encoding device in the above method embodiment. As shown in FIG14, the tone mapping device 140 may include a mapping unit 1401, a matching unit 1402, and a processing unit 1403.
[0286] The mapping unit 1401 is used to execute operation S601 in the method illustrated in FIG6. The matching unit 1402 is used to execute operation S602 in the method illustrated in FIG6. The processing unit 1403 is used to execute operations S603 and / or S604 in the method illustrated in FIG6.
[0287] On the other hand, this application provides a tone mapping device 150, which is used to implement the function of the decoding device in the above method embodiment. As shown in FIG15, the tone mapping device 150 may include a receiving unit 1501, a mapping unit 1502, a matching unit 1503, and a processing unit 1504.
[0288] The receiving unit 1501 is used to execute operation S701 in the method illustrated in FIG7. The mapping unit 1502 is used to execute operation S702 in the method illustrated in FIG7. The matching unit 1503 is used to execute operation S703 in the method illustrated in FIG7. The processing unit 1504 is used to execute operations S704 and / or S705 in the method illustrated in FIG7.
[0289] In another aspect, embodiments of this application provide an imaging system, including an encoding device and a decoding device. The encoding device is used to perform the method illustrated in FIG. 6, or the decoding device is used to perform the method illustrated in FIG. 7.
[0290] Furthermore, embodiments of this application also provide a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform some or all of the operational steps of the method in the above method embodiments.
[0291] Furthermore, embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform some or all of the operational steps of the method in the above-described method embodiments.
[0292] Furthermore, embodiments of this application also provide a chip, including a processor and a power supply circuit. The power supply circuit supplies power to the processor; the processor executes some or all of the operational steps of the method described in the method embodiments.
[0293] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a computing device. Of course, the processor and storage medium can also exist as discrete components in the computing device.
[0294] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD). The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A tone mapping method, characterized in that, Applied to an encoding device, the method includes: A second image is obtained by performing tone mapping on the first image according to the mapping parameters; the mapping parameters include the bit width of the first image and / or the metadata of the first image. If the second image does not meet the human eye's perception capability, the mapping parameters are adjusted, and the tone mapping is re-executed on the first image according to the adjusted mapping parameters, wherein the human eye's perception capability is the range of human eye's perception of image quality; If the second image satisfies human visual perception, a bitstream of the first image is generated. The bitstream includes the first image with an initial bit width and metadata mapped to the second image.
2. The method according to claim 1, characterized in that, The mapping parameter is the bit width of the first image; when the second image obtained by mapping the first image with the first bit width satisfies the human eye's perception ability, the method further includes: generating a code stream including difference information, the difference information being used to indicate the difference between the first image with the initial bit width and the first image with the first bit width.
3. The method according to claim 1, characterized in that, The mapping parameters are the metadata of the first image; when the second image obtained by performing tone mapping on the first image according to the first metadata satisfies the human eye's perception ability, the metadata of the second image obtained by mapping is the first metadata.
4. A tone mapping method, characterized in that, Applied to a decoding device, the method includes: Receive a bitstream from an encoding device, the bitstream including a first image and second metadata; Based on the second metadata, tone mapping is performed on the first image to obtain the second image; If the second image does not meet the human eye's perception ability, adjust the bit width of the first image, and re-perform tone mapping on the first image after adjusting the bit width according to the second metadata; If the second image satisfies the human eye's perception capabilities, then the second image is displayed.
5. The method according to claim 4, characterized in that, The method further includes: receiving a bitstream including difference information, the difference information being used to indicate the difference between the first image with an initial bit width and the first image with a first bit width; Adjusting the bit width of the first image includes: adjusting the first image to the first bit width based on the difference information.
6. The method according to claim 2 or 5, characterized in that, The difference information includes: The first image with the initial bit width and the mask image of the first image with the first bit width; or, Metadata extracted based on the first image with the initial bit width and the mask image of the first image with the first bit width.
7. The method according to any one of claims 1-6, characterized in that, The second image does not meet the requirements of human visual perception, including: The first quality index of the second image relative to the first image with the initial bit width does not meet the human eye's perception capability. and / or The second quality index of the second image does not meet the requirements of human visual perception.
8. The method according to claim 7, characterized in that, The first quality indicator is the peak signal-to-noise ratio (PSNR). The first quality indicator does not meet the human eye's perception ability, including: the PSNR ratio is less than or equal to a first threshold. or, The first quality indicator is the structural similarity index (SSIM). The first quality indicator does not meet the human eye's perception ability, including: the SSIM is less than or equal to a second threshold. or, The first quality indicator is the Feature Similarity Index (FSIM). The first quality indicator does not meet the human eye's perception ability, including: the FSIM is less than or equal to a third threshold. or, The first quality indicator is visual information fidelity (VIF). The first quality indicator does not meet the human eye's perception ability, including: the VIF is less than or equal to the fourth threshold. or, The first quality indicator is gradient magnitude similarity deviation (GMSD). The first quality indicator does not meet the human eye's perception ability, including: the GMSD is greater than or equal to the fifth threshold. or, The first quality metric is the learning-perceptual image patch similarity LPIPS. The first quality metric does not meet the human eye's perception ability, including: the LPIPS is greater than or equal to the sixth threshold.
9. The method according to claim 7, characterized in that, The second quality index is the first evaluation result obtained by the BRISQUE spatial quality evaluator without reference image. The second quality index does not meet the human eye perception ability, including: the first evaluation result is greater than or equal to the seventh threshold. or, The second quality index is the second evaluation result obtained by the Natural Image Quality Evaluator (NIQE). The second quality index does not meet the human eye's perception ability, including: the second evaluation result is greater than or equal to the eighth threshold. or, The second quality indicator is information entropy. The second quality indicator does not meet the human eye's perception ability, including: the information entropy is less than or equal to the ninth threshold. or, The second quality index is the sharpness index. The second quality index does not meet the human eye's perception ability, including: the sharpness index is less than or equal to the tenth threshold.
10. The method according to any one of claims 1-7, characterized in that, The tone mapping includes luminance mapping and / or chrominance mapping.
11. The method according to any one of claims 1-8, characterized in that, The metadata includes: luminance metadata for describing luminance characteristics, and / or chrominance metadata for describing color characteristics.
12. The method according to claim 9, characterized in that, The metadata also includes environmental data, which describes the environmental characteristics at the time the first image was acquired.
13. A tone mapping device, characterized in that, Applied to an encoding device, the device includes: A mapping unit is configured to perform tone mapping on a first image according to mapping parameters to obtain a second image; the mapping parameters include the bit width of the first image and / or the metadata of the first image. A matching unit is used to determine whether the second image meets the human eye's perception capability; wherein, the human eye's perception capability is the range of human eye's perception of image quality; The processing unit is configured to adjust the mapping parameters if the matching unit determines that the second image does not meet the human eye's perception capability; and to generate a bitstream of the first image if the matching unit determines that the second image meets the human eye's perception capability, wherein the bitstream includes the first image with an initial bit width and metadata mapped to the second image. The mapping unit is further configured to re-perform tone mapping on the first image according to the mapping parameters adjusted by the adjustment unit.
14. A tone mapping device, characterized in that, Applied to a decoding device, the device includes: A receiving unit is configured to receive a bitstream from an encoding device, the bitstream including a first image and second metadata; A mapping unit is configured to perform tone mapping on the first image based on the second metadata to obtain a second image; The matching unit is used to determine whether the second image meets the human eye's perceptual ability; The processing unit is configured to adjust the bit width of the first image if the matching unit determines that the second image does not meet the human eye's perception capability; and to display the second image if the matching unit determines that the second image meets the human eye's perception capability. The mapping unit is further configured to: re-perform tone mapping on the first image after adjusting the bit width based on the second metadata.
15. A computing device, characterized in that, The computing device includes a processor and a memory; The processor is configured to execute instructions stored in the memory to cause the computing device to perform the operational steps of the method as described in any one of claims 1 to 12.
16. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device, the computing device performs the operational steps of the method as described in any one of claims 1 to 12.
17. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a computing device, cause the computing device to perform the operational steps of the method as described in any one of claims 1 to 12.