Image encoding method, image decoding method, and related apparatus

By using optimization equations or deep learning network models to determine the sampling point position of the tone mapping curve at the image encoding end, the problem of poor image display effect after tone mapping in the prior art is solved, and a better image display effect is achieved.

WO2025092103A1PCT designated stage expired Publication Date: 2025-05-08HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/111449
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-08-12
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

In the prior art, when mapping HDR images tones, commonly used tone mapping curves such as Dolby sigmoidal curves and Bezier curves may cause the mapped image to be distorted, resulting in poor display effects.

Method used

By optimizing equations or deep learning network models, the position coordinates of multiple sampling points of the tone mapping curve corresponding to the image are determined, and a more suitable tone mapping curve is constructed to take into account image contrast, image brightness, and contrast human eye perception thresholds.

Benefits of technology

The display effect of the tone-mapping image in ordinary display devices is improved, so that the brightness and contrast of the image are consistent with the original image to the greatest extent, and the human eye's perception of image details is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024111449_08052025_PF_FP_ABST
    Figure CN2024111449_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing, and discloses an image encoding method, an image decoding method, and a related apparatus. The method comprises: acquiring a first image, wherein the first image is an image requiring tone mapping; on the basis of the first image, and by means of an optimization equation or a deep learning network model, determining position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image; and encoding the first image and the position coordinates of the plurality of sampling points into a bitstream. According to the present application, an optimization equation or a deep learning model is established based on at least one of an image contrast, image brightness and a contrast human eye perception threshold, and an image requiring tone mapping is encoded by means of the optimization equation or the deep learning model. In this way, a decoding end establishes a tone mapping curve on the basis of a bitstream, and achieves tone mapping of the image, so that the display effect of the image is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Image encoding, image decoding method and related device

[0001] This application claims priority to Chinese patent application No. 202311440178.3 filed on October 31, 2023, entitled “Image Coding, Image Decoding Methods and Related Devices,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of image processing, and in particular to an image encoding and decoding method and related devices. Background Art

[0003] High dynamic range (HDR) technology has developed rapidly in recent years. HDR technology can enhance the contrast between the extremes of brightness and darkness in an image, revealing rich details in bright and dark areas, and presenting images that are closer to the real human eye's perception. However, ordinary display devices have a limited dynamic range, so tone mapping is required to ensure that HDR images can be properly displayed on ordinary display devices.

[0004] In related technologies, tone mapping methods typically align the maximum and minimum image brightness of an image to the maximum and minimum screen brightness values ​​of the display. For image brightnesses between the maximum and minimum image brightnesses, global or local mapping is performed based on a tone mapping curve to map these image brightnesses to a range between the maximum and minimum screen brightnesses. However, the tone mapping curves currently used are typically Dolby sigmoidal curves, Bezier curves, etc., and tone mapping using these curves may distort the tone-mapped image, resulting in poor display quality.

[0005] Summary of the Invention

[0006] This application provides an image encoding and decoding method and related apparatus, which can improve the display effect of tone-mapped images on ordinary display devices. The technical solution is as follows:

[0007] In a first aspect, an image encoding method is provided, the method comprising:

[0008] Acquire a first image, the first image being an image to be tone mapped; determine, based on the first image, position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image using an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a human eye perception threshold of contrast, the plurality of sampling points including a starting point, an ending point, and a plurality of key points located between the starting point and the ending point; and encode the first image and the position coordinates of the plurality of sampling points into a bitstream.

[0009] That is to say, the present application determines multiple sampling points of the tone mapping curve corresponding to the first image at the encoding end, and the position coordinates of multiple key points among these sampling points are determined by an optimization equation or a deep learning network model, and the optimization equation or deep learning network model takes into account the factors affecting the display effect - at least one of the image contrast, image brightness and contrast human eye perception threshold. Therefore, the tone mapping curve constructed based on multiple sampling points at the decoding end can enable the mapped image to achieve the best display effect.

[0010] Image contrast refers to the ratio between the maximum and minimum brightness of a region. For a pixel in an image, the brightness of that pixel is the image brightness corresponding to that pixel. The contrast human eye perception threshold refers to the minimum contrast that the human eye can perceive for a pixel when the brightness of that pixel changes.

[0011] It should be noted that since the essence of tone mapping is to compress the high dynamic range into the low dynamic range, the first image can not only be an HDR / SDR image, but also other high dynamic range images, as long as the brightness range of the first image is greater than the brightness range that can be displayed by the decoding end.

[0012] Optionally, based on the first image, the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image are determined by optimizing an equation or a deep learning network model, including: generating a target histogram based on the first image, the target histogram including multiple histogram intervals, the multiple histogram intervals being divided based on the maximum image brightness and the minimum image brightness of the first image, the height of the histogram interval indicating the number of pixels in the first image whose brightness is within the histogram interval; determining the position coordinates of a starting point and an ending point, the position coordinates including a horizontal coordinate and a vertical coordinate; selecting a point from each of the multiple histogram intervals as one of multiple key points, and using the image brightness corresponding to each of the multiple key points as the horizontal coordinates of the multiple key points; determining the target probabilities corresponding to each of the multiple sampling points, wherein the target probabilities corresponding to the starting point and the ending point are specified probabilities, and the target probabilities corresponding to the key points indicate the probabilities that the image brightness of the key points falls within the histogram interval; determining the vertical coordinates of the multiple key points based on the position coordinates of the starting point and the ending point, the target probabilities corresponding to each of the multiple sampling points, and the horizontal coordinates of the multiple key points.

[0013] That is to say, there are two ways to determine the vertical coordinates of multiple key points. One is to determine them through optimization equations, and the other is to determine them through a deep learning network model.

[0014] Optionally, based on the position coordinates of the starting point and the ending point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points, the vertical coordinates of the multiple key points are determined by an optimization equation, including: determining a contrast human eye perception threshold of the multiple sampling points based on the horizontal coordinates of the multiple sampling points; based on the position coordinates of the starting point and the ending point, the number of the multiple sampling points, the target probabilities corresponding to the multiple sampling points, the horizontal coordinates of the multiple key points, and the contrast human eye perception threshold of the multiple sampling points, obtaining the vertical coordinates of the multiple key points by solving the optimization equation.

[0015] Optionally, based on the position coordinates of the starting point and the ending point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points, the vertical coordinates of the multiple key points are determined through a deep learning network model, including: inputting the maximum image brightness, the minimum image brightness and the target probabilities corresponding to the multiple sampling points into the deep learning network model to obtain the derivatives of the vertical coordinates of the multiple key points output by the deep learning network model; determining the vertical coordinates of the multiple key points based on the position coordinates of the starting point and the ending point, the horizontal coordinates of the multiple key points, and the derivatives of the vertical coordinates of the multiple key points.

[0016] Optionally, the method also includes: obtaining multiple training samples, each training sample includes the maximum sample image brightness, the minimum sample image brightness, the target probabilities corresponding to multiple sample sampling points, and the derivatives of the vertical coordinates of multiple sample key points, and the derivatives of the vertical coordinates of the multiple sample key points are determined by optimization equations; based on the multiple training samples, the initial network model is trained to obtain a deep learning network model.

[0017] Since the training samples of the deep learning network model are determined by the optimization equation, the deep learning network model can take into account at least one of the image brightness, image contrast and the contrast perception threshold of the human eye in the same way as the optimization equation, and can also enable the image after subsequent tone mapping to achieve better display effects.

[0018] Optionally, the above optimization equation includes but is not limited to any one of the following equations:

[0019] or

[0020] or

[0021] Among them, y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of multiple sampling points, and p k Indicates the target probability corresponding to the kth sampling point, x k Indicates the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

[0022] In the first equation above, This section takes into account image contrast and the human eye's contrast perception threshold. It uses the argmin function to make the slope of the resulting tone mapping curve as close to 1 as possible, thereby maximizing image contrast restoration. This means that the contrast of the tone-mapped image is kept as close to that of the original image as possible.

[0023] In the second equation above, This part takes into account the image contrast and the contrast perception threshold of the human eye. It achieves the maximum restoration of image contrast by minimizing the difference between the pixel gradient of the tone-mapped image and the pixel gradient of the original image, that is, through spatial consistency.

[0024] In the third equation above, This part considers the image contrast and the contrast threshold of human eyes, and applies the theory of Pearson correlation coefficient to make y k+1 -y k and t(x k )(x k+1 -x k ) is as small as possible to achieve the maximum restoration of image contrast.

[0025] At the same time, since the minimum contrast that the human eye can perceive is different for pixels with different brightness when the brightness of the pixel changes, the above three equations can be solved by adding t(x k ), that is, the contrast human eye perception threshold corresponding to the k-th sampling point, to improve the display effect of the image. This part takes the image brightness into consideration, with the goal of making the brightness value of each pixel in the tone-mapped image as consistent as possible with the brightness value of each pixel in the first image.

[0026] In addition, the weight coefficient λ is used to adjust the proportions of the two parts on the right side of the equation. The smaller λ is, the more the optimization equation emphasizes the impact of image contrast and the human eye's contrast perception threshold on the image display effect compared to image brightness. Conversely, the optimization equation emphasizes the impact of image brightness on the image display effect. This weight coefficient is set in advance and can be set according to actual needs. This embodiment of the application does not limit this.

[0027] When the weight coefficient is set to 0, the optimization equation only considers the two factors of image contrast and contrast human eye perception threshold; when the weight coefficient is set to 1, the optimization equation only considers the image brightness factor; when the weight coefficient is set to any value between 0 and 1, the optimization equation considers the three factors of image contrast, contrast human eye perception threshold and image brightness at the same time. Of course, when the weight coefficient is set to 0 and t(x k ), the optimization equation only considers the image contrast.

[0028] In a second aspect, an image decoding method is provided, the method comprising:

[0029] A reconstructed image is obtained based on a bitstream; position coordinates of a plurality of sampling points are parsed from the bitstream, the position coordinates of the plurality of sampling points being determined by an optimization equation or a deep learning network model, the optimization equation and the deep learning network model being determined based on at least one of image contrast, image brightness, and a human eye perception threshold of contrast, the plurality of sampling points including a starting point, an ending point, and a plurality of key points located between the starting point and the ending point; a tone mapping curve is obtained based on the position coordinates of the plurality of sampling points; and tone mapping is performed on the reconstructed image based on the tone mapping curve, a maximum screen brightness, and a minimum screen brightness.

[0030] The reconstructed image refers to an image constructed based on the first image data in the code stream; the maximum screen brightness is the maximum brightness value that can be displayed by the decoding end; and the minimum screen brightness is the minimum brightness value that can be displayed by the decoding end.

[0031] When the decoding end establishes the tone mapping curve, the position coordinates of the multiple key points are determined by an optimization equation or a deep learning network model, and the optimization equation or the deep learning network model takes into account at least one of the image contrast, image brightness and contrast human eye perception threshold, so that the brightness and contrast of the image can be kept consistent with the first image to the greatest extent, and the human eye can perceive the difference between light and dark at different image brightnesses, and thus perceive more details in the tone-mapped image. Therefore, after mapping the image through the tone mapping curve, a better display effect can be achieved.

[0032] Optionally, the tone mapping curve is determined by curve fitting, and the curve fitting method includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

[0033] Optionally, the above optimization equation includes but is not limited to any one of the following equations:

[0034] or

[0035] or

[0036] Among them, y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of multiple sampling points, and p k Indicates the target probability corresponding to the kth sampling point, x k Indicates the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

[0037] In the first equation above, This section takes into account image contrast and the human eye's contrast perception threshold. It uses the argmin function to make the slope of the resulting tone mapping curve as close to 1 as possible, thereby maximizing image contrast restoration. This means that the contrast of the tone-mapped image is kept as close to that of the original image as possible.

[0038] In the second equation above, This part takes into account the image contrast and the contrast perception threshold of the human eye. It achieves the maximum restoration of image contrast by minimizing the difference between the pixel gradient of the tone-mapped image and the pixel gradient of the original image, that is, through spatial consistency.

[0039] In the third equation above, This part considers the image contrast and the contrast threshold of human eyes, and applies the theory of Pearson correlation coefficient to make y k+1 -y k and t(x k )(xk +1 -x k ) is as small as possible to achieve the maximum restoration of image contrast.

[0040] At the same time, since the minimum contrast that the human eye can perceive is different for pixels with different brightness when the brightness of the pixel changes, the above three equations can be solved by adding t(x k ), that is, the contrast human eye perception threshold corresponding to the k-th sampling point, to improve the display effect of the image. This part takes the image brightness into consideration, with the goal of making the brightness value of each pixel in the tone-mapped image as consistent as possible with the brightness value of each pixel in the first image.

[0041] In addition, the weight coefficient λ is used to adjust the proportions of the two parts on the right side of the equation. The smaller λ is, the more the optimization equation emphasizes the impact of image contrast and the human eye's contrast perception threshold on the image display effect compared to image brightness. Conversely, the optimization equation emphasizes the impact of image brightness on the image display effect. This weight coefficient is set in advance and can be set according to actual needs. This embodiment of the application does not limit this.

[0042] When the weight coefficient is set to 0, the optimization equation only considers the two factors of image contrast and contrast human eye perception threshold; when the weight coefficient is set to 1, the optimization equation only considers the image brightness factor; when the weight coefficient is set to any value between 0 and 1, the optimization equation considers the three factors of image contrast, contrast human eye perception threshold and image brightness at the same time. Of course, when the weight coefficient is set to 0 and t(x k ), the optimization equation only considers the image contrast.

[0043] In a third aspect, an image coding apparatus is provided, wherein the apparatus has the function of implementing the image coding method in the first aspect. The image coding apparatus includes at least one module, wherein the at least one module is configured to implement the image coding method in the first aspect.

[0044] In a fourth aspect, an image decoding device is provided, wherein the device has the function of implementing the image decoding method in the second aspect. The image decoding device includes at least one module, wherein the at least one module is used to implement the image decoding method provided in the second aspect.

[0045] In a fifth aspect, an encoding end device is provided, comprising a processor and a memory, wherein the memory is configured to store a computer program for executing the image encoding method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the image encoding method described in the first aspect.

[0046] Optionally, the encoding end device may further include a communication bus, which is used to establish a connection between the processor and the memory.

[0047] In a sixth aspect, a decoding device is provided, comprising a processor and a memory, wherein the memory is configured to store a computer program for executing the image decoding method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the image decoding method provided in the second aspect.

[0048] Optionally, the decoding end device may further include a communication bus, which is used to establish a connection between the processor and the memory.

[0049] In the seventh aspect, a computer-readable storage medium is provided, wherein the storage medium stores instructions. When the instructions are executed on a computer, the computer executes the steps of the image encoding method described in the first aspect, or executes the steps of the image decoding method described in the second aspect.

[0050] In an eighth aspect, a computer program product comprising instructions is provided. When the instructions are executed on a computer, the computer is caused to execute the steps of the image encoding method described in the first aspect, or the steps of the image decoding method described in the second aspect. Alternatively, a computer program is provided. When the computer program is executed on a computer, the computer is caused to execute the steps of the image encoding method described in the first aspect, or the steps of the image decoding method described in the second aspect.

[0051] The technical effects obtained in the above-mentioned third, fourth, fifth, sixth, seventh and eighth aspects are similar to the technical effects obtained by the corresponding technical means in the first or second aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] FIG1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0053] FIG2 is a flowchart of an image encoding method provided by an embodiment of the present application;

[0054] FIG3 is a flow chart of a method for determining the position coordinates of multiple sampling points provided in an embodiment of the present application;

[0055] FIG4 is a schematic diagram of a target histogram and a process of constructing a tone mapping curve provided by an embodiment of the present application;

[0056] FIG5 is a schematic diagram of a process for determining the vertical coordinates of multiple key points by a deep learning network model according to an embodiment of the present application;

[0057] FIG6 is a flowchart of an image decoding method provided by an embodiment of the present application;

[0058] FIG7 is a schematic diagram of another process for constructing a tone mapping curve provided by an embodiment of the present application;

[0059] FIG8 is a schematic structural diagram of an image encoding device provided in an embodiment of the present application;

[0060] FIG9 is a schematic structural diagram of an image decoding device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0062] Before explaining in detail the image encoding method and image decoding method provided in the embodiments of the present application, the application scenarios and implementation environments involved in the embodiments of the present application are first introduced.

[0063] First, the application scenarios involved in the embodiments of the present application are introduced.

[0064] Natural scenes have an extremely wide range of colors and brightness, typically with a maximum brightness approaching 10^6 cd / m^2 and a minimum brightness approaching 10^(-3) cd / m^2. The human visual system can automatically adjust to this brightness variation over a range of nearly 10 orders of magnitude. However, images generated by digital cameras have a limited dynamic range. To reproduce the real world as closely as possible, it is necessary to capture a larger dynamic range, which has driven the rapid development of HDR imaging technology. HDR images are stored in floating-point format and have a large dynamic range. However, ordinary displays generally only have 8 bits, which has a very limited dynamic range. Therefore, if HDR images are to be displayed properly on ordinary displays, the dynamic range of the HDR image must be compressed to fit within the dynamic range of the display device. This process is called tone mapping.

[0065] In related technologies, the tone mapping method generally aligns the maximum image brightness and minimum image brightness of an image to the maximum screen brightness value and minimum screen brightness of the display. For image brightness between the maximum image brightness and the minimum image brightness, global or local mapping is performed based on the tone mapping curve to map these image brightnesses to the range between the maximum screen brightness and the minimum screen brightness. However, the tone mapping curves currently used do not consider the impact of image contrast on the display effect, so that the original image cannot achieve the best display effect. For example, tone mapping curves such as the Dolby sigmoidal curve and the Bezier curve will cause a certain loss of details in the mapped image, and the local contrast is not high, resulting in distortion of the mapped image. In addition, the tone mapping curves currently used do not consider the factors of human eye perception, resulting in a large visual difference between the mapped image and the real scene.

[0066] Based on this, the embodiments of the present application establish an optimization equation or a deep learning model based on at least one of image contrast, image brightness, and the contrast human eye perception threshold, and encode the image that needs to be tone mapped through the optimization equation or deep learning model. In this way, the decoding end establishes a tone mapping curve based on the bitstream to achieve image tone mapping, so that the same HDR image can obtain a corresponding tone mapping curve for different display devices, so that the contrast and brightness of the image after tone mapping are as consistent as possible with the HDR image before tone mapping, and the image display effect perceived by the human eye is improved.

[0067] Next, the implementation environment involved in the embodiments of this application is introduced.

[0068] Please refer to Figure 1, which is a schematic diagram of an implementation environment provided by an embodiment of the present application. The implementation environment includes a source device 10, a destination device 20, a link 30, and a storage device 40. The source device 10 can generate encoded images. Therefore, the source device 10 can also be referred to as an image encoding device or encoding end. The destination device 20 can decode the encoded images generated by the source device 10. Therefore, the destination device 20 can also be referred to as an image decoding device or decoding end. The link 30 can receive the encoded images generated by the source device 10 and transmit the encoded images to the destination device 20. The storage device 40 can receive the encoded images generated by the source device 10 and store them. Under such conditions, the destination device 20 can directly obtain the encoded images from the storage device 40. Alternatively, the storage device 40 can correspond to a file server or another intermediate storage device that can store the encoded images generated by the source device 10. Under such conditions, the destination device 20 can stream or download the encoded images stored by the storage device 40.

[0069] The source device 10 and the destination device 20 may each include one or more processors and a memory coupled to the one or more processors, wherein the memory may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium that can be used to store desired program code in the form of computer-accessible instructions or data structures. For example, the source device 10 and the destination device 20 may each include a mobile phone, a smartphone, a personal digital assistant (PDA), a wearable device, a pocket PC (PPC), a tablet computer, a smart car computer, a smart TV, a smart speaker, a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video game console, an in-vehicle computer, or the like.

[0070] Link 30 may include one or more media or devices capable of transmitting encoded images from source device 10 to destination device 20. In one possible implementation, link 30 may include one or more communication media that enable source device 10 to send encoded images directly to destination device 20 in real time. In an embodiment of the present application, source device 10 may modulate the encoded images based on a communication standard, such as a wireless communication protocol, and may transmit the modulated images to destination device 20. The one or more communication media may include wireless and / or wired communication media, such as radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, such as a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from source device 10 to destination device 20, although this embodiment of the present application does not specifically limit this.

[0071] In one possible implementation, the storage device 40 may store the received encoded image sent by the source device 10, and the destination device 20 may directly obtain the encoded image from the storage device 40. Under such conditions, the storage device 40 may include any of a variety of distributed or locally accessible data storage media, for example, any of the various distributed or locally accessible data storage media may be a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded images.

[0072] In one possible implementation, storage device 40 may correspond to a file server or another intermediate storage device that can store the encoded images generated by source device 10, and destination device 20 may stream or download the images stored on storage device 40. The file server may be any type of server capable of storing and transmitting the encoded images to destination device 20. In one possible implementation, the file server may include a network server, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Destination device 20 may obtain the encoded images via any standard data connection, including an internet connection. Any standard data connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of the two suitable for obtaining the encoded images stored on the file server. The transmission of the encoded images from storage device 40 may be streaming, downloading, or a combination of the two.

[0073] The implementation environment shown in FIG1 is only one possible implementation method, and the technology of the embodiment of the present application is applicable not only to the source device 10 that can encode images and the destination device 20 that can decode encoded images shown in FIG1 , but also to other devices that can encode images and decode encoded images, and the embodiment of the present application does not make specific limitations on this.

[0074] In the implementation shown in FIG1 , source device 10 includes a data source 120, an encoder 100, and an output interface 140. In some embodiments, output interface 140 may include a modem and / or a transmitter, where the transmitter may also be referred to as a transmitter. Data source 120 may include an image capture device (e.g., a camera), an archive containing previously captured images, a feed interface for receiving images from an image content provider, and / or a computer graphics system for generating images, or a combination of these sources of images.

[0075] The data source 120 may send an image to the encoder 100, and the encoder 100 may encode the image received from the data source 120 to generate an encoded image. The encoder may send the encoded image to an output interface. In some embodiments, the source device 10 directly sends the encoded image to the destination device 20 via the output interface 140. In other embodiments, the encoded image may also be stored on the storage device 40 for later retrieval by the destination device 20 for decoding and / or display.

[0076] In the implementation environment shown in FIG1 , the destination device 20 includes an input interface 240, a decoder 200, and a display device 220. In some embodiments, the input interface 240 includes a receiver and / or a modem. The input interface 240 may receive encoded images via the link 30 and / or from the storage device 40, and then transmit the encoded images to the decoder 200. The decoder 200 may decode the received encoded images to obtain decoded images. The decoder may transmit the decoded images to the display device 220. The display device 220 may be integrated with the destination device 20 or may be external to the destination device 20. Generally, the display device 220 displays the decoded images. The display device 220 may be any of a variety of types of display devices, for example, the display device 220 may be a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.

[0077] Although not shown in FIG1 , in some aspects, encoder 100 and decoder 200 can be integrated with an encoder and decoder, respectively, and can include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software for encoding both audio and video in a common data stream or in separate data streams. In some embodiments, the MUX-DEMUX units can conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP), if applicable.

[0078] The encoder 100 and the decoder 200 can each be any of the following circuits: one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology of the embodiments of the present application is implemented in part by software, the device can store instructions for the software in a suitable non-volatile computer-readable storage medium, and can use one or more processors to execute the instructions in hardware to implement the technology of the embodiments of the present application. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) can be regarded as one or more processors. Each of the encoder 100 and the decoder 200 can be included in one or more encoders or decoders, and any of the encoders or decoders can be integrated as part of a combined encoder / decoder (encoder / decoder) in the corresponding device.

[0079] Embodiments of the present application may generally refer to encoder 100 as "signaling" or "sending" certain information to another device, such as decoder 200. The terms "signaling" or "sending" may generally refer to the transmission of syntax elements and / or other data used to decode a compressed image. This transmission may occur in real time or near real time. Alternatively, this communication may occur over time, such as when syntax elements are stored in the encoded bitstream to a computer-readable storage medium during encoding, and a decoding device may then retrieve the syntax elements at any time after they are stored to this medium.

[0080] It should be noted that the application scenarios and implementation environments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of application scenarios and implementation environments, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0081] Next, the image encoding method and image decoding method provided in the embodiments of the present application are explained in detail.

[0082] FIG2 is a flowchart of an image encoding method provided by an embodiment of the present application, which can be applied to a source device in the above implementation environment, which is also called an encoding end. Referring to FIG2 , the method includes the following steps.

[0083] Step 201: Acquire a first image, where the first image is an image that needs to be tone mapped.

[0084] The first image can be an image obtained from an HDR / standard dynamic range (SDR) video stream, or an image obtained from an HDR / SDR image source, or an image from other sources. The first image can be in RGB color space format, YUV color space format, or any other color space format. The embodiment of the present application does not limit the source and color space format of the first image.

[0085] Since the essence of tone mapping is to compress a high dynamic range into a low dynamic range, the first image can be not only an HDR / SDR image, but also an image with other high dynamic ranges, as long as the brightness range of the first image is greater than the brightness range that can be displayed by the decoding end.

[0086] Step 202: Based on the first image, determine the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image through an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and contrast human eye perception threshold, and the multiple sampling points include a starting point, an end point, and multiple key points located between the starting point and the end point.

[0087] Image contrast refers to the ratio between the maximum and minimum brightness values ​​of a region. For a pixel in an image, the brightness value of that pixel is the image brightness corresponding to that pixel. The contrast human eye perception threshold refers to the minimum contrast that the human eye can perceive for a pixel when the brightness of that pixel changes.

[0088] In some embodiments, as shown in FIG3 , based on the first image, the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image are determined by optimizing an equation or a deep learning network model according to the following steps (1)-(4).

[0089] (1) A target histogram is generated based on a first image, where the target histogram includes a plurality of histogram intervals, where the plurality of histogram intervals are divided based on maximum image brightness and minimum image brightness of the first image, and where the height of the histogram interval indicates the number of pixels in the first image whose brightness is within the histogram interval.

[0090] In some embodiments, the brightness value of each pixel in the first image is obtained, the maximum brightness value is used as the maximum image brightness, and the minimum brightness value is used as the minimum image brightness; the difference between the maximum image brightness and the minimum image brightness is divided by the specified number of intervals to obtain the width of each histogram interval, and the brightness range contained in each histogram interval is determined; the number of pixels in the first image whose brightness is within each histogram interval is counted, and the number is used as the height of each histogram interval.

[0091] Among them, the number of specified intervals is set in advance. The more the number, the more accurate the tone mapping curve will be, but at the same time the system complexity will also become higher. In actual applications, reasonable settings can be made according to needs, and the embodiments of this application do not limit this.

[0092] For example, please refer to Figure 4. Assuming that the maximum image brightness of the first image is Xmax, the minimum image brightness is Xmin, and the number of specified intervals is M, the generated target histogram is shown in the dotted part of Figure 4. The width of each histogram interval is (Xmax-Xmin) / M, the starting point of the first histogram interval is Xmin, and the end point of the last histogram interval is Xmax.

[0093] (2) Determine the position coordinates of the starting point and the ending point, where the position coordinates include a horizontal coordinate and a vertical coordinate, select a point from each of the multiple histogram intervals as one of the multiple key points, and use the image brightness corresponding to the multiple key points as the horizontal coordinates of the multiple key points.

[0094] In some embodiments, the position coordinates of the starting point and the ending point are set default coordinates. The default coordinates are set through historical statistical data. That is, for the horizontal coordinate of the starting point, it can be set to the minimum brightness value of the image that needs to be tone mapped in the historical statistical data; for the vertical coordinate of the starting point, it can be set to the minimum brightness value that the decoding end can display in the historical statistical data; for the horizontal coordinate of the ending point, it can be set to the maximum brightness value of the image that needs to be tone mapped in the historical statistical data; for the vertical coordinate of the ending point, it can be set to the maximum brightness value that the decoding end can display in the historical statistical data; of course, technicians can also set the horizontal / vertical coordinates of the starting point / end point based on subjective experience, and the embodiment of the present application does not limit the values ​​of the horizontal / vertical coordinates of the starting point / end point.

[0095] In other embodiments, the position coordinates of the starting point and the ending point are determined based on the first selection interval, the second selection interval, the minimum reference brightness, and the maximum reference brightness. That is, the horizontal coordinate of the starting point can be set to the horizontal coordinate of any point in the first selection interval; the horizontal coordinate of the ending point can be set to the horizontal coordinate of any point in the second selection interval; the vertical coordinate of the starting point can be set to any brightness value within the fluctuation range of the minimum reference brightness, and the brightness value is greater than 0; the vertical coordinate of the ending point can be set to any brightness value within the fluctuation range of the maximum reference brightness.

[0096] Among them, the first selection interval is the interval from 0 to the right endpoint of the first histogram interval, and is a fully closed interval. The second selection interval is the interval from the left endpoint of the last histogram interval to positive infinity, and is a left-closed and right-open interval. The abscissa of the right endpoint of the first histogram interval is the sum of the minimum image brightness and the width of the histogram interval, and the abscissa of the left endpoint of the last histogram interval is the difference between the maximum image brightness and the width of the histogram interval.

[0097] The minimum reference brightness refers to the minimum brightness value that the reference decoding end can display, and the maximum reference brightness refers to the maximum brightness value that the reference decoding end can display. The reference decoding end is any decoding end. In other words, the embodiments of the present application determine the vertical coordinates of the starting and ending points by setting the minimum reference brightness and the maximum reference brightness in advance. In addition, the left and right fluctuation ranges of the minimum reference brightness and the maximum reference brightness are also set in advance to indicate the small fluctuation range near the minimum reference brightness and the maximum reference brightness. The left and right fluctuation ranges of the minimum reference brightness and the maximum reference brightness can be the same or different, and the embodiments of the present application do not limit this.

[0098] In some embodiments, multiple key points are determined according to a specified rule. The specified rule is set in advance. The specified rule can be to use the midpoint of the histogram interval as the key point, or to use the left endpoint or right endpoint of the histogram interval as the key point, or other rules. The embodiments of the present application do not limit this, as long as the selection method of the multiple key points is the same.

[0099] Based on the above description, the horizontal coordinate of the starting point may be outside or inside the first histogram interval. If the horizontal coordinate of the starting point is outside the first histogram interval, a point can be selected from the first histogram interval as a key point; if the horizontal coordinate of the starting point is inside the first histogram interval, a key point is not selected from the first histogram interval, or a key point is selected from the first histogram, but the horizontal coordinate of the key point is greater than that of the starting point.

[0100] Similarly, based on the above description, the abscissa of the end point may be outside the last histogram interval or inside the last histogram interval. If the abscissa of the end point is outside the last histogram interval, a point can be selected from the last histogram interval as a key point; if the abscissa of the end point is inside the last histogram interval, a key point is not selected from the last histogram interval, or a key point is selected from the last histogram interval, but the abscissa of the key point is smaller than the abscissa of the end point.

[0101] (3) Determine the target probabilities corresponding to the multiple sampling points, where the target probabilities corresponding to the starting point and the ending point are specified probabilities, and the target probability corresponding to the key point indicates the probability that the image brightness of the key point falls within the histogram interval.

[0102] Among them, the specified probability is set in advance and can be 0, or any decimal within a very small range greater than 0. For example, the specified probability can be 0.01, 0.001, 0.0001, etc., and the embodiment of the present application does not limit this.

[0103] In some embodiments, the target probability corresponding to a key point may be determined by: determining the sum of the heights of all histogram bins to obtain the total bin height; determining the ratio of the height of the histogram bin containing the first key point to the total bin height, and using this ratio as the target probability corresponding to the first key point. The target probabilities corresponding to the other key points may be determined in the same manner. The first key point is any one of the multiple key points.

[0104] (4) Based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points, the vertical coordinates of the multiple key points are determined by optimizing the equation or the deep learning network model.

[0105] In some embodiments, the methods of determining the vertical coordinates of the multiple key points through optimization equations or deep learning network models are different, which will be introduced below.

[0106] The first implementation method determines the ordinates of multiple key points using an optimization equation. Specifically, the contrast thresholds of the multiple sampling points are determined based on their abscissas. The ordinates of the key points are determined by solving the optimization equation based on the coordinates of the starting and ending points, the number of sampling points, the target probabilities corresponding to each sampling point, the abscissas of the key points, and the contrast thresholds of the multiple sampling points.

[0107] In some embodiments, the encoder stores a contrast human perception curve that indicates the contrast human perception thresholds corresponding to pixels of different brightnesses. The horizontal axis of the curve represents brightness, and the vertical axis represents the contrast human perception threshold. After determining the horizontal coordinate of the first sampling point, the vertical coordinate corresponding to the horizontal coordinate of the first sampling point can be found on the curve. This vertical coordinate is the contrast human perception threshold of the first sampling point. The contrast human perception thresholds corresponding to multiple other sampling points can be determined in the same manner. The first sampling point is any one of the multiple sampling points.

[0108] The optimization equation includes but is not limited to any one of the following equations (1)-(3):

[0109] In equation (1), y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of sampling points, and p k Indicates the target probability corresponding to the kth sampling point, x k Indicates the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 It represents the ordinate of the k+1th sampling point, δ represents the width of the histogram interval, the width of each histogram interval is equal, and λ represents the weight coefficient.

[0110] In equation (1) This section considers the image contrast and the contrast threshold of the human eye. Indicates the slope of the curve segment between two adjacent sampling points in the tone mapping curve (please refer to the marked ), the purpose of the accumulation operation is to comprehensively consider the slope of the entire curve. In this part, the argmin function is used to make the slope of the final tone mapping curve as close to 1 as possible to achieve the maximum restoration of the image contrast, that is, to make the contrast of the tone mapped image as consistent as possible with the contrast of the first image. At the same time, since the minimum contrast that the human eye can perceive is different for pixels with different brightness, when the brightness of the pixel changes, by adding t(x k ), that is, the contrast human eye perception threshold corresponding to the k-th sampling point, to improve the display effect of the image.

[0111] This part takes the image brightness into consideration, with the goal of making the brightness value of each pixel in the tone-mapped image as consistent as possible with the brightness value of each pixel in the first image.

[0112] In equation (2), y′, N-1 、 k, p k 、x k 、t(x k ),y k 、y k+1 , and λ have the same meanings as in equation (1).

[0113] In equation (2) This part is the same as equation (1), but in equation (2), the part that considers the image contrast and the contrast perception threshold of the human eye is The maximum image contrast is restored by making the difference between the pixel gradient of the tone-mapped image and the pixel gradient of the original image as small as possible, that is, by spatial consistency. At the same time, as in equation (1), t(x k ) to improve the display effect of the image.

[0114] In equation (3), y′, N-1, k, p k 、x k 、t(x k ),y k 、y k+1 , λ have the same meaning as in equation (1), and ρ represents the Pearson correlation coefficient. Therefore, equation (3) can also be expressed as:

[0115] Among them, P I represents y k+1 -y k, P J represents t(x k )(x k+1 -x k ).

[0116] In equation (3) This part is also the same as equation (1), but in equation (3), the part that considers the image contrast and the contrast perception threshold of the human eye is The theory of Pearson correlation coefficient is applied, by making y k+1 -y k and t(x k )(x k+1 -x k ) is as small as possible to achieve the maximum restoration of image contrast. At the same time, as in equation (1), t(x k ) to improve the display effect of the image.

[0117] In addition, the weight coefficient λ in equations (1)-(3) is used to adjust the proportion of the two parts on the right side of the equation. The smaller λ is, the more emphasis is placed on the effect of image contrast and the human eye perception threshold on the image display effect compared to image brightness. Conversely, the more emphasis is placed on the effect of image brightness on the image display effect. The weight coefficient is set in advance and its specific size can be set according to actual needs. This embodiment of the application does not limit this.

[0118] When the weight coefficient is set to 0, the optimization equation only considers the image contrast and the human eye perception threshold; when the weight coefficient is set to 1, the optimization equation only considers the image brightness; when the weight coefficient is set to any value between 0 and 1, the optimization equation considers the three factors of image contrast, human eye perception threshold and image brightness at the same time. Of course, when the weight coefficient is set to 0 and t(x k ), the optimization equation only considers the image contrast.

[0119] In the process of solving the vertical coordinate of the key point, it is necessary to combine the optimization equation with y k+1 =y k +y k+1 ′(x k+1 -x k ) and y k+1 ≥y k Solve the problem together. k Indicates the ordinate of the kth sampling point, y k+1 Indicates the ordinate of the k+1th sampling point, y k+1 ′ represents the derivative of the ordinate of the k+1th sampling point, x k Indicates the horizontal coordinate of the kth sampling point, x k+1 Represents the horizontal coordinate of k+1 sampling points.

[0120] The above y k+1 ≥y k It can ensure the monotonicity of the final generated tone mapping curve, and in the calculation process, it can find the solution that satisfies the optimization equation and y by gradual approximation. k+1 ≥y k The optimal solution includes the derivatives of the ordinates of multiple key points and the ordinates of multiple key points.

[0121] The second implementation method determines the vertical coordinates of multiple key points using a deep learning network model. Specifically, the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to the multiple sampling points are input into the deep learning network model, and the derivatives of the vertical coordinates of the multiple key points output by the deep learning network model are obtained. The vertical coordinates of the multiple key points are determined based on the position coordinates of the starting and ending points, the horizontal coordinates of the multiple key points, and the derivatives of the vertical coordinates of the multiple key points.

[0122] After obtaining the derivatives of the ordinates of multiple key points output by the deep learning network model, the ordinate of the first key point can be determined based on the position coordinates of the starting point, the abscissa of the first key point, and the derivative of the ordinate of the first key point according to the formula y1=y0+y1′(x1-x0), where y1 represents the ordinate of the first key point, y0 represents the ordinate of the starting point, y1′ represents the derivative of the ordinate of the first key point, x1 represents the abscissa of the first key point, and x0 represents the abscissa of the starting point. After determining the ordinate of the first key point, the ordinate of the second key point can be determined using the formula y2=y1+y2′(x2-x1), where y2 represents the ordinate of the second key point, y1 represents the ordinate of the first key point, y2′ represents the derivative of the ordinate of the second key point, x2 represents the abscissa of the second key point, and x1 represents the abscissa of the first key point. Similarly, the ordinates of multiple other key points can be determined.

[0123] For example, please refer to Figure 5, where Xmax represents the maximum image brightness, Xmin represents the minimum image brightness, and p k Represents the target probability corresponding to multiple sampling points. By inputting these three into the deep learning network model, we can get d k , that is, the derivative vector composed of the derivatives of the vertical coordinates of multiple key points, and then the derivative vector, combined with the position coordinates of the starting point and the end point, and the horizontal coordinates of multiple key points, can get Y k , that is, the ordinate vector composed of the ordinates of multiple key points.

[0124] In some embodiments, a deep learning network model can be obtained by: obtaining multiple training samples, each training sample including the maximum sample image brightness, the minimum sample image brightness, the target probabilities corresponding to multiple sample sampling points, and the derivatives of the vertical coordinates of multiple sample key points, and the derivatives of the vertical coordinates of the multiple sample key points are determined by optimization equations; based on multiple training samples, the initial network model is trained to obtain a deep learning network model.

[0125] Among them, the method of acquiring the sample image is the same as the method of acquiring the first image, and the method of determining the maximum sample image brightness, the minimum sample image brightness, and the target probabilities corresponding to multiple sample sampling points are also the same as the method of determining the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to multiple sampling points of the first image described above, and will not be repeated here.

[0126] After obtaining the sample image, the derivatives of the vertical coordinates of multiple sample key points of each sample image can be determined by the above optimization equation, combined with the above maximum sample image brightness, minimum sample image brightness, and the target probabilities corresponding to multiple sample sampling points, so as to obtain each training sample, and then input each training sample into the initial network model, and train the initial network model to obtain a deep learning network model.

[0127] In the above training process, the loss function used is Among them, L1 represents the loss function, n is the sum of the number of key points of multiple samples in each training sample, and y i Represents the ordinate of the i-th sample key point determined by the above optimization equation, Represents the vertical coordinate of the i-th sample key point predicted by the deep learning network model, y i and The corresponding horizontal coordinates are the same. This loss function can show the degree of difference between the vertical coordinate predicted by the deep learning network model and the vertical coordinate determined by the above optimization equation, and is used to guide the optimization of the deep learning network model.

[0128] Since the deep learning network model is trained using the training samples obtained through the above-mentioned optimization equation, the deep learning network model, like the above-mentioned optimization equation, also takes into account at least one of the image brightness, image contrast, and the contrast human eye perception threshold, and can also enable the image after subsequent tone mapping to achieve a better display effect.

[0129] After determining the vertical coordinates of multiple key points through the above two implementation methods, each vertical coordinate is combined with the corresponding horizontal coordinate to obtain the position coordinates of the multiple key points.

[0130] Step 203: Encode the first image and the position coordinates of the plurality of sampling points into a bitstream.

[0131] The position coordinates of the multiple sampling points are determined in step 202. After the position coordinates of the multiple sampling points are determined, they can be directly encoded into the bitstream or normalized before being encoded into the bitstream, thereby reducing the system complexity of the subsequent decoding end.

[0132] In some embodiments, the vertical coordinates of the position coordinates of the plurality of sampling points are normalized, while the horizontal coordinates remain unchanged, to obtain the normalized position coordinates of the plurality of sampling points.

[0133] Since the encoding end does not know the information of the decoding end, the position coordinates of the multiple sampling points determined by the encoding end are virtual position coordinates. The decoding end cannot directly generate a tone mapping curve through the virtual position coordinates, but needs to convert the virtual position coordinates into real position coordinates based on the information of the decoding end (mainly the maximum brightness that the decoding end can display, that is, the maximum screen brightness), and then determine the tone mapping curve based on the real position coordinates of the multiple sampling points. The vertical coordinates after normalization are all between 0-1, the smallest vertical coordinate is 0, and the largest vertical coordinate is 1. Therefore, it is only necessary to multiply each vertical coordinate by the maximum brightness that the decoding end can display, and then combine it with the corresponding horizontal coordinate to convert the virtual position coordinates into real position coordinates. If normalization is not performed, the conversion cannot be performed in the above manner, but must be performed in other ways including more steps, which will increase the system complexity of the decoding end.

[0134] In an embodiment of the present application, when determining the position coordinates of multiple sampling points of a tone mapping curve corresponding to a first image, an optimization equation or deep learning network model is used that takes into account at least one of image contrast, image brightness, and the contrast threshold of the human eye. This ensures that the position coordinates of the multiple key points determined are optimal solutions that satisfy the conditions of the optimization equation, maximizes consistency between the contrast of the subsequent tone-mapped image and the contrast of the first image, and ensures that the brightness of each pixel in the subsequent tone-mapped image is as close as possible to the brightness of the corresponding pixel in the first image. Furthermore, after the brightness of the mapped image changes, the human eye can still perceive the difference in brightness between each area of ​​the image, thereby perceiving more details in the tone-mapped image, thereby achieving a better display effect. Furthermore, before encoding the position coordinates of the multiple sampling points into the bitstream, the position coordinates are normalized, which makes it easier for the encoder to convert virtual position coordinates into real position coordinates, thereby reducing system complexity at the decoder.

[0135] FIG6 is a flowchart of an image decoding method provided by an embodiment of the present application, which can be applied to a target device in the above implementation environment, and the target device is also called a decoding end.

[0136] Step 601: Obtain a reconstructed image based on a code stream.

[0137] The reconstructed image refers to an image constructed based on the first image data in the code stream.

[0138] The decoding end parses the relevant image data from the code stream, processes the image data, and obtains a reconstructed image.

[0139] Step 602: Parse the position coordinates of multiple sampling points from the bitstream. The position coordinates of the multiple sampling points are determined by an optimization equation or a deep learning network model. The optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a human eye perception threshold of contrast. The multiple sampling points include a start point, an end point, and a plurality of key points located between the start point and the end point.

[0140] The optimization equation includes but is not limited to any one of equations (1)-(3) described above.

[0141] As can be seen from the above description, the position coordinates of multiple sampling points parsed from the code stream are virtual position coordinates. It is also necessary to convert the virtual position coordinates into real position coordinates based on the information of the decoding end (mainly the maximum brightness that the decoding end can display, that is, the maximum screen brightness). The obtained real position coordinates are used as the position coordinates for subsequent applications.

[0142] As can be seen from the above description, the virtual position coordinates obtained by parsing the code stream may be normalized position coordinates or may be non-normalized position coordinates.

[0143] When the virtual position coordinates are normalized position coordinates, the vertical coordinates in the virtual position coordinates are normalized vertical coordinates, and the horizontal coordinates are not normalized. In this case, the vertical coordinates in each virtual position coordinate are multiplied by the maximum screen brightness to obtain the converted vertical coordinates, and the converted vertical coordinates are combined with the corresponding horizontal coordinates to obtain the real position coordinates of multiple sampling points.

[0144] When the virtual position coordinates are unnormalized, the decoder can first normalize the virtual position coordinates to obtain normalized virtual coordinates. The maximum screen brightness value is then determined, and the ordinate in each normalized virtual position coordinate is multiplied by the value to obtain a converted ordinate. The converted ordinate is then combined with the corresponding abscissa to obtain the true position coordinates of the multiple sampling points. Of course, unnormalized virtual position coordinates can also be converted to true position coordinates using other methods, which are not limited in this embodiment of the present application.

[0145] In some embodiments, after the real position coordinates of the plurality of sampling points are determined, the data format of the real position coordinates may be transformed.

[0146] In the above process, since the maximum screen brightness value involved is generally in standard data format, such as 100nit, 1000nit, etc., the obtained real position coordinates are generally also in standard data format. In this case, the form of the real position coordinates can be changed into the log form of any base, that is, the horizontal and vertical coordinates in the standard data form are converted into the corresponding log form horizontal and vertical coordinates. For example, (10nit, 3nit) is converted into the log form with base 10 - (lg10 10 ,lg10 3 ). Of course, it can also be converted into other data forms, which is not limited in the embodiment of the present application.

[0147] Step 603: Obtain a tone mapping curve based on the position coordinates of the multiple sampling points.

[0148] It should be noted that the position coordinates used in this step are the real position coordinates in any data format determined in the above step 601 .

[0149] In some embodiments, the tone mapping curve is determined by curve fitting, which may include, but is not limited to, a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

[0150] For example, please refer to Figure 4. Each black dot in Figure 4 represents a sampling point, where the first black dot is the starting point with a position coordinate of (0, 0), and the last black dot is the ending point with a position coordinate of (x N ,y N ), (x2, y2) is the position coordinate of the second sampling point, which is also the position coordinate of the first key point, (x k ,y k ) is the position coordinate of the kth sampling point. Assuming that the curve fitting method is straight line connection, as shown in Figure 4, the tone mapping curve can be successfully established by connecting the N sampling points in sequence with straight line segments. The horizontal axis of this tone mapping curve is the brightness of the pixel in the reconstructed image, and the vertical axis is the screen display brightness, that is, the brightness of each pixel in the reconstructed image after tone mapping.

[0151] For example, please refer to Figure 7. Each black dot in Figure 7 represents a sampling point. Assuming that the curve fitting method is the cubic spline connection method, as shown in Figure 7, the position coordinates of every three sampling points can be used to uniquely determine the cubic spline curve connecting every three sampling points. By combining the cubic spline curves of every three sampling points, the tone mapping curve can be successfully established.

[0152] Step 604: Perform tone mapping on the reconstructed image based on the tone mapping curve, the maximum screen brightness, and the minimum screen brightness.

[0153] The maximum screen brightness is the maximum brightness value that the decoding end can display; the minimum screen brightness is the minimum brightness value that the decoding end can display.

[0154] As shown in Figure 4, the horizontal axis of the tone mapping curve is the brightness of the pixel in the reconstructed image, and the vertical axis is the screen display brightness. For each pixel in the reconstructed image, the brightness of the pixel is obtained as the horizontal coordinate, and the corresponding vertical coordinate is found on the tone mapping curve. The vertical coordinate is the screen display brightness corresponding to the pixel.

[0155] As can be seen from the above description, the horizontal coordinate of the end point determined by the encoding end may be within the last histogram interval or outside the last histogram interval. In some embodiments, the horizontal coordinate of the end point determined by the encoding end is within the last histogram interval, that is, the horizontal coordinate of the end point is less than the maximum image brightness, which will cause the brightness of individual pixels in the reconstructed image to be greater than the maximum horizontal coordinate of the tone mapping curve. In this case, when tone mapping is performed on the reconstructed image, the brightness of these pixels is directly mapped to the maximum screen brightness. Similarly, in some embodiments, the horizontal coordinate of the starting point determined by the encoding end is within the first histogram interval, that is, the horizontal coordinate of the starting point is greater than the minimum image brightness, which will cause the brightness of individual pixels in the reconstructed image to be less than the minimum horizontal coordinate of the tone mapping interval. In this case, when tone mapping is performed on the reconstructed image, the brightness of these pixels is directly mapped to the minimum screen brightness.

[0156] In an embodiment of the present application, when establishing a tone mapping curve based on the position coordinates of multiple sampling points parsed from a bitstream, the position coordinates of multiple key points among the multiple sampling points are determined using an optimization equation or a deep learning network model, and the optimization equation or deep learning network model takes into account at least one of image contrast, image brightness, and a human eye perception threshold of contrast, so that the position coordinates of the multiple key points determined are optimal solutions that satisfy the conditions of the optimization equation, so that the contrast of the subsequent tone-mapped image is kept consistent with the contrast of the first image to the greatest extent possible, the brightness of each pixel in the tone-mapped image is as close as possible to the brightness value of the corresponding pixel in the first image, and the human eye can still perceive the difference in brightness between each area of ​​the image after the brightness of the mapped image changes, thereby perceiving more details in the tone-mapped image, thereby achieving a better display effect.

[0157] FIG8 is a schematic diagram of the structure of an image encoding apparatus provided in an embodiment of the present application. The image encoding apparatus can be implemented by software, hardware, or a combination of both to form part or all of an encoding end device, which can be the source device shown in FIG1 . Referring to FIG8 , the apparatus includes an acquisition module 801, a first determination module 802, and an encoding module 803.

[0158] An acquisition module 801 is configured to acquire a first image, where the first image is an image that needs to be tone mapped.

[0159] A first determining module 802 is configured to determine, based on the first image, position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image using an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a human eye perception threshold of contrast, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points located between the starting point and the ending point;

[0160] The encoding module 803 is configured to encode the first image and the position coordinates of the plurality of sampling points into a bitstream.

[0161] Optionally, the first determining module 802 includes:

[0162] a generating submodule, configured to generate a target histogram based on the first image, the target histogram comprising a plurality of histogram intervals, the plurality of histogram intervals being divided based on maximum image brightness and minimum image brightness of the first image, the height of the histogram interval indicating the number of pixels in the first image whose brightness falls within the histogram interval;

[0163] A first coordinate determination submodule is configured to determine the position coordinates of a starting point and an ending point, wherein the position coordinates include a horizontal coordinate and a vertical coordinate, select a point from each of the multiple histogram intervals as one of the multiple key points, and use the image brightness corresponding to each of the multiple key points as the horizontal coordinate of the multiple key points;

[0164] The probability determination submodule is used to determine the target probabilities corresponding to the multiple sampling points, where the target probabilities corresponding to the starting point and the ending point are specified probabilities, and the target probability corresponding to the key point indicates the probability that the image brightness of the key point falls within the histogram interval;

[0165] The second coordinate determination submodule is used to determine the vertical coordinates of multiple key points based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points through an optimization equation or a deep learning network model.

[0166] Optionally, the second coordinate determination submodule is specifically configured to:

[0167] Determining a contrast human eye perception threshold of the plurality of sampling points based on the abscissas of the plurality of sampling points;

[0168] Based on the position coordinates of the starting point and the end point, the number of the multiple sampling points, the target probabilities corresponding to the multiple sampling points, the horizontal coordinates of the multiple key points and the contrast human eye perception threshold of the multiple sampling points, the vertical coordinates of the multiple key points are obtained by solving the optimization equation.

[0169] Optionally, the second coordinate determination submodule is specifically configured to:

[0170] Inputting the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to multiple sampling points into the deep learning network model, and obtaining the derivatives of the vertical coordinates of multiple key points output by the deep learning network model;

[0171] The vertical coordinates of the plurality of key points are determined based on the position coordinates of the starting point and the ending point, the horizontal coordinates of the plurality of key points, and the derivatives of the vertical coordinates of the plurality of key points.

[0172] Optionally, the device further comprises:

[0173] A sample acquisition module is used to acquire multiple training samples, each training sample includes a maximum sample image brightness, a minimum sample image brightness, target probabilities corresponding to multiple sample sampling points, and derivatives of the ordinates of multiple sample key points, where the derivatives of the ordinates of the multiple sample key points are determined by an optimization equation;

[0174] The model training module is used to train the initial network model based on multiple training samples to obtain a deep learning network model.

[0175] Optionally, the above optimization equation includes but is not limited to any one of the following equations:

[0176] or

[0177] or

[0178] Among them, y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of multiple sampling points, and p k Indicates the target probability corresponding to the kth sampling point , x k Indicates the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

[0179] In an embodiment of the present application, when determining the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image, the optimization equation or deep learning network model used takes into account at least one of the image contrast, image brightness and contrast human eye perception threshold, so that the position coordinates of the multiple key points determined are the optimal solution that meets the conditions of the optimization equation, so that the contrast of the subsequent tone-mapped image is kept consistent with the contrast of the first image to the greatest extent, the brightness of each pixel in the subsequent tone-mapped image is as similar as possible to the brightness value of the corresponding pixel in the first image, and it is ensured that after the brightness of the mapped image changes, the human eye can still perceive the difference in brightness and darkness in each area of ​​the image, and then perceive more details in the tone-mapped image, thereby achieving a better display effect.

[0180] It should be noted that the image coding apparatus provided in the above embodiments uses the division of the above functional modules as an example only. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the image coding apparatus provided in the above embodiments and the image coding method embodiments are based on the same concept. The specific implementation process is detailed in the method embodiments and will not be repeated here.

[0181] Figure 9 is a schematic diagram of the structure of an image decoding device provided in an embodiment of the present application. This image decoding device can be implemented as part or all of a decoding end device using software, hardware, or a combination of both. This decoding end device can be the target device shown in Figure 1 . Referring to Figure 9 , the device includes an image reconstruction module 901, a coordinate analysis module 902, a curve establishment module 903, and a mapping module 904.

[0182] An image reconstruction module 901 is configured to obtain a reconstructed image based on a code stream;

[0183] a coordinate parsing module 902 configured to parse the position coordinates of a plurality of sampling points from a bitstream, wherein the position coordinates of the plurality of sampling points are determined by an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a human eye perception threshold of contrast, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points located between the starting point and the ending point;

[0184] A curve building module 903 is configured to build a tone mapping curve by curve fitting based on the position coordinates of multiple sampling points;

[0185] The mapping module 904 is configured to perform tone mapping on the reconstructed image based on the tone mapping curve, the maximum screen brightness, and the minimum screen brightness. Optionally, the above optimization equation includes but is not limited to any one of the following equations:

[0186] or

[0187] or

[0188] Among them, y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of multiple sampling points, and p k Indicates the target probability corresponding to the kth sampling point, x k Indicates the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

[0189] Optionally, the tone mapping curve is determined by curve fitting, and the curve fitting method includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

[0190] In an embodiment of the present application, when establishing a tone mapping curve based on the position coordinates of multiple sampling points parsed from a bitstream, the position coordinates of multiple key points among the multiple sampling points are determined using an optimization equation or a deep learning network model, and the optimization equation or deep learning network model takes into account at least one of image contrast, image brightness, and a human eye perception threshold of contrast, so that the position coordinates of the multiple key points determined are optimal solutions that satisfy the conditions of the optimization equation, so that the contrast of the subsequent tone-mapped image is kept consistent with the contrast of the first image to the greatest extent possible, the brightness of each pixel in the tone-mapped image is as close as possible to the brightness value of the corresponding pixel in the first image, and the human eye can still perceive the difference in brightness between each area of ​​the image after the brightness of the mapped image changes, thereby perceiving more details in the tone-mapped image, thereby achieving a better display effect.

[0191] It should be noted that the image decoding device provided in the above embodiments uses the division of the above functional modules as an example only. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image decoding device provided in the above embodiments and the image decoding method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0192] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.

[0193] It should be understood that the "plurality" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.

[0194] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0195] The above description is an embodiment provided for this application and is not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.

Claims

1. An image coding method, characterized in that: The method comprises: Acquire a first image, where the first image is an image that needs to be tone mapped; Based on the first image, determine position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image by using an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point; The first image and the position coordinates of the plurality of sampling points are encoded into a bitstream.

2. The method according to claim 1, characterized in that The determining, based on the first image, position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image by optimizing an equation or a deep learning network model comprises: Generate a target histogram based on the first image, the target histogram comprising a plurality of histogram intervals, the plurality of histogram intervals being divided based on a maximum image brightness and a minimum image brightness of the first image, the height of the histogram interval indicating the number of pixels in the first image whose brightness is within the histogram interval; Determine the position coordinates of the starting point and the ending point, wherein the position coordinates include a horizontal coordinate and a vertical coordinate; Selecting one point from each of the multiple histogram intervals as one of the multiple key points, and using the image brightness corresponding to each of the multiple key points as the horizontal coordinates of the multiple key points; Determine the target probabilities corresponding to the multiple sampling points respectively, wherein the target probabilities corresponding to the starting point and the ending point are specified probabilities, and the target probability corresponding to the key point indicates the probability that the image brightness of the key point falls within the histogram interval; Based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points respectively, and the horizontal coordinates of the multiple key points, the vertical coordinates of the multiple key points are determined through the optimization equation or the deep learning network model.

3. The method according to claim 2, characterized in that The determining the vertical coordinates of the multiple key points through the optimization equation based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points includes: Determine a contrast human eye perception threshold of the plurality of sampling points based on the abscissas of the plurality of sampling points; Based on the position coordinates of the starting point and the end point, the number of the multiple sampling points, the target probabilities respectively corresponding to the multiple sampling points, the horizontal coordinates of the multiple key points and the contrast human eye perception threshold of the multiple sampling points, the vertical coordinates of the multiple key points are obtained by solving the optimization equation.

4. The method according to claim 2, characterized in that The determining the vertical coordinates of the multiple key points through the deep learning network model based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points includes: Inputting the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to the multiple sampling points into the deep learning network model, and obtaining the derivatives of the ordinates of the multiple key points output by the deep learning network model; The vertical coordinates of the plurality of key points are determined based on the position coordinates of the starting point and the ending point, the horizontal coordinates of the plurality of key points, and the derivatives of the vertical coordinates of the plurality of key points.

5. The method according to claim 4, characterized in that The method further comprises: Acquire multiple training samples, each training sample includes a maximum sample image brightness, a minimum sample image brightness, target probabilities corresponding to multiple sample sampling points, and derivatives of the ordinates of multiple sample key points, wherein the derivatives of the ordinates of the multiple sample key points are determined by the optimization equation; Based on the multiple training samples, the initial network model is trained to obtain the deep learning network model.

6. The method according to any one of claims 1 to 5, characterized in that: The optimization equation includes but is not limited to any one of the following equations: or or Wherein, y′ represents a derivative vector composed of the derivatives of the ordinates of the multiple key points, N represents the number of the multiple sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k represents the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

7. An image decoding method, characterized in that: The method comprises: Obtain a reconstructed image based on the code stream; parsing position coordinates of a plurality of sampling points from the bitstream, the position coordinates of the plurality of sampling points being determined by an optimization equation or a deep learning network model, the optimization equation and the deep learning network model being determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, the plurality of sampling points comprising a starting point, an ending point, and a plurality of key points between the starting point and the ending point; Obtaining a tone mapping curve based on the position coordinates of the plurality of sampling points; The reconstructed image is tone mapped based on the tone mapping curve, the maximum screen brightness and the minimum screen brightness.

8. The method according to claim 7, characterized in that The tone mapping curve is determined by curve fitting, and the curve fitting method includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

9. The method according to claim 7, characterized in that The optimization equation includes but is not limited to any one of the following equations: or or Wherein, y′ represents a derivative vector composed of the derivatives of the ordinates of the multiple key points, N represents the number of the multiple sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k represents the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

10. An image encoding device, characterized in that: The device comprises: An acquisition module, used for acquiring a first image, where the first image is an image that needs to be tone mapped; a first determination module, for determining, based on the first image, position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image by means of an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an end point, and a position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image. a plurality of key points between the starting point and the end point; The encoding module is used to encode the first image and the position coordinates of the multiple sampling points into a bit stream.

11. The device according to claim 10, characterized in that The first determining module comprises: A generating submodule, configured to generate a target histogram based on the first image, wherein the target histogram includes a plurality of histogram intervals, wherein the plurality of histogram intervals are obtained by dividing based on a maximum image brightness and a minimum image brightness of the first image, and the height of the histogram interval indicates the number of pixels in the first image whose brightness is within the histogram interval; A first coordinate determination submodule is used to determine the position coordinates of the starting point and the ending point, wherein the position coordinates include a horizontal coordinate and a vertical coordinate, and select a point from each of the multiple histogram intervals as one of the multiple key points, and use the image brightness corresponding to each of the multiple key points as the horizontal coordinate of the multiple key points; A probability determination submodule, used to determine the target probabilities corresponding to the plurality of sampling points, wherein the target probabilities corresponding to the starting point and the ending point are designated probabilities, and the target probability corresponding to the key point indicates the probability that the image brightness of the key point falls within the histogram interval; The second coordinate determination submodule is used to determine the vertical coordinates of the multiple key points based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points through the optimization equation or the deep learning network model.

12. The device according to claim 11, characterized in that The second coordinate determination submodule is specifically used for: Determine a contrast human eye perception threshold of the plurality of sampling points based on the abscissas of the plurality of sampling points; Based on the position coordinates of the starting point and the end point, the number of the multiple sampling points, the target probabilities respectively corresponding to the multiple sampling points, the horizontal coordinates of the multiple key points and the contrast human eye perception threshold of the multiple sampling points, the vertical coordinates of the multiple key points are obtained by solving the optimization equation.

13. The device according to claim 11, characterized in that The second coordinate determination submodule is specifically used for: Inputting the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to the multiple sampling points into the deep learning network model, and obtaining the derivatives of the ordinates of the multiple key points output by the deep learning network model; The vertical coordinates of the plurality of key points are determined based on the position coordinates of the starting point and the ending point, the horizontal coordinates of the plurality of key points, and the derivatives of the vertical coordinates of the plurality of key points.

14. The device according to claim 13, characterized in that The device also includes: A sample acquisition module, used to acquire multiple training samples, each training sample includes a maximum sample image brightness, a minimum sample image brightness, target probabilities corresponding to multiple sample sampling points, and derivatives of the ordinates of multiple sample key points, wherein the derivatives of the ordinates of the multiple sample key points are determined by the optimization equation; The model training module is used to train the initial network model based on the multiple training samples to obtain the deep learning network model.

15. The device according to any one of claims 10 to 14, characterized in that: The optimization equation includes but is not limited to any one of the following equations: or or Wherein, y′ represents a derivative vector composed of the derivatives of the ordinates of the multiple key points, N represents the number of the multiple sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast of the kth sampling point. Perception threshold, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

16. An image decoding device, characterized in that: The device comprises: An image reconstruction module, used for obtaining a reconstructed image based on a bit stream; a coordinate parsing module, configured to parse position coordinates of a plurality of sampling points from the bitstream, wherein the position coordinates of the plurality of sampling points are determined by an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point; A curve building module, used for obtaining a tone mapping curve based on the position coordinates of the plurality of sampling points; A mapping module is used to perform tone mapping on the reconstructed image based on the tone mapping curve, the maximum screen brightness and the minimum screen brightness.

17. The device according to claim 16, characterized in that The tone mapping curve is determined by curve fitting, and the curve fitting method includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

18. The device according to claim 16, characterized in that The optimization equation includes but is not limited to any one of the following equations: or or Wherein, y′ represents a derivative vector composed of the derivatives of the ordinates of the multiple key points, N represents the number of the multiple sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k represents the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

19. A computer-readable storage medium, characterized in that: The storage medium stores instructions, and when the instructions are executed on the computer, the computer executes the steps of the method according to any one of claims 1 to 9.

20. A computer program product, characterized in that The computer program product comprises instructions, and when the instructions are executed on a computer, the computer is caused to execute the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image encoding method, image decoding method and related device

    CN119919329A

  • High-fidelity full reference and high-efficiency reduced reference encoding in end-to-end single-layer backward compatible encoding pipeline

    CN112106357A

  • Image processing method and device

    CN112215760A

  • Image processing method and device

    CN112686810A

  • HDR image representations using neural network mappings

    US20210150812A1