Image encoding method, image decoding method and related device

By using optimization equations or deep learning models to determine the location of the sampling point and constructing a suitable tone mapping curve during the tone mapping process, the problem of poor image display effect in the prior art is solved and a better image display effect is achieved.

CN119919329APending Publication Date: 2025-05-02HUAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311440178.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

In the prior art, in the process of tone mapping, commonly used methods such as Dolby sigmoidal curves and Bezier curves may cause image distortion and lead to poor display effects.

Method used

By optimizing equations or deep learning network models, the position coordinates of multiple sampling points of the tone mapping curve corresponding to the image are determined, and a more suitable tone mapping curve is constructed to take into account image contrast, image brightness, and contrast human eye perception thresholds.

Benefits of technology

The display effect of the tone-mapping image in ordinary display devices is improved, so that the brightness and contrast of the image are maximized to the original image, and the human eye's perception of image details is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919329A_ABST
    Figure CN119919329A_ABST
Patent Text Reader

Abstract

The invention discloses an image coding method, an image decoding method and related devices, and belongs to the field of image processing. The method comprises the following steps: acquiring a first image, wherein the first image is an image needing tone mapping; based on the first image, determining position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image through an optimization equation or a deep learning network model; and compiling the first image and the position coordinates of the plurality of sampling points into a code stream. According to the method, the optimization equation or the deep learning model is established based on at least one of the image contrast, the image brightness and the contrast human eye perception threshold, and the image needing tone mapping is coded through the optimization equation or the deep learning model. Therefore, the decoding end establishes the tone mapping curve based on the code stream and realizes tone mapping of the image, so that the display effect of the image is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to an image encoding and decoding method and related devices. Background Art

[0002] High dynamic range (HDR) technology has developed rapidly in recent years. HDR technology can make the contrast between the extreme brightness and darkness of an image higher, and show rich details of the bright and dark areas, presenting a picture closer to the real human eye perception. However, the dynamic range that ordinary display devices can display is relatively low, so it is necessary to perform tone mapping on HDR images so that they can be displayed normally on ordinary display devices.

[0003] In the related art, the tone mapping method generally aligns the maximum image brightness and the minimum image brightness of the image to the maximum screen brightness value and the minimum screen brightness of the display, and performs global or local mapping based on the tone mapping curve for the image brightness between the maximum image brightness and the minimum image brightness, so as to map these image brightnesses to the range between the maximum screen brightness and the minimum screen brightness. However, the tone mapping curves currently used are generally Dolby sigmoidal curves, Bezier curves, etc., and tone mapping through these curves may distort the tone-mapped image, resulting in poor display effects. Summary of the invention

[0004] The present application provides an image encoding and decoding method and related devices, which can improve the display effect of tone-mapped images in ordinary display devices. The technical solution is as follows:

[0005] In a first aspect, an image encoding method is provided, the method comprising:

[0006] A first image is obtained, where the first image is an image that needs to be tone mapped; based on the first image, position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image are determined by an optimization equation or a deep learning network model, where the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point; and the position coordinates of the first image and the plurality of sampling points are encoded into a bitstream.

[0007] That is to say, the present application determines multiple sampling points of the tone mapping curve corresponding to the first image at the encoding end, and the position coordinates of multiple key points among these sampling points are determined by an optimization equation or a deep learning network model, and the optimization equation or deep learning network model takes into account the factors affecting the display effect - at least one of the image contrast, image brightness and contrast human eye perception threshold. Therefore, the tone mapping curve constructed based on multiple sampling points at the decoding end can enable the mapped image to achieve the best display effect.

[0008] Among them, image contrast refers to the ratio between the maximum and minimum brightness of a certain area; for a certain pixel in the image, the brightness value of the pixel is the image brightness corresponding to the pixel; the contrast human eye perception threshold refers to the minimum contrast that the human eye can perceive for a certain pixel when the brightness of the pixel changes.

[0009] It should be noted that since the essence of tone mapping is to compress the high dynamic range into the low dynamic range, the first image can not only be an HDR / SDR image, but also other high dynamic range images, as long as the brightness range of the first image is greater than the brightness range that can be displayed by the decoding end.

[0010] Optionally, based on the first image, the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image are determined by optimizing an equation or a deep learning network model, including: generating a target histogram based on the first image, the target histogram including multiple histogram intervals, the multiple histogram intervals are divided based on the maximum image brightness and the minimum image brightness of the first image, and the height of the histogram interval indicates the number of pixels in the first image whose brightness is within the histogram interval; determining the position coordinates of the starting point and the ending point, the position coordinates including a horizontal coordinate and a vertical coordinate; selecting a point from the multiple histogram intervals as one of the multiple key points, and using the image brightness corresponding to the multiple key points as the horizontal coordinates of the multiple key points; determining the target probabilities corresponding to the multiple sampling points, wherein the target probabilities corresponding to the starting point and the ending point are specified probabilities, and the target probabilities corresponding to the key points indicate the probability that the image brightness of the key points falls within the histogram interval; based on the position coordinates of the starting point and the ending point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points, determining the vertical coordinates of the multiple key points by optimizing an equation or a deep learning network model.

[0011] That is to say, there are two ways to determine the vertical coordinates of multiple key points, one is to determine them through optimization equations, and the other is to determine them through a deep learning network model.

[0012] Optionally, based on the position coordinates of the starting point and the ending point, the target probabilities corresponding to the multiple sampling points respectively, and the horizontal coordinates of the multiple key points, the vertical coordinates of the multiple key points are determined by optimizing equations, including: determining the contrast human eye perception threshold of the multiple sampling points based on the horizontal coordinates of the multiple sampling points; based on the position coordinates of the starting point and the ending point, the number of the multiple sampling points, the target probabilities corresponding to the multiple sampling points respectively, the horizontal coordinates of the multiple key points and the contrast human eye perception threshold of the multiple sampling points, the vertical coordinates of the multiple key points are obtained by solving the optimization equations.

[0013] Optionally, based on the position coordinates of the starting point and the ending point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points, the vertical coordinates of the multiple key points are determined through a deep learning network model, including: inputting the maximum image brightness, the minimum image brightness and the target probabilities corresponding to the multiple sampling points into the deep learning network model to obtain the derivatives of the vertical coordinates of the multiple key points output by the deep learning network model; determining the vertical coordinates of the multiple key points based on the position coordinates of the starting point and the ending point, the horizontal coordinates of the multiple key points, and the derivatives of the vertical coordinates of the multiple key points.

[0014] Optionally, the method also includes: obtaining multiple training samples, each training sample includes the maximum sample image brightness, the minimum sample image brightness, the target probabilities corresponding to multiple sample sampling points, and the derivatives of the vertical coordinates of multiple sample key points, and the derivatives of the vertical coordinates of the multiple sample key points are determined by optimization equations; based on the multiple training samples, training the initial network model to obtain a deep learning network model.

[0015] Since the training samples of the deep learning network model are determined by the optimization equation, the deep learning network model can take into account at least one of the image brightness, image contrast and the contrast human eye perception threshold in the same way as the optimization equation, and can also enable the image to achieve better display effect after subsequent tone mapping.

[0016] Optionally, the above optimization equation includes but is not limited to any one of the following equations:

[0017] or

[0018] or

[0019]

[0020] Where y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(xk ) represents the contrast human eye perception threshold of the kth sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

[0021] In the first equation above, This part takes into account the image contrast and the contrast perception threshold of the human eye. This part uses the argmin function to make the slope of the final generated tone mapping curve as close to 1 as possible to achieve the maximum restoration of the image contrast, that is, to make the contrast of the tone mapped image as consistent as possible with the contrast of the first image.

[0022] In the second equation above, This part takes into account the image contrast and the contrast perception threshold of the human eye. It achieves the maximum restoration of image contrast by making the difference between the pixel gradient of the tone-mapped image and the pixel gradient of the original image as small as possible, that is, through spatial consistency.

[0023] In the third equation above, This part considers the image contrast and the contrast perception threshold of the human eye, and applies the theory of Pearson correlation coefficient. k+1 -y k and t(x k )(x k+1 -x k ) as small as possible to achieve the maximum restoration of image contrast.

[0024] At the same time, since for pixels with different brightness, the minimum contrast that the human eye can perceive is different when the brightness of the pixel changes, the above three equations can be added by adding t(x k ), that is, the contrast human eye perception threshold corresponding to the k-th sampling point, to improve the display effect of the image. This part takes the image brightness into consideration, with the goal of making the brightness value of each pixel in the tone-mapped image as consistent as possible with the brightness value of each pixel in the first image.

[0025] In addition, the weight coefficient λ is used to adjust the proportion of the two parts on the right side of the equation. The smaller λ is, the more emphasis is placed on the effect of image contrast and the contrast human eye perception threshold on the image display effect compared to image brightness. Conversely, the more emphasis is placed on the effect of image brightness on the image display effect. The weight coefficient is set in advance, and the specific size can be set according to actual needs, and the embodiment of the present application does not limit this.

[0026] When the weight coefficient is set to 0, the optimization equation only considers the image contrast and the contrast human eye perception threshold; when the weight coefficient is set to 1, the optimization equation only considers the image brightness; when the weight coefficient is set to any value between 0 and 1, the optimization equation considers the image contrast, the contrast human eye perception threshold and the image brightness at the same time. Of course, when the weight coefficient is set to 0 and t(x k ), the optimization equation only considers the image contrast.

[0027] In a second aspect, an image decoding method is provided, the method comprising:

[0028] A reconstructed image is obtained based on a bitstream; position coordinates of a plurality of sampling points are parsed from the bitstream, the position coordinates of the plurality of sampling points are determined by an optimization equation or a deep learning network model, the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, the plurality of sampling points include a starting point, an end point, and a plurality of key points between the starting point and the end point; a tone mapping curve is obtained based on the position coordinates of the plurality of sampling points; and tone mapping is performed on the reconstructed image based on the tone mapping curve, the maximum screen brightness, and the minimum screen brightness.

[0029] Among them, the reconstructed image refers to an image constructed based on the first image data in the code stream; the maximum screen brightness is the maximum brightness value that can be displayed by the decoding end; and the minimum screen brightness is the minimum brightness value that can be displayed by the decoding end.

[0030] When the decoding end establishes the tone mapping curve, the position coordinates of the multiple key points are determined by an optimization equation or a deep learning network model, and the optimization equation or the deep learning network model takes into account at least one of the image contrast, image brightness and contrast human eye perception threshold, so that the brightness and contrast of the image can be kept consistent with the first image to the greatest extent, and the human eye can perceive the difference between light and dark at different image brightnesses, and then perceive more details in the tone mapped image. Therefore, after mapping the image through the tone mapping curve, a better display effect can be achieved.

[0031] Optionally, the tone mapping curve is determined by curve fitting, and the curve fitting method includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

[0032] Optionally, the above optimization equation includes but is not limited to any one of the following equations:

[0033] or

[0034] or

[0035]

[0036] Where y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the kth sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

[0037] In the first equation above, This part takes into account the image contrast and the contrast perception threshold of the human eye. This part uses the argmin function to make the slope of the final generated tone mapping curve as close to 1 as possible to achieve the maximum restoration of the image contrast, that is, to make the contrast of the tone mapped image as consistent as possible with the contrast of the first image.

[0038] In the second equation above, This part takes into account the image contrast and the contrast perception threshold of the human eye. It achieves the maximum restoration of image contrast by making the difference between the pixel gradient of the tone-mapped image and the pixel gradient of the original image as small as possible, that is, through spatial consistency.

[0039] In the third equation above, This part considers the image contrast and the contrast perception threshold of the human eye, and applies the theory of Pearson correlation coefficient. k+1 -y k and t(x k )(x k+1 -x k ) as small as possible to achieve the maximum restoration of image contrast.

[0040] At the same time, since for pixels with different brightness, the minimum contrast that the human eye can perceive is different when the brightness of the pixel changes, the above three equations can be added by adding t(x k ), that is, the contrast human eye perception threshold corresponding to the k-th sampling point, to improve the display effect of the image. This part takes the image brightness into consideration, with the goal of making the brightness value of each pixel in the tone-mapped image as consistent as possible with the brightness value of each pixel in the first image.

[0041] In addition, the weight coefficient λ is used to adjust the proportion of the two parts on the right side of the equation. The smaller λ is, the more emphasis is placed on the effect of image contrast and the contrast human eye perception threshold on the image display effect compared to image brightness. Conversely, the more emphasis is placed on the effect of image brightness on the image display effect. The weight coefficient is set in advance, and the specific size can be set according to actual needs, and the embodiment of the present application does not limit this.

[0042] When the weight coefficient is set to 0, the optimization equation only considers the image contrast and the contrast human eye perception threshold; when the weight coefficient is set to 1, the optimization equation only considers the image brightness; when the weight coefficient is set to any value between 0 and 1, the optimization equation considers the image contrast, the contrast human eye perception threshold and the image brightness at the same time. Of course, when the weight coefficient is set to 0 and t(x k ), the optimization equation only considers the image contrast.

[0043] In a third aspect, an image coding device is provided, the device having the function of implementing the image coding method in the first aspect. The image coding device comprises at least one module, the at least one module being used to implement the image coding method provided in the first aspect.

[0044] In a fourth aspect, an image decoding device is provided, wherein the device has the function of implementing the image decoding method in the second aspect. The image decoding device comprises at least one module, wherein the at least one module is used to implement the image decoding method provided in the second aspect.

[0045] In a fifth aspect, a coding end device is provided, the coding end device comprising a processor and a memory, the memory being used to store a computer program for executing the image coding method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the image coding method described in the first aspect.

[0046] Optionally, the encoding end device may further include a communication bus, and the communication bus is used to establish a connection between the processor and the memory.

[0047] In a sixth aspect, a decoding end device is provided, the decoding end device comprising a processor and a memory, the memory being used to store a computer program for executing the image decoding method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the image decoding method described in the second aspect.

[0048] Optionally, the decoding end device may further include a communication bus, and the communication bus is used to establish a connection between the processor and the memory.

[0049] In the seventh aspect, a computer-readable storage medium is provided, wherein the storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the steps of the image encoding method described in the first aspect, or executes the steps of the image decoding method described in the second aspect.

[0050] In an eighth aspect, a computer program product comprising instructions is provided, and when the instructions are executed on a computer, the computer executes the steps of the image encoding method described in the first aspect, or executes the steps of the image decoding method described in the second aspect. In other words, a computer program is provided, and when the computer program is executed on a computer, the computer executes the steps of the image encoding method described in the first aspect, or executes the steps of the image decoding method described in the second aspect.

[0051] The technical effects obtained by the above-mentioned third, fourth, fifth, sixth, seventh and eighth aspects are similar to the technical effects obtained by the corresponding technical means in the first or second aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0053] Figure 2 is a flowchart of an image encoding method provided by an embodiment of the present application;

[0054] Figure 3 is a flow chart of a method for determining position coordinates of multiple sampling points provided in an embodiment of the present application;

[0055] Figure 4 is a schematic diagram of a target histogram and a process of constructing a tone mapping curve provided in an embodiment of the present application;

[0056] Figure 5 It is a schematic diagram of a process of determining the vertical coordinates of multiple key points by a deep learning network model provided in an embodiment of the present application;

[0057] Figure 6 is a flowchart of an image decoding method provided by an embodiment of the present application;

[0058] Figure 7 is a schematic diagram of another process of constructing a tone mapping curve provided in an embodiment of the present application;

[0059] Figure 8 is a structural diagram of an image encoding device provided in an embodiment of the present application;

[0060] Fig. 9 It is a structural schematic diagram of an image decoding device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0062] Before explaining in detail the image encoding method and the image decoding method provided in the embodiments of the present application, the application scenarios and implementation environments involved in the embodiments of the present application are first introduced.

[0063] First, the application scenarios involved in the embodiments of the present application are introduced.

[0064] Scenes in nature have an extremely wide range of colors and brightness. Usually, the maximum brightness is close to 10^6cd / m^2 and the minimum brightness is close to 10^(-3)cd / m^2. The human visual system can automatically adjust to adapt to the brightness changes of nearly 10 orders of magnitude. However, the images generated by digital cameras have only a limited dynamic range. In order to restore the real world as much as possible, it is necessary to be able to capture a larger dynamic range, which has prompted the rapid development of HDR imaging technology. HDR images are stored in floating-point format and have a large dynamic range, but ordinary display devices generally only have 8 bits and a very limited dynamic range. Therefore, if HDR images are to be displayed normally on ordinary display devices, the dynamic range of the HDR image needs to be compressed to the dynamic range of the display device. This process is called tone mapping.

[0065] In the related art, the tone mapping method usually aligns the maximum image brightness and the minimum image brightness of the image to the maximum screen brightness value and the minimum screen brightness of the display. For the image brightness between the maximum image brightness and the minimum image brightness, global or local mapping is performed based on the tone mapping curve to map these image brightnesses to the range between the maximum screen brightness and the minimum screen brightness. However, the tone mapping curves commonly used at present do not consider the impact of image contrast on the display effect, so that the original image cannot achieve the best display effect. For example, tone mapping curves such as Dolby sigmoidal curves and Bezier curves will cause a certain loss of details in the mapped image, and the local contrast is not high, resulting in distortion of the mapped image. In addition, the tone mapping curves commonly used at present do not consider the factors of human eye perception, resulting in a large visual difference between the mapped image and the real scene.

[0066] Based on this, the embodiment of the present application establishes an optimization equation or a deep learning model based on at least one of the image contrast, image brightness, and the contrast human eye perception threshold, and encodes the image that needs to be tone mapped through the optimization equation or the deep learning model. In this way, the decoding end establishes a tone mapping curve based on the bitstream to achieve the tone mapping of the image, so that the same HDR image can obtain the corresponding tone mapping curve for different display devices, so that the contrast and brightness of the image after tone mapping are as consistent as possible with the HDR image before tone mapping, and the image display effect perceived by the human eye is improved.

[0067] Secondly, the implementation environment involved in the embodiments of the present application is introduced.

[0068] Please refer to Figure 1 , Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application. The implementation environment includes a source device 10, a destination device 20, a link 30, and a storage device 40. Among them, the source device 10 can generate an encoded image. Therefore, the source device 10 can also be referred to as an image encoding device or an encoding end. The destination device 20 can decode the encoded image generated by the source device 10. Therefore, the destination device 20 can also be referred to as an image decoding device or a decoding end. The link 30 can receive the encoded image generated by the source device 10, and can transmit the encoded image to the destination device 20. The storage device 40 can receive the encoded image generated by the source device 10, and can store the encoded image. Under such conditions, the destination device 20 can directly obtain the encoded image from the storage device 40. Alternatively, the storage device 40 can correspond to a file server or another intermediate storage device that can store the encoded image generated by the source device 10. Under such conditions, the destination device 20 can stream or download the encoded image stored in the storage device 40.

[0069] The source device 10 and the destination device 20 may each include one or more processors and a memory coupled to the one or more processors, the memory may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, any other medium that can be used to store desired program code in the form of instructions or data structures that can be accessed by a computer, etc. For example, the source device 10 and the destination device 20 may each include a mobile phone, a smart phone, a personal digital assistant (PDA), a wearable device, a pocket PC (PPC), a tablet computer, a smart car machine, a smart TV, a smart speaker, a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video game console, a car computer, or the like.

[0070] The link 30 may include one or more media or devices capable of transmitting the encoded image from the source device 10 to the destination device 20. In one possible implementation, the link 30 may include one or more communication media that enable the source device 10 to directly send the encoded image to the destination device 20 in real time. In an embodiment of the present application, the source device 10 may modulate the encoded image based on a communication standard, which may be a wireless communication protocol, etc., and may send the modulated image to the destination device 20. The one or more communication media may include wireless and / or wired communication media, for example, the one or more communication media may include a radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, which may be a local area network, a wide area network, or a global network (e.g., the Internet), etc. The one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from the source device 10 to the destination device 20, etc., and the embodiment of the present application does not specifically limit this.

[0071] In a possible implementation, the storage device 40 may store the received encoded image sent by the source device 10, and the destination device 20 may directly obtain the encoded image from the storage device 40. Under such conditions, the storage device 40 may include any of a variety of distributed or locally accessible data storage media, for example, any of the multiple distributed or locally accessible data storage media may be a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded images.

[0072] In one possible implementation, the storage device 40 may correspond to a file server or another intermediate storage device that can store the encoded image generated by the source device 10, and the destination device 20 may stream or download the image stored in the storage device 40. The file server may be any type of server that can store the encoded image and send the encoded image to the destination device 20. In one possible implementation, the file server may include a network server, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive, etc. The destination device 20 may obtain the encoded image through any standard data connection (including an Internet connection). Any standard data connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of the two suitable for obtaining the encoded image stored on the file server. The transmission of the encoded image from the storage device 40 may be streaming transmission, download transmission, or a combination of the two.

[0073] Figure 1 The implementation environment shown is only one possible implementation method, and the technology of the embodiment of the present application can be applied not only to Figure 1 The source device 10 that can encode an image and the destination device 20 that can decode the encoded image shown can also be applicable to other devices that can encode an image and decode the encoded image, and the embodiments of the present application do not specifically limit this.

[0074] exist Figure 1In the illustrated implementation, the source device 10 includes a data source 120, an encoder 100, and an output interface 140. In some embodiments, the output interface 140 may include a modulator / demodulator (modem) and / or a transmitter, where the transmitter may also be referred to as a transmitter. The data source 120 may include an image capture device (e.g., a camera, etc.), an archive containing previously captured images, a feed interface for receiving images from an image content provider, and / or a computer graphics system for generating images, or a combination of these sources of images.

[0075] The data source 120 may send an image to the encoder 100, and the encoder 100 may encode the image received from the data source 120 to obtain an encoded image. The encoder may send the encoded image to the output interface. In some embodiments, the source device 10 directly sends the encoded image to the destination device 20 via the output interface 140. In other embodiments, the encoded image may also be stored in the storage device 40 for the destination device 20 to obtain and use for decoding and / or display later.

[0076] exist Figure 1 In the illustrated implementation environment, the destination device 20 includes an input interface 240, a decoder 200, and a display device 220. In some embodiments, the input interface 240 includes a receiver and / or a modem. The input interface 240 may receive an encoded image via the link 30 and / or from the storage device 40, and then send it to the decoder 200. The decoder 200 may decode the received encoded image to obtain a decoded image. The decoder may send the decoded image to the display device 220. The display device 220 may be integrated with the destination device 20 or may be external to the destination device 20. In general, the display device 220 displays the decoded image. The display device 220 may be any of a variety of types of display devices, for example, the display device 220 may be a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.

[0077] although Figure 11 and 2. Although not shown in the drawings, in some aspects, encoder 100 and decoder 200 may be integrated with an encoder and decoder, respectively, and may include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software for encoding both audio and video in a common data stream or in separate data streams. In some embodiments, if applicable, the MUX-DEMUX units may conform to the ITU H.223 multiplexer protocol, or other protocols such as the user datagram protocol (UDP).

[0078] The encoder 100 and the decoder 200 may each be any of the following circuits: one or more microprocessors, digital signal processors (digital signal processing, DSP), application specific integrated circuits (application specific integrated circuits, ASIC), field programmable gate arrays (field-programmable gate array, FPGA), discrete logic, hardware or any combination thereof. If the technology of the embodiment of the present application is partially implemented in software, the device may store instructions for the software in a suitable non-volatile computer-readable storage medium, and may use one or more processors to execute the instructions in hardware to implement the technology of the embodiment of the present application. Any of the foregoing contents (including hardware, software, a combination of hardware and software, etc.) may be regarded as one or more processors. Each of the encoder 100 and the decoder 200 may be included in one or more encoders or decoders, and any of the encoders or decoders may be integrated as part of a combined encoder / decoder (encoder-decoder) in a corresponding device.

[0079] The embodiments of the present application may generally refer to the encoder 100 as "signaling" or "sending" certain information to another device, such as the decoder 200. The term "signaling" or "sending" may generally refer to the transmission of syntax elements and / or other data used to decode the compressed image. This transmission may occur in real time or near real time. Alternatively, this communication may occur over a period of time, such as when the syntax elements are stored in the encoded bitstream to a computer-readable storage medium at the time of encoding, and the decoding device may then retrieve the syntax elements at any time after the syntax elements are stored to this medium.

[0080] It should be noted that the application scenarios and implementation environments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of application scenarios and implementation environments, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0081] Next, the image encoding method and image decoding method provided in the embodiments of the present application are explained in detail.

[0082] Figure 2 is a flowchart of an image encoding method provided by an embodiment of the present application. The method can be applied to a source device in the above implementation environment, and the source device is also called an encoding end. Figure 2 , the method comprises the following steps.

[0083] Step 201: Acquire a first image, where the first image is an image that needs to be tone mapped.

[0084] The first image may be an image obtained from an HDR / standard dynamic range (SDR) video stream, or an image obtained from an HDR / SDR image source, or an image from other sources. The first image may be in RGB color space format, in YUV color space format, or in any other color space format. The embodiment of the present application does not limit the source and color space format of the first image.

[0085] Since the essence of tone mapping is to compress a high dynamic range into a low dynamic range, the first image can be not only an HDR / SDR image, but also other high dynamic range images, as long as the brightness range of the first image is greater than the brightness range that can be displayed by the decoding end.

[0086] Step 202: Based on the first image, determine the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image through an optimization equation or a deep learning network model, the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness and contrast human eye perception threshold, and the multiple sampling points include a starting point, an end point and a plurality of key points located between the starting point and the end point.

[0087] Among them, image contrast refers to the ratio between the maximum and minimum brightness of a certain area; for a certain pixel in the image, the brightness value of the pixel is the image brightness corresponding to the pixel; the contrast human eye perception threshold refers to the minimum contrast that the human eye can perceive for a certain pixel when the brightness of the pixel changes.

[0088] In some embodiments, Figure 3 As shown, based on the first image, the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image are determined according to the following steps (1)-(4) by optimizing the equation or the deep learning network model.

[0089] (1) A target histogram is generated based on a first image, wherein the target histogram includes a plurality of histogram intervals, wherein the plurality of histogram intervals are divided based on a maximum image brightness and a minimum image brightness of the first image, and the height of the histogram interval indicates the number of pixels in the first image whose brightness is within the histogram interval.

[0090] In some embodiments, the brightness value of each pixel in the first image is obtained, and the maximum brightness value is used as the maximum image brightness, and the minimum brightness value is used as the minimum image brightness; the difference between the maximum image brightness and the minimum image brightness is divided by the specified number of intervals to obtain the width of each histogram interval, and the brightness range included in each histogram interval is determined; the number of pixels in the first image whose brightness is within each histogram interval is counted, and the number is used as the height of each histogram interval.

[0091] Among them, the number of specified intervals is set in advance. The larger the number, the more accurate the tone mapping curve will be, but at the same time the system complexity will also increase. In practical applications, reasonable settings can be made according to needs, and the embodiments of the present application do not limit this.

[0092] For example, please refer to Figure 4 , assuming that the maximum image brightness of the first image is Xmax, the minimum image brightness is Xmin, and the number of specified intervals is M, then the generated target histogram is as follows Figure 4 As shown in the dotted line part, the width of each histogram interval is (Xmax-Xmin) / M, the starting point of the first histogram interval is Xmin, and the end point of the last histogram interval is Xmax.

[0093] (2) Determine the position coordinates of the starting point and the ending point, the position coordinates include a horizontal coordinate and a vertical coordinate, select a point from each of the multiple histogram intervals as one of the multiple key points, and use the image brightness corresponding to the multiple key points as the horizontal coordinates of the multiple key points.

[0094] In some embodiments, the position coordinates of the starting point and the ending point are set default coordinates. The default coordinates are set through historical statistical data. That is, for the horizontal coordinate of the starting point, it can be set to the minimum brightness value of the image that needs to be tone mapped in the historical statistical data; for the vertical coordinate of the starting point, it can be set to the minimum brightness value that the decoding end can display in the historical statistical data; for the horizontal coordinate of the ending point, it can be set to the maximum brightness value of the image that needs to be tone mapped in the historical statistical data; for the vertical coordinate of the ending point, it can be set to the maximum brightness value that the decoding end can display in the historical statistical data; of course, the technician can also set the horizontal / vertical coordinates of the starting point / end point based on subjective experience, and the embodiment of the present application does not limit the values ​​of the horizontal / vertical coordinates of the starting point / end point.

[0095] In other embodiments, the position coordinates of the starting point and the ending point are determined based on the first selection interval, the second selection interval, the minimum reference brightness, and the maximum reference brightness. That is, the horizontal coordinate of the starting point can be set to the horizontal coordinate of any point in the first selection interval; the horizontal coordinate of the ending point can be set to the horizontal coordinate of any point in the second selection interval; the vertical coordinate of the starting point can be set to any brightness value within the left and right fluctuation range of the minimum reference brightness, and the brightness value is greater than 0; the vertical coordinate of the ending point can be set to any brightness value within the left and right fluctuation range of the maximum reference brightness.

[0096] Among them, the first selection interval is the interval from 0 to the right endpoint of the first histogram interval, and is a fully closed interval. The second selection interval is the interval from the left endpoint of the last histogram interval to positive infinity, and is a left-closed and right-open interval. The abscissa of the right endpoint of the first histogram interval is the sum of the minimum image brightness and the width of the histogram interval, and the abscissa of the left endpoint of the last histogram interval is the difference between the maximum image brightness and the width of the histogram interval.

[0097] Among them, the minimum reference brightness refers to the minimum brightness value that the reference decoding end can display, and the maximum reference brightness refers to the maximum brightness value that the reference decoding end can display; the reference decoding end is any decoding end; that is, the embodiment of the present application determines the vertical coordinates of the starting point and the ending point by setting the minimum reference brightness and the maximum reference brightness in advance. In addition, the left and right fluctuation ranges of the minimum reference brightness and the maximum reference brightness are also set in advance to indicate the small fluctuation range near the minimum reference brightness and the maximum reference brightness, and the left and right fluctuation ranges of the minimum reference brightness and the maximum reference brightness can be the same or different, and the embodiment of the present application does not limit this.

[0098] In some embodiments, multiple key points are determined according to a specified rule. The specified rule is set in advance, and the specified rule can be to use the midpoint of the histogram interval as the key point, or the left endpoint or the right endpoint of the histogram interval as the key point, or other rules. The embodiment of the present application does not limit this, as long as the selection method of the multiple key points is the same.

[0099] Based on the above description, the horizontal coordinate of the starting point may be outside the first histogram interval or inside the first histogram interval. If the horizontal coordinate of the starting point is outside the first histogram interval, a point can be selected from the first histogram interval as a key point; if the horizontal coordinate of the starting point is within the first histogram interval, a key point is not selected from the first histogram interval, or a key point is selected from the first histogram, but the horizontal coordinate of the key point is greater than the horizontal coordinate of the starting point.

[0100] Similarly, based on the above description, the horizontal coordinate of the end point may be outside the last histogram interval or inside the last histogram interval. If the horizontal coordinate of the end point is outside the last histogram interval, a point can be selected from the last histogram interval as a key point; if the horizontal coordinate of the end point is inside the last histogram interval, a key point is not selected from the last histogram interval, or a key point is selected from the last histogram interval, but the horizontal coordinate of the key point is smaller than the horizontal coordinate of the end point.

[0101] (3) Determine the target probabilities corresponding to the multiple sampling points, where the target probabilities corresponding to the start point and the end point are specified probabilities, and the target probability corresponding to the key point indicates the probability that the image brightness of the key point falls within the histogram interval.

[0102] Among them, the specified probability is set in advance and can be 0, or any decimal within a very small range greater than 0. For example, the specified probability can be 0.01, 0.001, 0.0001, etc., and the embodiment of the present application does not limit this.

[0103] In some embodiments, the target probability corresponding to the key point may be determined by: determining the sum of the heights of all histogram intervals to obtain the total height of the intervals; determining the ratio of the height of the histogram interval where the first key point is located to the total height of the intervals, using the ratio as the target probability corresponding to the first key point, and obtaining the target probabilities corresponding to other key points in the same manner. The first key point is any one of the multiple key points.

[0104] (4) Based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points, the vertical coordinates of the multiple key points are determined by optimizing the equation or the deep learning network model.

[0105] In some embodiments, the methods of determining the vertical coordinates of the multiple key points by optimizing equations or deep learning network models are different, which will be introduced below respectively.

[0106] The first implementation method is to determine the ordinates of multiple key points by optimizing equations. That is, the contrast human eye perception threshold of multiple sampling points is determined based on the abscissas of multiple sampling points. Based on the position coordinates of the starting point and the end point, the number of multiple sampling points, the target probabilities corresponding to the multiple sampling points, the abscissas of multiple key points, and the contrast human eye perception threshold of multiple sampling points, the ordinates of multiple key points are obtained by solving the optimization equations.

[0107] In some embodiments, the encoder stores a contrast human eye perception curve for indicating the contrast human eye perception threshold corresponding to pixels of different brightness, the horizontal axis of the curve is brightness, and the vertical axis is the contrast human eye perception threshold. After determining the horizontal coordinate of the first sampling point, the vertical coordinate corresponding to the horizontal coordinate of the first sampling point can be found on the curve, and the vertical coordinate is the contrast human eye perception threshold of the first sampling point; in the same way, the contrast human eye perception thresholds corresponding to the other multiple sampling points can be determined. The first sampling point is any one of the multiple sampling points.

[0108] The optimization equation includes but is not limited to any one of the following equations (1)-(3):

[0109]

[0110]

[0111]

[0112] In equation (1), y′ represents a derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the kth sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of the histogram interval, the width of each histogram interval is equal, and λ represents the weight coefficient.

[0113] In equation (1) This section considers image contrast and the contrast perception threshold of the human eye. Indicates the slope of the curve segment between two adjacent sampling points in the tone mapping curve (see Figure 4 Marked in ), the purpose of the accumulation operation is to comprehensively consider the slope of the entire curve. In this part, the argmin function is used to make the slope of the final tone mapping curve as close to 1 as possible, so as to restore the image contrast to the maximum extent, that is, to make the contrast of the tone mapped image as consistent as possible with the contrast of the first image. At the same time, for pixels with different brightness, the minimum contrast that the human eye can perceive is different when the brightness of the pixel changes. Therefore, by adding t(x k ), that is, the contrast human eye perception threshold corresponding to the k-th sampling point, to improve the display effect of the image.

[0114] This part takes the image brightness into consideration, with the goal of making the brightness value of each pixel in the tone-mapped image as consistent as possible with the brightness value of each pixel in the first image.

[0115] In equation (2), y′, N-1, k, p k 、x k , t(x k ), y k ,y k+1 , λ have the same meanings as in equation (1).

[0116] In equation (2) This part is the same as equation (1), but in equation (2), the part that considers the image contrast and the contrast perception threshold of the human eye is The maximum image contrast is restored by making the difference between the pixel gradient of the tone-mapped image and the pixel gradient of the original image as small as possible, that is, by spatial consistency. At the same time, as in equation (1), t(x k ) to improve the display effect of the image.

[0117] In equation (3), y′, N-1, k, p k 、x k , t(x k ), y k ,y k+1 , λ have the same meanings as in equation (1), and ρ represents the Pearson correlation coefficient. Therefore, equation (3) can also be expressed as:

[0118] Among them, P I Represents yk+1 -y k , P J It means t(x k )(x k+1 -x k ).

[0119] In equation (3) This part is also the same as equation (1), but in equation (3), the part that considers the image contrast and the contrast perception threshold of the human eye is The theory of Pearson correlation coefficient is applied, by making y k+1 -y k and t(x k )(x k+1 -x k ) is as small as possible to achieve the maximum restoration of image contrast. At the same time, as in equation (1), t(x k ) to improve the display effect of the image.

[0120] In addition, the weight coefficient λ in equations (1)-(3) is used to adjust the proportion of the two parts on the right side of the equation. The smaller λ is, the more emphasis is placed on the effect of image contrast and the contrast human eye perception threshold on the image display effect compared to image brightness. Conversely, the more emphasis is placed on the effect of image brightness on the image display effect. The weight coefficient is set in advance, and the specific size can be set according to actual needs, and the embodiment of the present application does not limit this.

[0121] When the weight coefficient is set to 0, the optimization equation only considers the image contrast and the human eye perception threshold; when the weight coefficient is set to 1, the optimization equation only considers the image brightness; when the weight coefficient is set to any value between 0 and 1, the optimization equation considers the three factors of image contrast, human eye perception threshold and image brightness at the same time. Of course, when the weight coefficient is set to 0 and t(x k ), the optimization equation only considers the image contrast.

[0122] In the process of solving the ordinate of the key point, it is necessary to combine the optimization equation with y k+1 =y k +y k+1 ′(x k+1 -x k ) and y k+1 ≥y k Solve the problem together. Among them, y k Indicates the ordinate of the kth sampling point, y k+1 Indicates the ordinate of the k+1th sampling point, y k+1 ′ represents the derivative of the ordinate of the k+1th sampling point, xk Indicates the horizontal coordinate of the kth sampling point, x k+1 Represents the horizontal coordinate of k+1 sampling points.

[0123] The above k+1 ≥y k It can ensure the monotonicity of the final generated tone mapping curve, and gradually approximate the solution that satisfies the optimization equation and y in the calculation process. k+1 ≥y k The optimal solution includes the derivatives of the ordinates of multiple key points and the ordinates of multiple key points.

[0124] The second implementation method is to determine the ordinates of multiple key points through a deep learning network model. That is, the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to the multiple sampling points are input into the deep learning network model to obtain the derivatives of the ordinates of the multiple key points output by the deep learning network model. The ordinates of the multiple key points are determined based on the position coordinates of the starting point and the end point, the horizontal coordinates of the multiple key points, and the derivatives of the ordinates of the multiple key points.

[0125] After obtaining the derivatives of the ordinates of multiple key points output by the deep learning network model, the ordinate of the first key point can be determined based on the position coordinates of the starting point, the abscissa of the first key point, and the derivative of the ordinate of the first key point according to the formula y1=y0+y1′(x1-x0), where y1 represents the ordinate of the first key point, y0 represents the ordinate of the starting point, y1′ represents the derivative of the ordinate of the first key point, x1 represents the abscissa of the first key point, and x0 represents the abscissa of the starting point. After determining the ordinate of the first key point, the ordinate of the second key point can be determined by the formula y2=y1+y2′(x2-x1), where y2 represents the ordinate of the second key point, y1 represents the ordinate of the first key point, y2′ represents the derivative of the ordinate of the second key point, x2 represents the abscissa of the second key point, and x1 represents the abscissa of the first key point. By analogy, the ordinates of multiple other key points can be determined.

[0126] For example, please refer to Figure 5 , Figure 5 Where Xmax represents the maximum image brightness, Xmin represents the minimum image brightness, and p k Represents the target probability corresponding to multiple sampling points. By inputting these three into the deep learning network model, we can get d k , that is, the derivative vector composed of the derivatives of the ordinates of multiple key points, and then the derivative vector, combined with the position coordinates of the starting point and the end point, and the abscissas of multiple key points, can be used to obtain Y k , which is the ordinate vector composed of the ordinates of multiple key points.

[0127] In some embodiments, a deep learning network model can be obtained in the following manner: obtaining multiple training samples, each training sample including the maximum sample image brightness, the minimum sample image brightness, the target probabilities corresponding to multiple sample sampling points, and the derivatives of the vertical coordinates of multiple sample key points, wherein the derivatives of the vertical coordinates of the multiple sample key points are determined by an optimization equation; based on multiple training samples, the initial network model is trained to obtain a deep learning network model.

[0128] Among them, the method of acquiring the sample image is the same as the method of acquiring the first image, and the method of determining the maximum sample image brightness, the minimum sample image brightness, and the target probabilities corresponding to multiple sample sampling points are also the same as the method of determining the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to multiple sampling points of the first image described above, and will not be repeated here.

[0129] After obtaining the sample image, the derivatives of the vertical coordinates of multiple sample key points of each sample image can be determined by the above optimization equation, combined with the above maximum sample image brightness, minimum sample image brightness, and the target probabilities corresponding to multiple sample sampling points, so as to obtain each training sample, and then input each training sample into the initial network model, train the initial network model, and thus obtain a deep learning network model.

[0130] In the above training process, the loss function used is Among them, L1 represents the loss function, n is the sum of the number of key points of multiple samples in each training sample, and y i represents the ordinate of the i-th sample key point determined by the above optimization equation, Represents the ordinate of the i-th sample key point predicted by the deep learning network model, y i and The corresponding horizontal axes are the same. This loss function can show the difference between the vertical axis predicted by the deep learning network model and the vertical axis determined by the above optimization equation, and is used to guide the optimization of the deep learning network model.

[0131] Since the deep learning network model is trained by the training samples obtained by the above optimization equation, the deep learning network model, like the above optimization equation, also takes into account at least one of the image brightness, image contrast and contrast human eye perception threshold, and can also enable the image to achieve better display effect after subsequent tone mapping.

[0132] After the ordinates of the multiple key points are determined by the above two implementation methods, each ordinate is combined with the corresponding abscissa to obtain the position coordinates of the multiple key points.

[0133] Step 203: Encode the first image and the position coordinates of the plurality of sampling points into a bitstream.

[0134] The position coordinates of the multiple sampling points are determined by the above step 202. After the position coordinates of the multiple sampling points are determined, the position coordinates of the multiple sampling points can be directly encoded into the bitstream, or can be normalized before being encoded into the bitstream, so that the system complexity of the subsequent decoding end can be reduced.

[0135] In some embodiments, the vertical coordinates of the position coordinates of the multiple sampling points are normalized, and the horizontal coordinates remain unchanged, thereby obtaining the position coordinates of the multiple sampling points after the normalization process.

[0136] Since the encoder does not know the information of the decoder, the position coordinates of the multiple sampling points determined by the encoder are virtual position coordinates. The decoder cannot directly generate a tone mapping curve through the virtual position coordinates, but needs to convert the virtual position coordinates into real position coordinates based on the information of the decoder (mainly the maximum brightness that the decoder can display, that is, the maximum screen brightness), and then determine the tone mapping curve based on the real position coordinates of the multiple sampling points. The normalized vertical coordinates are between 0-1, the smallest vertical coordinate is 0, and the largest vertical coordinate is 1. Therefore, it is only necessary to multiply each vertical coordinate by the maximum brightness that the decoder can display, and then combine it with the corresponding horizontal coordinate to convert the virtual position coordinates into real position coordinates. If normalization is not performed, the conversion cannot be performed in the above manner, but must be performed in other ways including more steps, which will increase the system complexity of the decoder.

[0137] In an embodiment of the present application, when determining the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image, the optimization equation or deep learning network model used takes into account at least one of the image contrast, image brightness and contrast human eye perception threshold, so that the position coordinates of the multiple key points determined are the optimal solution that satisfies the optimization equation conditions, so that the contrast of the subsequent tone-mapped image is consistent with the contrast of the first image to the greatest extent, the brightness of each pixel in the subsequent tone-mapped image is as close as possible to the brightness value of the corresponding pixel in the first image, and it is ensured that after the brightness of the mapped image changes, the human eye can still perceive the brightness difference of each area of ​​the image, and then perceive more details in the tone-mapped image, so as to achieve a better display effect. In addition, before the position coordinates of the multiple sampling points are encoded into the bitstream, the position coordinates are normalized, so that the encoding end can more conveniently convert the virtual position coordinates into real position coordinates, thereby reducing the system complexity of the decoding end.

[0138] Figure 6is a flowchart of an image decoding method provided by an embodiment of the present application. The method can be applied to a target device in the above implementation environment, and the target device is also called a decoding end. Figure 6 , the method comprises the following steps.

[0139] Step 601: Obtain a reconstructed image based on a code stream.

[0140] The reconstructed image refers to an image constructed based on the first image data in the code stream.

[0141] The decoding end parses the relevant image data from the bit stream, processes the image data, and obtains a reconstructed image.

[0142] Step 602: parsing position coordinates of a plurality of sampling points from a bitstream, wherein the position coordinates of the plurality of sampling points are determined by an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point.

[0143] The optimization equation includes but is not limited to any one of equations (1)-(3) described above.

[0144] From the above description, it can be seen that the position coordinates of multiple sampling points parsed from the bitstream are virtual position coordinates. It is also necessary to convert the virtual position coordinates into real position coordinates based on the information of the decoding end (mainly the maximum brightness that the decoding end can display, that is, the maximum screen brightness), and use the obtained real position coordinates as the position coordinates for subsequent application.

[0145] As can also be seen from the above description, the virtual position coordinates obtained by parsing the bitstream may be normalized position coordinates or may be unnormalized position coordinates.

[0146] When the virtual position coordinates are normalized position coordinates, the vertical coordinates in the virtual position coordinates are normalized vertical coordinates, and the horizontal coordinates have not been normalized. In this case, the vertical coordinates in each virtual position coordinate are multiplied by the maximum screen brightness to obtain the converted vertical coordinates, and the converted vertical coordinates are combined with the corresponding horizontal coordinates to obtain the real position coordinates of multiple sampling points.

[0147] When the virtual position coordinates are position coordinates that have not been normalized, the virtual position coordinates can be normalized at the decoding end to obtain normalized virtual coordinates, and then the value of the maximum screen brightness is determined, and the ordinate in each normalized virtual position coordinate is multiplied by the value to obtain the converted ordinate, and the converted ordinate is combined with the corresponding abscissa to obtain the real position coordinates of multiple sampling points. Of course, the virtual position coordinates that have not been normalized can also be converted into real position coordinates by other methods, and the embodiments of the present application are not limited to this.

[0148] In some embodiments, after the real position coordinates of the plurality of sampling points are determined, the data format of the real position coordinates may be transformed.

[0149] In the above process, since the maximum screen brightness value involved is generally in standard data form, such as 100nit, 1000nit, etc., the obtained real position coordinates are generally also in standard data form. In this case, the form of the real position coordinates can be changed into a log form with an arbitrary base, that is, the horizontal coordinate and vertical coordinate in the standard data form are converted into the corresponding horizontal coordinate and vertical coordinate in log form. For example, (10nit, 3nit) is converted into a log form with a base of 10 - (lg10 10 ,lg10 3 ). Of course, it can also be converted into other data forms, which is not limited in the embodiments of the present application.

[0150] Step 603: Obtain a tone mapping curve based on the position coordinates of the multiple sampling points.

[0151] It should be noted that the position coordinates used in this step are real position coordinates in any data format determined in the above step 601 .

[0152] In some embodiments, the tone mapping curve is determined by curve fitting, which includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

[0153] For example, please refer to Figure 4 , Figure 4 Each black dot in the figure represents a sampling point. The first black dot is the starting point with a position coordinate of (0, 0), and the last black dot is the ending point with a position coordinate of (x N ,y N ), (x2, y2) is the position coordinate of the second sampling point, which is also the position coordinate of the first key point, (x k ,y k ) is the position coordinate of the kth sampling point. Assuming that the curve fitting method is a straight line connection method, then Figure 4As shown, by connecting N sampling points in sequence with straight line segments, a tone mapping curve can be successfully established. The horizontal axis of the tone mapping curve is the brightness of the pixel in the reconstructed image, and the vertical axis is the screen display brightness, that is, the brightness of each pixel in the reconstructed image after tone mapping.

[0154] For example, please refer to Figure 7 , Figure 7 Each black dot in represents a sampling point. Assuming that the curve fitting method is cubic spline connection, then Figure 7 As shown, through the position coordinates of every three sampling points, the cubic spline curve connecting every three sampling points can be uniquely determined, and the tone mapping curve can be successfully established by combining the cubic spline curves of every three sampling points.

[0155] Step 604: Perform tone mapping on the reconstructed image based on the tone mapping curve, the maximum screen brightness, and the minimum screen brightness.

[0156] The maximum screen brightness is the maximum brightness value that can be displayed by the decoding end; the minimum screen brightness is the minimum brightness value that can be displayed by the decoding end.

[0157] like Figure 4 As shown, the horizontal axis of the tone mapping curve is the brightness of the pixel in the reconstructed image, and the vertical axis is the screen display brightness. For each pixel in the reconstructed image, the brightness of the pixel is obtained as the horizontal coordinate, and the corresponding vertical coordinate is found on the tone mapping curve. The vertical coordinate is the screen display brightness corresponding to the pixel.

[0158] As can be seen from the above description, the horizontal coordinate of the end point determined by the encoding end may be within the last histogram interval or outside the last histogram interval. In some embodiments, the horizontal coordinate of the end point determined by the encoding end is within the last histogram interval, that is, the horizontal coordinate of the end point is less than the maximum image brightness, which will cause the brightness of individual pixels in the reconstructed image to be greater than the maximum horizontal coordinate of the tone mapping curve. In this case, when tone mapping is performed on the reconstructed image, the brightness of these pixels is directly mapped to the maximum screen brightness. Similarly, in some embodiments, the horizontal coordinate of the starting point determined by the encoding end is within the first histogram interval, that is, the horizontal coordinate of the starting point is greater than the minimum image brightness, which will cause the brightness of individual pixels in the reconstructed image to be less than the minimum horizontal coordinate of the tone mapping interval. In this case, when tone mapping is performed on the reconstructed image, the brightness of these pixels is directly mapped to the minimum screen brightness.

[0159] In an embodiment of the present application, when a tone mapping curve is established based on the position coordinates of multiple sampling points parsed from a bitstream, the position coordinates of multiple key points among the multiple sampling points are determined by an optimization equation or a deep learning network model, and the optimization equation or the deep learning network model takes into account at least one of image contrast, image brightness, and contrast human eye perception threshold, so that the position coordinates of the determined multiple key points are the optimal solution that satisfies the optimization equation conditions, so that the contrast of the subsequent tone-mapped image is kept consistent with the contrast of the first image to the greatest extent, the brightness of each pixel in the tone-mapped image is as close as possible to the brightness value of the corresponding pixel in the first image, and it is ensured that after the brightness of the mapped image changes, the human eye can still perceive the difference in brightness and darkness in each area of ​​the image, and then perceive more details in the tone-mapped image, thereby achieving a better display effect.

[0160] Figure 8 is a structural diagram of an image encoding device provided in an embodiment of the present application. The image encoding device can be implemented by software, hardware, or a combination of both to become part or all of an encoding end device. The encoding end device can be Figure 1 Source device shown. Figure 8 The device includes: an acquisition module 801, a first determination module 802 and an inclusion module 803.

[0161] An acquisition module 801 is used to acquire a first image, where the first image is an image that needs to be tone mapped;

[0162] A first determination module 802 is used to determine, based on the first image, position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image by using an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point;

[0163] The encoding module 803 is used to encode the first image and the position coordinates of multiple sampling points into the bitstream.

[0164] Optionally, the first determining module 802 includes:

[0165] A generating submodule, configured to generate a target histogram based on the first image, wherein the target histogram includes a plurality of histogram intervals, wherein the plurality of histogram intervals are obtained by dividing the maximum image brightness and the minimum image brightness of the first image, and the height of the histogram interval indicates the number of pixels in the first image whose brightness is within the histogram interval;

[0166] A first coordinate determination submodule is used to determine the position coordinates of the starting point and the ending point, the position coordinates including a horizontal coordinate and a vertical coordinate, select a point from each of the multiple histogram intervals as one of the multiple key points, and use the image brightness corresponding to the multiple key points as the horizontal coordinates of the multiple key points;

[0167] The probability determination submodule is used to determine the target probabilities corresponding to the multiple sampling points, wherein the target probabilities corresponding to the starting point and the end point are the specified probabilities, and the target probability corresponding to the key point indicates the probability that the image brightness of the key point falls within the histogram interval;

[0168] The second coordinate determination submodule is used to determine the vertical coordinates of multiple key points based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points through an optimization equation or a deep learning network model.

[0169] Optionally, the second coordinate determination submodule is specifically used for:

[0170] Determine the contrast human eye perception threshold of the plurality of sampling points based on the abscissas of the plurality of sampling points;

[0171] Based on the position coordinates of the starting point and the end point, the number of multiple sampling points, the target probabilities corresponding to the multiple sampling points, the horizontal coordinates of the multiple key points and the contrast human eye perception threshold of the multiple sampling points, the vertical coordinates of the multiple key points are obtained by solving the optimization equation.

[0172] Optionally, the second coordinate determination submodule is specifically used for:

[0173] Inputting the maximum image brightness, the minimum image brightness and the target probabilities corresponding to the multiple sampling points into the deep learning network model, and obtaining the derivatives of the ordinates of the multiple key points output by the deep learning network model;

[0174] The vertical coordinates of the plurality of key points are determined based on the position coordinates of the starting point and the ending point, the horizontal coordinates of the plurality of key points, and the derivatives of the vertical coordinates of the plurality of key points.

[0175] Optionally, the device further comprises:

[0176] A sample acquisition module is used to acquire multiple training samples, each of which includes a maximum sample image brightness, a minimum sample image brightness, target probabilities corresponding to multiple sample sampling points, and derivatives of the ordinates of multiple sample key points, where the derivatives of the ordinates of the multiple sample key points are determined by an optimization equation;

[0177] The model training module is used to train the initial network model based on multiple training samples to obtain a deep learning network model.

[0178] Optionally, the above optimization equation includes but is not limited to any one of the following equations:

[0179] or

[0180] or

[0181]

[0182] Where y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the kth sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

[0183] In an embodiment of the present application, when determining the position coordinates of multiple sampling points of the tone mapping curve corresponding to the first image, the optimization equation or deep learning network model used takes into account at least one of the image contrast, image brightness and contrast human eye perception threshold, so that the position coordinates of the multiple key points determined are the optimal solution that satisfies the conditions of the optimization equation, so that the contrast of the subsequent tone-mapped image is kept consistent with the contrast of the first image to the greatest extent, the brightness of each pixel in the subsequent tone-mapped image is as similar as possible to the brightness value of the corresponding pixel in the first image, and it is ensured that after the brightness of the mapped image changes, the human eye can still perceive the difference in brightness in each area of ​​the image, and then perceive more details in the tone-mapped image, thereby achieving a better display effect.

[0184] It should be noted that: the image encoding device provided in the above embodiment only uses the division of the above functional modules as an example for illustration during encoding. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the image encoding device provided in the above embodiment and the image encoding method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0185] Fig. 9 is a structural diagram of an image decoding device provided in an embodiment of the present application. The image decoding device can be implemented by software, hardware, or a combination of both to become part or all of a decoding end device. The decoding end device can be Figure 1 The destination device shown. Fig. 9 The device includes: an image reconstruction module 901, a coordinate analysis module 902, a curve establishment module 903 and a mapping module 904.

[0186] An image reconstruction module 901, used to obtain a reconstructed image based on a code stream;

[0187] A coordinate parsing module 902 is used to parse the position coordinates of multiple sampling points from the bitstream, where the position coordinates of the multiple sampling points are determined by an optimization equation or a deep learning network model, where the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the multiple sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point;

[0188] A curve establishing module 903, used to establish a tone mapping curve by curve fitting based on the position coordinates of multiple sampling points;

[0189] The mapping module 904 is configured to perform tone mapping on the reconstructed image based on the tone mapping curve, the maximum screen brightness, and the minimum screen brightness.

[0190] Optionally, the above optimization equation includes but is not limited to any one of the following equations:

[0191] or

[0192] or

[0193]

[0194] Where y′ represents the derivative vector composed of the derivatives of the ordinates of multiple key points, N represents the number of sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the kth sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

[0195] Optionally, the tone mapping curve is determined by curve fitting, and the curve fitting method includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

[0196] In an embodiment of the present application, when a tone mapping curve is established based on the position coordinates of multiple sampling points parsed from a bitstream, the position coordinates of multiple key points among the multiple sampling points are determined by an optimization equation or a deep learning network model, and the optimization equation or the deep learning network model takes into account at least one of image contrast, image brightness, and contrast human eye perception threshold, so that the position coordinates of the determined multiple key points are the optimal solution that satisfies the optimization equation conditions, so that the contrast of the subsequent tone-mapped image is kept consistent with the contrast of the first image to the greatest extent, the brightness of each pixel in the tone-mapped image is as close as possible to the brightness value of the corresponding pixel in the first image, and it is ensured that after the brightness of the mapped image changes, the human eye can still perceive the difference in brightness and darkness in each area of ​​the image, and then perceive more details in the tone-mapped image, thereby achieving a better display effect.

[0197] It should be noted that: the image decoding device provided in the above embodiment only uses the division of the above functional modules as an example for explanation during decoding. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the image decoding device provided in the above embodiment and the image decoding method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0198] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiment of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.

[0199] It should be understood that the "multiple" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate the clear description of the technical solution of the embodiments of the present application, in the embodiments of the present application, the words "first", "second" and the like are used to distinguish between the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the words "first", "second" and the like do not limit the quantity and execution order, and the words "first", "second" and the like do not limit them to be necessarily different.

[0200] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0201] The above-mentioned embodiments are provided for the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An image coding method, characterized in that: The method comprises: Acquire a first image, where the first image is an image that needs to be tone mapped; Based on the first image, determine position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image by using an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point; The first image and the position coordinates of the plurality of sampling points are encoded into a bitstream.

2. The method according to claim 1, characterized in that The determining, based on the first image, position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image by optimizing an equation or a deep learning network model includes: Generate a target histogram based on the first image, the target histogram comprising a plurality of histogram intervals, the plurality of histogram intervals being divided based on a maximum image brightness and a minimum image brightness of the first image, the height of the histogram interval indicating the number of pixels in the first image whose brightness is within the histogram interval; Determine the position coordinates of the starting point and the ending point, wherein the position coordinates include a horizontal coordinate and a vertical coordinate; Selecting one point from each of the multiple histogram intervals as one of the multiple key points, and using the image brightness corresponding to each of the multiple key points as the horizontal coordinates of the multiple key points; Determine the target probabilities corresponding to the multiple sampling points respectively, wherein the target probabilities corresponding to the starting point and the ending point are specified probabilities, and the target probability corresponding to the key point indicates the probability that the image brightness of the key point falls within the histogram interval; Based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points respectively, and the horizontal coordinates of the multiple key points, the vertical coordinates of the multiple key points are determined through the optimization equation or the deep learning network model.

3. The method according to claim 2, characterized in that The determining the vertical coordinates of the multiple key points through the optimization equation based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points includes: Determine a contrast human eye perception threshold of the plurality of sampling points based on the abscissas of the plurality of sampling points; Based on the position coordinates of the starting point and the end point, the number of the multiple sampling points, the target probabilities respectively corresponding to the multiple sampling points, the horizontal coordinates of the multiple key points and the contrast human eye perception threshold of the multiple sampling points, the vertical coordinates of the multiple key points are obtained by solving the optimization equation.

4. The method according to claim 2, characterized in that The determining the vertical coordinates of the multiple key points through the deep learning network model based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points includes: Inputting the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to the multiple sampling points into the deep learning network model, and obtaining the derivatives of the ordinates of the multiple key points output by the deep learning network model; The vertical coordinates of the plurality of key points are determined based on the position coordinates of the starting point and the ending point, the horizontal coordinates of the plurality of key points, and the derivatives of the vertical coordinates of the plurality of key points.

5. The method according to claim 4, characterized in that The method further comprises: Acquire multiple training samples, each training sample includes a maximum sample image brightness, a minimum sample image brightness, target probabilities corresponding to multiple sample sampling points, and derivatives of the ordinates of multiple sample key points, wherein the derivatives of the ordinates of the multiple sample key points are determined by the optimization equation; Based on the multiple training samples, the initial network model is trained to obtain the deep learning network model.

6. The method according to any one of claims 1 to 5, characterized in that: The optimization equation includes but is not limited to any one of the following equations: or or Wherein, y′ represents a derivative vector composed of the derivatives of the ordinates of the multiple key points, N represents the number of the multiple sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k represents the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

7. An image decoding method, characterized in that: The method comprises: Obtain a reconstructed image based on the code stream; parsing position coordinates of a plurality of sampling points from the bitstream, the position coordinates of the plurality of sampling points being determined by an optimization equation or a deep learning network model, the optimization equation and the deep learning network model being determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, the plurality of sampling points comprising a starting point, an ending point, and a plurality of key points between the starting point and the ending point; Obtaining a tone mapping curve based on the position coordinates of the plurality of sampling points; The reconstructed image is tone mapped based on the tone mapping curve, the maximum screen brightness and the minimum screen brightness.

8. The method according to claim 7, characterized in that The tone mapping curve is determined by curve fitting, and the curve fitting method includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

9. The method according to claim 7, characterized in that The optimization equation includes but is not limited to any one of the following equations: or or Wherein, y′ represents a derivative vector composed of the derivatives of the ordinates of the multiple key points, N represents the number of the multiple sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k represents the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

10. An image encoding device, characterized in that: The device comprises: An acquisition module, used for acquiring a first image, where the first image is an image that needs to be tone mapped; A first determination module is configured to determine, based on the first image, position coordinates of a plurality of sampling points of a tone mapping curve corresponding to the first image by using an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point; The encoding module is used to encode the first image and the position coordinates of the multiple sampling points into a bit stream.

11. The device according to claim 10, characterized in that The first determining module comprises: A generating submodule, configured to generate a target histogram based on the first image, wherein the target histogram includes a plurality of histogram intervals, wherein the plurality of histogram intervals are obtained by dividing based on a maximum image brightness and a minimum image brightness of the first image, and the height of the histogram interval indicates the number of pixels in the first image whose brightness is within the histogram interval; A first coordinate determination submodule is used to determine the position coordinates of the starting point and the ending point, wherein the position coordinates include a horizontal coordinate and a vertical coordinate, and select a point from each of the multiple histogram intervals as one of the multiple key points, and use the image brightness corresponding to each of the multiple key points as the horizontal coordinate of the multiple key points; A probability determination submodule, used to determine the target probabilities corresponding to the plurality of sampling points, wherein the target probabilities corresponding to the starting point and the ending point are designated probabilities, and the target probability corresponding to the key point indicates the probability that the image brightness of the key point falls within the histogram interval; The second coordinate determination submodule is used to determine the vertical coordinates of the multiple key points based on the position coordinates of the starting point and the end point, the target probabilities corresponding to the multiple sampling points, and the horizontal coordinates of the multiple key points through the optimization equation or the deep learning network model.

12. The device according to claim 11, characterized in that The second coordinate determination submodule is specifically used for: Determine a contrast human eye perception threshold of the plurality of sampling points based on the abscissas of the plurality of sampling points; Based on the position coordinates of the starting point and the end point, the number of the multiple sampling points, the target probabilities respectively corresponding to the multiple sampling points, the horizontal coordinates of the multiple key points and the contrast human eye perception threshold of the multiple sampling points, the vertical coordinates of the multiple key points are obtained by solving the optimization equation.

13. The device according to claim 11, characterized in that The second coordinate determination submodule is specifically used for: Inputting the maximum image brightness, the minimum image brightness, and the target probabilities corresponding to the multiple sampling points into the deep learning network model, and obtaining the derivatives of the ordinates of the multiple key points output by the deep learning network model; The vertical coordinates of the plurality of key points are determined based on the position coordinates of the starting point and the ending point, the horizontal coordinates of the plurality of key points, and the derivatives of the vertical coordinates of the plurality of key points.

14. The device according to claim 13, characterized in that The device also includes: A sample acquisition module, used to acquire multiple training samples, each training sample includes a maximum sample image brightness, a minimum sample image brightness, target probabilities corresponding to multiple sample sampling points, and derivatives of the ordinates of multiple sample key points, wherein the derivatives of the ordinates of the multiple sample key points are determined by the optimization equation; The model training module is used to train the initial network model based on the multiple training samples to obtain the deep learning network model.

15. The device according to any one of claims 10 to 14, characterized in that: The optimization equation includes but is not limited to any one of the following equations: or or Wherein, y′ represents a derivative vector composed of the derivatives of the ordinates of the multiple key points, N represents the number of the multiple sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the kth sampling point, y k Indicates the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

16. An image decoding device, characterized in that: The device comprises: An image reconstruction module, used for obtaining a reconstructed image based on a bit stream; a coordinate parsing module, configured to parse position coordinates of a plurality of sampling points from the bitstream, wherein the position coordinates of the plurality of sampling points are determined by an optimization equation or a deep learning network model, wherein the optimization equation and the deep learning network model are determined based on at least one of image contrast, image brightness, and a contrast human eye perception threshold, and the plurality of sampling points include a starting point, an ending point, and a plurality of key points between the starting point and the ending point; A curve building module, used for obtaining a tone mapping curve based on the position coordinates of the plurality of sampling points; A mapping module is used to perform tone mapping on the reconstructed image based on the tone mapping curve, the maximum screen brightness and the minimum screen brightness.

17. The device according to claim 16, characterized in that The tone mapping curve is determined by curve fitting, and the curve fitting method includes but is not limited to: a straight line connection method, a cubic spline connection method, and a polynomial fitting method.

18. The device according to claim 16, characterized in that The optimization equation includes but is not limited to any one of the following equations: or or Wherein, y′ represents a derivative vector composed of the derivatives of the ordinates of the multiple key points, N represents the number of the multiple sampling points, and p k represents the target probability corresponding to the kth sampling point, x k represents the horizontal coordinate of the kth sampling point, t(x k ) represents the contrast human eye perception threshold of the k-th sampling point, y k represents the ordinate of the kth sampling point, y k+1 represents the ordinate of the k+1th sampling point, δ represents the width of each histogram interval, λ represents the weight coefficient, and ρ represents the Pearson correlation coefficient.

19. A computer-readable storage medium, characterized in that: The storage medium stores instructions, and when the instructions are executed on the computer, the computer executes the steps of the method according to any one of claims 1 to 9.

20. A computer program product, characterized in that The computer program product comprises instructions, and when the instructions are executed on a computer, the computer is caused to execute the steps of the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Image encoding method, image decoding method, and related apparatus

    EP4779558A1

  • Image encoding method, image decoding method, and related apparatus

    WO2025092103A1