Image processing method and device, equipment, storage medium and computer program product

Through the image processing method of electronic devices, multi-frame images are generated and tone mapping and processing are performed, the problem of insufficient dynamic range of traditional electronic devices is solved, the display needs of high and low dynamic range images are realized, and the visual effect and detail retention ability of the image are improved.

CN120378549APending Publication Date: 2025-07-25GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446259.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The images captured by traditional electronic devices have low dynamic range and cannot meet the display needs of high dynamic range images.

Method used

By processing the first image, a second image of multiple frames is obtained, and tone mapping is performed on the second image of multiple frames and the first image, the bit depth of the third image is smaller than the bit depth of the first image, and then the third image is processed to obtain that the dynamic range of the fourth image is greater than the dynamic range of the fourth image, and finally the third image and the fourth image with different dynamic ranges are output.

Benefits of technology

Meet the display needs of electronic devices for content with different dynamic ranges at different heights, improve the visual effect of the image and retain more high-frequency details of the original image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378549A_ABST
    Figure CN120378549A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method, device and equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: processing a first image to obtain multiple frames of second images; performing tone mapping on the multiple frames of second images and the first image to obtain a third image; the bit depth of the third image is smaller than that of the first image; processing the third image to obtain a fourth image; the dynamic range of the third image is larger than that of the fourth image; and outputting the third image and the fourth image. And the third image and the fourth image with different dynamic ranges are output, so that the display requirements of the electronic equipment on contents with different dynamic ranges can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and particularly to an image processing method, apparatus, device, computer-readable storage medium, and computer program product. Background Art

[0002] With the popularization and intelligence of electronic devices, more and more users need to display high-quality images when using electronic devices.

[0003] In traditional technologies, the dynamic range of the images output after an electronic device captures images is relatively low, and cannot meet the display requirements of the electronic device for images with a relatively high dynamic range. Summary of the Invention

[0004] This application provides an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can meet the display requirements of an electronic device for images with a relatively high dynamic range.

[0005] In a first aspect, this application provides an image processing method, the method comprising:

[0006] Processing a first image to obtain multiple frames of second images;

[0007] Performing tone mapping on the multiple frames of the second images and the first image to obtain a third image; the bit depth of the third image is less than the bit depth of the first image;

[0008] Processing the third image to obtain a fourth image; the dynamic range of the third image is greater than the dynamic range of the fourth image;

[0009] Outputting the third image and the fourth image.

[0010] In a second aspect, this application further provides an image processing apparatus, the apparatus comprising:

[0011] A first processing module, configured to process a first image to obtain multiple frames of second images;

[0012] A tone mapping module, configured to perform tone mapping on the multiple frames of the second images and the first image to obtain a third image; the bit depth of the third image is less than the bit depth of the first image;

[0013] A second processing module, configured to process the third image to obtain a fourth image; the dynamic range of the third image is greater than the dynamic range of the fourth image;

[0014] An output module, configured to output the third image and the fourth image.

[0015] In a third aspect, the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the image processing method provided in the first aspect are implemented.

[0016] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image processing method provided in the first aspect are implemented.

[0017] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the image processing method provided in the first aspect are implemented.

[0018] The above-mentioned image processing method, device, electronic device, computer-readable storage medium, and computer program product process the first image to obtain multiple frames of second images, then perform tone mapping processing on the multiple frames of first images and the first image to obtain a third image, and then process the third image to obtain a fourth image. The dynamic range of the third image is greater than that of the fourth image. Finally, the third image and the fourth image with different dynamic ranges are output, which can meet the display requirements of the electronic device for content with different high and low dynamic ranges. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for describing the embodiments of the present application or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart of an image processing method in an embodiment;

[0021] Figure 2 It is a schematic flowchart of processing the first image to obtain multiple frames of second images in an embodiment;

[0022] Figure 3 It is a schematic flowchart of processing the first image to obtain multiple frames of second images in another embodiment;

[0023] Figure 4 It is a schematic diagram of the training and application of a tone mapping network model in an embodiment;

[0024] Figure 5 It is a schematic flowchart of performing tone transformation processing on the third image to obtain the fourth image in an embodiment;

[0025] Figure 6Schematic diagram of the process of performing tone transformation processing on a third image to obtain a fourth image in another embodiment;

[0026] Figure 7 Schematic diagram of a tone mapping system in one embodiment;

[0027] Figure 8 Schematic diagram of a model selection interface in one embodiment;

[0028] Figure 9A Schematic diagram of a model management system in one embodiment;

[0029] Figure 9B Schematic diagram of a model management system in another embodiment;

[0030] Figure 10 Flowchart of an implementation scheme for the test process of a neural network module in one embodiment;

[0031] Figure 11 Flowchart of the process of implementing portrait detection by a neural network module on a device in one embodiment;

[0032] Figure 12 Block diagram of the structure of an image processing apparatus in one embodiment;

[0033] Figure 13 Internal structure diagram of an electronic device in one embodiment. Detailed implementation manners

[0034] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0035] The image processing method provided by the embodiments of the present application can be applied to an electronic device. Among them, the electronic device can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. It should be noted that the electronic device can be a terminal or a server.

[0036] In some exemplary embodiments, as Figure 1 shown, an image processing method is provided, which can be applied to an electronic device and includes the following steps 102 to step 108. Among them:

[0037] Step 102: Process the first image to obtain multiple frames of second images.

[0038] The first image can be an image captured by an electronic device, an image obtained from a network or other devices, a synthesized image, etc. Multiple frames can include two or more frames. The specific number of multiple frames can be configured as needed, such as 2 frames, 3 frames, 4 frames, 5 frames, etc., which is not limited to this. Exemplarily, at least two of the multiple frames of second images have different brightness. When including at least two second images with different brightness, richer brightness information can be obtained during the tone mapping process, resulting in a higher contrast.

[0039] Exemplarily, the processor of the electronic device can perform sampling processing on the first image to obtain multiple frames of second images. Among them, the resolution of the first image can be greater than the resolution of the second image. When the resolution of the first image is greater than the resolution of the second image, and then using the second image for tone mapping processing later can save processing resources and computational volume. The sampling processing can include downsampling, and the sampling ratio of downsampling can be configured as needed, such as a sampling ratio of 2 times, 3 times, 4 times, etc.

[0040] Step 104: Perform tone mapping on the multiple frames of second images and the first image to obtain a third image; the bit depth of the third image is less than the bit depth of the first image.

[0041] Tone mapping is a technique that compresses the brightness information of a high-dynamic-range image into the low-dynamic range supported by an electronic device, aiming to retain image details while making the visual effect closer to the true perception of the human eye. The first image is used to help the second image after tone mapping processing restore more original image high-frequency details to obtain the third image. The size of the third image can be the same as or different from the size of the first image. The resolution of the third image can be the same as or different from the resolution of the first image.

[0042] The bit depth of an image, also known as color depth or pixel depth, represents the number of binary digits used to store color information for each pixel in the image, and is represented by bit. It directly determines the number of colors that the image can represent and the color accuracy. The higher the bit depth, the more delicate the colors that can be represented, and the smoother the transition. A high bit depth can retain more brightness levels, especially in the details of dark and bright parts. A pixel with a bit depth of 8bit can represent 2 8 = 256 color values. The higher the bit depth, the larger the image data volume and the more storage space it occupies. A low bit depth is suitable for network transmission, etc.

[0043] Exemplarily, the processor of the electronic device can use a tone mapping network model to perform tone mapping processing on the multiple frames of second images and the first image to obtain the third image.

[0044] Step 106: Process the third image to obtain a fourth image. The dynamic range of the third image is greater than that of the fourth image.

[0045] The dynamic range of an image refers to the ratio of the brightness between the brightest part and the darkest part that can be recorded or displayed in the image. The dynamic range of the third image is greater than that of the fourth image, that is, the brightness ratio corresponding to the third image is greater than the brightness ratio corresponding to the fourth image.

[0046] Exemplarily, the processor of the electronic device performs a tone transformation process on the third image to obtain a fourth image. The resolution of the fourth image and the third image may be the same or different. The bit depth of the fourth image and the third image may be the same or different. The bit depth of the third image may be greater than the bit depth of the fourth image.

[0047] Step 108: Output the third image and the fourth image.

[0048] Exemplarily, the processor of the electronic device outputs the third image and the fourth image, and stores the third image and the fourth image. The electronic device may directly store the third image and the fourth image, or perform a format conversion on the third image and the fourth image to obtain images that conform to the Ultra High Dynamic Range (UHDR) format.

[0049] In this embodiment, the first image is processed to obtain multiple frames of second images, and then tone mapping processing is performed on the multiple frames of first images and the first image to obtain a third image. Then, the third image is processed to obtain a fourth image. The dynamic range of the third image is greater than that of the fourth image. Finally, the third image and the fourth image with different dynamic ranges are output, which can meet the display requirements of the electronic device for content with different high and low dynamic ranges.

[0050] In some exemplary embodiments, the processing of the first image to obtain multiple frames of second images includes: performing downsampling and brightness adjustment on the first image to obtain multiple frames of second images, where the brightness of each frame of the second image is different, and the resolution of the first image is greater than the resolution of the second image.

[0051] Downsampling is to reduce the resolution of a signal or an image, such as reducing the image size. The sampling ratio of downsampling can be configured as needed. Brightness adjustment can be achieved through brightness gain. Multiple brightness gains can be obtained. Multiple can include two or more. At least two different brightness gains are included in the multiple brightness gains.

[0052] Downsample the first image and adjust its brightness to obtain multiple frames of second images. The brightness of each frame of the second images is different, and the resolution of the first image is higher than that of the second images. Since the resolution of the second images is low, the computational load can be saved, and the hardware resources can be conserved. The different brightness of the multiple frames of second images can effectively prompt the neural network to output the required brightness. At the same time, each image frame has content with normal exposure, and the optimal exposure regions of each frame can be adaptively selected for fusion, resulting in a relatively high contrast of the result.

[0053] In some exemplary embodiments, the downsampling and brightness adjustment of the first image to obtain multiple frames of second images includes:

[0054] Step 202: Perform downsampling on the first image to obtain an intermediate image.

[0055] Downsampling is achieved by reducing the resolution of a signal or an image, such as shrinking the image size. The downsampling ratio can be configured as needed. A downsampling ratio of 2X means that the image is reduced to 1 / 2 of its original size in each dimension, and the total number of pixels becomes 1 / 4 of the original. The downsampling method can be subsampling. For example, if the first image is an image with a resolution of 4K, and the resolution of the intermediate image obtained after downsampling is 1K, then the downsampling ratio is one-fourth.

[0056] Step 204: Adjust the brightness of the intermediate image to obtain multiple frames of second images, where the brightness of each frame of the second images is different, and the resolution of the first image is higher than that of the second images.

[0057] The brightness adjustment can be achieved through brightness gain. Multiple brightness gains can be obtained. "Multiple" can include two or more. At least two different brightness gains are included among the multiple brightness gains.

[0058] Exemplarily, the processor of the electronic device can obtain multiple brightness gains, and multiply the brightness of the intermediate image by the multiple brightness gains in sequence to obtain multiple frames of second images. For example, 5 brightness gains are obtained, namely gain0, gain1, gain2, gain3, and gain4. The intermediate image is multiplied by gain0, gain1, gain2, gain3, and gain4 respectively to obtain ev0, ev1, ev2, ev3, and ev4.

[0059] Downsample the first image to obtain an intermediate image, then adjust the brightness of the intermediate image to obtain multiple frames of second images. The brightness of each frame of the second image is different, and the resolution of the first image is greater than that of the second image. Since the resolution of the second image is low, the computational load can be saved and the hardware resources can be saved. The different brightnesses of the multiple frames of second images can effectively prompt what kind of brightness the neural network output needs. At the same time, each image frame has content with normal exposure, and the optimal exposure area of each frame can be adaptively selected and fused, so that the result has a higher contrast.

[0060] In some exemplary embodiments, step 204 may include: adjusting the brightness of the intermediate image to obtain multiple frames of adjusted images, and then performing truncation processing and gamma processing on each frame of the adjusted image to obtain the corresponding second image.

[0061] Performing truncation processing on each frame of the image may include: the part of the adjusted image that exceeds 1. By performing truncation processing and gamma processing on the adjusted image, a more accurate image can be obtained.

[0062] Such as Figure 3 As shown, the first image is a 4K HDR linear map. Downsample the 4K high dynamic range (HDR) linear map to obtain an intermediate image, then multiply the intermediate image by the ev0 brightness gain, ev1 brightness gain, ev2 brightness gain, ev3 brightness gain, and ev4 brightness gain respectively, then truncate to 0-1, and then perform gamma processing to obtain 5 frames of 1K standard dynamic range (SDR) images, namely ev0, ev1, ev2, ev3, and ev4.

[0063] In some exemplary embodiments, performing tone mapping on the multiple frames of the second image and the first image to obtain a third image includes: obtaining a portrait mask map corresponding to the first image; performing tone mapping on the multiple frames of the second image, the first image, and the portrait mask map corresponding to the first image to obtain a third image.

[0064] Exemplarily, obtaining a portrait mask map corresponding to the first image may include: performing portrait recognition on the first image to obtain a portrait area, and then performing binarization processing on the first image to obtain a portrait mask map, that is, setting the pixel values of the portrait area to 1 and the non-portrait area to 0 to obtain a portrait mask map.

[0065] Perform tone mapping on the multiple frames of the second image, the first image, and the portrait mask map corresponding to the first image through a tone mapping network model to obtain a third image.

[0066] In some exemplary embodiments, tone mapping is performed on multiple frames of the second image and the first image through a tone mapping network model to obtain a third image; wherein, the tone mapping network model is trained based on a high-dynamic range sample image set.

[0067] Input multiple frames of the second image and the first image into the tone mapping network model. After tone mapping, a third image is output. The third image can be a high-dynamic range encoding curve (Hybrid Log-Gamma, abbreviated as HLG), and the color gamut is P3. HLG is an HDR encoding standard jointly developed by the BBC and NHK. Its core is to combine the gamma curve with the logarithmic curve. In the low-brightness area (dark part), the gamma curve is adopted, which is compatible with the gamma curve of SDR to ensure that the basic picture is retained when displayed on traditional devices. In the high-brightness area (bright part), it switches to the logarithmic curve to extend the dynamic range to a maximum brightness of 4000 nits (nits) to avoid the loss of high-brightness details. P3 (DCI-P3) is a wide color gamut standard.

[0068] The tone mapping network model is trained based on a high-dynamic range sample image set. Each sample image pair in the high-dynamic range sample image set includes an HDR sample image and a corresponding HDR annotation image. The HDR annotation image can be obtained by manual retouching or machine retouching of the HDR sample image. Using the HDR sample image set as training data, since the HDR domain has a higher brightness range and color performance than the SDR domain, with more amazing tone and color results, and electronic devices can display HDR content exceeding 1000 nits, the tone mapping network model can output images with a higher dynamic range.

[0069] As Figure 4 shown, the tone mapping network model includes a training and an application process. The training process of the tone mapping network model includes:

[0070] (1) Obtain an HDR sample image set. The HDR sample image set includes HDR sample image pairs, and each HDR sample image pair includes an HDR sample image and an HDR annotation image.

[0071] The HDR sample image can be a 4K HDR sample linear graph. The bit depth of the HDR annotation image is less than the bit depth of the HDR sample linear graph. For example, the bit depth of the HDR annotation image is 10 bit, and the bit depth of the HDR sample linear graph can be 16 - 20 bit, or can also be 10 - 16 bit, etc.

[0072] (2) Perform preprocessing on the HDR sample image, that is, perform downsampling, brightness adjustment, truncation processing, and gamma processing to obtain multiple frames of sample brightness images.

[0073] At least two of the multi-frame sample luminance images have different luminances. The multi-frames can be configured as needed. The resolution of the sample luminance image is less than that of the HDR sample image. The HDR sample image is split into 5 frames of 1K SDR images with different exposures by means of virtual exposure.

[0074] (3)Input the multi-frame sample luminance images and the HDR sample image into an initial tone mapping network model for tone mapping to obtain a predicted HDR image.

[0075] The predicted HDR image can be a 10-bit to 16-bit HDR image. The resolution of the predicted HDR image can be the same as or different from that of the HDR sample image. The resolution of the predicted HDR image is the same as that of the HDR annotation image, and the bit depth of the predicted HDR image is the same as that of the HDR annotation image.

[0076] (4)Calculate the loss value based on the predicted HDR image and the HDR annotation image, and adjust the parameters of the initial tone mapping network model according to the loss value until the training end condition is met to obtain a trained tone mapping network model.

[0077] Training the tone mapping network model with HDR training data can enable the tone mapping network model to have richer color performance, rich tonal levels, and sufficient luminance range.

[0078] The application process of the tone mapping network model includes:

[0079] (5)Obtain a 4K HDR linear image, and perform pre-processing on the 4K HDR linear image, that is, perform downsampling, luminance adjustment, truncation processing, and gamma processing to obtain 5 frames of 1K SDR images.

[0080] (6)Input the 5 frames of 1K SDR images and the 4K HDR linear image into the tone mapping network model for tone mapping to obtain a 10-bit HDR image.

[0081] (7)Perform tone transformation on the 10-bit HDR image to obtain an 8-bit SDR image.

[0082] The 4K HDR linear image can be tone-mapped through a tone mapping network model to obtain a 10-bit HDR image, and then an 8-bit SDR image can be obtained based on the 10-bit HDR image, thereby obtaining a 10-bit HDR image and an 8-bit SDR image. The 10-bit HDR image is used for display on electronic devices such as mobile phones, and the 8-bit SDR image is used for traditional display devices and network sharing. Compared with the related technology where the SDR domain is the direct output target effect, the dynamic range is very low, the effect space is small, and through the post-processing algorithm, it is reversed back to the pseudo-HDR domain, and the dynamic loss cannot be compensated. Through the technical solution of this application, the display requirements of electronic devices for high-quality HDR image content can be met.

[0083] In some exemplary embodiments, the tone mapping network model includes a Transformer structure, a diffusion structure, or a mamba structure.

[0084] In some exemplary embodiments, processing the third image to obtain a fourth image includes: performing a tone transformation on the third image to obtain the fourth image; wherein, the bit depth of the fourth image is less than the bit depth of the third image.

[0085] The tone transformation can be downsampling processing. Performing downsampling processing on the third image to obtain the fourth image. The resolution of the third image can be the same as or different from the resolution of the fourth image. In possible implementation manners, the resolution of the third image is greater than the resolution of the fourth image. The third image can be an HDR image, and the fourth image can be an SDR image, etc. The bit depth of the fourth image is less than the bit depth of the third image. Performing a tone transformation on the third image to obtain the fourth image can obtain support for the display of traditional display devices and network transmission.

[0086] In some exemplary embodiments, as Figure 5 shown, processing the third image to obtain the fourth image includes steps 502 to 510. Among them:

[0087] Step 502, downsample the third image to obtain a sampled image.

[0088] The downsampling ratio can be configured as needed. The downsampling ratio can be 2, 3, 4, 5, 6, 7, 8, etc. For example, if the downsampling ratio is 2X, it means that the image is reduced to 1 / 2 of the original in each dimension (width and height), and the total number of pixels becomes 1 / 4 of the original.

[0089] Step 504, obtain the luminance information and color information of the sampled image.

[0090] Convert the sampled image to a color space with separated luminance and chrominance, such as YUV or LAB. Separate the luminance component from the color space as luminance information, set the luminance component to a neutral value, and combine the original chrominance component to convert back to RGB to obtain color information. The luminance information can be the luminance component, and the color information can be the color component.

[0091] Step 506: Perform tone mapping transformation on the luminance information to obtain the transformed luminance information.

[0092] Perform global tone mapping and local tone mapping on the luminance information to obtain the transformed luminance information.

[0093] Step 508: Reconstruct a color image based on the transformed luminance information and color information.

[0094] Reconstruct the transformed luminance information and color information to obtain an RGB color image.

[0095] Step 510: Upsample the color image to obtain a fourth image; the fourth image has the same resolution as the third image.

[0096] The upsampling ratio is the same as the downsampling ratio. The upsampling ratio can be 2, 3, 4, etc. An upsampling ratio of 2X means that the image is enlarged to 2 times the original in each dimension (width and height), and the total number of pixels becomes 4 times the original. The upsampling method can include bilinear interpolation, transposed convolution (Transposed Convolution), subpixel convolution (Subpixel Convolution), etc.

[0097] By downsampling the third image, then separating the luminance information and color information, performing tone mapping on the luminance information, then reconstructing with the color information to obtain a color image, and then upsampling the color image to obtain a fourth image, the third image and the fourth image have the same resolution, and the bit depth of the third image is greater than that of the fourth image to meet the display requirements of traditional devices.

[0098] In some exemplary embodiments, step 504 includes: processing the sampled image through a decoding curve to obtain a linear domain image; obtaining the luminance information and color information of the linear domain image. Correspondingly, step 510 includes: converting the color image to a curve domain image through an encoding curve, and upsampling the curve domain image to obtain a fourth image.

[0099] By converting the image to a linear domain image, then performing luminance and color decomposition to obtain luminance information and color information, and then converting the color image to a curve domain image and upsampling the curve domain image to obtain a fourth image, it is convenient for calculation.

[0100] In some exemplary embodiments, as Figure 6 shown, performing a tone transformation on the third image to obtain a fourth image includes:

[0101] (1) Downsampling a 10-bit HDR image to obtain a 1K HDR image.

[0102] (2) Applying a Hybrid Log-Gamma Electro-Optical Transfer Function (HLG EOTF) to the 1K HDR image to return it to the linear domain, obtaining a linear HDR image; the EOTF (Electro-Optical Transfer Function) is to restore the electronic signal to the display optical signal.

[0103] (3) Decomposing the linear HDR image into luminance and color components, obtaining the luminance component and color component of the linear HDR image.

[0104] (4) Performing a tone transformation on the luminance component, passing through a global tone transformation mapping and a local tone mapping respectively, and outputting the transformed luminance component.

[0105] (5) Reconstructing the transformed luminance component and color component together into a downscaled pseudo-linear image.

[0106] (6) Converting the pseudo-linear image to the SDR curve domain through the SRGB encoding curve SRGB Opto-Electronic Transfer Function (SRGB OETF), obtaining a 1K SDR image; the OETF (Opto-Electronic Transfer Function) is to convert the optical signal to the electronic signal. The OETF and the EOTF are inverse processes, jointly ensuring the accurate transmission of the image from acquisition to display.

[0107] (7) Upsampling the 1K SDR image to obtain a 4K SDR image.

[0108] Upsampling the 1K SDR image through an upsampling module to obtain a 4K SDR image.

[0109] In some exemplary embodiments, the third image is subjected to a hue transformation through a hue transformation network model to obtain a fourth image; wherein, the hue transformation network model is trained based on a hue transformation training set, and each sample pair in the hue transformation training set includes a high-dynamic-range (HDR) sample image and a standard-dynamic-range (SDR) annotation image. The SDR annotation image serves as the label data corresponding to the HDR sample image, and the bit depth of the SDR annotation image is less than that of the HDR sample image.

[0110] Each sample pair in the hue transformation training set includes an HDR sample image and an SDR annotation image, and the SDR annotation image can be obtained by retouching the corresponding HDR sample image. The HDR sample image can be a 10-bit HDR image, and the SDR annotation image can be an 8-bit SDR image. The color and brightness of the SDR annotation image have a corresponding relationship with the color and brightness of the corresponding HDR sample image. In terms of color, the subjective hue perception of HDR and SDR by the human eye is similar. In terms of brightness, the brightness of HDR and SDR in the medium and low brightness ranges is basically the same, but there are significant differences in the highlight parts. The SDR annotation image serves as the label data corresponding to the HDR sample image. That is, the HDR sample image is used as the input of the hue transformation network model, and the SDR annotation image is used as a reference for the result output after the hue transformation network model transforms the HDR sample image, so as to calculate the loss value.

[0111] The training process of the hue transformation network model includes: obtaining a sample pair, where the sample pair includes an HDR sample image and an SDR annotation image, inputting the HDR sample image into the initial hue transformation network model to output a predicted SDR image, calculating the transformation loss value between the predicted SDR image and the SDR annotation image, and adjusting the parameters of the initial hue transformation network model according to the transformation loss value until the training end condition is met, thereby obtaining the trained hue transformation network model.

[0112] In some exemplary embodiments, the bit depth of the third image is greater than that of the fourth image. For example, the bit depth of the third image is 10 bits, and the bit depth of the fourth image is 8 bits, etc. The bit depth of the third image being greater than that of the fourth image allows the electronic device to obtain images of different qualities to support different display requirements.

[0113] In some exemplary embodiments, the resolution of at least one of the third image and the fourth image is the same as that of the first image. In one embodiment, the bit depth of the third image is less than that of the first image, and the bit depth of the fourth image is less than that of the third image.

[0114] That the resolution of at least one of the third image and the fourth image is the same as that of the first image may include that the resolution of the third image is the same as that of the first image, or the resolution of the fourth image is the same as that of the first image, or the resolutions of both the third image and the fourth image are the same as that of the first image. The bit depth of the third image is less than that of the first image, and the third image can support the display requirements of the electronic device. The bit depth of the fourth image is less than that of the third image, and the fourth image can support the display requirements of the traditional device.

[0115] In some exemplary embodiments, the first image and the third image are high-dynamic range images, the second image and the fourth image are standard-dynamic range images, the bit depth of the first image is greater than that of the third image, and the bit depth of the second image is greater than that of the fourth image.

[0116] The first image can be a 16-bit to 20-bit HDR image, the second image can be a 10-bit SDR image, the third image can be a 10-bit HDR image, and the fourth image can be an 8-bit SDR image. The bit depth of the first image is greater than that of the third image, and the bit depth of the second image is greater than that of the fourth image, thereby obtaining the third image and the fourth image with different display requirements.

[0117] In some exemplary embodiments, the output third image and the fourth image include a standard-dynamic range image and a high-dynamic range gain map; the standard-dynamic range image represents the fourth image, and the standard-dynamic range image and the high-dynamic range gain map represent the third image.

[0118] Convert a 10-bit HDR image and an 8-bit SDR image together into a standard image that conforms to the UHDR format. The UHDR format includes an 8-bit SDR image and an 8-bit HDR gain map. When it is displayed on the screen of the electronic device, the HDR image will be displayed. When the image is transmitted and shared on the network, the SDR image will be transmitted.

[0119] As Figure 7 shown, obtain a linear HDR map, input the linear HDR map into an artificial intelligence (AI) tone mapping system (i.e., a tone mapping network model) for tone mapping to obtain a 10-bit HDR image, perform a down-conversion on the 10-bit HDR image to obtain an 8-bit SDR image, and convert the 10-bit HDR image and the 8-bit SDR image to obtain a UHDR image.

[0120] Figure 8Schematic diagram of the model selection interface 800 in some exemplary embodiments. In some embodiments, the model selection interface 800 may be presented within a web browser. For example, a user may access a website, and the model selection interface 800 is presented to the user within the web browser of the website. As depicted, the model selection interface 800 may provide a variety of menus 802-810, options within these menus, and sub-options or categories 812-814 that the user may select to achieve the AI functions required for their application. It should be noted that the user interface shown here is for illustrative purposes and is not limited in any way to the menus, options, and sub-options shown in this document. A variety of other menus, options, and sub-options are possible and within the scope of this application.

[0121] As Figure 8 shown, the model selection interface 800 may include a task menu 802, a device type menu 804, and constraint menus 806-810. The task menu 802 may include a variety of AI tasks from which the user may select a desired task to incorporate into their application. By way of example and not limitation, these tasks may include scene recognition, image tagging, object detection, object tracking, object segmentation, human pose estimation, image enhancement, behavior recognition, human emotion recognition, speech recognition, text recognition, natural language understanding, tone mapping, and / or any kind of data processing task. In possible embodiments, one or more categories 812 may be included within one or more of these tasks. The one or more categories within a task may include specific objects or items of interest that the user may wish to focus on particularly. For example, the categories within the object detection task may include people, trees, cars, bicycles, dogs, buildings, etc. Also, the categories within the scene recognition task may include beaches, snow-capped mountains, flowers, buildings, autumn leaves, waterfalls, night scenes, etc. Again, the categories within the image enhancement task may include denoising, super-resolution, etc.

[0122] The device type menu 804 may include different types of user devices on which a given task (e.g., the task selected from the task menu 802) may be deployed. By way of example and not limitation, user devices may include smart phones, tablets, smart watches, ARM-based platforms, embedded sensors, cameras, Intel-based platforms, drones, Snapdragon-based platforms, NVIDIA-based platforms (e.g., various GPUs), X-86 series, Ambarella platforms, etc. From the list of user devices, the user may select the device on which they wish to deploy the selected task.

[0123] The constraint menus 806 - 810 may include a memory constraint menu 806, a latency constraint menu 808, and a power constraint menu 810. The memory constraint menu 806 may enable a user to specify the amount of memory to be used for their selected task. For example, as shown in the figure, the user may choose to allocate 100 MB of memory for an object detection task. The latency constraint menu 808 may enable a user to specify the refresh rate or frames per second (FPS) at which they desire their task to run. For example, as shown in the figure, the user may choose to run the object detection task at 10 FPS. The power constraint menu 810 may enable a user to specify the amount of power to be utilized by their selected task. For example, as shown in the figure, the user may choose to allocate 1 JPS (Joules per second) of power for the object detection task.

[0124] Figure 9A Schematic diagram of the various components of a model management system 900 and their associated functions for allocating an appropriate AI model according to user specifications in some exemplary embodiments. Once the user has defined their specifications (e.g., task, device type, memory constraint, power constraint, latency constraint) using the Figure 8 interface discussed, the user specifications can be sent to the query generator server 902. The query generator server 902 receives the user specifications and generates a query 904, which can be used to retrieve one or more raw models 908 (or device - level models) from the database 906. The device - level models can be defined in certain programming languages. For example, the device - level models can be defined in the Python language. The database 906 may include multiple AI models that can be defined and stored based on previous user specifications. The previous user specifications may or may not be related to the current user specifications. For example, a previous specification may be related to an image tagging task for an application designed to operate on an Android phone, and the model management system 900 generates a model according to this specification and stores the model in the database 906 for future access and / or retrieval. Storing previous models or having a database of such models is advantageous because it can avoid creating models from scratch and can reuse and adapt existing models to match the current user specifications.

[0125] The original or advanced model 908 retrieved from the database 906 can be processed by the model compiler 910 to make it suitable according to the current user specifications. For example, the model compiler 910 can compile the model to remove any unnecessary modules / components and add any required components to match the user specifications. For example, the original model can be pre-generated based on user specifications such as task: object detection; device type: Android; memory constraint: 1GB; latency constraint: 100FPS; power constraint: 1JPS; and some additional constraints. The model compiler 910 can compile the original model to remove the additional constraints, and can change the device type to a mobile phone and the memory constraint to 100MB to make the model suitable for the current user specifications. In a possible implementation, the output of the model compiler 910 is a compact model package or model binary 912. In a possible implementation, compiling the original model essentially converts the advanced or original model into a model binary, which may also include user queries, such as Figure 9A as shown

[0126] The interface binder 914 can bind the compact model / model binary with one or more interface functions. The interface function can be a function that the actual user will use in the application or software once the model is deployed to the application. For example, to perform the task of image tagging, the interface function can process the input image, and the output of the function will be a set of tags within the image.

[0127] Once one or more interface functions are bound to the model binary, the final software binary (a combination of the model binary and the interface functions) can be sent to the distributor 918, which creates a link 920 for the user to download the final software binary 916 on their device. Upon download, the user can incorporate the received binary into their application and then can finally release the binary for use by the end user. The end user can use the developer's software to perform AI functions.

[0128] In a possible implementation, the built-in automatic benchmarking component in the model can automatically evaluate the performance of the model based on, for example, the extent to which the provided model serves the user's needs and various performance constraints associated with the model (such as the amount of memory and power usage, speed, etc.). The automatic benchmarking component can report the benchmarked data to the model management system 900, which can store the data in the database 906 of the AI model and use it for training the model. In some implementations, model training can also be performed based on feedback provided by the user at query time.

[0129] Figure 9BSchematic diagram of additional components of a model management system 900 for training an AI model pool in some exemplary embodiments and their associated functions. Training of the model may include improving the performance of an existing model, adding one or more new features to an existing model, and / or creating a new model. The user may provide training data at query time, or in other words, when the user defines their specifications using the model selection interface 800. By way of example and not limitation, the training data may include a request from the user to perform object detection on objects that may not currently be included in the model selection interface 800, but the user may have data corresponding to those objects. The user may then provide their data along with their desired specifications to the model management system 900, which may use the data to refine its existing model related to object detection to include the new object categories provided by the user. As another example, the user may provide performance feedback for the model at query time (e.g., the extent to which the previous model served their needs in terms of speed, memory, and power usage), and the model management system 900 may use the performance feedback to train the AI model pool (indicated by reference numeral 922) on the server. Other possible ways of training the model are also possible and within the scope of the present disclosure.

[0130] The trained AI model 924 may be sent to a ranker 926 for ranking. The ranker 926 may rank the models based on various factors. These factors may include, for example, performance metrics associated with each of the models, benchmark data, user feedback on the models, etc. Based on this ranking, the top-level model 928 may be determined. For example, the top five ranked models may be selected and sent to the model converter 930 of the device. The model converter 930 may convert the highest ranked model into a device-level model 932, as discussed above. Automatic benchmarking (indicated by reference numeral 934) may be performed on these models in terms of speed, memory, and power constraints, and finally these device-level models may be stored in the database 906 of AI models together with the benchmark data for future access and / or retrieval. In a possible implementation, Figure 9B The steps 922-934 shown may be performed Figure 9A asynchronously with the steps shown.

[0131] Figure 10 Flowchart of an implementation of a test process for a neural network module in some exemplary embodiments. During the test process, a sample image input may be provided to the neural network module along with operation parameters. The neural network module may provide sample output data by processing the sample input image using the operation parameters. The sample output data may be compared with the known data of the sample image to see if the data matches in the matching data. The neural network module may be a tone mapping network model or a tone transformation network model.

[0132] If the sample output data matches the known data of the sample image, the operating parameters are set. If the sample output data does not match the known data of the sample image (within the desired tolerance), the training process can be fine-tuned. Fine-tuning the training process can include providing additional training images to the training process and / or other adjustments during the training process to optimize the operating parameters of the neural network module (or generate new operating parameters).

[0133] Once the operating parameters for the neural network module are set in setting the operating parameters, the operating parameters can be applied to the device by providing the operating parameters to the neural network module in the electronic device or server. In some embodiments, the operating parameters of the neural network module for training and the operating parameters of the neural network module applied on the device are in different numerical representation modes. For example, the neural network module for training may use floating-point numbers, while the neural network module applied on the device uses integers. Thus, in such embodiments, the operating parameters of the neural network module for training are converted from floating-point operating parameters to integer operating parameters of the neural network module applied on the device.

[0134] After providing the operating parameters to the neural network module on the device, the neural network module can operate on the device to implement a human detection process on the device. Figure 11 A flowchart for implementing a human detection process by the neural network module on the device in some exemplary embodiments. The image input can include an image captured using a camera on the device. The captured image can be a flood infrared illumination image or a depth map image. The human detection process can be used to detect whether a human exists in the image (e.g., placing a bounding box around the human), and if a human is detected, evaluate the values of the attributes of the human (e.g., position, pose, and / or distance).

[0135] The captured image from the image input can be provided to an encoder process. The encoder process can be performed by an encoder. In some embodiments, the encoder module is a multi-scale convolutional neural network. In the encoder process, the encoder module can encode the image input to represent the features in the image as feature vectors in a feature space. The encoder process can output the feature vectors. The feature vectors can be, for example, the encoded image features represented as vectors.

[0136] The feature vector can be provided to a decoder process. The decoder process can be performed by an encoder module. In some embodiments, the decoder module can be a recurrent neural network. In the decoder process, the decoder module can decode the feature vector to evaluate one or more attributes of the image input to determine (e.g., extract) output data from the image input. Decoding the feature vector can include classifying the feature vector using classification parameters determined during a training process. Classifying the feature vector can include operating on the feature vector using one or more classifiers or networks that support classification.

[0137] In some embodiments, the decoder process includes decoding the feature vectors for each region in the feature space. The feature vectors from each region of the feature space can be decoded into non-overlapping boxes in the output data. In some embodiments, decoding a feature vector for a region (e.g., extracting information from the feature vector) includes determining (e.g., detecting) whether there is a human figure in the region. Since the decoder process operates on each region in the feature space, the decoder module can provide a human figure detection score for each region in the feature space (e.g., based on a prediction of a confidence score regarding whether a human figure or a part of a human figure is detected / present in the region). In some embodiments, using an RNN, multiple predictions regarding the presence of a human figure (or part of a human figure) can be provided for each region of the feature space, where the prediction includes predictions regarding both the part of the human figure within the region and the part of the human figure around the region (e.g., in adjacent regions). These predictions can be folded into a final decision regarding the presence of a human figure in the image input (e.g., detection of a human figure in the image input). In some embodiments, the predictions are used to form (e.g., place) a bounding box around the detected human figure in the image input. The output data can include a decision regarding the presence of a human figure in the image input (e.g., in the captured image) and the bounding box formed around the human figure.

[0138] In some embodiments, the human figure detection process detects the presence of a human figure in the image input regardless of the orientation of the human figure in the image input. For example, a neural network module applied in an electronic device can operate the human figure detection process using the operating parameters implemented by a training process that is developed to detect human figures in any orientation in the image input, as described above. Thus, the human figure detection process can detect human figures in any orientation in the image input without rotating the image and / or receiving any other sensor input data that can provide information regarding the orientation of the image (e.g., accelerometer or gyroscope data). Detecting the face in any orientation also increases the range of pose estimation in the bounding box. For example, the roll estimation in the bounding box ranges from -180° to +180° (all roll orientations).

[0139] In some embodiments, the human portrait detection process detects the presence of a partial human portrait in the image input. The partial human portrait detected in the image input may include any part of the human portrait present in the image input. The amount of the human portrait required to be present in the image input for detection and / or the ability of the neural network module to detect the partial human portrait in the image input may depend on the operating parameters used to generate the neural network module and / or the training of the detectable human portrait features in the image input.

[0140] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0141] Based on the same inventive concept, the embodiments of the present application also provide an image processing apparatus for implementing the above-mentioned image processing method. The solution provided by this apparatus for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following image processing apparatus can refer to the limitations on the image processing method in the above text, and will not be repeated here.

[0142] In some exemplary embodiments, as Figure 12 shown, an image processing apparatus 1200 is provided, including a first processing module 1202, a tone mapping module 1204, a second processing module 1206, and an output module 1208.

[0143] The first processing module 1202 is configured to perform image processing on the first image to obtain multiple frames of second images.

[0144] The tone mapping module 1204 is configured to perform tone mapping on the multiple frames of the second images and the first image to obtain a third image; the bit depth of the third image is less than the bit depth of the first image.

[0145] The second processing module 1206 is configured to process the third image to obtain a fourth image; the dynamic range of the third image is greater than the dynamic range of the fourth image.

[0146] The output module 1208 is configured to output the third image and the fourth image.

[0147] In some exemplary embodiments, the first processing module 1202 is further configured to downsample and adjust the brightness of the first image to obtain multiple frames of second images, where the brightness of each frame of the second images is different, and the resolution of the first image is greater than that of the second images.

[0148] In some exemplary embodiments, the first processing module 1202 includes a downsampling module and a brightness adjustment module.

[0149] The downsampling module is configured to perform downsampling processing on the first image to obtain an intermediate image;

[0150] The brightness adjustment module adjusts the brightness of the intermediate image to obtain multiple frames of second images, where the brightness of each frame of the second images is different, and the resolution of the first image is greater than that of the second images.

[0151] In some exemplary embodiments, the above device further includes: a portrait mask map acquisition module. The portrait mask map acquisition module is configured to acquire the portrait mask map corresponding to the first image.

[0152] The tone mapping module 1204 is further configured to perform tone mapping on multiple frames of the second images, the first image, and the portrait mask map corresponding to the first image to obtain a third image.

[0153] In some exemplary embodiments, the tone mapping module 1204 is further configured to perform tone mapping on multiple frames of the second images and the first image through a tone mapping network model to obtain a third image; wherein, the tone mapping network model is trained based on a high dynamic range sample image set.

[0154] In some exemplary embodiments, the second processing module 1206 is further configured to perform tone transformation on the third image to obtain a fourth image; wherein, the bit depth of the fourth image is less than that of the third image.

[0155] In some exemplary embodiments, the second processing module 1206 is further configured to downsample the third image to obtain a sampled image; acquire the brightness information and color information of the sampled image; perform tone mapping transformation on the brightness information to obtain the transformed brightness information; reconstruct a color image according to the transformed brightness information and the color information; upsample the color image to obtain a fourth image; the fourth image has the same resolution as the third image.

[0156] In some exemplary embodiments, the second processing module 1206 is further configured to perform a hue transformation on the third image through a hue transformation network model to obtain a fourth image; wherein, the hue transformation network model is trained based on a hue transformation training set, and each sample in the hue transformation training set includes a high-dynamic range sample image and a standard-dynamic range annotation image, and the standard-dynamic range annotation image serves as the label data corresponding to the high-dynamic range sample image, and the bit depth of the standard-dynamic range annotation image is less than the bit depth of the high-dynamic range sample image.

[0157] In some exemplary embodiments, the bit depth of the third image is greater than the bit depth of the fourth image.

[0158] In some exemplary embodiments, the resolution of at least one of the third image and the fourth image is the same as the resolution of the first image; the bit depth of the fourth image is less than the bit depth of the third image.

[0159] In some exemplary embodiments, the first image and the third image are high-dynamic range images, the second image and the fourth image are standard-dynamic range images, and the bit depth of the second image is greater than the bit depth of the fourth image.

[0160] In some exemplary embodiments, the output third image and the fourth image include a standard-dynamic range image and a high-dynamic range gain map; the standard-dynamic range image represents the fourth image, and the standard-dynamic image and the high-dynamic range gain map represent the third image.

[0161] It can be understood that the functions implemented by each module in the above device are consistent with the corresponding steps in the above image processing method, and the implementation process is the same, which will not be elaborated here. For the specific implementation, refer to the implementation of the corresponding steps in the above image processing method.

[0162] Each module in the above image processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0163] In an exemplary embodiment, an electronic device is provided. The electronic device can be a terminal, and its internal structure diagram can be as Figure 13As shown in the figure. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and external devices. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an image processing method. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.

[0164] Those skilled in the art can understand that Figure 13 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0165] In an exemplary embodiment, an electronic device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the steps of the image processing method in the above embodiment.

[0166] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of the image processing method in the above embodiment.

[0167] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps of the image processing method in the above embodiment.

[0168] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0169] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.

[0170] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0171] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several variations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. An image processing method, characterized in that, The method includes: Processing a first image to obtain multiple frames of second images; Performing tone mapping on the multiple frames of the second images and the first image to obtain a third image; the bit depth of the third image is less than the bit depth of the first image; Processing the third image to obtain a fourth image; the dynamic range of the third image is greater than the dynamic range of the fourth image; Outputting the third image and the fourth image.

2. The method according to claim 1, characterized in that The processing of the first image to obtain multiple frames of second images includes: Downsampling and brightness adjustment are performed on the first image to obtain multiple frames of second images, wherein the brightness of each frame of the second image is different, and the resolution of the first image is greater than the resolution of the second image.

3. The method according to claim 1, characterized in that, The performing tone mapping on the multiple frames of the second images and the first image to obtain a third image includes: Obtaining a portrait mask image corresponding to the first image; Performing tone mapping on the multiple frames of the second images, the first image, and the portrait mask image corresponding to the first image to obtain a third image.

4. The method according to claim 1, wherein Performing tone mapping on the multiple frames of the second images and the first image through a tone mapping network model to obtain a third image; wherein the tone mapping network model is trained based on a high dynamic range sample image set.

5. The method according to claim 1, characterized in that, The processing of the third image to obtain a fourth image includes: Performing tone transformation on the third image to obtain a fourth image; wherein the bit depth of the fourth image is less than the bit depth of the third image.

6. The method according to claim 5, characterized in that, The performing tone transformation on the third image to obtain a fourth image includes: Downsampling the third image to obtain a sampled image; Obtaining the brightness information and color information of the sampled image; Performing tone mapping transformation on the brightness information to obtain transformed brightness information; Reconstructing a color image according to the transformed brightness information and the color information; Upsampling the color image to obtain a fourth image; the resolution of the fourth image is the same as the resolution of the third image.

7. The method according to claim 5, wherein Performing tone transformation on the third image through a tone transformation network model to obtain a fourth image; wherein the tone transformation network model is trained based on a tone transformation training set, and each sample in the tone transformation training set includes a high dynamic range sample image and a standard dynamic range annotation image, and the standard dynamic range annotation image serves as the label data corresponding to the high dynamic range sample image, and the bit depth of the standard dynamic range annotation image is less than the bit depth of the high dynamic range sample image.

8. The method according to any one of claims 1 to 7, characterized in that, The bit depth of the third image is greater than the bit depth of the fourth image.

9. The method according to any one of claims 1 to 7, characterized in that, The resolution of at least one of the third image and the fourth image is the same as the resolution of the first image.

10. The method according to any one of claims 1 to 7, characterized in that, The first image and the third image are high dynamic range images, the second image and the fourth image are standard dynamic range images, the bit depth of the first image is greater than the bit depth of the third image, and the bit depth of the second image is greater than the bit depth of the fourth image.

11. The method according to any one of claims 1 to 7, characterized in that, The output third image and the fourth image include a standard dynamic range image and a high dynamic range gain map; the standard dynamic range image represents the fourth image, and the standard dynamic image and the high dynamic range gain map represent the third image.

12. An image processing apparatus, characterized in that, The device includes: A first processing module, configured to process a first image to obtain multiple frames of second images; A tone mapping module, configured to perform tone mapping on the multiple frames of the second images and the first image to obtain a third image; the bit depth of the third image is less than the bit depth of the first image; A second processing module, configured to process the third image to obtain a fourth image; the dynamic range of the third image is greater than the dynamic range of the fourth image; An output module, configured to output the third image and the fourth image.

13. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.