High dynamic range image processing method, electronic equipment and readable storage medium

By inserting the first intermediate frame between the short exposure frame and the mid-exposure frame, inserting the second intermediate frame between the mid-exposure frame and the long-exposure frame, adjusting the exposure span, the problem of difficulty in fusion between the short exposure image and the long-exposure image in the prior art is solved, and a wider exposure dynamic range and better HDR imaging effect are achieved.

CN120238747APending Publication Date: 2025-07-01HONOR DEVICE CO LTD

Patent Information

Application Number
CN202311777333.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the high dynamic range imaging, the exposure span between the short-exposure image and the long-exposure image is too large, resulting in too low image similarity and difficult to successfully fuse, which in turn affects the HDR imaging effect.

Method used

By inserting at least one first intermediate frame between the short exposure frame and the mid exposure frame, and inserting at least one second intermediate frame between the mid exposure frame and the long exposure frame, the exposure span of each adjacent frame is adjusted to make it appropriate, thereby improving the similarity of the image and the alignment and registration success rate of feature points.

Benefits of technology

It achieves a wider dynamic range of exposure, improves the imaging effect of HDR images, and provides users with a better HDR imaging experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238747A_ABST
    Figure CN120238747A_ABST
Patent Text Reader

Abstract

The invention discloses a high dynamic range image processing method, electronic equipment and a readable storage medium, and relates to the technical field of image.The method comprises the steps that when shooting is carried out in an HDR mode, the exposure amounts and the frame numbers corresponding to a short exposure frame, a middle exposure frame, a long exposure frame, a first middle frame and a second middle frame are determined, the exposure amount of the first intermediate frame is greater than the exposure amount of the short exposure frame and less than the exposure amount of the intermediate exposure frame, and the exposure amount of the second intermediate frame is greater than the exposure amount of the intermediate exposure frame and less than the exposure amount of the long exposure frame; and generating a high dynamic range image according to the short exposure frame, the middle exposure frame, the long exposure frame, the first middle frame and the second middle frame. Based on the scheme of the invention, a relatively large exposure span can be set, and the HDR image is generated by fusing the original images with more than three exposures, so that a relatively wide exposure dynamic range is realized, the HDR image with a relatively wide dynamic range is obtained, and a relatively good HDR imaging effect is provided for a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image technology, and particularly to a high dynamic range image processing method, an electronic device, and a readable storage medium. Background Art

[0002] Scenes in the real world often have a very high dynamic range in terms of brightness, that is, a very large difference in light and darkness or contrast. However, the dynamic range sensed by sensors in ordinary digital imaging devices (such as cameras or mobile phones) is much smaller, resulting in difficulty in showing all details in a single exposure image when shooting scenes with a large dynamic range, and overexposure or underexposure will occur in places that are too bright or too dark. In order to make the imaging result present rich color details and light and dark levels, and better match the cognitive characteristics of the human eye to real-world scenes, high dynamic range (HDR) imaging has become an increasingly popular imaging technology in digital imaging devices.

[0003] High dynamic range imaging usually obtains long exposure images, short exposure images, and medium exposure images by using different exposure amounts for the same scene, and then combines the long exposure images, short exposure images, and medium exposure images into an HDR image. However, when the exposure amount span between the short exposure image and the long exposure image is large, the similarity between the images is too low, which will cause the synthesis to fail, resulting in poor HDR imaging effects. Summary of the Invention

[0004] This application provides a high dynamic range image processing method, an electronic device, and a readable storage medium, which can achieve a wider dynamic range and provide better HDR imaging effects for users. The technical solutions are as follows:

[0005] In a first aspect, a high dynamic range image processing method is provided, which is applied to an electronic device. The method includes:

[0006] When the high dynamic range imaging mode of the shooting function of the electronic device is turned on, upon detecting a shooting instruction, the shooting instruction requests the electronic device to take a picture; in response to the shooting instruction, determine the exposure amounts and the number of frames corresponding to the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame, where the exposure amount of the first intermediate frame is greater than the exposure amount of the short exposure frame and less than the exposure amount of the medium exposure frame, and the exposure amount of the second intermediate frame is greater than the exposure amount of the medium exposure frame and less than the exposure amount of the long exposure frame; according to the exposure amounts and the number of frames corresponding to the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame, output the corresponding short exposure frame, medium exposure frame, long exposure frame, first intermediate frame, and second intermediate frame; generate a high dynamic range image based on the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame.

[0007] Based on the above technical solution, the present application can set a relatively large exposure amount span between the short exposure frame and the long exposure frame, then insert at least one first intermediate frame between the short exposure frame and the medium exposure frame, and insert at least one second intermediate frame between the medium exposure frame and the long exposure frame. The exposure amount span between each adjacent two frames of images is appropriate, so that the similarity between each adjacent two frames of images is high, and a larger number of feature points can be found for alignment and registration, so that the short exposure frame, the medium short frame, the medium exposure frame, the medium long frame, and the long exposure frame can be successfully fused. The method provided by the embodiments of the present application generates an HDR image by fusing frames with more than three exposure amounts, can achieve a wider dynamic range of exposure amounts, obtain an HDR image with a wider dynamic range, and provide a better HDR imaging effect for users.

[0008] In combination with the first aspect, in some implementation manners of the first aspect, determining the number of frames of the first intermediate frame and the second intermediate frame includes: determining the total number of frames M taken; according to the total number of frames M taken and the number of frames corresponding to the short exposure frame, the medium exposure frame, and the long exposure frame, determining the number of frames of the first intermediate frame and the second intermediate frame, and the number of frames corresponding to the short exposure frame, the medium exposure frame, and the long exposure frame is all 1.

[0009] In combination with the first aspect, in some implementation manners of the first aspect, determining the exposure amount of the first intermediate frame and the second intermediate frame includes: determining a first difference between the exposure amount of the medium exposure frame and the exposure amount of the short exposure frame, and a second difference between the exposure amount of the long exposure frame and the exposure amount of the medium exposure frame; according to the first difference and the number of frames of the first intermediate frame, determining the exposure amount of each first intermediate frame; according to the second difference and the number of frames of the second intermediate frame, determining the exposure amount of each second intermediate frame.

[0010] In combination with the first aspect, in some implementation manners of the first aspect, when the number of frames of the first intermediate frame is greater than 1, the exposure amounts of each first intermediate frame are all different; when the number of frames of the second intermediate frame is greater than 1, the exposure amounts of each second intermediate frame are all different. Since the image details obtained under different exposure amounts are different, this helps to obtain richer and more differentiated picture detail information. Based on information theory, after fusing more and richer information, the imaging quality of the image output by the deep learning network is better.

[0011] In combination with the first aspect, in some implementation manners of the first aspect, the exposure amount is determined by the exposure time and the sensitivity. The exposure times corresponding to the medium exposure frame, the first intermediate frame, and the second intermediate frame are different, and the sensitivities corresponding to the medium exposure frame, the first intermediate frame, and the second intermediate frame are the same. In this way, while obtaining an HDR image with a wider dynamic range, the total exposure time required for shooting remains unchanged, avoiding the extension of the shooting time when inserting different exposure frames, and can improve the user experience.

[0012] In combination with the first aspect, in some implementation manners of the first aspect, generating a high-dynamic range image according to a short-exposure frame, a medium-exposure frame, a long-exposure frame, a first intermediate frame, and a second intermediate frame, including: processing the short-exposure frame, the medium-exposure frame, the long-exposure frame, the first intermediate frame, and the second intermediate frame by using a target network model to obtain a high-dynamic range image, where the target network model is used to perform feature fusion on two adjacent exposure frames, and continuously perform feature fusion on the fused features step by step. In this way, the target network model can better adapt to the method provided in this application for fusing frames with more than three exposure amounts, achieve a better HDR imaging effect, and has a relatively simple network structure, which can improve the processing efficiency.

[0013] In combination with the first aspect, in some implementation manners of the first aspect, the target network model is trained based on a transformer network and a feature extractor, and the feature extractor is used to perform feature fusion on two adjacent frames. The transformer network can extract global information for feature fusion, can more fully extract the effective information of the reference frame and the feature fusion frame, can make the processed image achieve better enhancement and noise reduction effects, and can greatly improve the image quality of the HDR image.

[0014] In combination with the first aspect, in some implementation manners of the first aspect, the feature extractor includes an attention module and a plurality of convolutional layers. The attention module is used to implement an attention mechanism. The feature extractor selectively suppresses the weights of the information of two branches through the attention module, so that more attention is paid to important feature information during the processing, the weight of the important feature information is greater during the processing, and the rejection of redundant information and the enhancement of effective information are realized, which can greatly improve the image quality of the HDR image.

[0015] In combination with the first aspect, in some implementation manners of the first aspect, processing the short-exposure frame, the medium-exposure frame, the long-exposure frame, the first intermediate frame, and the second intermediate frame by using a target network model to obtain a high-dynamic range image, including:

[0016] Starting from the short exposure frame and the first intermediate frame in ascending order of exposure amount, perform feature fusion on two adjacent frames with different exposure amounts through a feature extractor to obtain at least one first-level feature fusion frame; perform feature fusion on two adjacent first-level feature fusion frames to obtain at least one second-level feature fusion frame, and so on. Through progressive feature fusion, obtain the n-level feature fusion frame; starting from the long exposure frame and the second intermediate frame in descending order of exposure amount, perform feature fusion on two adjacent frames with different exposure amounts through a feature extractor to obtain at least one first-level feature fusion frame; perform feature fusion on two adjacent first-level feature fusion frames to obtain at least one second-level feature fusion frame, and so on. Through progressive feature fusion, obtain the m-level feature fusion frame; perform feature fusion on the n-level feature fusion frame and the m-level feature fusion frame with the medium exposure frame respectively to obtain the (n + 1)-level feature fusion frame and the (m + 1)-level feature fusion frame; perform feature fusion on the (n + 1)-level feature fusion frame and the (m + 1)-level feature fusion frame, and pass through the transformer network to obtain the high dynamic range image.

[0017] Combined with the first aspect, in some implementation manners of the first aspect, using the target network model to process the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame and the second intermediate frame to obtain the high dynamic range image, including:

[0018] Perform feature fusion on the short exposure frame and the adjacent first intermediate frame to obtain a feature fusion frame; repeatedly perform feature fusion on the feature fusion frame and the adjacent first intermediate frame until the obtained feature fusion frame is fused with the medium exposure frame to obtain the (n + 1)-level feature fusion frame; perform feature fusion on the long exposure frame and the adjacent second intermediate frame to obtain a feature fusion frame; repeatedly perform feature fusion on the feature fusion frame and the adjacent second intermediate frame until the obtained feature fusion frame is fused with the medium exposure frame to obtain the (m + 1)-level feature fusion frame; perform feature fusion on the (n + 1)-level feature fusion frame and the (m + 1)-level feature fusion frame, and pass through the transformer network to obtain the high dynamic range image.

[0019] By progressively fusing frames with multiple exposure amounts, the two frames are not affected by other frames during fusion, which can make full use of the effective features in each frame and is more conducive to the fusion of frames with multiple exposure amounts.

[0020] In a second aspect, an embodiment of the present application provides an electronic device, including: one or more processors; one or more memories; the memory stores one or more programs, and when the one or more programs are executed by the processor, the electronic device is enabled to execute any one of the methods in the first aspect above.

[0021] In a third aspect, an embodiment of the present application provides a device, which is included in an electronic device and has a function of implementing the behavior of the electronic device in the above aspects and the possible implementation manners of the above aspects. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a display module or unit, a detection module or unit, a processing module or unit, etc.

[0022] In a fourth aspect, a chip system is provided. The chip system is applied to an electronic device and includes one or more processors for calling computer instructions to cause the electronic device to execute any method in the first aspect.

[0023] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium storing instructions, which when running on a computer, cause the computer to execute any method in the first aspect.

[0024] In a sixth aspect, an embodiment of the present application provides a computer program product including instructions, which when running on a computer, cause the computer to execute any method in the first aspect.

[0025] The technical effects obtained in the second, third, fourth, fifth, and sixth aspects are similar to those obtained by the corresponding technical means in the first aspect and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 FIG. shows a schematic diagram of an application scenario provided by an embodiment of the present application;

[0027] Figure 2 FIG. shows a schematic diagram of the principle of synthesizing an HDR image provided by an embodiment of the present application;

[0028] Figure 3 FIG. shows another schematic diagram of the principle of synthesizing an HDR image provided by an embodiment of the present application;

[0029] Figure 4 FIG. shows another schematic diagram of the principle of synthesizing an HDR image provided by an embodiment of the present application;

[0030] Figure 5 FIG. shows a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;

[0031] Figure 6 FIG. shows a schematic diagram of the software structure of an electronic device provided by an embodiment of the present application;

[0032] Figure 7Shows a schematic flowchart of synthesizing an HDR image provided by an embodiment of the present application;

[0033] Figure 8 Shows a schematic flowchart of a high-dynamic range image processing method provided by an embodiment of the present application;

[0034] Figure 9 Shows a schematic diagram of the principle of a high-dynamic range image processing method provided by an embodiment of the present application;

[0035] Figure 10 Shows a schematic flowchart of synthesizing an HDR image through a network provided by an embodiment of the present application;

[0036] Figure 11 Shows another schematic flowchart of synthesizing an HDR image through a network provided by an embodiment of the present application;

[0037] Figure 12 Shows another schematic flowchart of synthesizing an HDR image through a network provided by an embodiment of the present application;

[0038] Figure 13 Shows another schematic flowchart of synthesizing an HDR image through a network provided by an embodiment of the present application;

[0039] Figure 14 Shows a schematic diagram of the structure of a feature extractor provided by an embodiment of the present application;

[0040] Figure 15 Shows a schematic diagram of the structure of a device provided by an embodiment of the present application;

[0041] Figure 16 Shows a schematic diagram of the structure of a chip provided by an embodiment of the present application. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe in detail the embodiments of the present application in conjunction with the accompanying drawings. Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise stated, the meaning of "a plurality" is two or more.

[0043] 1. Exposure Value (EV for short)

[0044] The exposure refers to the intensity of light sensed by the camera and the duration of time. When taking a photo with a camera, due to the limitation of the dynamic range, the exposure amount may be too high or too low, which will directly cause overexposure or underexposure of the subject and the background. If overexposed, the captured image will be too bright to show the details of the bright parts; if underexposed, the captured image will be too dark to show the details of the dark parts.

[0045] Exposure parameters include aperture, exposure time, and sensitivity. Among them, by controlling the shutter speed, the exposure time can be adjusted. The faster the shutter speed, the shorter the exposure time, and the less the exposure amount; on the contrary, the slower the shutter speed, the longer the exposure time, and the more the exposure amount, and the image brightness increases. It should be noted that the length of the exposure time here is relative, and the specific value of the exposure time can be determined according to actual usage requirements.

[0046] Among them, the larger the aperture (the smaller the value), such as F2.8, the more the exposure amount, and the image brightness increases. The smaller the aperture (the larger the value), such as F16, the less the exposure amount, and the image brightness decreases.

[0047] Among them, sensitivity is used to measure the sensitivity of the photosensitive component to light. Specifically, the higher the sensitivity, the stronger the ability to analyze light, the more light is sensed, and the image brightness increases; on the contrary, the lower the sensitivity, the weaker the ability to analyze light, the less light is sensed, and the image brightness decreases. Sensitivity (International Standards Organization, ISO) is usually represented by the ISO sensitivity value. The ISO sensitivity value can be divided into several levels, such as 50, 100, 200, 400, 800, 1600, 3200,.... The higher the ISO sensitivity value, the stronger the photosensitive ability of the photosensitive component.

[0048] In practical applications, the exposure amount during shooting can be controlled by controlling parameters such as aperture, exposure time, and sensitivity, so as to control the shooting effect. By reasonably adjusting the exposure parameters, the generated HDR image can have more vivid colors, higher contrast, and clearer image details.

[0049] 2. Automatic Exposure (AE) algorithm, an exposure method that automatically adjusts the exposure amount according to the intensity of light.

[0050] 3. Long-frame images, medium-frame images, and short-frame images

[0051] Normal exposure image: Also known as the middle frame image, middle exposure frame, N-frame image, it is an image captured by a camera under the condition of an exposure value of 0 EV. That is to say, the exposure value of the normal exposure image is 0 EV. Here, 0 EV is a relative value, not that the exposure amount is 0. Exemplarily, the aperture size of a mobile phone camera is fixed, so the brightness of mobile phone photography is controlled by the exposure time and ISO. Exposure amount = exposure time * ISO. Assuming that the normal exposure image is captured under the condition that the ISO is 200 and the exposure time is 50 milliseconds, the actual exposure amount corresponding to 0 EV is the product of 200 and 50 milliseconds.

[0052] Short frame image: Also known as the short exposure frame, S-frame image, it is an image captured by a camera under the condition of an exposure value less than 0 EV. That is to say, the exposure value of the short frame image is less than 0 EV, such as the exposure value of the short frame image is -2 EV, -4 EV, etc.

[0053] Long frame image: Also known as the long exposure frame, L-frame image, it is an image captured by a camera under the condition of an exposure value greater than 0 EV. That is to say, the exposure value of the long frame image is greater than 0 EV, such as the exposure value of the long frame image is +2 EV, +4 EV, etc.

[0054] 4. RAW Image

[0055] A RAW image is the original image captured by a camera in an electronic device. A RAW image generally does not directly appear on the display screen of the electronic device because the human eye usually cannot directly obtain the scene information from the RAW image. The electronic device needs to perform a series of operations on the RAW image, including white balance correction, color space conversion, tone mapping, etc., to obtain the image that finally appears on the display screen of the electronic device.

[0056] 5. High-Dynamic Range (HDR)

[0057] High-dynamic range imaging is a technology that synthesizes an image with a high dynamic range through a series of consecutive frames of images with different exposure values. Compared with ordinary images, the bright parts of a high-dynamic range image are not overexposed, and the dark details are clearly visible, capable of providing more dynamic range and image details.

[0058] In some shooting scenarios, such as backlit shooting, strong sunlight, shooting indoor and outdoor scenery at the same time, etc., the brightness difference in the captured picture may be too large, and the captured image is prone to the situation of the bright part being too bright or the dark part being too dark, affecting the image quality. To improve the image quality, the electronic device can use the HDR function to capture images and generate HDR images.

[0059] Exemplarily, taking a mobile phone as the electronic device, Figure 1The figure shows a schematic diagram of the interface for enabling the HDR mode in the camera application in an embodiment of the present application.

[0060] When the user lights up the screen of the electronic device and controls the electronic device to be in an unlocked state, the mobile phone can display an interface as shown in Figure 1 (a) below. Among them, this interface can be the desktop of the electronic device, and icons of multiple installed applications are displayed on the desktop of the electronic device, such as file management application icons, email application icons, weather application icons, calculator application icons, clock application icons, recorder application icons, music application icons, settings application icons, address book application icons, phone application icons, message application icons, and camera application icon 10, etc.

[0061] The user can perform a touch operation on the camera application icon 10. This touch operation can be a click operation, a long - press operation, etc., so that the electronic device receives the touch operation of the user on the camera application icon 10. In response to the touch operation of the user on the camera application, the electronic device starts the camera application.

[0062] After the camera application is started, the electronic device can display a shooting interface as shown in Figure 1 (b) below. This shooting interface includes a currently captured preview image, a shooting control 11, and function controls corresponding to various shooting modes. For example, the function controls corresponding to various shooting modes can include a night - mode control, a portrait - mode control, a photo - taking mode control, a video - recording mode control, an aperture - mode control, and more controls 12 for enabling more functions in the camera application, etc. The shooting control 11 is used to trigger the shooting operation of the electronic device.

[0063] When the user needs to use the HDR function of the electronic device to capture an image, the user can click the more controls 12. In response to the user's click operation, the electronic device can display a corresponding menu bar as shown in Figure 1 (c) below. An HDR control 13 is displayed in the menu bar. In response to the trigger operation of the user on the HDR control 13, the mobile phone enables the HDR mode. As shown in Figure 1 (d) below, after the electronic device enables the HDR mode, an HDR logo 14 and a shooting control 11 can be displayed in the camera shooting interface. The HDR logo 14 can prompt the user that the current camera is in the HDR mode. Then, in response to the user's click operation on the shooting control 11 (an example of a shooting instruction), the electronic device can use the HDR function to capture image frames with different exposure amounts, and then the electronic device can fuse the image frames with different exposure amounts to generate an HDR image.

[0064] Currently, a multiple - exposure fusion (MEF) scheme is generally adopted to achieve high - dynamic - range imaging. AsFigure 2 As shown in the figure, the multi-exposure fusion scheme usually means that the camera exposes multiple times in a short period of time, collects three types of RAW images: short-exposure frames, medium-exposure frames, and long-exposure frames, and then generates an HDR image by fusing these three types of RAW images. Among them, different exposure amounts provide picture information of different brightness regions. The short-exposure frame is underexposed and can provide high-brightness information (such as bright area 21), the long-exposure frame is overexposed and can provide dark area information (such as dark area 22), and the medium-exposure frame is used as a reference frame and can provide medium-brightness information.

[0065] However, it is currently difficult to achieve a good HDR imaging effect in scenes with extremely high dynamic range. Since during fusion, the three types of exposure frames need to perform feature point matching between different frames, and these feature points are relied on to compensate and align some problems brought about when taking pictures of different frames (such as the picture offset caused by the user's hand shaking and moving). In order to ensure that the short-exposure frame, medium-exposure frame, and long-exposure frame can be successfully fused, the span between the exposure amounts of the short-exposure frame and the long-exposure frame cannot be set too large. As Figure 3 shown in the figure, when the span between the exposure amounts of the short-exposure frame and the long-exposure frame is too large, such as the exposure amount of the medium-exposure frame is 0EV, the exposure amount of the short-exposure frame is -8EV, and the exposure amount of the long-exposure frame is +8EV, the picture differences between the medium-exposure frame and the short-exposure frame, and the medium-exposure frame and the long-exposure frame are too large, it is difficult to find similar points, and thus the number of feature points for matching is too small, making it difficult to register and align, and further resulting in problems such as registration failure and fusion failure, and it is difficult to obtain a good HDR imaging effect (as Figure 3 shown in the figure, the HDR imaging effect is close to the medium frame). That is to say, in a scene with an extremely high dynamic range (the span between the exposure amounts of the short-exposure frame and the long-exposure frame is large), such as a night scene, when synthesizing an HDR image based on the RAW images of the three exposure types: short-exposure frame, medium-exposure frame, and long-exposure frame, the HDR imaging effect is not ideal. In addition, the repeated medium-exposure frames are used to denoise the image. However, limited by the deep learning network architecture, the repeated medium-frame data is often difficult to be efficiently utilized, and the image quality effect fails to meet the user's requirements.

[0066] In view of this, the present application proposes a high dynamic range image processing method, which can start the camera application by performing a touch operation on the camera application icon. After the camera application is started, when a shooting instruction input by the user is received in the HDR mode, it indicates that the electronic device needs to shoot an HDR image, and the electronic device can determine the exposure amounts corresponding to the short exposure frame, the medium exposure frame, and the long exposure frame respectively. Then, the electronic device can determine the number of frames and the exposure amount of the first intermediate frame, and the number of frames and the exposure amount of the second intermediate frame, where the exposure amount of the first intermediate frame is greater than the exposure amount of the short exposure frame and less than the exposure amount of the medium exposure frame, the exposure amount of the second intermediate frame is greater than the exposure amount of the medium exposure frame and less than the exposure amount of the long exposure frame, the number of frames of the first intermediate frame is greater than or equal to 1, and the number of frames of the second intermediate frame is also greater than or equal to 1. Then, the electronic device can shoot the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame according to different exposure amounts, and input the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame into a deep learning network, and the deep learning network outputs a synthesized HDR image.

[0067] In the present application, the exposure amount span between the short exposure frame and the long exposure frame captured (e.g., -8EV to +8EV) is greater than that obtained by the traditional method (such as Figure 2 ) Figure 4 As shown, since a medium short frame (the first intermediate frame) is inserted between the short exposure frame and the medium exposure frame, and a medium long frame (the second intermediate frame) is inserted between the medium exposure frame and the long exposure frame, when synthesizing the HDR image, the exposure amount spans between the short exposure frame and the medium short frame, the medium short frame and the medium exposure frame, the medium exposure frame and the medium long frame, and the medium long frame and the long exposure frame are all appropriate. The similarity between the short exposure frame and the medium short frame is relatively high, the similarity between the medium short frame and the medium exposure frame is relatively high, the similarity between the medium exposure frame and the medium long frame is relatively high, and the similarity between the medium long frame and the long exposure frame is relatively high, that is, the similarity between each adjacent two frames of images is high, and more feature points can be found for alignment and registration. Therefore, the short exposure frame, the medium short frame, the medium exposure frame, the medium long frame, and the long exposure frame can be successfully fused to obtain a better HDR imaging effect. It should be understood that Figure 4 This is only an exemplary illustration. In fact, multiple medium long frames and multiple medium short frames can be inserted, and the exposure amounts of each medium short frame and each medium long frame can be different. It should be understood that Figures 2 to 4 The images with different exposure amounts in

[0068] Therefore, based on the solution of this application, a relatively large exposure amount span can be set. Then, medium-short frames are inserted between short-exposure frames and medium-exposure frames, and medium-long frames are inserted between medium-exposure frames and long-exposure frames. HDR images are generated by fusing original images with more than three exposure amounts, which can solve the problem that the span is too large to be registered when there are only three exposure amounts of short, medium, and long, achieve a wider dynamic range of exposure amounts, obtain HDR images with a wider dynamic range, and provide better HDR imaging effects for users. In addition, since the image details obtained under different exposure amounts are different, when the total number of frames is the same, inserting medium-short frames and medium-long frames is equivalent to replacing the repeated frames with medium-short frames and medium-long frames, which helps to obtain richer and more diverse picture detail information. Based on information theory, after fusing more and richer information, the imaging quality of the images output by the deep learning network is better.

[0069] This application also provides a new deep learning network structure, which can be called a super-dynamic multi-exposure frame RAW image fusion network. This super-dynamic multi-exposure frame RAW image fusion network can better adapt to the method provided by this application, fuse RAW images with more than three exposures, and achieve better HDR imaging effects.

[0070] The high-dynamic range image processing method provided by the embodiments of this application can be applied to various electronic devices with HDR shooting functions. The electronic device can be, but is not limited to, mobile phones, tablet computers, handheld computers, laptop computers, vehicle-mounted devices, ultra-mobile personal computers (UMPCs), netbooks, cellular phones, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, etc. The embodiments of this application do not make any limitations in this regard.

[0071] Exemplarily, Figure 5 FIG. 11 shows a schematic hardware structure diagram of an electronic device 100 provided by the embodiments of this application.

[0072] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0073] It can be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0074] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0075] Among them, the controller may generate operation control signals according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0076] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0077] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0078] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0079] The charging management module 140 is used to receive a charging input from a charger. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.

[0080] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc.

[0081] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0082] Antenna 1 and Antenna 2 are used for transmitting and receiving electromagnetic wave signals. The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The wireless communication module 160 can provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc.

[0083] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).

[0084] Electronic device 100 implements a display function through a GPU, display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or change display information.

[0085] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0086] Electronic device 100 can implement a shooting function through an ISP, camera 193, video codec, GPU, display screen 194, and an application processor, etc.

[0087] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin tone of the image through algorithms. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.

[0088] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0089] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0090] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0091] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process the input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.

[0092] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0093] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.

[0094] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, music playback, recording, voice calls, video calls, etc. The headphone interface 170D is used to connect a wired headphone. The headphone interface 170D can be a USB interface 130.

[0095] Among them, the sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0096] The keys 190 include a power-on key, a volume key, etc.

[0097] The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback.

[0098] The indicator 192 can be an indicator light, which can be used to indicate the charging state, the change in battery power, and can also be used to indicate messages, missed calls, notifications, etc.

[0099] The SIM card interface 195 is used to connect to the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1.

[0100] The hardware system of the electronic device 100 has been described in detail above. Next, the software system of the image electronic device 100 will be introduced.

[0101] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present invention, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 100. It should be noted that in the embodiments of the present application, the operating system of the electronic device may include, but is not limited to (Symbian), (Andriod), (iOS), (Blackberry), HarmonyOS (HarmonyOS), etc. The operating system is not limited in this application.

[0102] Figure 6 Shows a schematic structural diagram of the software system of the electronic device 100 provided in the embodiments of the present application.

[0103] As Figure 6 shown, the system architecture may include an application layer 610, an application framework layer 620, a hardware abstraction layer 630, a driver layer 640, and a hardware layer 650.

[0104] The application layer 610 may include applications such as a camera, music, gallery, calendar, short message, call, navigation, and video. The camera application provides an HDR mode. Of course, when other applications need to use the shooting function, they can also call the camera application to implement the shooting function.

[0105] The application framework layer 620 can provide application programming interfaces (application programming interfaces, APIs) and programming frameworks to the applications in the application layer; the application framework layer may include some predefined functions.

[0106] For example, the application framework layer 620 may include a camera access interface; the camera access interface may include camera management and a camera device; among them, the camera management can be used to provide an access interface for managing the camera; the camera device can be used to provide an interface for accessing the camera.

[0107] The application framework layer 620 may further include a window manager, a content provider, a view system, a resource manager, a notification manager, etc. (not shown in the figure), and the embodiments of the present application do not impose any restrictions thereon.

[0108] The hardware abstraction layer 630 is used to abstract the hardware. By calling the hardware abstraction layer interface in the hardware abstraction layer, the connection between the application layer and the application framework layer above the hardware abstraction layer and the driver layer and the hardware layer below can be realized, and the shooting data transmission and function control can be achieved.

[0109] For example, the hardware abstraction layer 630 may include a camera hardware abstraction module (camera hardware abstraction layer, camera HAL) and other hardware device abstraction modules; the camera hardware abstraction module may call the algorithms in the camera algorithm library. The camera algorithm library may include software algorithms for image processing.

[0110] Exemplarily, the camera hardware abstraction module may include an AE module, an interpolation module, and an image processing module. The AE module encapsulates an AE algorithm, which can determine different shooting scenarios according to the preview frame, and determine the total number of shooting frames M and the exposure amounts corresponding to the short exposure frame, the medium exposure frame, and the long exposure frame respectively with reference to the shooting scenario. Among them, the shooting scenario may include daytime shooting, night shooting, backlight shooting, night scene shooting, low light shooting, point light source shooting, etc., and the shooting algorithms, the total number of shooting frames, and the EV span corresponding to different shooting scenarios are different. In the embodiments of the present application, the EV span between the short exposure frame and the long exposure frame corresponding to each scenario in the AE module can be modified in advance, and the modified EV span range is greater than the EV span range in the traditional method. The interpolation module is used to calculate the exposure amount and the number of frames of the medium short frame, and the exposure amount and the number of frames of the medium long frame. The image processing module is used to process the original image.

[0111] The driver layer 640 is used to provide drivers for different hardware devices. For example, the driver layer may include a camera driver, a display driver, and an image signal processor driver.

[0112] The hardware layer 650 may include a camera, an image signal processor, a display screen, an image sensor, and other hardware devices.

[0113] Refer to Figure 7The process of denoising RAW images and HDR fusion implemented by Artificial Intelligence (AI) is as follows. After the user triggers the shooting instruction, the camera application sends the shooting instruction to the camera hardware abstraction module through the camera access interface. After receiving the shooting instruction, the AE module determines different shooting scenarios based on the preview frame, determines the total number of shooting frames M with reference to the shooting scenario, and calculates the exposure amounts corresponding to the short-exposure frame, medium-exposure frame, and long-exposure frame required for synthesizing the HDR image using the corresponding shooting algorithm. The AE module sends the total number of shooting frames M, and the exposure amounts corresponding to the short-exposure frame, medium-exposure frame, and long-exposure frame respectively to the frame interpolation module. Further, the frame interpolation module calculates the exposure amounts of the medium-short frame and medium-long frame, as well as the number of frames of various exposure frames, and determines the exposure time and photosensitive value corresponding to the exposure amount. Then the camera hardware abstraction module sends the exposure information (the exposure time, photosensitive value, and number of frames of each frame) to the camera through the underlying driver. The sensor in the camera outputs the raw images with more than three exposure amounts. After that, the image processing module receives the raw image data output by the sensor, and performs image preprocessing on the raw image (including white balance, registration, noise transformation, etc.). After preprocessing, the raw image is fused according to the multi-exposure fusion model (the model structure is the super-dynamic multi-exposure frame RAW image fusion network), and then post-processing is performed (including brightness / sharpness / color enhancement, etc.) to obtain the HDR image. After that, the camera hardware abstraction module returns the HDR image to the upper-layer camera application through the camera access interface.

[0114] The following combines Figure 6 、 Figure 7 and Figure 8 to elaborate in detail on the high dynamic range image processing method provided by this application. Figure 8 It is a flowchart of a high dynamic range image processing method provided by an embodiment of this application. Figure 8 The high dynamic range image method shown can be executed by the electronic device shown in Figure 5 , or by the chip configured in the electronic device shown in Figure 5 ; Figure 8 The high dynamic range image processing method shown includes S801 to S805, and the following will describe S801 to S805 in detail respectively.

[0115] S801, the electronic device starts the camera application and enables the HDR mode.

[0116] When the user wants to start the camera application, the user can click on the camera application icon 10 shown in (a) of Figure 1 , so that the electronic device can receive the touch operation on the camera application icon 10. The electronic device responds to the touch operation, starts the camera application, and after the camera application starts, it displays asFigure 1 In the shooting interface shown in (b) in Figure 1 the interface shown, the user can turn on the HDR mode for shooting.

[0117] It can be understood that there are multiple ways to operate the camera application. In addition to the above-mentioned touch operation on the camera application icon to start the camera application, it can also be triggered by voice or sliding, etc., to start the camera application. For example, when the electronic device is in the locked screen state, the user can use a gesture of swiping right on the display screen of the electronic device to instruct the electronic device to start the camera application. Or, when the electronic device is in the locked screen state and the locked screen interface includes the icon of the camera application, the user clicks the icon of the camera application to instruct the electronic device to start the camera application. Or, when the electronic device is running other applications and the application has the permission to call the camera application; the user can click the corresponding control to instruct the electronic device to start the camera application program. For example, when the electronic device is running an instant messaging application program, the user can use the control of the camera function to instruct the electronic device to start the camera application program, etc. The specific operation method for starting the camera application in the embodiments of the present application is not limited.

[0118] In some embodiments, after the electronic device starts the camera application, the HDR mode of the camera application is already in the on state, and the electronic device can display a shooting interface as shown in (d) in Figure 1 the figure.

[0119] S802. The electronic device detects a shooting instruction, and the shooting instruction requests the electronic device to perform shooting.

[0120] After the HDR mode of the shooting function of the electronic device is turned on, the user can issue a shooting instruction to request the electronic device to shoot in the HDR mode to obtain an HDR image. Exemplarily, the shooting instruction can be a click operation of the user on the shooting control, and the shooting control can be the shooting control 11 as shown in (d) in Figure 1 the figure. It should be understood that the user can also trigger the electronic device to shoot through a voice instruction or a gesture instruction, etc., and the present application does not make any limitation in this regard.

[0121] S803. In response to the shooting instruction, the electronic device determines the exposure amount and the number of frames corresponding to the short exposure frame, the medium exposure frame, the long exposure frame, the medium short frame, and the medium long frame respectively, where the exposure amount of the medium short frame is greater than the exposure amount of the short exposure frame and less than the exposure amount of the medium exposure frame, and the exposure amount of the medium long frame is greater than the exposure amount of the medium exposure frame and less than the exposure amount of the long exposure frame.

[0122] It should be understood that after starting the camera application, during the preview process, the camera can obtain multiple frames of preview images and the original data of each frame of preview image, and the original data includes the exposure time, the brightness of the pixels, the ISO sensitivity value, etc.

[0123] The electronic device can determine the shooting scene based on the preview image, and then call the algorithm corresponding to the shooting scene to determine the exposure amount of the short-exposure frame, the exposure amount of the medium-exposure frame, and the exposure amount of the long-exposure frame. Among them, determining the exposure amount includes determining the exposure time and ISO sensitivity value corresponding to each frame. It should be noted that the EV span is a variable that can be set. Since exposure amount interpolation will be performed later, the EV span of each scene can be set to a relatively large value in advance. Therefore, in the embodiments of the present application, the exposure amount spans of the short-exposure frame and the long-exposure frame determined are larger, so as to ensure that the dynamic range of the shooting scene is covered as completely as possible.

[0124] Specifically, the electronic device can determine which shooting scene the current shooting belongs to according to the ratio of the exposure time and the ISO sensitivity of the current preview image, as well as the preview image. For example, when the ratio R of the exposure time and the ISO is less than a certain preset ratio (such as 0.9), it indicates that the current shooting scene is a daytime shooting. On the contrary, it is a night scene shooting; if it is a night scene shooting, and the proportion value of overexposed pixels in the preview image exceeds a certain threshold T1 (the threshold is small) and is less than a certain threshold T2 (the threshold is large), then the current shooting scene may be a point light source shooting; if it is a night scene shooting, and the proportion value of overexposed pixels in the preview image is less than the above T1, then it can be considered a low-light shooting.

[0125] Furthermore, the electronic device can determine the total number of shooting frames M with reference to the shooting scene. The total number of shooting frames M is the sum of the number of frames corresponding to the short-exposure frame, the medium-exposure frame, the long-exposure frame, the medium-short frame, and the medium-long frame respectively, and M>3. The brighter the current shooting scene, the smaller the set number of shooting frames M. On the contrary, because the noise level in the image is high, M is larger. For example, in the night scene mode, M is an odd number greater than 3. In some other implementation manners, the average brightness value of the over-dark area in the preview image can be calculated, and the total number of shooting frames M can be determined according to the corresponding relationship between the average brightness value of the over-dark area and the number of shooting frames. The present application does not limit the determination method of the total number of shooting frames M.

[0126] After that, the electronic device can first determine the number of frames i of the medium-short frame and the number of frames r of the medium-long frame according to the exposure amount of the short-exposure frame, the exposure amount of the medium-exposure frame, the exposure amount of the long-exposure frame, and the total number of shooting frames M, and then determine the exposure amount of the medium-short frame and the exposure amount of the medium-long frame.

[0127] Among them, the sum of the number of medium - short frames \(i\) and the number of medium - long frames \(r\) is equal to the total number of frames \(M\) captured minus the number of frames corresponding to short - exposure frames, medium - exposure frames, and long - exposure frames. \(i\geq1\), \(r\geq1\). In one implementation, the number of short - exposure frames, medium - exposure frames, and long - exposure frames is all 1, and \(i + r = M - 3\). When \(i + r\) is an even number, \(i\) and \(r\) can be evenly divided, that is, \(i=r=(M - 3) / 2\). When \(i + r\) is an odd number, the first difference between the exposure amount of the medium - exposure frame and the exposure amount of the short - exposure frame, and the second difference between the exposure amount of the long - exposure frame and the exposure amount of the medium - exposure frame can be determined. When the first difference is greater than the second difference, \(i = r + 1\); when the first difference is less than the second difference, \(r = i + 1\); when the first difference is equal to the second difference, \(i = r + 1\) or \(r = i + 1\) can be randomly determined.

[0128] In another implementation, the correspondence between the EV span range and the number of interpolated frames is pre - stored in the electronic device. When the first difference is within the first span range (less than the first threshold), the number of medium - short frames \(i\) is within the first number range, for example, \([1,3]\); when the first difference is within the second span range (greater than or equal to the first threshold and less than or equal to the second threshold), the number of medium - short frames \(i\) is within the second number range, for example, \((3,5]\); when the first difference is within the third span range (greater than the second threshold), the number of medium - short frames \(i\) is within the third number range, for example, \((5,10]\). The same applies to the second difference and the number of medium - long frames \(r\). The electronic device can first determine the first difference between the exposure amount of the medium - exposure frame and the exposure amount of the short - exposure frame, and the second difference between the exposure amount of the long - exposure frame and the exposure amount of the medium - exposure frame. Then, according to the first difference, the second difference, the total number of frames \(M\) captured, and the correspondence between the EV span range and the number of interpolated frames, the number of medium - short frames \(i\) and the number of medium - long frames \(r\) are determined, and the number of frames corresponding to the short - exposure frame, medium - exposure frame, and long - exposure frame are determined respectively. In this implementation, the number of short - exposure frames, medium - exposure frames, and long - exposure frames can be greater than 1.

[0129] It should be understood that the above number ranges are only exemplary descriptions and should not limit this application.

[0130] Furthermore, determining the exposure amount of the medium - short frames and the exposure amount of the medium - long frames may include: determining the exposure amount of each medium - short frame according to the first difference and the number of medium - short frames \(i\); determining the exposure amount of each medium - long frame according to the second difference and the number of medium - long frames \(r\). It should be noted that when \(i>1\), the exposure amounts of each medium - short frame are different, and when \(r>1\), the exposure amounts of each medium - long frame are different.

[0131] Specifically, the exposure amounts corresponding to the medium - short frames and the medium - long frames can be determined by arithmetic interpolation. That is, the exposure amount of the medium - short frame is equal to (E1 - E0) / (i + 1), and the exposure amount of the medium - long frame is equal to (E2 - E0) / (r + 1). Here, E0 is the exposure amount of the medium - exposure frame, E1 is the exposure amount of the short - exposure frame, and E2 is the exposure amount of the long - exposure frame.

[0132] In other implementation manners, the exposure amounts of each medium - short frame and each medium - long frame can also be determined randomly. The random rule is that the exposure amount span between two adjacent frames is greater than the fourth threshold and less than the fifth threshold, so that the content details between two adjacent frames are neither overly repetitive nor misaligned.

[0133] For example, see Figure 9 , the electronic device determines that the exposure amount of the short - exposure frame is - 6EV, the exposure amount of the medium - exposure frame is 0EV, the exposure amount of the long - exposure frame is 4EV, and the total number of frames M to be captured is 5. The number of frames of the short - exposure frame, the medium - exposure frame, and the long - exposure frame is all 1. By arithmetic interpolation, it can be determined that i = r=(M - 3) / 2 = 1, that is, one medium - short frame is inserted between the short - exposure frame and the medium - exposure frame, and one medium - long frame is inserted between the medium - exposure frame and the long - exposure frame. The exposure amount of the medium - short frame is equal to (E1 - E0) / (i + 1)= - 3EV, and the exposure amount of the medium - long frame is equal to (E2 - E0) / (r + 1)=2EV.

[0134] It should be understood that the exposure amounts of the medium - short frame and the medium - long frame are achieved by adjusting the exposure time and / or the sensitivity value. In the embodiments of the present application, whether to adjust the exposure time or the sensitivity value is set in advance according to the scene information. Exemplarily, in some scenes, there are generally multiple repeated medium - exposure frames when using traditional methods. In this scene, it can be set to obtain the exposure amounts of the medium - short frame and the medium - long frame by adjusting the length of the exposure time based on the exposure time and the sensitivity of the medium - exposure frame; in other scenes, there are generally multiple repeated short - exposure frames when using traditional methods. In this scene, it can be set to obtain the exposure amounts of the medium - short frame and the medium - long frame by adjusting the sensitivity value based on the exposure time and the sensitivity of the short - exposure frame. Thus, compared with the traditional method, the method provided by the embodiments of the present application can obtain an HDR image with a wider dynamic range while keeping the total exposure time required for shooting unchanged, avoiding the extension of the shooting time when inserting different exposure frames, and improving the user experience.

[0135] S804, the electronic device outputs the corresponding short - exposure frame, medium - exposure frame, long - exposure frame, medium - short frame, and medium - long frame according to the exposure amounts and the number of frames corresponding to the short - exposure frame, medium - exposure frame, long - exposure frame, medium - short frame, and medium - long frame respectively.

[0136] The electronic device sends the exposure amounts and frame numbers corresponding to the short exposure frame, medium exposure frame, long exposure frame, medium - short frame, and medium - long frame to the image sensor respectively. The image sensor outputs the corresponding short exposure frame, medium exposure frame, long exposure frame, medium - short frame, and medium - long frame according to the exposure amounts and frame numbers respectively. Exemplarily, referring to Figure 9 , the image sensor outputs one short exposure frame, one medium exposure frame, one long exposure frame, one medium - short frame, and one medium - long frame. The corresponding exposure amounts are - 6EV, 0EV, 4EV, - 3EV, and 2EV respectively. An HDR image is generated from the above - mentioned multiple exposure frames.

[0137] S805, the electronic device generates an HDR image based on the short exposure frame, medium exposure frame, long exposure frame, medium - short frame, and medium - long frame.

[0138] An embodiment of the present application provides a super - dynamic multi - exposure - frame raw map fusion network. The electronic device selects the network corresponding to the frame number according to the total number of shooting frames, and inputs the short exposure frame, medium exposure frame, long exposure frame, medium - short frame, and medium - long frame into the network, and fuses the features of different frames to obtain an HDR image.

[0139] As Figures 10 to 12As shown, in one implementation, taking the medium-exposure frame (reference frame) as the dividing line, for the frames with exposure amounts lower than the medium-exposure frame, starting from the lowest exposure amount to the highest, pairwise adjacent frames with different exposure amounts are subjected to feature fusion through a feature extractor to obtain first-level feature fusion frames; pairwise adjacent first-level feature fusion frames are subjected to feature fusion through a feature extractor to obtain second-level feature fusion frames, and so on. Among them, the medium-short frames or h-level feature fusion frames that are not pairwise combined can be subjected to feature fusion with the (n - 1)-level feature fusion frames, where h < n - 1; through gradually progressive feature fusion, n-level feature fusion frames can be obtained. On the other side, for the frames with exposure amounts higher than the medium-exposure frame, starting from the highest exposure amount to the lowest, pairwise adjacent frames with different exposure amounts are subjected to feature fusion through a feature extractor to obtain first-level feature fusion frames; pairwise adjacent first-level feature fusion frames are subjected to feature fusion through a feature extractor to obtain second-level feature fusion frames, and so on. Among them, the medium-long frames or k-level feature fusion frames that are not pairwise combined can be subjected to feature fusion with the (m - 1)-level feature fusion frames, where k < m - 1; through gradually progressive feature fusion, m-level feature fusion frames can be obtained. The n-level feature fusion frames and the m-level feature fusion frames are respectively fused with the medium-exposure frame to obtain (n + 1)-level feature fusion frames and (m + 1)-level feature fusion frames; the (n + 1)-level feature fusion frames and the (m + 1)-level feature fusion frames are subjected to feature fusion and input into a Transformer decoder, and an HDR image is output. Among them, n and m can be equal or not equal. The medium-exposure frame can also pass through a Transformer encoder before fusion. By using the Transformer network to extract global information for feature fusion, the effective information of the reference frame and the feature fusion frames can be more fully extracted, which can make the processed image achieve better enhancement and noise reduction effects, and can greatly improve the image quality of the HDR image.

[0140] Specifically, it is illustrated through several examples. Figure 10 It shows a case where the total number of captured frames is 5. The sum of the number of short-exposure frames and medium-short frames is even, and the sum of the number of medium-long frames and long-exposure frames is even. Sorted by exposure amount from low to high, they are short-exposure frame, medium-short frame 1, medium-exposure frame, medium-long frame 1, and long-exposure frame. When multiple exposure frames are input into the super-dynamic multi-exposure frame raw image fusion network for feature fusion, the short-exposure frame and the medium-short frame are pairwise fused to obtain first-level feature fusion frames (n = 1), the medium-long frame and the long-exposure frame are pairwise fused to obtain first-level feature fusion frames (m = 1), the medium-exposure frame passes through a Transformer encoder and then is fused with the two first-level feature fusion frames respectively to obtain two second-level feature fusion frames. The two second-level feature fusion frames are successively subjected to feature extraction and feature fusion through a feature extractor and a Transformer decoder to obtain an HDR image.

[0141] Figure 11 A case where the total number of captured frames is 8 is shown. The number of medium - short frames is 3, the number of medium - long frames is 2, and there is 1 short - exposure frame, 1 medium - exposure frame, and 1 long - exposure frame each. The sum of the number of short - exposure frames and medium - short frames is an even number, and the sum of the number of medium - long frames and long - exposure frames is an odd number. When multiple exposure frames are input into the super - dynamic multi - exposure frame raw image fusion network for feature fusion, starting from the lowest exposure amount, the short - exposure frame and medium - short frame 1 are fused pairwise to obtain a first - level feature fusion frame, the medium - short frame 2 and medium - short frame 3 are fused pairwise to obtain a first - level feature fusion frame, and the two first - level feature fusion frames are fused to obtain a second - level feature fusion frame (n = 2). On the other side, starting from the highest exposure amount, the long - exposure frame and medium - long frame 2 are fused pairwise to obtain a first - level feature fusion frame, and the medium - long frame 1 and the first - level feature fusion frame are fused to obtain a second - level feature fusion frame (m = 2). The medium - exposure frame is fused with the two second - level feature fusion frames respectively after passing through the transformer encoder to obtain two third - level feature fusion frames. The two third - level feature fusion frames are successively passed through the feature extractor and the transformer decoder for feature extraction and feature fusion to obtain the HDR image.

[0142] Figure 12 A case where the total number of captured frames is 9 is shown. The number of medium - short frames is 5, the number of medium - long frames is 1, and there is 1 short - exposure frame, 1 medium - exposure frame, and 1 long - exposure frame each. The sum of the number of short - exposure frames and medium - short frames is an even number, and the sum of the number of medium - long frames and long - exposure frames is an even number. When multiple exposure frames are input into the super - dynamic multi - exposure frame raw image fusion network for feature fusion, on the side where the exposure amount is lower than the medium - exposure frame, starting from the lowest exposure amount, the short - exposure frame and medium - short frame 1 are fused pairwise to obtain a first - level feature fusion frame, the medium - short frame 2 and medium - short frame 3 are fused pairwise to obtain a first - level feature fusion frame, the medium - short frame 4 and medium - short frame 5 are fused pairwise to obtain a first - level feature fusion frame. The first two first - level feature fusion frames are fused to obtain a second - level feature fusion frame, and the latter first - level feature fusion frame and the second - level feature fusion frame are fused to obtain a third - level feature fusion frame (n = 3, h = 1). On the other side where the exposure amount is greater than the medium - exposure frame, starting from the highest exposure amount, the long - exposure frame and medium - long frame 1 are fused pairwise to obtain a first - level feature fusion frame (m = 1). The medium - exposure frame is fused with the third - level feature fusion frame and the first - level feature fusion frame respectively after passing through the transformer encoder to obtain a fourth - level feature fusion frame and a second - level feature fusion frame. Then the fourth - level feature fusion frame and the second - level feature fusion frame are successively passed through the feature extractor and the transformer decoder for feature extraction and feature fusion to obtain the HDR image.

[0143] In another implementation, as Figure 13As shown, taking the medium exposure frame (reference frame) as the dividing line, for the frames with exposure amounts lower than the medium exposure frame, starting from the lowest to the highest according to the exposure amount, the short exposure frame and the adjacent medium short frame (medium short frame 1) are subjected to feature fusion through a feature extractor to obtain a first-level feature fusion frame; the first-level feature fusion frame and the next medium short frame (medium short frame 2) are subjected to feature fusion through a feature extractor to obtain a second-level feature fusion frame; the second-level feature fusion frame and the next medium short frame (medium short frame 3) are subjected to feature fusion through a feature extractor to obtain a third-level feature fusion frame, and so on. By progressively performing feature fusion level by level, an n-level feature fusion frame can be obtained; the medium exposure frame is subjected to feature fusion with the n-level feature fusion frame after passing through a transformer encoder to obtain an (n + 1)-level feature fusion frame. On the other side, for the frames with exposure amounts higher than the medium exposure frame, starting from the highest to the lowest according to the exposure amount, the long exposure frame and the adjacent medium long frame (medium long frame 3) are subjected to feature fusion through a feature extractor to obtain a first-level feature fusion frame; the first-level feature fusion frame and the next medium long frame (medium long frame 2) are subjected to feature fusion through a feature extractor to obtain a second-level feature fusion frame; the second-level feature fusion frame and the next medium long frame (medium long frame 1) are subjected to feature fusion through a feature extractor to obtain a third-level feature fusion frame, and so on. By progressively performing feature fusion level by level, an m-level feature fusion frame can be obtained; the medium exposure frame is fused with the m-level feature fusion frame after passing through a transformer encoder to obtain an (m + 1)-level feature fusion frame. The (n + 1)-level feature fusion frame and the (m + 1)-level feature fusion frame are subjected to feature fusion and input into a transformer decoder to output an HDR image.

[0144] Preferably, the structure of the feature extractor is as Figure 14As shown in the figure, the F1 frame and the F2 frame are input into the feature extractor. Among them, the exposure of F2 is closer to the medium-exposure frame than that of F1. F1 and F2 can be raw image data. For example, F1 is a short-exposure frame and F2 is a medium-short frame 1; F1 and F2 can also be the feature layer results after the previous-level feature extraction. For example, F1 is a first-level feature fusion frame obtained by fusing a short-exposure frame and a medium-short frame 1, and F2 is a first-level feature fusion frame obtained by fusing a medium-short frame 2 and a medium-short frame 3. On the one hand, F1 and F2 each pass through a residual convolution block (Res Block) to extract shallow features. On the other hand, F1 and F2 are input into a convolution block (Covn Block) together to obtain a brightness compensation value (gain map), and the pixels in the F1 frame are adjusted in brightness through the brightness compensation value to make it closer to the brightness of F2, which is closer to the exposure of the reference frame. Among them, the structure of the convolution block in this application is not limited. After that, both the brightness-compensated F1 frame and the F2 frame enter the attention module (attention block), and the effective information in the F1 frame and the F2 frame is extracted through the attention module. The attention module can determine the weights of the effective information and the invalid information in the F1 frame, as well as the weights of the effective information and the invalid information in the F2 frame; the brightness-compensated F1 frame is multiplied by the weight (attention map) output by the attention module, and the shallow features of the F2 frame are also multiplied by the weight output by the attention module. Then, the features of the two branches are fused by a convolutional layer to achieve feature fusion and obtain a feature fusion frame.

[0145] In the embodiment of this application, through brightness compensation, the brightness of the feature fusion frame output by the feature extractor is closer to the medium-exposure frame, which can avoid problems caused by excessive data differences during subsequent fusion. The attention module is used to implement the attention mechanism. The feature extractor selectively suppresses the weights of the information of the two branches through the attention module, so that more attention is paid to the important feature information during the processing process, making the important feature information have a greater weight during the processing process, realizing the rejection of duplicate information and the enhancement of effective information, which can greatly improve the image quality effect of the HDR image. For example, for the short-exposure frame and the medium-short frame 1, if it is necessary to pay more attention to the bright area information of the short-exposure frame and the dark area information of the medium-short frame 1 during processing, the weight of the bright area information of the short-exposure frame can be increased, the weight of the dark area information of the short-exposure frame can be reduced, the weight of the dark area information of the medium-short frame 1 can be increased, and the weight of the bright area information of the medium-short frame 1 can be reduced.

[0146] Optionally, before performing the above method, the super-dynamic multi-exposure frame raw image fusion network provided in the embodiment of this application can be obtained through the following training method, that is, the method provided in this application can also include: combining the training image and the label image to train the initial network model to obtain the super-dynamic multi-exposure frame raw image fusion network.

[0147] The training images are original images with different exposure levels collected in an ultra-high dynamic range scene, or images stored in an electronic device. Of course, they may also be images downloaded from a server or received from other electronic devices. The embodiments of the present application do not impose any restrictions on this.

[0148] The label images are used to indicate images with the same content as the training images and relatively better quality. The label images can carry labels manually or machine-annotated, and the labels are used to indicate images with a wide dynamic range, accurate colors, and no noise.

[0149] The initial network model is the model before training, and the ultra-dynamic multi-exposure frame raw image fusion network is the model generated after training. The ultra-dynamic multi-exposure frame raw image fusion network corresponds to the aforementioned target network model.

[0150] Specifically, input the training images, use the initial network model to extract the feature information of the training images and process them to obtain predicted images. Based on the predicted images and the label images, determine the loss function between the two, and adjust the parameters of the initial network model according to the loss function, and iterate to obtain the target network model.

[0151] The network structure of the above ultra-dynamic multi-exposure frame raw image fusion network is relatively simple, which is more conducive to the fusion of frames with multiple exposure levels, can improve the processing efficiency, and the more frames there are, the more obvious the effect. By gradually progressing to fuse frames with multiple exposure levels, the two frames are not affected by other frames during fusion, and the effective features in each frame can be fully utilized. In addition, the ultra-dynamic multi-exposure frame raw image fusion network incorporates mechanisms such as transformers and attention, which can greatly improve the image quality effect.

[0152] Above in combination with Figures 1 to 14 , the high-dynamic range image processing method provided by the embodiments of the present application has been described. Next, the apparatus for executing the above method provided by the embodiments of the present application will be described. It should be understood that the apparatus in the embodiments of the present application can execute various methods of the foregoing embodiments of the present application, that is, the specific working processes of the following various products can refer to the corresponding processes in the foregoing method embodiments.

[0153] As Figure 15 shown, Figure 15 is a schematic structural diagram of an apparatus provided by an embodiment of the present application. The apparatus may be the electronic device in the embodiments of the present application, or a chip or chip system within the electronic device. As Figure 15 shown, the apparatus 1500 may include: a display unit 1501 and a processing unit 1502. Among them, the display unit 1501 is used to support the apparatus 1500 to execute the above display step; the processing unit 1502 is used to support the apparatus 1500 to execute the above processing step.

[0154] In a possible implementation, the device 1500 further includes a storage unit 1503. The storage unit 1503 and the processing unit 1502 are connected by a line. The storage unit 1503 may include one or more memories, and the memory may be one or more devices or components in a circuit for storing programs or data. The storage unit 1503 may exist independently and be connected to the processing unit 1502 through a communication bus. The storage unit 1503 may also be integrated with the processing unit 1502.

[0155] The storage unit 1503 may store computer-executable instructions of the method in the electronic device, so that the processing unit 1502 executes the method in the above embodiments. The storage unit 1503 may be a register, a cache, a random access memory (RAM), etc., or may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions.

[0156] Figure 16 Schematic diagram of the structure of a chip provided by an embodiment of the present application. As Figure 16 shown, the chip 1600 includes one or more than two (including two) processors 1601, a communication line 1602, and a communication interface 1603. Optionally, the chip 1600 further includes a memory 1604.

[0157] In some embodiments, the memory 1604 stores the following elements: executable modules or data structures, or subsets thereof, or extended sets thereof.

[0158] The method described in the above embodiments of the present application may be applied to or implemented by the processor 1601. The processor 1601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor 1601 or instructions in software form. The above-mentioned processor 1601 may be a general-purpose processor (e.g., a microprocessor or a conventional processor), a digital signal processor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate, transistor logic devices, or discrete hardware components. The processor 1601 may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0159] The steps of the method disclosed in the embodiments of the present application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. Among them, the software module can be located in a mature storage medium in the art such as a random access memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable read-only memory (EEPROM). This storage medium is located in the memory 1604, and the processor 1601 reads the information in the memory 1604 and combines its hardware to complete the steps of the above method.

[0160] Communication can be carried out among the processor 1601, the memory 1604, and the communication interface 1603 through the communication line 1602.

[0161] In the above embodiments, the instructions stored in the memory for the processor to execute can be implemented in the form of a computer program product. Among them, the computer program product can be pre-written in the memory, or downloaded and installed in the memory in the form of software.

[0162] The embodiments of the present application also provide a computer program product, which includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can store, or a data storage device such as a server or a data center that includes one or more available media integrated. For example, the available medium can include a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a digital versatile disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).

[0163] The embodiments of the present application provide an electronic device, which includes a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the computer program to execute the above method.

[0164] An embodiment of the present application provides a chip. The chip includes a processor, which is used to call a computer program in a memory to execute the technical solutions in the above embodiments. Its implementation principle and technical effects are similar to those of the above related embodiments, and will not be elaborated here.

[0165] In the embodiments provided in the present application, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form. In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0166] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above various method embodiments can be implemented. The computer-readable storage medium stores a computer program or instruction. When the computer program or instruction is executed by a processor, the above method is implemented. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. The computer-readable medium can include a computer storage medium and a communication medium, and can also include any medium that can transmit a computer program from one place to another place. The storage medium can be any target medium accessible by a computer.

[0167] As a possible design, a computer-readable medium may include a compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; the computer-readable medium may include magnetic disk storage or other magnetic disk storage devices. Moreover, any connecting wire may also be appropriately referred to as a computer-readable medium. For example, if software is transmitted using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave from a website, server, or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, magnetic disks and optical discs include optical discs (CDs), laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs, where magnetic disks typically reproduce data magnetically, while optical discs use lasers to optically reproduce data. Combinations of the above should also be included within the scope of computer-readable media.

[0168] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processing unit of the computer or other programmable data processing device generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0169] In the above description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0170] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0171] It should also be understood that the "plurality" mentioned in the specification of this application and the appended claims means two or more. In the description of this application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; the "and / or" herein is merely a description of the relationship between associated objects, referring to any combination and all possible combinations of one or more of the associated listed items, and including these combinations. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, these three situations.

[0172] As used in the specification of this application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0173] In addition, for the convenience of clearly describing the technical solution of this application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, cannot be understood as indicating or implying relative importance, and words such as "first" and "second" do not necessarily mean different.

[0174] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0175] The above-described embodiments are only used to illustrate the technical solution of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A high dynamic range image processing method, characterized in that The method includes: Detecting a shooting instruction that requests the electronic device to take a picture, where the high dynamic range imaging mode of the shooting function of the electronic device is turned on; In response to the shooting instruction, determining the exposure amounts and the number of frames corresponding to the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame respectively, where the exposure amount of the first intermediate frame is greater than the exposure amount of the short exposure frame and less than the exposure amount of the medium exposure frame, and the exposure amount of the second intermediate frame is greater than the exposure amount of the medium exposure frame and less than the exposure amount of the long exposure frame; Outputting the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame respectively according to the exposure amounts and the number of frames corresponding to the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame; Generating a high dynamic range image according to the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame.

2. The method according to claim 1, characterized in that, Determining the number of frames of the first intermediate frame and the second intermediate frame includes: Determining the total number of frames M for shooting; Determining the number of frames of the first intermediate frame and the second intermediate frame according to the total number of frames M for shooting and the number of frames corresponding to the short exposure frame, the medium exposure frame, and the long exposure frame, and the number of frames corresponding to the short exposure frame, the medium exposure frame, and the long exposure frame is 1 each.

3. The method according to claim 2, wherein Determining the exposure amounts of the first intermediate frame and the second intermediate frame includes: Determining a first difference between the exposure amount of the medium exposure frame and the exposure amount of the short exposure frame, and a second difference between the exposure amount of the long exposure frame and the exposure amount of the medium exposure frame; Determining the exposure amount of each first intermediate frame according to the first difference and the number of frames of the first intermediate frame; Determining the exposure amount of each second intermediate frame according to the second difference and the number of frames of the second intermediate frame.

4. The method according to claim 2, wherein When the number of frames of the first intermediate frame is greater than 1, the exposure amounts of the first intermediate frames are all different; When the number of frames of the second intermediate frame is greater than 1, the exposure amounts of the second intermediate frames are all different.

5. The method according to any one of claims 1 to 4, characterized in that, The exposure amount is determined by the exposure time and the sensitivity. The exposure times corresponding to the medium exposure frame, the first intermediate frame, and the second intermediate frame are different, and the sensitivities corresponding to the medium exposure frame, the first intermediate frame, and the second intermediate frame are the same.

6. The method according to any one of claims 1 to 5, characterized in that The generating a high dynamic range image according to the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame includes: Processing the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame by using a target network model to obtain the high dynamic range image, where the target network model is used to perform feature fusion on two adjacent exposure frames and continue to perform feature fusion on the fused features step by step.

7. The method according to claim 6, characterized in that, The target network model is trained based on a transformer network and a feature extractor, and the feature extractor is used to perform feature fusion on two adjacent frames.

8. The method according to claim 7, wherein The feature extractor includes an attention module and a plurality of convolutional layers.

9. The method according to claim 7, wherein Processing the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame using the target network model to obtain the high-dynamic range image includes: Starting from the short exposure frame and the first intermediate frame in ascending order of exposure amount, performing feature fusion on two adjacent frames with different exposure amounts through a feature extractor to obtain at least one first-level feature fusion frame; performing feature fusion on two adjacent first-level feature fusion frames to obtain at least one second-level feature fusion frame,..., and through progressive feature fusion, obtaining an n-level feature fusion frame; Starting from the long exposure frame and the second intermediate frame in descending order of exposure amount, performing feature fusion on two adjacent frames with different exposure amounts through a feature extractor to obtain at least one first-level feature fusion frame; performing feature fusion on two adjacent first-level feature fusion frames to obtain at least one second-level feature fusion frame,..., and through progressive feature fusion, obtaining an m-level feature fusion frame; Performing feature fusion on the n-level feature fusion frame and the m-level feature fusion frame with the medium exposure frame respectively to obtain an n + 1-level feature fusion frame and an m + 1-level feature fusion frame; Performing feature fusion on the n + 1-level feature fusion frame and the m + 1-level feature fusion frame, and passing through the transformer network to obtain the high-dynamic range image.

10. The method according to claim 7, characterized in that, Processing the short exposure frame, the medium exposure frame, the long exposure frame, the first intermediate frame, and the second intermediate frame using the target network model to obtain the high-dynamic range image includes: Performing feature fusion on the short exposure frame and the adjacent first intermediate frame to obtain a feature fusion frame; repeatedly performing feature fusion on the feature fusion frame and the adjacent first intermediate frame until the obtained feature fusion frame is fused with the medium exposure frame to obtain an n + 1-level feature fusion frame; Performing feature fusion on the long exposure frame and the adjacent second intermediate frame to obtain a feature fusion frame; repeatedly performing feature fusion on the feature fusion frame and the adjacent second intermediate frame until the obtained feature fusion frame is fused with the medium exposure frame to obtain an m + 1-level feature fusion frame; Performing feature fusion on the n + 1-level feature fusion frame and the m + 1-level feature fusion frame, and passing through the transformer network to obtain the high-dynamic range image.

11. An electronic device, characterized in that, Including: One or more processors; one or more memories; The memory stores one or more programs, and when the one or more programs are executed by the processor, the electronic device is caused to execute the high-dynamic range image processing method according to any one of claims 1 to 10.

12. A chip system, characterized in that, The chip system is applied to an electronic device. The chip system includes one or more processors, and the processors are used to call computer instructions to cause the electronic device to execute the high-dynamic range image processing method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when executed on a computer, cause the computer to perform the high-dynamic-range image processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Vehicle Camera System

    CN108965726A

  • Method and device for identifying object from video

    CN116030387A

  • High dynamic range multi-exposure image fusion model and method based on attention mechanism

    CN116152128A

  • Method for generating HDR image based on LDR image of Transform

    CN116245968A

  • Image pickup device

    JP2014171146A

Cited By

  • Exposure parameter adjustment method and device, computer equipment and storage medium

    CN120602789A