Image generation method and electronic equipment
By fusing long and short frames in Stagger HDR mode, the problem of blurring in static areas when shooting moving targets is solved, improving image quality and achieving high signal-to-noise ratio and rich detail in image presentation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2026-03-10
AI Technical Summary
When the target object is in motion, if the electronic device uses the Stagger HDR shooting mode, the captured image may have problems such as blurred static areas, which will affect the user's visual experience.
By acquiring frames of different frame types (long and short frames) from the preview frame sequence in Stagger HDR mode and performing inter-frame fusion, the high signal-to-noise ratio of long frames is used to compensate for the insufficient signal-to-noise ratio in static areas, and the detailed features of moving areas in short frames are combined to improve the quality of the image fusion frames.
It improves the signal-to-noise ratio of still areas and the quality of moving areas in image fusion frames, solves image quality anomalies, and improves the overall quality of captured images.
Smart Images

Figure CN121645005A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to an image generation method and an electronic device. Background Technology
[0002] High Dynamic Range (HDR) technology is a technique that uses an image sensor to capture frames with different exposure times and then fuses these frames to create an HDR image. These frames are specifically categorized as long frames and short frames, with long frames having a longer exposure time than short frames. Long frames, due to their longer exposure time, capture more light and have a wider brightness range, while short frames, due to their shorter exposure time, capture fleeting moments and rapidly changing image details. Therefore, HDR images generated using HDR technology more realistically reflect changes in ambient light, exhibiting richer color gradations and a wider brightness range, providing users with a more lifelike and detailed visual experience.
[0003] Stagger High Dynamic Range (Stagger HDR) is an HDR technique that captures both long and short frame sequences simultaneously in a single shot. It then fuses these long frames into a single image, creating an HDR image. By reducing the time intervals between frames, Stagger HDR reduces the probability of ghosting in HDR images, further improving image quality. Currently, some electronic devices have integrated HDR functionality based on Stagger HDR technology to enhance captured image quality.
[0004] When electronic devices are shooting, they often need to capture the fleeting image of a subject, such as news events, social activities, or exciting moments in sports competitions. However, when the subject is in motion and the device is in Stagger HDR shooting mode, the captured image may exhibit problems such as blurred still areas, affecting the user's visual experience. Summary of the Invention
[0005] This application provides an image generation method and an electronic device that can improve the image quality of images captured by the electronic device when the target object is in motion.
[0006] To achieve the above objectives, this application provides the following technical solution:
[0007] In a first aspect, embodiments of this application provide an image processing method applied to an electronic device, wherein the electronic device enables an interleaved high dynamic range (Stagger HDR) mode, and in response to receiving an operation that triggers the electronic device to capture an image, the method acquires a first frame and a second frame from a preview frame sequence in the Stagger HDR mode, the preview frame sequence including multiple latest preview frames cached during the preview process; the frame type of the second frame is different from that of the first frame, and the frame type includes: long frames in the Stagger HDR mode and short frames in the Stagger HDR mode; the first frame, the captured long frame sequence, and the short frame sequence are fused to obtain a first fused frame; the long frame sequence includes multiple captured long frames, and the short frame sequence includes multiple captured short frames; the second frame and the first fused frame are fused to obtain a second fused frame; and the captured image is obtained based on the second fused frame. In other words, when electronic devices fuse short and long frame sequences acquired during shooting, the long and short frames in Stagger HDR mode are used as benchmarks. The signal-to-noise ratio (SNR) of the long frames in stagger mode is significantly higher than that of the static areas in the X-frame, compensating for the insufficient SNR in the static areas of the fused frame. This improves the SNR of the static areas in the fused image, thus resolving image quality anomalies such as various textures caused by insufficient SNR in the static areas of the fused image, and improving the overall image quality. Furthermore, the rich motion region details in the short frames of stagger mode can be utilized to improve the quality of the motion regions in the fused image.
[0008] In one specific implementation, a reference preview frame is determined from the preview frame sequence. If the proportion of pixels in the moving region to the total number of pixels in the preview frame is greater than a preset proportion threshold, the long frame of the Stagger HDR mode is determined as the first frame; otherwise, the short frame of the Stagger HDR mode is determined as the first frame. The motion amount of the pixel data in the moving region is greater than or equal to a first motion amount threshold. Therefore, in shooting scenarios where the moving region accounts for a relatively high proportion, using the short frame as the first frame can improve the image quality of a single fusion.
[0009] In another specific implementation, based on a first size specification, the reference preview frame is divided into multiple image blocks, each with the same size specification. For each image block, the following steps are performed: determining the motion of at least one feature point in the image block within the preview frame sequence; determining the motion of the image block data based on the motion of the feature point data; and determining the motion region of the reference preview frame based on the motion of multiple image block data. Obtaining low-resolution motion data based on image blocks can reduce computational complexity, improve image processing speed, and enhance robustness.
[0010] In another specific implementation, feature points among the feature points of adjacent image blocks that satisfy a preset distance condition from the pixel to be reconstructed in the image block are selected as candidate feature points. The pixel to be reconstructed in each image block has the same relative position, and the preset distance condition is: closest to the pixel to be reconstructed in the image block, farthest from the pixel to be reconstructed in the image block, or a predetermined distance. The motion amount of the pixel to be reconstructed in the image block is determined based on the motion amount of the candidate feature point data in adjacent image blocks. The motion amount of the image block data is then determined based on the motion amount of the pixel to be reconstructed in the image block. Thus, by using the pixel to be reconstructed, the motion amount of pixels within the frame is evenly distributed, thereby enabling the motion amount of the image block data to accurately describe the overall motion trend of the moving object in the reference preview frame.
[0011] In another specific implementation, intra-frame weights corresponding to candidate feature points in adjacent image blocks are determined. These intra-frame weights are negatively correlated with the distance from the candidate feature points in adjacent image blocks to the pixel to be reconstructed. The motion magnitude of the pixel to be reconstructed is determined based on the product of the intra-frame weights and the motion magnitudes of the candidate feature point data in adjacent image blocks. Since the motion intensity values of feature points closer to the pixel to be reconstructed have a more significant impact on the pixel being reconstructed, increasing the intra-frame weights can improve the accuracy of the motion amplitude values at the reconstructed location.
[0012] In another specific implementation, the motion magnitude and confidence level of feature point data in adjacent preview frames of the reference preview frame are obtained. Based on the confidence level of the feature point data in adjacent preview frames, the inter-frame weights of the feature point data in adjacent preview frames are determined, and the inter-frame weights are positively correlated with the motion magnitude of the feature point data in adjacent preview frames. Based on the product of the inter-frame weights and the motion magnitudes of the feature points in adjacent preview frames, the motion magnitude of the feature points in the reference preview frame is determined. This fully considers the inter-frame weights of the motion data in adjacent frames, improving the accuracy of the obtained motion data of the reference preview frame.
[0013] In another specific implementation, the long frame sequence is globally aligned with the denoised short frame sequence; using the first frame as a reference, the globally aligned long frame sequence and the short frame sequence are fused to obtain the first fused frame. Through denoising and full alignment, the signal-to-noise ratio is further improved.
[0014] In another specific implementation, the first frame, a long frame sequence, and a short frame sequence are input into a fusion network to obtain a first fused frame. The loss function used to train the fusion network represents the noise level of the first fused frame in the static region. Using the trained fusion network for fusion helps improve the signal-to-noise ratio in the static region.
[0015] In this system, the loss function value corresponds one-to-one with the target difference data, which is the difference between the first data and the second data. The first data is the product of multiple pixel data and the fusion mask value of the same pixel in the fusion frame output by the fusion network during training. The second data is the product of multiple pixel data and the fusion mask value of the same pixel in the standard frame of the fusion frame during training. The fusion mask value is the first mask value when the motion amount of the pixel data is less than the first motion amount threshold. The fusion mask value is the second mask value when the motion amount of the pixel data is greater than or equal to the first motion amount threshold.
[0016] In another specific implementation, if the second frame is a long frame in Stagger HDR mode, the motion of multiple pixel data points in the second frame relative to the same pixel data points in the first fused frame is obtained. Pixels with motion values less than a second motion value threshold are designated as pixels in the stationary region. If the signal-to-noise ratio (SNR) metric of the pixels in the stationary region of the second frame is greater than the SNR metric of the same pixels in the first fused frame, the second fused frame is obtained. The data of the stationary region in the second fused frame is the sum of the data of the stationary region in the second frame and the data of the stationary region in the first fused frame. The SNR metric is used to measure the magnitude of the SNR. Thus, the SNR of the stationary region in the long frame is used to compensate for the SNR of the stationary region in the first fused frame, thereby improving the SNR of the stationary region in the second fused frame.
[0017] In another specific implementation, the data of the still region in the second fused frame is the sum of the first product data and the second product data. The first product data is the product of the data of the still region in the second frame and the first fusion weight, and the second product data is the product of the data of the still region in the first fused frame and the second fusion weight. The sum of the first fusion weight and the second fusion weight is 1. The difference between the signal-to-noise ratio (SNR) of the pixels in the still region of the second frame and the SNR of the same pixels in the first fused frame is the first difference value. The first fusion weight and the first difference value are positively correlated. Different fusion weights are assigned to the still regions of different frames to further improve the SNR of the still regions in the fused frames.
[0018] In another specific implementation, the second frame includes at least the first pixel. The method further includes: obtaining a first window from the second frame, the first window being an image block including the current pixel and having a size of a second dimension; determining a first structural similarity index (SSIM) between the first window and its adjacent windows, the first window having the same size as its adjacent windows; obtaining a signal-to-noise ratio (SNR) metric for the current pixel in the second frame; the SNR metric for the current pixel in the second frame is positively correlated with the first SSIM. That is, the electronic device utilizes the principle that a lower SNR indicates higher noise and more significant damage to the image structure, thus accurately obtaining the SNR of the first fused frame.
[0019] In another specific implementation, for each pixel in the first fused frame, a second window is obtained from the first fused frame. The second window is an image block including the current pixel and with a size of a third dimension. A second SSIM (Signal-to-Noise Ratio) is determined between the second window and its adjacent windows, ensuring that the second window and its adjacent windows have the same size. The signal-to-noise ratio (SNR) metric of the current pixel in the first fused frame is obtained. The SNR metric of the current pixel in the first fused frame is positively correlated with the second SSIM. Thus, the SNR of the first fused frame is accurately obtained. Furthermore, the first fused frame and the second frame use the same method to obtain their SNR, and their metrics are on the same dimension, which is beneficial for controlling the SNR of static areas during fusion.
[0020] In another specific implementation, the fusion mask value of pixels corresponding to motion values less than a second motion value threshold is set as the first mask value; the fusion mask value of pixels corresponding to motion values greater than or equal to the second motion value threshold is set as the second mask value; pixels corresponding to motion amplitude values less than a second motion amplitude threshold are designated as static region pixels, including: pixels with a fusion mask value of the first mask value are designated as static region pixels. The mask values accurately distinguish between static and moving regions of the image frame.
[0021] In another specific implementation, if the frame type of the second frame is a short frame in Stagger HDR mode, the motion of multiple pixel data in the second frame relative to the same pixel data in the first fused frame is obtained; the pixels corresponding to the motion of the multiple motion values that are greater than or equal to the second motion value threshold are taken as the pixels of the motion region; the second fused frame is obtained, and the data of the motion region of the second fused frame is the sum of the data of the motion region of the second frame and the data of the motion region of the first fused frame.
[0022] In another specific implementation, short frames from multiple Stagger HDR modes in the preview frame sequence are fused to obtain a short-frame fused frame; the short-frame sequence includes the short-frame fused frame. By fusing multiple short frames, the signal-to-noise ratio in static areas is further improved.
[0023] In another specific implementation, the long frame gain and the exposure reduction ratio are determined; multiple short frames of Stagger HDR mode in the preview frame sequence are fused to obtain a short frame fused frame, including: if the long frame gain is less than or equal to a first gain threshold and the exposure reduction ratio is less than a first ratio threshold, n1 short frames of Stagger HDR mode in the preview frame sequence are fused to obtain a short frame fused frame, where n1 is an integer greater than or equal to 1.
[0024] In another specific implementation, if the long frame gain is less than or equal to the second gain threshold and greater than the first gain threshold; and the exposure reduction ratio is greater than or equal to the first ratio threshold and less than the second ratio threshold, then the short frames of the Stagger HDR mode in the n2 frames of the preview frame sequence are merged to obtain the short frame fused frame; n2 is an integer greater than n1.
[0025] In another specific implementation, if the long frame gain is less than or equal to the third gain threshold and greater than the second gain threshold; and the exposure reduction ratio is greater than or equal to the second ratio threshold and less than the third ratio threshold, then the short frames of the Stagger HDR mode in the n3 frames of the preview frame sequence are merged to obtain the short frame fused frame, where n3 is an integer greater than n2.
[0026] In another specific implementation, if the maximum gain of the short frame in the Stagger HDR mode of the preview frame sequence is less than the first short frame gain threshold, a first noise reduction algorithm is used for fusion noise reduction processing; if the maximum gain of the short frame in the Stagger HDR mode of the preview frame sequence is less than the second short frame gain threshold but greater than or equal to the first short frame gain threshold, a second noise reduction algorithm is used for fusion noise reduction processing; if the maximum gain of the short frame in the Stagger HDR mode of the preview frame sequence is less than the third short frame gain threshold but greater than or equal to the second short frame gain threshold, a third noise reduction algorithm is used for fusion noise reduction processing; wherein, the third short frame gain threshold is greater than the second short frame gain threshold, which is greater than the first short frame gain threshold, the fusion noise reduction processing capability of the third noise reduction algorithm is better than that of the second noise reduction algorithm, and the fusion noise reduction processing capability of the second noise reduction algorithm is better than that of the first noise reduction algorithm. This application embodiment, for the image signal-to-noise ratio levels of S frames under different gain conditions, adopts a hierarchical, progressively increasing complexity multi-frame noise reduction processing algorithm, thereby achieving the best noise reduction effect while maintaining the processing effect and obtaining a higher signal-to-noise ratio in still areas.
[0027] In another specific implementation, the third denoising algorithm includes a first sub-denoising algorithm for stationary regions and a second sub-denoising algorithm for moving regions. The first sub-denoising algorithm is more effective at denoising stationary regions than the second sub-denoising algorithm, and the second multi-frame denoising algorithm is more effective at denoising moving regions than the first multi-frame denoising algorithm. Further improving the signal-to-noise ratio in stationary regions and employing different denoising algorithms for different regions helps improve processing efficiency.
[0028] Secondly, embodiments of this application provide an electronic device, including a processor and a memory; wherein,
[0029] Memory is used to store programs;
[0030] The processor shown is used to execute a program stored in memory, and when the program stored in memory is executed, it performs any of the methods described in the first aspect.
[0031] Thirdly, this application provides a computer storage medium for storing a computer program, which, when executed, implements the image generation method provided in any one of the first to second aspects of this application.
[0032] Fourthly, this application provides a computer program product containing instructions that, when run on at least one computing device, causes the at least one computing device to implement the image generation method provided in any one of the first to second aspects of this application. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of multiple frames captured in a shooting mode, provided as an embodiment of this application.
[0034] Figure 2 This is a schematic diagram of multiple frames captured under another shooting mode provided in an embodiment of this application;
[0035] Figure 3 This is a schematic diagram of a method for capturing moving objects.
[0036] Figure 4 A flowchart of an image generation method provided in an embodiment of this application;
[0037] Figure 5 A flowchart illustrating a method for implementing motion estimation of a preview path, as provided in an embodiment of this application;
[0038] Figure 6 This is a schematic diagram illustrating a method for obtaining low-resolution motion data based on a preview frame, as provided in an embodiment of this application.
[0039] Figure 7 A schematic diagram of intra-frame reconstructed motion provided in an embodiment of this application;
[0040] Figure 8 This is a schematic diagram illustrating a method for reconstructing inter-frame motion data according to an embodiment of this application;
[0041] Figure 9 A schematic diagram of an image frame sequence provided in an embodiment of this application;
[0042] Figure 10 This is a schematic diagram illustrating an embodiment of obtaining an image frame sequence.
[0043] Figure 11 This is a schematic diagram illustrating another method of acquiring an image sequence according to an embodiment of this application;
[0044] Figure 12 A flowchart of a one-time fusion processing method provided in an embodiment of this application;
[0045] Figure 13 A flowchart of a secondary fusion processing method provided in an embodiment of this application;
[0046] Figure 14 A flowchart illustrating another secondary fusion processing method provided in this application embodiment;
[0047] Figure 15 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0048] Figure 16 A schematic diagram of the software structure of an electronic device provided for the implementation of this application. Detailed Implementation
[0049] It should be noted that the embodiments described in this application are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0050] To make the following embodiments clear, the technical terms involved in this application will be introduced first.
[0051] (1) Shooting Mode
[0052] In this embodiment, the shooting mode is also called the snapshot mode, which refers to a camera shooting mode that captures a fleeting image of a target object. When an electronic device runs a camera application and the camera application is in snapshot mode, the shooting button (e.g., shutter) of the electronic device is triggered, and the electronic device can capture an image of the target object. In particular, when the target object is in motion, the electronic device can capture an image of the target object in motion. Hereinafter, a target object in motion will be referred to as a moving object.
[0053] (2) High Dynamic Range (HDR):
[0054] HDR technology refers to the technique of fusing multiple image frames (referred to as multi-frames) captured by an image sensor within different acquisition periods into a single frame. These multiple frames have different exposure times. Compared to non-HDR technology, images captured by electronic devices using HDR technology can capture more detailed information. Among HDR technologies, StaggerHDR is the most widely used.
[0055] Stagger HDR achieves this by increasing the frame rate of an interleaved sensor, enabling the capture of multiple frames with different exposure times within a single acquisition cycle. The interleaved sensor is an image sensor. In capture mode, Stagger HDR can acquire multiple frames simultaneously in a single shot. Compared to other HDR techniques that complete one frame before starting the next, this helps reduce the time interval between frames, thus reducing the probability of ghosting in the captured image.
[0056] Exemplary illustration: Appendix Figure 1 This illustration shows a multi-frame capture method in a shooting mode, as provided in an embodiment of this application. When the user presses the shutter button, the electronic device can capture multiple frames with different exposure times in a single shot. For example, the multiple frames can be 4 frames, 5 frames, or more; this embodiment of the application does not specifically limit the number of frames. Figure 1 This demonstrates how an electronic device can capture six frames with different exposure times in a single shot. The first column shows a sequence of multiple N-frames, and the second column shows a sequence of multiple X-frames. N-frames and X-frames in the same row represent frames captured simultaneously, with the exposure time of the N-frame being longer than that of the X-frame.
[0057] Zero Shutter Lag Buffer (ZSL Buffer) is a technology that caches a number of spare frames with different exposure times during the preview process and provides the spare frame closest to the moment the shutter is pressed (referred to as the preferred spare frame) when the shutter button is pressed. When the electronic device triggers the shutter button, it continues to shoot multiple frames with different exposure times. By fusing the preferred spare frame and the continued shooting frames, the captured image can be obtained. Based on the ZSL Buffer, the delay between pressing the shutter button and the actual image is reduced, improving the shooting experience and the ability to capture fleeting moments, achieving an instant shooting effect.
[0058] Exemplary illustration: Appendix Figure 2 This is a schematic diagram illustrating multiple frames captured under another shooting mode provided in an embodiment of this application. The electronic device buffers multiple sets of N-frames and S-frames as spare frames during the preview process. Since the preview process simulates the actual shooting process, the N-frames and S-frames in the preview process are also image frames with different exposure times, where the exposure time of the N-frame is longer than that of the S-frame. (See attached diagram) Figure 2 Two sets of backup frames are shown, where the exposure time of frame N is longer than that of frame S. The two frames (N and S) closest to the press time are used as backup frames (see attached). Figure 2 The gray area indicates the preferred spare frame provided by the electronic device. When the user presses the shutter, the electronic device will continue to capture multiple frames with different exposure times, as shown in the attached image. Figure 1As shown, at this time, the electronic device captures multiple frames in shooting mode: 2 frames (N frames) and 1 frame (S frames) before the press time, and 6 frames after the press time.
[0059] (3) Exercise volume
[0060] Motion quantity refers to the relative motion information between two adjacent image frames, including the magnitude of the displacement of pixel data from one image frame to the corresponding pixel data in another image frame. It should be noted that in the embodiments of this application, information such as motion speed can also be used instead of motion quantity, and the processing method is the same as that for motion quantity, which will not be discussed further here.
[0061] In this embodiment, the motion data includes low-resolution motion data and high-resolution motion data. The accuracy of low-resolution motion data is lower than that of high-resolution motion data.
[0062] Low-resolution motion data refers to the motion data obtained by processing preview frames cached during the preview process. The image resolution of the preview frames is lower than that of the image frames acquired during shooting. Due to the low resolution of the object being processed, the accuracy of the extracted motion data is low. However, because the amount of preview frame data cached during the preview process is relatively small, the computational complexity of the extracted motion data is relatively low, which helps to achieve real-time processing.
[0063] High-resolution motion measurement refers to the motion measured by processing image frames acquired during the shooting process. Because the resolution of image frames acquired during the capture process is higher than that of the preview frames, the accuracy of the extracted motion measurement is higher than that of low-resolution motion measurement.
[0064] (4) Long frames and short frames.
[0065] When an electronic device is in shooting mode, and the target object is a moving object, the acquired image includes a static region and a moving region. The static region refers to an area in the image where there is no significant motion, or an image region where the data of pixels or image blocks remains relatively stable within a preset time period. The moving region refers to an area in the image where there is significant motion, or an image region where the data of pixels or image blocks changes significantly within a preset time period. In this embodiment, each frame of the multiple frames acquired when the electronic device captures a moving object includes both a static region and a moving region.
[0066] Long frames and short frames: Long frames, as mentioned above (N frames), refer to frames with an exposure time exceeding a preset time threshold. Short frames, as mentioned above (S frames and X frames), refer to frames with an exposure time not exceeding a preset time threshold. Because long frames have longer exposure times, they capture more light and reduce random noise; therefore, the signal-to-noise ratio (SNR) in still areas of long frames is higher than that of short frames. Still areas in long frames are also clearer. Here, SNR refers to the ratio of signal power to noise power.
[0067] However, long exposure times can cause motion blur in moving objects, resulting in low image sharpness in moving areas of long frames. Short frame exposure times avoid overexposure, thus capturing the instantaneous state of moving objects and rapidly changing details.
[0068] In summary, the image quality of moving areas in short frames is higher than that in long frames. However, the image quality of still areas in short frames is lower than that in long frames. Therefore, in the Stagger HDR-based capture mode, after the electronic device acquires both long and short frames simultaneously, it can fuse the still areas of the long frames with the moving areas of the short frames to obtain images with high dynamic range and rich detail. This captured image clearly presents both the rich details of still areas and the details of moving areas.
[0069] Currently, if the moving object is moving rapidly, electronic devices can improve the sharpness of the moving area by further reducing the exposure time of short frames during the capture process. In other words, the exposure time of the short frames acquired by the electronic device during the capture process is lower than the exposure time of the short frames during the preview process, i.e., [further details needed]. Figure 2 The exposure time of the S-frame shown is greater than that of the X-frame. Furthermore, to distinguish between short frames during the preview process and short frames during the capture process, in this embodiment, the short frames during the capture process are referred to as down-exposure frames (or X-frames), and the short frames during the preview process are referred to as short frames (or S-frames). Therefore, the X-frame can also be called a down-exposure frame. Because the X-frame has a lower exposure time, it captures less light compared to the S-frame, thus further reducing the signal-to-noise ratio of the X-frame.
[0070] However, a decrease in the signal-to-noise ratio (SNR) of the X-frame will cause a decrease in the SNR of the fused frame obtained by fusing the N-frame and X-frame. This will lead to image quality abnormalities in the captured images obtained by electronic devices, such as the generation of various pseudo-textures. For example, Figure 3 This is a schematic diagram of a method for capturing moving objects. Figure 3 The shooting scenes include both static and moving objects, with the moving objects being people and the static objects being numbers and houses. Figure 3 In the image, (a) represents the image the user expects to capture, while Figure 3 Image (b) is a captured image obtained when the user enabled Stagger HDR mode. At this time, the still area is relative to... Figure 3 The static area shown in (a) is blurred.
[0071] In view of the above problems, this application provides an image generation method. When an electronic device enables the Stagger High Dynamic Range (Stagger HDR) mode, in response to receiving an operation that triggers the electronic device to capture an image, a first frame and a second frame are obtained from the preview frame sequence in the Stagger HDR mode. The preview frame sequence includes multiple latest preview frames cached during the preview process. The frame type of the second frame is different from that of the first frame, including: long frames in the Stagger HDR mode and short frames in the Stagger HDR mode. The first frame, the captured long frame sequence, and the short frame sequence are fused to obtain a first fused frame. The long frame sequence includes multiple captured long frames, and the short frame sequence includes multiple captured short frames. The second frame and the first fused frame are fused to obtain a second fused frame. Based on the second fused frame, the captured image is obtained. In other words, when electronic devices fuse short and long frame sequences acquired during shooting, the long and short frames in Stagger HDR mode are used as benchmarks. The signal-to-noise ratio (SNR) of the long frames in stagger mode is significantly higher than that of the static areas in the X-frame, compensating for the insufficient SNR in the static areas of the fused frame. This improves the SNR of the static areas in the fused image, thus resolving image quality anomalies such as various textures caused by insufficient SNR in the static areas of the fused image, and improving the overall image quality. Furthermore, the rich motion region details in the short frames of stagger mode can be utilized to improve the quality of the motion regions in the fused image.
[0072] The image generation method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of this application.
[0073] Appendix Figure 4 This application provides a flowchart of an image generation method, which includes a preview path processing method and a photo capture path processing method.
[0074] The preview path processing method refers to the data processing method performed during the preview process when an electronic device is in Stagger HDR mode. The preview path refers to the path from which an image sensor (e.g., a camera) captures raw image data based on the principle of snapshot before shooting, and then the image processor performs preliminary processing to generate a preview image. In this embodiment, the preview path processing method is used to perform low-resolution motion estimation and combine the low-resolution motion estimation to determine the reference frame.
[0075] The image capture path processing method refers to the method of data processing in the image capture path. The image capture path refers to the process by which the image sensor captures an image and generates the final captured image after shooting (e.g., pressing the shutter button or triggering the shutter button) in Stagger HDR mode. In the embodiments of this application, after the electronic device captures an image sequence, it performs fusion processing on the image sequence to obtain a captured image with a high signal-to-noise ratio in still areas.
[0076] The following sections provide detailed explanations of the preview path processing method and the image capture path processing method. (See attached...) Figure 4 As shown, the preview path processing method includes:
[0077] S410: Perform preview path motion estimation on the preview frame sequence to obtain low-resolution motion.
[0078] A preview frame sequence refers to the sequence of original image frames captured by the image sensor based on the capture principle before the capture action. In this embodiment, the electronic device adopts a capture mode based on Stagger HDR technology, and the preview frame sequence can be a sequence of spare frames buffered in ZSLBuffer.
[0079] In this embodiment, to achieve an instant capture effect, the electronic device needs to obtain a preferred backup frame from the preview frame sequence. Starting with the preferred backup frame, the electronic device performs fusion processing on the image frames acquired during the capture process to obtain a fused frame. Furthermore, to ensure the quality of the acquired fused image, the electronic device uses the complete preview frame acquired closest to the moment the capture button is pressed as the preferred backup frame. For example, regarding the attached... Figure 2 As shown, the electronic device uses frames N2 and S2 as preferred backup frames.
[0080] Furthermore, when fusing long and short frame sequences during the image capture process, to ensure the image quality of the fused frame, the electronic device first acquires a high-quality reference frame (also known as the first frame), and then uses this reference frame as a benchmark to fuse the long and short frame sequences. For example, if the image quality of N frames is better than that of S frames, then N frames are selected as the reference frame; if the image quality of N frames is worse than that of S frames, then S frames are selected as the reference frame.
[0081] In one example, for capturing a moving object, the electronic device can compare the image quality of two preview frames in a backup frame and select the best-quality preview frame (i.e., the backup frame) as the reference frame. Since the preview frame includes both moving and stationary areas, for N frames in a preview frame sequence, the image quality of the stationary area is better than that of the moving area, while for S frames, the image quality of the moving area is better than that of the stationary area. If the moving area occupies a large proportion of the image in a preview frame, for example, 80%, meaning the image quality of the preview frame is significantly affected by the image quality of the moving area, then the image quality of the S frame is better than that of the N frame, and the electronic device selects the shorter S frame as the reference frame. Conversely, if the moving area occupies a small proportion of the image in a preview frame, for example, 10%, meaning the image quality of the preview frame is significantly affected by the image quality of the stationary area, then the image quality of the N frame is better than that of the S frame, and the electronic device selects the N frame as the reference frame.
[0082] In another example, to obtain a high-quality reference frame, the electronic device first performs preview path motion estimation on the preview frame sequence to obtain low-resolution motion data of the pixel data in the reference preview frame within the preview frame sequence. The electronic device utilizes the characteristic that the low-resolution motion data can quickly estimate and track the motion state of moving objects to quickly and accurately locate the reference frame.
[0083] The following section provides a detailed explanation of the preview pathway estimation.
[0084] In one example, preview path estimation can be achieved by the electronic device determining a frame from the preview frame sequence as a reference preview frame. The reference preview frame can be the preview frame closest to the moment the shutter was pressed, or it can be a preview frame at a preset time; this embodiment does not specifically limit the reference frame. Multiple feature points are extracted from the reference preview frame, and then the motion of these feature points to corresponding feature points in adjacent preview frames is determined. Corresponding feature points in adjacent preview frames refer to feature points whose feature point data (e.g., color, structure, etc., representing image information) is the same as those in the reference preview frame. However, this method has high computational complexity, which affects the image processing speed. Furthermore, individual feature points are significantly affected by the environment, resulting in poor robustness of the estimation.
[0085] In another example, the preview path can acquire low-resolution motion data in a patch-based manner to reduce computational complexity, improve image processing speed, and enhance robustness. (See attached diagram.) Figure 5 -Appendix Figure 8 A detailed explanation is provided below. For ease of understanding, this application uses a preview frame as an example for illustration. It is understood that other preview frames can also be obtained based on the above method.
[0086] Appendix Figure 5A flowchart illustrating a method for implementing motion estimation in a preview path, provided as an embodiment of this application. The method includes the following:
[0087] S510. Divide the reference preview frame into multiple image blocks and obtain the motion data of each image block.
[0088] The electronic device divides the reference preview frame into multiple image blocks, for example, the reference preview frame is divided into P × Q image blocks. Here, P is an integer greater than 1, and Q is also an integer greater than 1, for example, P is 10 and Q is 8. It should be noted that the values of P and Q can be adjusted by those skilled in the art as needed. Each image block has the same size specifications.
[0089] The electronic device performs feature extraction and tracking on each image patch, acquiring feature points for each patch and the motion magnitude of each feature point. Then, based on the motion magnitude of each feature point, the electronic device estimates the low-resolution motion magnitude of the entire image patch data. Example illustration: (See attached image) Figure 6 This is a schematic diagram illustrating a method for acquiring low-resolution motion data based on a preview frame, provided as an embodiment of this application. The electronic device divides the preview frame into 8×10 image blocks, totaling 80 blocks. Feature points are extracted from each image block to obtain the data as shown in the attached diagram. Figure 6 The black dots shown represent feature points for each image patch. Electronic devices can obtain low-resolution motion data for each image patch based on the motion estimation vectors of the black dot positions within that patch.
[0090] This application does not specifically limit the method for extracting feature points for each image patch. For example, electronic devices can use feature point detection algorithms to identify and extract feature points in image patches, such as corner points and edge points in the image patch. Feature point detection algorithms can be Scale-Invariant Feature Transform (SIFT) algorithms, or Speeded-Up Robust Features (SURF) algorithms, etc.
[0091] Feature tracking refers to comparing the positional changes of feature points in adjacent preview frames to obtain the motion data of each feature point. This application does not specifically limit the implementation method of feature tracking; for example, electronic devices can perform feature tracking based on optical flow or background subtraction methods to obtain the motion data of each feature point.
[0092] The electronic device estimates the low-resolution motion of the entire image patch data based on the motion of each feature point. For example, the electronic device can estimate the low-resolution motion of the entire image patch data by averaging the motion of each feature point in the image patch, or by selecting the motion of the most representative feature point as the low-resolution motion of the image patch data. Furthermore, the electronic device can also perform the estimation in other ways, which are not specifically limited in the embodiments of this application.
[0093] Furthermore, in this embodiment, the electronic device can also acquire the confidence level of low-resolution motion in the image patch and adjust the motion of the image patch data. The confidence level is used to evaluate the reliability of the acquired low-resolution motion; it is understood that the higher the reliability, the higher the confidence level, and the more accurate the acquired low-resolution motion. In one possible implementation, to ensure the reliability of the acquired low-resolution motion of the image patch, image patches with a confidence level below a confidence threshold are deleted, and no further processing is performed.
[0094] In this embodiment, the low-resolution motion data of the image block obtained in S410 is obtained by extracting feature points in the image block. The feature points in the image block are unevenly distributed. For example, some image blocks have sparse feature points and some have dense feature points. This results in uneven distribution of motion data of all image block data in the reference preview frame. The motion area determined based on the uneven motion data of the image block has low accuracy.
[0095] In one specific implementation, the electronic device can execute S520 to reconstruct low-resolution motion data that is uniformly distributed within the frame.
[0096] S520. Reconstruct the motion data of the image patch data to obtain the low-resolution motion data that is uniformly distributed within the frame.
[0097] Uniformly distributed low-resolution motion within a frame is used to indicate the relatively uniform distribution of low-resolution motion between image blocks within the same preview frame.
[0098] In one specific implementation, the electronic device can reconstruct the low-resolution motion of image patch data based on the following method.
[0099] Step 1: The electronic device pre-determines the pixels to be reconstructed in the image block.
[0100] This application does not specifically limit the pixels to be reconstructed. For example, an electronic device can use the center point of each image block as the pixel to be reconstructed. Exemplarily, see the attached... Figure 7 This is a schematic diagram illustrating intra-frame motion reconstruction as provided in an embodiment of this application. The electronic device will determine the center position of an image patch, such as... The pixel to be reconstructed is located at a certain point. The relative positions of the pixels to be reconstructed in each image block are the same, thus ensuring a uniform distribution of motion of feature points within the frame.
[0101] In addition, electronic devices can select other specific locations within an image block as pixels to be reconstructed based on the preview frame content, feature point distribution, and motion trajectory.
[0102] Step 2: The electronic device calculates the distances between the current image block and the pixel to be reconstructed in the neighboring image blocks that meet the preset distance conditions, and uses the feature points corresponding to these distances as candidate feature points, and the low-resolution motion of the candidate feature points as candidate motion.
[0103] The preset distance conditions can be: closest to the pixel to be reconstructed in the image block, farthest from the pixel to be reconstructed in the image block, or a set distance from the pixel to be reconstructed in the image block.
[0104] For example, Appendix Figure 7 This is a schematic diagram of intra-frame motion reconstruction provided in an embodiment of this application. The adjacent image blocks of the current image block Block0 are divided into 8 parts, namely Block1 to Block8. The electronic device calculates the feature point in each of Blocks 1 to 8 that is closest to the pixel to be reconstructed, and the candidate position corresponding to the feature point is indicated by the black solid circle.
[0105] The electronic device calculates the distance between the pixel to be reconstructed and the candidate positions. For example, the distance between the candidate position of Block 1 and the pixel to be reconstructed is S1, the distance between the candidate position of Block 2 and the pixel to be reconstructed is S2, ..., and the distance between the candidate position of Block 8 and the pixel to be reconstructed is S8.
[0106] Electronic devices, based on a pre-calibrated Look-Up Table (LUP) and the distance between the pixel to be reconstructed and candidate positions, can obtain the interpolation reconstruction weights (also known as intra-frame weights) for each candidate position within a reference preview frame. The Look-Up Table is a pre-built mapping table between the distance between the pixel to be reconstructed and the candidate position and the interpolation reconstruction weights. The interpolation reconstruction weights are negatively correlated with the distance between the pixel to be reconstructed and the candidate position; the greater the distance, the smaller the interpolation reconstruction weight. Example: The Look-Up Table is shown in Table 1. Figure 7 The interpolation weights (Reconstruct_Weight, RW) of the candidate positions corresponding to the pixels to be reconstructed are RW(S1), RW(S2), ..., RW(S8).
[0107] Table 1
[0108]
[0109] In this embodiment, the electronic device utilizes a pre-calibrated position weight lookup table to quickly obtain the interpolation reconstruction weight of each candidate position based on the distance between the given candidate position and the pixel to be reconstructed, simplifying computational complexity and improving processing speed. Furthermore, the position weight lookup table can be adjusted and optimized according to actual needs, thus facilitating maintenance and expansion.
[0110] It should be noted that the distance between the pixel to be reconstructed and the candidate position provided in the embodiments of this application can be Euclidean distance, Manhattan distance, or Chebyshev distance, etc., and is not specifically limited in the embodiments of this application.
[0111] Step 3: The electronic device uses the sum of the interpolation reconstruction weights of all candidate positions and the corresponding motion quantities to obtain the motion quantity of the pixel to be reconstructed.
[0112] The electronic device uses the motion of the pixel to be reconstructed as the low-resolution motion of the current image patch.
[0113] For example, for the attached Figure 7 As shown, if the motion amount of the candidate position of Block 1 is MV_Block1, the motion amount of the candidate position of Block 2 is MV_Block2, ..., the motion amount of the candidate position of Block 8 is MV_Block8, then the motion amount RV_S of the pixel to be reconstructed is:
[0114] RV_S=RW(S1)*MV_Block1+...+RW(S8)*MV_Block8 (1)
[0115] The motion of the current image block Block0 is the motion of the pixel to be reconstructed, RV_S.
[0116] Using the above method, the electronic device can acquire the motion data of all image blocks within the preview frame. Therefore, the electronic device only needs to determine the pixels to be reconstructed and the reconstruction rules for other image blocks using the same method to ensure a uniform distribution of low-resolution motion data among image blocks within the frame. Furthermore, the electronic device can improve the accuracy of motion estimation by assigning appropriate weights to make candidate positions closer to the reconstruction location and more reliable to contribute more to the final motion data through reasonable weight allocation.
[0117] Furthermore, compared to local interpolation methods or median filtering methods using known feature points within an image block, the motion data obtained by the above-mentioned methods is less susceptible to the influence of abnormal motion within the block, resulting in higher accuracy of the obtained intra-frame motion data.
[0118] Furthermore, if the distribution of low-resolution motion between adjacent preview frames in the reference preview frame sequence is uneven, it will also affect the accuracy of obtaining the overall motion trend of the moving object in the preview frame. Therefore, in another specific implementation, the electronic device can execute S530 after executing S520.
[0119] S530: Reconstruct inter-frame motion and adjust the low-resolution motion of image patch data in the reference image frame.
[0120] Inter-frame motion refers to the motion of corresponding feature points between preview frames. In this embodiment, the electronic device reconstructs the inter-frame motion to obtain spatially consistent inter-frame motion between preview frame sequences. Therefore, the electronic device can accurately obtain the motion distribution and magnitude of moving objects, thereby accurately identifying preview frames with higher image quality and better suited to the current motion scene requirements as reference frames.
[0121] In this embodiment, the electronic device reconstructs the motion data of multiple pixels in the current frame based on the characteristics that the motion amount between adjacent frames is continuous and the motion path of the moving object is smooth in time, using the previous frame or the previous N1 frame, or the next frame or the next N2 frame of the reference preview frame.
[0122] For ease of description, in this application embodiment, the frame preceding the current frame or the previous N1 frames are referred to as historical frames, and the frame following the current frame or the next N2 frames are referred to as future frames. N1 is an integer greater than 1, and N2 is also an integer greater than 1. Exemplary description: Appendix Figure 8 A schematic diagram illustrating a method for reconstructing inter-frame motion data as provided in an embodiment of this application. (Attached) Figure 8 It displays the two preceding historical frames and the next future frame of the current frame. Historical frame 1 is the adjacent frame of the current frame, future frame 2 is the adjacent frame of the current frame, and historical frame 2 is the adjacent frame of historical frame 1.
[0123] For ease of description, in this application embodiment, the motion amount of pixels in historical frames is collectively referred to as RV-Pre, the motion amount of pixels in future frames is referred to as RV_Fut, the confidence level corresponding to the motion amount of pixels in historical frames is referred to as Conf_Pre, and the confidence level corresponding to the motion amount of pixels in future frames is referred to as Conf_Fut.
[0124] In one possible approach, the electronic device first uses S520 to acquire the motion data of mid-pixel points in historical and future frames from the preview frame sequence. Example illustration: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] Figure 8As shown, the electronic device first obtains the motion amount RV_Pre1 of historical frame 1 in image block 1, the motion amount RV_Pre2 of historical frame 2 in image block 2, and the motion amount RV_Fut of future frame in image block 4 using step 1. Image block 1, image block 2, and image block 4 refer to the positions of the same image block on different preview frames.
[0125] Then, the electronic device uses historical and future frames to reconstruct the inter-frame motion through a time-domain filtering method, thereby ensuring the continuity and smoothness of the inter-frame motion. The motion RV_T of the target image patch at its position in the current frame can be expressed as:
[0126] RV_T= RV_Fut*CW_Fut+ RV_Pre*CW_Pre (2)
[0127] Wherein, CW_Pre indicates the inter-frame reconstruction weight value (also known as inter-frame weight) assigned by the electronic device to the motion of pixels in historical frames, and CW_Fut indicates the inter-frame reconstruction weight value assigned by the electronic device to the motion of pixels in future frames. The inter-frame reconstruction weight value can be determined based on the time distance between adjacent frames, or it can be determined based on other factors, such as the electronic device's recognition result of the motion region in the preview frame, etc., which is not specifically limited in this application embodiment.
[0128] Furthermore, the inter-frame reconstruction weight value can also be determined based on the confidence level corresponding to the intra-frame motion. In one example, the electronic device first obtains the inter-frame reconstruction weight value corresponding to the motion of each pixel in the future frame and the inter-frame reconstruction weight value corresponding to the motion of each pixel in the historical frame, based on the confidence level of the intra-frame motion of the future frame and the confidence level of the intra-frame motion of the historical frame, combined with a pre-calibrated confidence weight lookup table. The confidence weight lookup table is a positive mapping mechanism between confidence level and inter-frame reconstruction weight value, constructed by those skilled in the art as needed. The higher the confidence level, the larger the inter-frame reconstruction weight value. For example, Table 2 is a confidence weight lookup table provided in an embodiment of this application. The electronic device can determine the inter-frame reconstruction weight value CW(Conf_Fut) of the future frame and the inter-frame reconstruction weight value CW(Conf_Pre) of the historical frame based on Table 2, according to the confidence levels of the future frame and the historical frame. This improves the speed of obtaining the reconstructed inter-frame motion.
[0129] Table 2
[0130] Input (confidence level) Output (confidence weight) Conf1 CW(Conf1) Conf2 CW(Conf2) …… …… Confi CW(Confi) …… ……
[0131] Example description: For the appendix Figure 8As shown, the confidence levels of the motion of the pixels to be reconstructed in the two historical frames are Conf_Pre1 and Conf_Pre2, respectively. The confidence level of the motion of the pixel to be reconstructed in the one future frame is Conf_Fut. The motion of the pixel to be reconstructed in the current frame obtained by the electronic device is:
[0132] RV_T=RV_Fut*CW(Conf_Fut)+RV_Pre1*CW(Conf_Pre1)+RV_Pre2*CW(Conf_Pre2)(3)
[0133] Where CW_Pre1 is the inter-frame reconstruction weight value corresponding to the motion of the pixel to be reconstructed in historical frame 1, and CW_Pre2 is the inter-frame reconstruction weight value corresponding to the motion of the pixel to be reconstructed in historical frame 2. At this time, the motion of the target image block data in the current frame is the motion corresponding to the small "☆". In addition, [the following is attached...] Figure 8 The other small "☆" symbols indicate that the electronic device uses the S520 method to acquire the motion amount RV_Pre1 of historical frame 1 in image block 1, the motion amount RV_Pre2 of historical frame 2 in image block 2, and the motion amount RV_Fut of future frame in image block 4. Among them, image block 1, image block 2, and image block 4 are the positions of the target image block on different preview frames.
[0134] Electronic devices use the confidence level of intra-frame motion to determine inter-frame reconstruction weight values, assigning higher weights to motion with high confidence levels. This results in more accurate and reliable inter-frame motion data for the current frame. Example illustration: (See attached diagram) Figure 8 As shown, the low-resolution motion data of the target image block in the current frame, such as the motion data corresponding to the large "☆", is directly calculated based on the method described in S520. Comparative analysis reveals that reconstructing the low-resolution motion data of the current frame using historical and future frames is more conducive to the electronic device understanding the overall motion trend of the moving object.
[0135] Furthermore, the embodiments of this application do not specifically limit the time-domain filtering method. For example, the weighted average method, Kalman filtering, Wiener filtering and other methods can be used to construct the time-domain filtering model and reconstruct the inter-frame motion using the constructed time-domain filtering model.
[0136] In this embodiment, the electronic device can increase the number of historical frames and future frames, thereby obtaining more accurate motion data of pixels in the current frame. It is understood that the more historical and future frames there are, the smoother the inter-frame motion becomes, and the stronger the interdependence, thus helping to reduce errors and jitter in motion estimation and improve reconstruction accuracy. If the number of historical and future frames is small, the smoothness of inter-frame motion decreases, and the independence becomes stronger, leading to reduced reconstruction accuracy.
[0137] S540. Obtain the low-resolution motion of the reference preview frame in the preview frame sequence.
[0138] In this embodiment of the application, the electronic device can acquire low-resolution motion data of the reference preview frame based on S510 and S520.
[0139] In summary, electronic devices adopt Figure 5 The low-resolution motion data obtained in the manner shown satisfies the following condition: the motion data of the image patch data in the reference preview frame is uniformly and accurately distributed. Therefore, by using the low-resolution motion data, the motion of image patches can be accurately located, and the motion region in the preview frame can be accurately identified, laying the groundwork for image fusion.
[0140] S420, Reference Frame Selection.
[0141] In this embodiment, the electronic device determines the motion range and motion intensity of the preview frame in the motion scene based on the low-resolution motion of the reference preview frame obtained from the preview frame sequence.
[0142] The motion range of a preview frame refers to the area occupied by the motion region within the entire preview frame. In this embodiment, the electronic device can determine the motion range of the preview frame based on the number of image blocks in the motion region. For example, if the number of image blocks in the motion region of the preview frame is 50, then the motion range of the preview frame is 50.
[0143] The motion intensity of a preview frame refers to the amount of displacement of a moving object or pixel in the preview frame over time.
[0144] In this embodiment, the electronic device determines whether an image block is located in a moving region based on whether the motion intensity of the image block is greater than a motion intensity threshold (i.e., Motion_th). If the motion intensity of the image block is greater than Motion_th, then the image block is located in a moving region. Motion_th is a value set by those skilled in the art as needed; for example, Motion_th is set to 20.
[0145] If the number of image blocks with motion intensity greater than Motion_th in the reference preview frame exceeds the number threshold (Number_th), it indicates that the moving regions in the preview frame occupy a large proportion of the image, and the image quality of the preview frame is significantly affected by the image quality of the moving regions. In this case, the image quality of frame S is better than that of frame N, and the electronic device selects frame X as the reference frame. Conversely, if the moving regions occupy a small proportion of the image in the preview frame, meaning that the image quality of the preview frame is significantly affected by the image quality of the stationary regions, then the image quality of frame N is better than that of frame S, and the electronic device selects frame N as the reference frame.
[0146] Here, Number_th is a dynamically changing value set by those skilled in the art as needed, for example, Number_th is 50% of the total number of image blocks.
[0147] As a result, electronic devices can acquire reference frames with high image quality.
[0148] The following describes the methods for processing the image path. See the appendix for further details. Figure 4 As shown, the image processing method includes:
[0149] S430, Obtain the image frame sequence.
[0150] An image frame sequence includes both long frame sequences and short frame sequences acquired during the same capture process. The long frame sequence includes long frames acquired during the capture of a moving object, while the short frame sequence includes short frames acquired during the capture of the object.
[0151] In one example, to achieve the instant capture effect, the long frame sequence also includes N frames of spare frames, and the short frame sequence also includes S frames of spare frames. Example illustration: Appendix Figure 9 This is a schematic diagram of an image frame sequence provided in an embodiment of this application. (Attached) Figure 9 This is a sequence of image frames acquired using Stagger HDR capture technology. The corresponding image frames are pairs of long and short frames acquired simultaneously. (The last sentence appears to be incomplete and possibly contains errors.) Figure 9 The long frame sequence shown includes a preferred spare frame N2, and the short frame sequence includes a preferred spare frame S2. The long frame sequence can be {N2, N3, N4, N5}, and the short frame sequence can be {S2, X1, X2, X3}.
[0152] In another example, because the signal-to-noise ratio (SNR) of the static area in short frames is low, in order to ensure that the fused frame acquired during fusion has a higher SNR in the static area, the electronic device can fuse multiple S-frames in the preview frame sequence to obtain a short-frame fused frame S_Fusion. The long frame sequence acquired during the capture process includes at least N frames from one preview frame sequence, and the long frame sequence includes S_Fusion.
[0153] Exemplary illustration: Appendix Figure 10 This is a schematic diagram illustrating the acquisition of an image frame sequence according to an embodiment of this application. The long frame sequence includes {N0, N1, N2, N3, N4, N5}, and the short frame sequence includes {S_Fusion, X1, X2, X3}. S_Fusion is an appendix... Figure 9 The image shows the fused frame obtained after fusing the four S-frames shown.
[0154] For example, appendix Figure 11This is a schematic diagram illustrating another method for acquiring an image sequence according to an embodiment of this application. The long frame sequence includes {N0, N1, N2, N3, N4, N5}, and the short frame sequence includes {S_Fusion, S1, X1, X2, X3}. Wherein, S_Fusion is an appendix... Figure 9 The image shows the fused frames of the three S-frames preceding frame S1.
[0155] In S_Fusion, the signal-to-noise ratio (SNR) of the stationary region is higher than that of the stationary region in each of the S frames. This is because the location and intensity of random noise typically vary across different frames, while the information in the stationary region remains consistent. Through fusion, noise can be effectively suppressed, while the image information in the stationary region is enhanced.
[0156] In one example, when fusing multiple S-frames, the electronic device first presets a standard preview frame. The standard preview frame indicates N frames in which the moving object has been correctly captured. In this embodiment, the standard preview frame can be N frames containing complete image information closest to the moment the shutter was pressed. (Regarding the appendix...) Figure 9 As shown, the standard preview frame can be an N2 frame.
[0157] In this embodiment, the electronic device sets a standard preview frame, causing multiple S-frames to be fused to be fused based on the S-frame corresponding to the standard preview frame. Since the S-frame corresponding to the standard preview frame contains the latest and most complete moving object data, it helps ensure that the short frame fused frame can reflect the scene of capturing the moving object at that moment, thus improving the image quality of the subsequently acquired image fused frame.
[0158] Furthermore, to ensure that moving objects are captured correctly, the electronic device dynamically adjusts the number of S-frames output based on the gain level and exposure reduction ratio (DER) of the current N frames. The number of S-frames output refers to the number of S_Fusion frames that need to be obtained through multi-frame fusion processing in the imaging path. Correct capture means that the moving object is recorded correctly, completely, and clearly.
[0159] The exposure reduction ratio refers to adjusting the camera application's exposure settings to reduce the amount of exposure, by a percentage of the original exposure. For example, if adjusting the camera application's exposure compensation value reduces the exposure by twice the original exposure (i.e., x), then the exposure reduction ratio is 2x.
[0160] In this embodiment, the electronic device can reduce the exposure ratio based on the current scene motion information and determine the number of frames to ensure that the moving object is captured correctly. The motion information includes information representing the direction and magnitude of motion, such as the moving object's speed, direction of movement, and acceleration.
[0161] When considering the number of frames output, electronic devices also take into account the current N-frame gain level to further improve the accuracy of capturing moving objects. The current N-frame gain level also directly affects the brightness and signal-to-noise ratio of the captured image. A higher gain value will increase the brightness of the captured image but will increase the noise level, while a lower gain will result in an overly dark image.
[0162] In Example 1, if the current N-frame gain is less than or equal to the first gain threshold (Gain_th1) and DER is less than the first scaling threshold (De_th1), the electronic device adjusts the number of short frame preview frames (S-frames) output to n1 frames. Here, n1 is an integer greater than or equal to 1. Otherwise, the number of S-frames output is n1' frames, where n1' > n1 and n1' is an integer.
[0163] Gain_th1 and De_th1 are the N-frame gain threshold and exposure reduction threshold set by those skilled in the art as needed to ensure correct capture of moving objects. For example, Gain_th1 = 2x, De_th1 = 2x. n1 frames and n1' frames are the minimum number of frames obtained by those skilled in the art based on experience, at which point correct capture of moving objects is possible. For example, n1 = 1, n1' = 2.
[0164] In Example 2, if the current N-frame gain is less than or equal to the first gain threshold (Gain_th1) and DER is less than the first proportional threshold (De_th1), the electronic device adjusts the number of short frame preview frames (S-frames) to n1 frames. If the second gain threshold (Gain_th2) is greater than or equal to the current N-frame gain and greater than Gain_th1, and the second proportional threshold (De_th2) is greater than DER and greater than De_th1, then the number of S-frames is n2 frames, where n2 is an integer greater than n1'. For example, n1' = 2, n2 = 3. In other cases, the number of S-frames is n3 frames, where n3 is an integer greater than n2, for example, n2 = 3, n3 = 4.
[0165] Gain_th2 and De_th2 are N-frame gain thresholds and exposure reduction thresholds that can be set by those skilled in the art to ensure that moving objects are captured correctly, for example, Gain_th2 = 4x and De_th3 = 3x.
[0166] In Example 3, if the current N-frame gain is less than or equal to the first gain threshold (Gain_th1) and DER is less than the first scaling threshold (De_th1), the electronic device adjusts the number of short-frame preview frames (S-frames) to n1 frames. If the second gain threshold (Gain_th2) is greater than or equal to the current N-frame gain and greater than Gain_th1, and the second scaling threshold (De_th2) is greater than DER and greater than De_th1, the number of S-frames is n2 frames. If the third gain threshold (Gain_th3) is greater than or equal to the current N-frame gain and greater than Gain_th2, and the third scaling threshold (De_th3) is greater than DER and greater than De_th2, the number of S-frames is n3 frames. In other cases, the exposure reduction ratio is lowered to avoid excessive exposure reduction exacerbating noise problems and causing image quality degradation.
[0167] Among them, Gain_th3 and De_th3 are N-frame gain thresholds and exposure reduction thresholds that can be correctly captured by those skilled in the art as needed, such as Gain_th3 = 5x and De_th3 = 4x.
[0168] In another example, to further improve the signal-to-noise ratio in the static region of the short-frame fusion frame, a global alignment of multiple S-frames to be fused can be performed based on a standard preview frame to obtain a globally aligned S-frame sequence. The electronic device then fuses the globally aligned S-frame sequence to obtain S_Fusion.
[0169] For example, as shown in the appendix Figure 9 As shown, if the four S-frame sequences are {S-1, S0, S1, S2}, the globally aligned S-frame sequence obtained after global alignment is {S-1_GA, S0_GA, S1_GA, S2_GA}. Merging {S-1_GA, S0_GA, S1_GA, S2_GA} yields the attached... Figure 10 The S_Fusion shown.
[0170] The embodiments of this application do not specifically limit the global alignment method. For example, the global alignment method can be a feature matching method or a sparse optical flow method, etc.
[0171] In another example, the electronic device can also perform noise reduction processing during the short frame fusion process.
[0172] In another example, the electronic device can also determine different noise reduction methods based on the maximum gain of the preview frame, thereby improving noise reduction speed and memory resource usage while reducing noise during the short frame fusion process. The maximum gain value of the preview frame is determined by the maximum gain of the current N frames and the maximum current underexposure ratio.
[0173] Example explanation: (1) If the current N-frame gain ≤ Gain_th1 and DER < De_th1, the number of output S-frames is n1 frames. At this time, the maximum gain of the S-frame is Gain_th1 * De_th1, and the image signal-to-noise ratio of the S-frame is the first signal-to-noise ratio, which is relatively good. Electronic devices can perform noise reduction processing based on the multi-frame noise reduction algorithm of the first-level S-frame. At this time, the short frame fusion frame is:
[0174] S_Fusion=λ1*S1_GA+……+λ n1 *Sn1_GA (4)
[0175] Where, λ1+……+λ n1 =1.
[0176] The embodiments of this application do not specifically limit the multi-frame noise reduction algorithm of the first-level S-frame; for example, it can be Gaussian noise reduction, etc.
[0177] (2) If Gain_th2 ≥ the current N-frame gain > Gain_th1, and De_th2 > DER ≥ De_th1, then the maximum gain of the S-frame is Gain_th2 * De_th2, and the image signal-to-noise ratio of the S-frame is the second signal-to-noise ratio, which is less than the first signal-to-noise ratio. In this case, the electronic device can perform noise reduction processing based on the multi-frame noise reduction algorithm of the second-level S-frame. The noise processing capability of the multi-frame noise reduction algorithm of the second-level S-frame is stronger than that of the multi-frame noise reduction algorithm of the first-level S-frame, and the multi-frame noise reduction algorithm of the second-level S-frame comprehensively considers noise reduction in the moving region.
[0178] The embodiments of this application do not specifically limit the multi-frame denoising algorithm for the second-level S-frame. For example, it can be a multi-scale spatial domain and frequency domain denoising algorithm, or a non-local means (NLM) transform domain denoising algorithm.
[0179] (3) If Gain_th3 ≥ the current N-frame gain > Gain_th2, and De_th3 > DER ≥ De_th2, then the maximum gain of the S-frame is Gain_th3 * De_th3, and the image signal-to-noise ratio (SNR) of the S-frame is the third SNR, which is less than the second SNR. At this point, the overall preview frame's SNR is relatively poor, and the electronic device can perform noise reduction processing based on the multi-frame noise reduction algorithm of the third-level S-frame. The noise processing capability of the multi-frame noise reduction algorithm of the third-level S-frame is greater than that of the multi-frame noise reduction algorithm of the second-level S-frame.
[0180] This application does not specifically limit the multi-frame denoising algorithm for third-level S-frames. For example, the multi-frame denoising algorithm for third-level S-frames can be a denoising algorithm combining multi-scale spatial domain and frequency domain denoising algorithms with AL denoising algorithms. For instance, the multi-frame denoising algorithm for third-level S-frames can be a partitioning algorithm, where static regions are processed using multi-scale spatial domain and frequency domain denoising algorithms, and moving regions are processed using a motion-compensated spatiotemporal filtering algorithm. It should be noted that using a partitioning algorithm for denoising can balance the signal-to-noise ratio between moving and static regions, thereby obtaining short-frame fused frames with higher image quality.
[0181] In summary, the embodiments of this application employ a hierarchical, multi-frame noise reduction algorithm with progressively increasing complexity to address the signal-to-noise ratio (SNR) levels of S-frames under different gain conditions. This achieves optimal noise reduction while maintaining processing performance, resulting in a higher SNR in stationary regions.
[0182] The embodiments of this application can also perform S-frame fusion in other ways, and the embodiments of this application are not specifically limited.
[0183] S440. Based on the low-resolution motion and the reference frame, the image frame sequence is fused once to obtain the first fused frame.
[0184] In this process, the electronic device can perform fusion processing on the long frame sequence, the short frame sequence and the reference frame based on the image fusion network to obtain the first fused frame.
[0185] In this embodiment, to ensure the fusion effect, the long frame sequence and the short frame sequence need to be globally aligned to obtain a globally aligned long and short frame sequence. In the globally aligned long and short frame sequence, the long frames and short frames are synchronized in time and space. The following refers to the appendix... Figure 12 A flowchart of a one-time fusion processing method provided in this application embodiment is described in detail.
[0186] S441: Denoise the short frame sequence to obtain a denoised short frame sequence.
[0187] In one example, the electronic device performs noise reduction on each short frame in the short frame sequence to obtain a noise-reduced short frame sequence. In another example, the electronic device may first perform global image alignment on the short frame sequence to synchronize the short frames in the sequence in time and space, and then perform noise reduction on the globally image-aligned short frame sequence to obtain a noise-reduced short frame sequence.
[0188] The embodiments of this application do not specifically limit the use of more advanced image registration algorithms, such as feature point-based matching algorithms or deep learning-based registration methods, to improve alignment accuracy.
[0189] S442: Perform global image alignment between the long frame sequence and the short frame denoised sequence to obtain the long and short frame global alignment sequence.
[0190] Because in an image frame sequence, long frames in a long frame sequence are used to capture scene changes over a longer time interval, while short frames in a short, denoised sequence are used to capture detail changes over a short period, even if the long and short frames are acquired at the same time, they are not synchronized in time and space. Furthermore, due to lens distortion, scale differences, parallax, and the influence of shooting position, alignment errors exist between the long and short frames, affecting the fusion effect. Therefore, to improve image fusion, electronic devices can first perform global image alignment processing on the long frame sequence and the short, denoised sequence to obtain a globally aligned long-short frame sequence.
[0191] The embodiments of this application do not specifically limit the use of more advanced image registration algorithms, such as feature point-based matching algorithms or deep learning-based registration methods, to improve alignment accuracy.
[0192] S443: The first fused frame is obtained by processing the global alignment sequence of long and short frames based on the image fusion network.
[0193] Electronic devices can use image fusion networks to fuse a globally aligned sequence of long and short frames with a reference frame to obtain a first fused frame. However, this alignment method does not take into account the signal-to-noise ratio difference between stationary and moving regions, which limits the image quality of the moving regions in the first fused frame.
[0194] In this embodiment, the electronic device adds a noise model to the image fusion network for noise estimation. The noise model includes the cost estimation function Loss_Function_update of the image fusion network. The loss value of Loss_Function_update measures the difference between the fusion result Fusion_Cur and the ground truth. If the loss value is lower than a preset loss threshold, the noise of the image fusion network is minimized, and the signal-to-noise ratio in the static area is highest. That is, the electronic device optimizes the image fusion network using Loss_Function_update to make the fusion result closer to the ground truth.
[0195] In this embodiment, Loss_Function_update fully considers the signal-to-noise ratio (SNR) difference between static and moving regions. Since the SNR of static regions tends to be more similar to that of frame X, insufficient SNR in frame X can lead to image quality degradation. Therefore, this embodiment optimizes the SNR of static regions by introducing static region information into Loss_Function_update.
[0196] Specifically, in the embodiments of the present application, the electronic device introduces a fusion mask loss Mask_For_Loss in Loss_Function_update. Mask_For_Loss is used to enable Loss_Function_update to only control the signal-to-noise ratio of the static region and not calculate the signal-to-noise ratio of the moving region.
[0197] In one example, for the static region, Mask_For_Loss is set to a fixed value g. For example, g = 1. For the moving region, the loss value of Loss_Function_update is not calculated. Exemplary illustration: For the pixel point (i, j) in the image, Mask_For_Loss(i, j) = 1 if RV_T(i,j) < low-resolution motion threshold (LRM_th). That is, in the static region, RV_T(i,j) < LRM_th, Mask_For_Loss(i, j) = 1; in the moving region, RV_T(i,j) ≥ LRM_th, and the loss value is not calculated.
[0198] The embodiments of the present application do not specifically limit Loss_Function_update. For example:
[0199] Loss_Function_update = Loss_Function[(Fusion_Cur * Mask_For_Loss), (Ground_Truth * Mask_For_Loss)](5)
[0200] In the cost estimation function provided by the embodiments of the present application, Loss_Function[(Fusion_Cur * Mask_For_Loss), (Ground_Truth * Mask_For_Loss)] refers to the loss function obtained by taking the difference after multiplying the fusion result Fusion_Cur and Ground_Truth by Mask_For_Loss respectively. Compared with the Loss_Function(Fusion_Cur, Ground_Truth) function, it realizes optimization in the static region and ensures that the control of the signal-to-noise ratio in the static region is not affected by the characteristics of the moving region.
[0201] Furthermore, the electronic device can continue to optimize Loss_Function_update to balance the signal-to-noise ratio of the static region and the moving region. This enables the first fusion frame to have a higher signal-to-noise ratio in the static region and better clarity in the moving region. Exemplarily:
[0202] Loss_Function_update=(1-α)Loss_Function[(Fusion_Cur*Mask_For_Loss), (Ground_Truth*Mask_For_Loss)]+αLoss_Function(Fusion_Cur, Ground_Truth)(6)
[0203] α is a random parameter, where 1 > α > 0.
[0204] In this embodiment, the image fusion network uses a reference frame as a benchmark to fuse the global image alignment sequence to obtain a first fused frame. The first fused frame is then input into a noise model for noise estimation. If the loss value of the cost estimation function is less than a preset loss threshold, the first fused frame is considered the first fused frame. If the loss value of the cost estimation function is not less than the preset loss threshold, the image parameters of the image fusion network are adjusted so that the loss value corresponding to the fused frame obtained by the adjusted image fusion network is less than the preset loss threshold.
[0205] In this embodiment, in areas where the motion region accounts for a large proportion, the quality of the captured image is significantly affected by the image quality of the motion region. To ensure better quality of the acquired image, a short preview frame is used as the reference frame. That is, using the short preview frame as the reference, the global image alignment sequence is fused to obtain the first fused frame.
[0206] In areas where the moving portion of the image is relatively small, the quality of the captured image is significantly affected by the quality of the stationary portion. To ensure better image quality, a long preview frame is used as the reference frame. Specifically, the global image alignment sequence is fused using the long preview frame as the reference to obtain the first fused frame. Because the long preview frame has a longer exposure time and a higher signal-to-noise ratio (SNR) in the stationary region, it effectively compensates for the insufficient SNR in the stationary region of the short frame when fusing long and short global alignment sequences, thus improving the SNR in the stationary region of the first fused frame.
[0207] In summary, by introducing a cost estimation function into the noise model and using the loss value of this function to measure the interpolation between the fusion result and the true value, a fusion frame with better fusion quality can be obtained. Furthermore, by introducing a fusion mask loss into the cost estimation function, the signal-to-noise ratio (SNR) of static regions can be precisely controlled during image fusion, thus helping to improve the SNR of static regions and the image quality of the fusion frame. Moreover, using a high-quality reference frame can further enhance the image quality of the fusion frame.
[0208] S450: Displays local motion estimation and obtains high-resolution motion data.
[0209] Display local motion estimation refers to the technique of estimating the motion of local regions in an image frame. In the embodiments of this application, display local motion estimation is used to compare the motion differences of two input image frames in the motion region to obtain high-resolution motion data.
[0210] In the embodiments of this application, the local motion estimation can be motion estimation based on optical flow, motion estimation based on feature point tracking, or motion estimation based on multi-scale pyramid block matching algorithm or dense optical flow tracing algorithm, etc. The embodiments of this application are not specifically limited.
[0211] In this embodiment, the electronic device is used to perform display local motion estimation on key frames and a first fused frame to obtain high-resolution motion data. The key frames are time-synchronized with the reference frame. If the reference frame is N, then the key frame is an S-frame (also known as a short-frame key frame); if the reference frame is an S-frame, then the key frame is an N-frame (also known as a long-frame key frame).
[0212] The following is in conjunction with the appendix Figure 13 ~Attached Figure 14 A detailed analysis will be conducted.
[0213] Appendix Figure 13 This is a flowchart of a secondary fusion processing method provided in an embodiment of this application. The electronic device performs display local motion estimation on the first fused frame and the long frame keyframe to obtain high-resolution motion data. Therefore, by accurately capturing and quantifying the motion differences between the long frame keyframe and the first fused frame in the motion region, the electronic device achieves more accurate compensation for motion differences between frames, reducing blurring and artifacts caused by inaccurate motion estimation.
[0214] Appendix Figure 14 This is a flowchart of another secondary fusion processing method provided in an embodiment of this application. The electronic device performs display local motion estimation on the first fused frame and the short frame keyframe to obtain high-resolution motion data. Due to the short frame exposure time, more details can be captured, but they may also be blurred due to motion. By estimating and compensating the actual local motion with the first fused frame, motion blur can be reduced while maintaining details in static areas.
[0215] In summary, the purpose of displaying local motion estimation is to obtain high-quality motion data, thereby improving the accuracy of motion compensation.
[0216] S460. Based on high-resolution motion data, the key preview frame and the first fused frame are fused a second time to obtain the second fused frame.
[0217] The second fused frame is obtained by fusing the keyframe and the first fused frame again. The first fused frame is obtained by fusing the reference frame with the image frame sequence, while the second fused frame is obtained by fusing the keyframe and the first fused frame again. The keyframe and the reference frame are time-synchronized long and short preview frame pairs. Therefore, the second fused frame obtained in this embodiment is obtained by fusing the image frame sequence with N frames from the preview frame sequence. That is, using long frames can compensate for the insufficient signal-to-noise ratio in the static area of short frames, resulting in a higher signal-to-noise ratio in the static area of the second fused frame. Furthermore, the second fused frame is obtained by motion compensation of S frames, and this fused frame has rich details and clarity in the moving area.
[0218] In one scenario, if frame S is the reference frame for the first fusion frame, indicating a high overall motion intensity in the scene, the electronic device can filter out stationary areas based on high-resolution motion data and optimize these stationary areas using long keyframes with high signal-to-noise ratios, as shown in the attached figure. Figure 13 As shown. The specific fusion method is as follows:
[0219] Step 1: Obtain the high-resolution motion threshold.
[0220] High-resolution motion threshold refers to a parameter that accurately distinguishes between moving and stationary regions in an image frame. The high-resolution motion threshold can be a fixed value set by those skilled in the art based on experience.
[0221] However, in snapshot mode, the image sensor gain value of an electronic device directly affects the noise level and contrast of the image, thus impacting the accuracy of distinguishing between moving and stationary areas. In this embodiment, the electronic device defines a high-resolution motion threshold as a dynamically changing value related to the image sensor gain value (Sensor_Gain).
[0222] In one example, the electronic device pre-constructs a motion threshold lookup table. The input to the motion threshold lookup table is the image sensor gain value, and the output is the high-resolution motion threshold. The motion threshold lookup table is a positive mapping mechanism that maps the image sensor gain value to the high-resolution motion threshold; as the image sensor gain value increases, the high-resolution motion threshold increases. The electronic device obtains `Sensor_Gain`, queries the motion threshold lookup table, and obtains the high-resolution motion threshold `LM_th`.
[0223] Step 2: Obtain a local motion mask based on high-resolution motion data and high-resolution motion threshold.
[0224] A local mask is a binary image that marks different regions of an image with different identifiers. An electronic device uses a first identifier to mark moving regions and a second identifier to mark stationary regions. For example, the first identifier can be "1" or "True", and the second identifier can be "0" or "False".
[0225] In this embodiment, the electronic device obtains a local motion mask based on the relationship between high-resolution motion quantity and high-resolution motion threshold (also known as the second motion intensity threshold). Specifically, for a pixel (i, j), where i is an integer greater than 0 and j is an integer greater than 0, if the high-resolution motion quantity (ELMV) is greater than or equal to the high-resolution motion threshold LM_th, the local motion mask information Local_Motion_Mask(i,j) = 1. If ELMV is less than LM_th, then Local_Motion_Mask(i,j) = 0.
[0226] That is: Local_Motion_Mask(i,j)=1if ELMV(i,j)≥LM_th(7)
[0227] Local_Motion_Mask(i,j) = 0 if ELMV(i,j)<LM_th (8)
[0228] Therefore, electronic devices can accurately distinguish between stationary and moving areas using local motion masks.
[0229] Step 3: Based on the local motion mask, obtain the fusion weights of the long frame keyframe and the first fusion frame.
[0230] The embodiments of this application may employ the following non-limiting method for obtaining the signal-to-noise ratio of the first fused frame and the long frame keyframe of the static region.
[0231] In one possible implementation, the electronic device can measure the structural similarity between different windows in an image frame based on the Structural Similarity Index (SSIM), and use this structural similarity as the signal-to-noise ratio (SNR). That is, the electronic device utilizes the principle that a lower SNR indicates greater noise and more significant damage to the image structure, thus accurately obtaining the SNR. Furthermore, the SSIM-based method provided in this application embodiment can still be used even when no noise is available, demonstrating good universality.
[0232] In one example, to determine the signal-to-noise ratio of pixel (i, j), the electronic device first divides the frame into multiple windows centered on pixel (i, j), each window being m*n in size. Here, m and n are values defined by those skilled in the art based on the serial port size; m > 0 and is an integer, and n > 0 and is an integer. For example, m = 9, n = 9. For ease of description, the electronic device refers to the m*n window at the pixel (i, j) position of the first fused frame as Fusion_Frame_Window, and the m*n window at the pixel (i, j) position of the long frame keyframe as Normal_Frame_Window.
[0233] The electronic device then calculates the SSIM (Signal-to-Noise Ratio) of the central m*n region and the surrounding m*n regions. Based on multiple SSIMs of the central m*n region and the surrounding m*n regions, the electronic device determines the SNR of pixel (i, j). For example, the electronic device determines the SNR of pixel (i, j) based on the distribution metrics of multiple SSIMs. For example, the electronic device takes the average of multiple SSIMs and uses this average as the SNR of pixel (i, j). Another example is the electronic device taking the weighted average of multiple SSIMs and using this weighted average as the SNR of pixel (i, j). Yet another example is the electronic device taking the median or mode value from multiple SSIMs and using the median or mode value as the SNR of pixel (i, j).
[0234] That is, the signal-to-noise ratio of the first fused frame:
[0235] Noise_Evaluation_Fu(i,j) = SSIM(Fusion_Frame_Window) (9)
[0236] Signal-to-noise ratio of long keyframes:
[0237] Noise_Evaluation_N(i,j) = SSIM(Normal_Frame_Window) (10)
[0238] In another possible implementation, the electronic device assesses the noise situation through transform domain evaluation and obtains the signal-to-noise ratio of the first fused frame and the long frame key frame.
[0239] For example, for Fusion_Frame_Window and Normal_Frame_Window, a Fourier transform is first performed to convert the spatial domain to the frequency domain, obtaining the first frequency domain corresponding to Fusion_Frame_Window and the second frequency domain corresponding to Normal_Frame_Window. The first and second frequency domains are complex arrays.
[0240] Calculate the amplitude in the first and second frequency domains. The amplitude is obtained by taking the square root of the sum of the squares of the real and imaginary parts of the complex number. Based on the amplitudes in the first and second frequency domains, obtain the first frequency energy corresponding to Fusion_Frame_Window and the second frequency energy corresponding to Normal_Frame_Window. The first frequency energy is the square of the amplitude in the first frequency domain, and the second frequency energy is the square of the amplitude in the second frequency domain.
[0241] Electronic devices identify noise based on frequency energy levels below a noise threshold (Frequency_Domain_Noise_Th). Frequency energy levels below Frequency_Domain_Noise_Th are considered to have a low signal-to-noise ratio (SNR), while those below are considered to have a high SNR.
[0242] In addition, the signal-to-noise ratio of the first fused frame and the long frame key frame can also be obtained through other means in the embodiments of this application, and the embodiments of this application are not specifically limited.
[0243] Electronic devices can acquire static regions based on local motion masks and determine fusion weights based on the signal-to-noise ratio of the first fused frame and the long frame keyframe of the static region.
[0244] In one example, after acquiring Noise_Evaluation_Fu and Noise_Evaluation_N, the electronic device, for a stationary region, if the signal-to-noise ratio of the long frame keyframe is greater than that of the first fused frame, uses the signal-to-noise ratio of the long frame keyframe to compensate for the insufficient signal-to-noise ratio of the first fused frame, thereby improving the signal-to-noise ratio of the first fused frame.
[0245] If Local_Motion_Mask(i,j) = 0, the current region is determined to be a stationary region.
[0246] For static regions, if Noise_Evaluation_Fu(i,j) > Noise_Evaluation_N(i,j), then the long frame keyframe is used to fuse the first fusion frame:
[0247] S_Fusion_Frame(i,j) =α* Norma_Fame(i,j) + (1-α) * Fusion_Frame(i,j) (11)
[0248] α is the fusion weight. In one example, α can be a fixed value set by those skilled in the art based on experience. In another example, α can be determined by the signal-to-noise ratio difference between the first fused frame and the long frame keyframe.
[0249] Exemplary illustration:
[0250] α=LUT_Fusion_Weigh(abs(Noise_Evaluation_N(i,j) - Noise_Evaluation_Fu(i,j)) (12)
[0251] LUT_Fusion_Weigh is a fusion weight lookup table. This lookup table is a pre-calibrated positive mapping mechanism between signal-to-noise ratio (SNR) differences and fusion weights; the greater the SNR difference, the greater the fusion weight.
[0252] Step 4: The first fused frame and the long frame keyframe are fused based on the fusion weight to obtain the second fused frame.
[0253] For specific fusion methods, please refer to formula (12), which will not be discussed here.
[0254] In this embodiment, the electronic device obtains a local motion mask based on high-resolution motion magnitude and a high-resolution motion threshold. Based on the local motion mask, it obtains fusion weights for a long-frame keyframe and a first fused frame. Based on these fusion weights, it performs image fusion to obtain a second fused frame. At this point, the long-frame keyframe can improve the signal-to-noise ratio (SNR) of the static region in the first fused frame, resulting in a higher SNR in the static region of the second fused frame.
[0255] In another case, if N frames in the preview frame sequence are the reference frames for the first fused frame, it indicates that the overall motion intensity of the scene is relatively low.
[0256] In one example, the electronic device can filter out motion regions based on high-resolution motion data, and then fuse short keyframes with a first fusion frame to obtain a second fusion frame. The second fusion frame then includes the high signal-to-noise ratio of the static region described in the first fusion frame. Furthermore, the rich motion region details of the short keyframes can be utilized to further enhance the image quality of the motion region in the fusion frame, resulting in a second fusion frame with both a high static region signal-to-noise ratio and good motion region details.
[0257] In one example, the electronic device stops performing secondary fusion and directly uses the first fused frame as the second fused frame to acquire the captured image. Since the first fused frame has a higher signal-to-noise ratio in the still area, the image quality of the captured image based on the first fused frame is relatively good.
[0258] This application provides an image generation method. When an electronic device enables Stagger HDR mode, in response to receiving an operation that triggers the electronic device to capture an image, a first frame and a second frame are obtained from a preview frame sequence in Stagger HDR mode. The preview frame sequence includes multiple latest preview frames cached during the preview process. The frame type of the second frame is different from that of the first frame, including: long frames in Stagger HDR mode and short frames in Stagger HDR mode. The first frame, the captured long frame sequence, and the short frame sequence are fused to obtain a first fused frame. The long frame sequence includes multiple captured long frames, and the short frame sequence includes multiple captured short frames. The second frame and the first fused frame are fused to obtain a second fused frame. Based on the second fused frame, the captured image is obtained. In other words, when electronic devices fuse short and long frame sequences acquired during shooting, they use the long and short frames in Stagger HDR mode as a benchmark. By leveraging the fact that the signal-to-noise ratio (SNR) of the long frames in stagger mode is much higher than that of the static area of the X-frame, the problem of insufficient SNR in the static area of the fused frame can be compensated for, thereby improving the SNR of the static area of the image fused frame. This solves the problem of image quality abnormalities such as various textures caused by insufficient SNR in the static area of the image fused frame, and improves the image quality of the captured image.
[0259] Furthermore, the rich motion region details of short frames in stagger mode can be utilized to improve the quality of motion regions in image fusion frames.
[0260] Furthermore, embodiments of this application also provide an electronic device, which, as an implementation, may have the following features: Figure 15 The structure of the electronic device 100 shown is illustrated. Figure 15 As shown, the electronic device 100 may include a processor 310, an external memory interface 320, an internal memory 321, a mobile communication module 350, a wireless communication module 360, a camera 340, a display screen 330, etc.
[0261] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0262] Processor 310 may include one or more processing units, such as: application processor (AP) [710A], modem, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0263] The processor 310 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 310 may be a cache memory. This memory can store instructions or data that the processor 310 has used or that are used frequently. If the processor 310 needs to use the instruction or data, it can directly retrieve it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 310, and thus improves the efficiency of the system.
[0264] Internal memory 321 can be used to store executable program code, including instructions. Internal memory 321 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 700 (such as audio data, phonebook, etc.). Furthermore, internal memory 321 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 310 executes various functional methods or data processing of electronic device 100 by running instructions stored in internal memory 321 and / or instructions stored in memory located in the processor.
[0265] The display screen 330 is used to display images, videos, etc. The display screen 330 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 700 may include one or more display screens 330.
[0266] Electronic device 100 can realize camera function through camera 340, ISP, video codec, GPU, display 330, AP, NPU, etc.
[0267] Camera 340 can be used to acquire color image data and depth data of the subject. An ISP can be used to process the color image data acquired by camera 340. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then transmitted to the ISP for processing, transforming it into a visible image. The ISP can also perform algorithmic optimizations on image noise, brightness, etc. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be integrated into camera 340.
[0268] In some embodiments, the electronic device 100 may include one or more cameras 340. Specifically, the electronic device 100 may include one front-facing camera 340 and one rear-facing camera 340. The front-facing camera 340 is typically used to capture images of the person facing the display screen 330, while the rear-facing camera 340 is used to capture images of the subject (such as a person, landscape, etc.) in front of the person.
[0269] The electronic device described in this application embodiment may also have a layered architecture. This layered architecture includes several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces.
[0270] For example, Figure 16 A schematic diagram of the composition of an electronic device 100 is shown. For example... Figure 11As shown, the layered architecture in this electronic device, from top to bottom, consists of the application layer, application framework layer, system library, hardware abstraction layer (HAL), kernel layer, and hardware layer.
[0271] The application layer can include a series of application packages.
[0272] like Figure 16 As shown, the application layer can include applications such as music, video, calls, ringtones, alarm clocks, Bluetooth, navigation, camera, and gallery. Of course, the application layer can also include other application packages; this application is not limited thereto. Among them, the camera application can provide users with photo-taking and image preview functions. The gallery application has the function of saving images, as well as providing users with functions such as viewing, sharing, and editing images.
[0273] The application framework layer provides application programming interfaces and programming frameworks for applications in the application layer. The camera application programming interface of the application framework is used to provide a camera server.
[0274] The kernel layer is the layer between hardware and software. The kernel layer includes at least Linux base drivers and device drivers, such as camera drivers.
[0275] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0276] Furthermore, embodiments of this application also provide a computer-readable storage medium storing instructions that, when executed on one or more computing devices, cause the one or more computing devices to perform the image generation method described in the above embodiments.
[0277] Furthermore, this application also provides a computer program product, which, when executed by one or more computing devices, allows the computing devices to execute any of the aforementioned image generation methods. This computer program product can be a software installation package; when any of the aforementioned image generation methods needs to be used, the computer program product can be downloaded and executed on a computer.
[0278] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0279] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0280] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0281] The system architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
Claims
1. An image generation method characterized by, The method is applied to an electronic device enabled with a stagger high dynamic range (Stagger HDR) mode, and the method comprises the following steps: In response to receiving an operation of triggering the electronic device to capture an image, a first frame and a second frame are obtained from a preview frame sequence in the Stagger HDR mode, the preview frame sequence comprising a plurality of latest preview frames cached in a preview process; the frame type of the second frame is different from the frame type of the first frame, and the frame type comprises a long frame of the Stagger HDR mode and a short frame of the Stagger HDR mode; Fusion is performed on the first frame, a long frame sequence and a short frame sequence captured and obtained, to obtain a first fusion frame; the long frame sequence comprises a plurality of long frames captured and obtained, and the short frame sequence comprises a plurality of short frames captured and obtained; The second fusion frame is obtained by fusing the second frame and the first fusion frame; The captured image is obtained based on the second fusion frame.
2. The method of claim 1, wherein, The method further comprises the following steps: A reference preview frame is determined from the preview frame sequence; If the proportion of the number of pixel points in a motion region in the reference preview frame to the total number of pixel points in the preview frame is greater than a preset proportion threshold, the long frame of the Stagger HDR mode is determined as the first frame, otherwise, the short frame of the Stagger HDR mode is determined as the first frame; The motion amount of the pixel point data in the motion region is greater than or equal to a first motion amount threshold.
3. The method of claim 2, wherein, The method further comprises the following steps: Based on a first size specification, the reference preview frame is divided into a plurality of image blocks, and the size specification of each image block in the plurality of image blocks is the same; For each image block, the following steps are respectively performed: determining the motion amount of at least one feature point data in the image block in the preview frame sequence, and determining the motion amount of the image block data based on the motion amount of the feature point data; The motion region of the reference preview frame is determined based on the motion amount of the plurality of image block data.
4. The method of claim 3, wherein, The determination of the motion amount of the image block data based on the motion amount of the feature point data comprises the following steps: The feature points in the adjacent image blocks of the image block, which satisfy a preset distance condition with the distance to the to-be-reconstructed pixel points in the image block, are taken as candidate feature points; Wherein, the relative positions of the to-be-reconstructed pixel points in each image block are the same, and the preset distance condition is that the distance to the to-be-reconstructed pixel points in the image block is the closest, the farthest, or a set distance; The motion amount of the to-be-reconstructed pixel point data in the image block is determined according to the motion amount of the candidate feature point data in the adjacent image blocks; The motion amount of the image block data is determined based on the motion amount of the to-be-reconstructed pixel point data in the image block.
5. The method of claim 4, wherein, The method further comprises the following steps: Intra-frame weights corresponding to the candidate feature points in the adjacent image blocks are determined; the intra-frame weights are in a negative correlation relationship with the distances from the candidate feature points in the adjacent image blocks to the to-be-reconstructed pixel points; The determination of the motion amount of the to-be-reconstructed pixel point data in the image block according to the motion amount of the candidate feature point data in the adjacent image blocks comprises the following steps: Determine the motion amount of the pixel point data to be reconstructed based on the product of the intra-frame weight and the motion amount of the candidate feature point data in the adjacent image block.
6. The method according to any one of claims 2 to 5, characterized in that, The method further comprises: Obtaining the motion amount and confidence of the feature point data in the adjacent preview frame of the reference preview frame; Determine the inter-frame weight of the feature point data in the adjacent preview frame based on the confidence of the feature point data in the adjacent preview frame, the inter-frame weight being positively correlated with the motion amount of the feature point data in the adjacent preview frame; Determine the motion amount of the feature point in the reference preview frame based on the product of the inter-frame weight and the motion amount of the feature point in the adjacent preview frame.
7. The method according to any one of claims 2 to 6, characterized in that, The fusion of the first frame, the long frame sequence and the short frame sequence obtained by shooting comprises: Perform global alignment processing on the long frame sequence and the short frame sequence after noise reduction processing; Take the first frame as a reference, and fuse the long frame sequence and the short frame sequence after global alignment processing to obtain the first fusion frame.
8. The method of claims 1-7, wherein, The fusion of the first frame, the long frame sequence and the short frame sequence obtained by shooting comprises: Input the first frame, the long frame sequence and the short frame sequence into a fusion network to obtain the first fusion frame, wherein the function value of the loss function during training of the fusion network represents the noise value size of the first fusion frame in the static region.
9. The method of claim 8, wherein, The function value of the loss function corresponds to the target difference data one by one, and the target difference data is the difference between the first data and the second data; The first data is the product of the pixel point data and the fusion mask value of the same pixel point in the fusion frame output by the fusion network during the training process, and the second data is the product of the data of multiple pixel points and the fusion mask value of the same pixel point in the standard frame of the fusion frame during the training process; The fusion mask value when the motion amount of the pixel point data is less than the first motion amount threshold is the first mask value; The fusion mask value when the motion amount of the pixel point data is greater than or equal to the first motion amount threshold is the second mask value.
10. The method according to any one of claims 1 to 9, characterized in that, If the frame type of the second frame is the long frame of the Stagger HDR mode, the fusion of the second frame and the first fusion frame comprises: Obtain the motion amount of the multiple pixel point data in the second frame relative to the same pixel point data in the first fusion frame; the pixel points corresponding to the motion amounts less than the second motion amount threshold in the multiple motion amounts are taken as the pixel points in the static region; If the signal-to-noise ratio metric value of the pixel points in the static region in the second frame is greater than the signal-to-noise ratio metric value of the same pixel points in the first fusion frame, obtain the second fusion frame, the data of the static region of the second fusion frame is the sum of the data of the static region in the second frame and the data of the static region in the first fusion frame; the signal-to-noise ratio metric value is used to measure the size of the signal-to-noise ratio.
11. The method of claim 10, wherein, Data of a static region of the second fusion frame is a sum of the first product data and the second product data, the first product data is a product of data of a static region of the second frame and a first fusion weight, and the second product data is a product of data of a static region of the first fusion frame and a second fusion weight; The first fusion weight and the second fusion weight are positively correlated with a first difference value, the first difference value is a difference between a signal-to-noise ratio of a pixel point of the static region of the second frame and a signal-to-noise ratio of the same pixel point in the first fusion frame.
12. The method of claim 10 or 11, wherein, For each pixel point in the second frame, the method further comprises: obtaining a first window from the second frame, the first window being an image block including the current pixel point and having a second size specification; determining a first structural similarity index (SSIM) of the first window and a neighboring window of the first window, the first window and the neighboring window of the first window having the same size specification; obtaining a signal-to-noise ratio metric value of the current pixel point in the second frame, the signal-to-noise ratio metric value of the current pixel point in the second frame being positively correlated with the first SSIM.
13. The method of claim 12, wherein, For each pixel point in the first fusion frame, the method further comprises: obtaining a second window from the first fusion frame, the second window being an image block including the current pixel point and having a third size specification; determining a second SSIM of the second window and a neighboring window of the second window, the second window and the neighboring window of the second window having the same size specification; obtaining a signal-to-noise ratio metric value of the current pixel point in the first fusion frame, the signal-to-noise ratio metric value of the current pixel point in the first fusion frame being positively correlated with the second SSIM.
14. The method of claim 10, wherein, The method further comprises: setting a fusion mask value of a pixel point corresponding to a motion value less than a second motion threshold to a first mask value, and setting a fusion mask value of a pixel point corresponding to a motion value greater than or equal to the second motion threshold to a second mask value; The pixel point corresponding to the motion value less than the second motion threshold is set as a static region pixel point.
15. The method of claim 12, wherein, If the frame type of the second frame is a short frame of the Stagger HDR mode, the fusing the second frame and the first fusion frame to obtain a second fusion frame comprises: obtaining a plurality of motion values of pixel point data in the second frame relative to the same pixel point data in the first fusion frame, and setting a pixel point corresponding to a motion value greater than or equal to a second motion threshold as a pixel point of a motion region; obtaining the second fusion frame, data of a motion region of the second fusion frame being a sum of data of a motion region of the second frame and data of a motion region of the first fusion frame.
16. The method of any one of claims 1-15, wherein, The method further comprises: fusing a plurality of short frames of the Stagger HDR mode in the preview frame sequence to obtain a short frame fusion frame, the short frame sequence including the short frame fusion frame.
17. The method of claim 16, wherein, The method further comprises: determine a long frame gain and a reduced exposure ratio, the reduced exposure ratio being a ratio of a reduced exposure amount to a previous exposure amount before adjustment; the fusing the short frames of the Stagger HDR mode in the preview frame sequence comprises: if the long frame gain is less than or equal to a first gain threshold and the reduced exposure ratio is less than a first ratio threshold, fusing short frames of the Stagger HDR mode in n1 consecutive frames in the preview frame sequence to obtain the short frame fusion frame, the n1 being an integer greater than or equal to 1.
18. The method of claim 17, wherein, The method further comprises: if the long frame gain is less than or equal to a second gain threshold, greater than the first gain threshold, and the reduced exposure ratio is greater than or equal to the first ratio threshold and less than a second ratio threshold, fusing short frames of the Stagger HDR mode in n2 consecutive frames in the preview frame sequence to obtain the short frame fusion frame, the n2 being an integer greater than n1.
19. The method of claim 18, wherein, The method further comprises: if the long frame gain is less than or equal to a third gain threshold and greater than the second gain threshold, and the reduced exposure ratio is greater than or equal to the second ratio threshold and less than a third ratio threshold, fusing short frames of the Stagger HDR mode in n3 consecutive frames in the preview frame sequence to obtain the short frame fusion frame, the n3 being an integer greater than n2.
20. The method of claim 17, wherein, In the process of fusing the short frames of the Stagger HDR mode in the preview frame sequence, the method further comprises: if the maximum gain of the short frames of the Stagger HDR mode in the preview frame sequence is less than a first short frame gain threshold, a first noise reduction algorithm is used for fusion noise reduction processing; if the maximum gain of the short frames of the Stagger HDR mode in the preview frame sequence is less than a second short frame gain threshold and greater than or equal to the first short frame gain threshold, a second noise reduction algorithm is used for fusion noise reduction processing; if the maximum gain of the short frames of the Stagger HDR mode in the preview frame sequence is less than a third short frame gain threshold and greater than or equal to the second short frame gain threshold, a third noise reduction algorithm is used for fusion noise reduction processing; wherein the third short frame gain threshold is greater than the second short frame gain threshold, which is greater than the first short frame gain threshold, the fusion noise reduction processing capability of the third noise reduction algorithm is better than that of the second noise reduction algorithm, and the fusion noise reduction processing capability of the second noise reduction algorithm is better than that of the first noise reduction algorithm.
21. The method of claim 20, wherein, The third noise reduction algorithm comprises a first sub-noise reduction algorithm for a static region and a second sub-noise reduction algorithm for a motion region. Wherein the noise reduction capability of the first sub-noise reduction algorithm for a static region is better than that of the second sub-noise reduction algorithm for a static region, and the noise reduction capability of the second multi-frame noise reduction algorithm for a motion region is better than that of the first multi-frame noise reduction algorithm for a motion region.
22. An electronic device, comprising: comprise a processor and a memory; wherein The memory is used to store a program; The processor is used to execute the program stored in the memory, and when the program stored in the memory is executed, the method of any one of claims 1 to 21 is executed.