Shooting method, electronic equipment, storage medium and chip system
By using multiple cameras to capture images with different exposure levels and performing image denoising and fusion processing, the quality problem caused by large differences between frames of night fireworks images was solved, improving image quality and user experience.
Patent Information
- Application Number
- CN202411126237.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-08-15
AI Technical Summary
When shooting fireworks images at night, due to insufficient light input to the night camera, each frame of the fireworks image requires a longer exposure time, resulting in large differences in spark content between frames, and causing large areas of fusion holes and false color problems in the fused fireworks image, affecting image quality.
Multiple cameras are used to simultaneously capture images with different exposures. Cameras with better photosensitivity capture images with smaller exposures, while cameras with poorer photosensitivity capture images with larger exposures. These images are then processed using an image denoising and fusion model to ensure that the image's field of view, size, and resolution are consistent. Finally, image fusion is performed.
It improves the overall quality of fireworks images, reduces inter-frame differences, improves image registration accuracy and fusion efficiency, shortens shooting waiting time, and enhances user experience.
Smart Images

Figure CN120751267A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to a shooting method, electronic equipment, storage medium and chip system. Background Art
[0002] In daily life, people sometimes use electronic devices such as mobile phones to capture images of fireworks. Because the burst of fireworks is fleeting, and the brightness contrast between the bright sparks and the dark night sky is quite pronounced, electronic devices typically use a single-camera, multi-frame, multi-exposure fusion capture strategy to preserve details in both the brightest and darkest parts of the fireworks image, thereby improving the overall quality of the fireworks image. This strategy involves sequentially capturing multiple frames of fireworks using a single camera, fusing these frames, and outputting the fused image.
[0003] However, due to insufficient light entering the image sensor in night cameras, a relatively long exposure time is required to ensure that each frame of fireworks captured by a single camera meets the required exposure. Since the sparks produced by fireworks move quickly, a long exposure time results in significant differences in the spark content between adjacent frames. Consequently, when fusing multiple frames of fireworks with significantly different spark content, the resulting fused fireworks image will contain large fusion holes, resulting in poor overall quality in the final fused fireworks image. Summary of the Invention
[0004] Embodiments of the present application provide a shooting method, electronic device, storage medium, and chip system, which can improve the overall quality of fireworks images shot by the electronic device.
[0005] In a first aspect, an embodiment of the present application provides a shooting method, comprising: upon receiving a shooting instruction and the scene to be shot in the shooting preview box is a target scene, shooting at least one group of images through multiple cameras; the target scene includes a fireworks scene; each group of images includes: multiple images with the same exposure time and different exposure amounts shot simultaneously by multiple cameras for the scene to be shot; in each group of images, images with a smaller exposure amount are shot by a camera with better photosensitivity among the multiple cameras, and images with a larger exposure amount are shot by a camera with a worse photosensitivity among the multiple cameras; different groups of images have different shooting times; each group of images is subjected to a first processing so that the corresponding field of view angles, sizes, and resolutions of all images in each group of images are the same; each group of images after the first processing is a group of images to be fused; multiple groups of images to be fused with adjacent shooting times are input into a trained image denoising fusion model for processing to obtain a fused image.
[0006] The electronic device receives a shooting instruction, which may include but is not limited to:
[0007] The electronic device detects a click or continuous press operation on a shooting control in a shooting preview interface;
[0008] Alternatively, the electronic device detects a quick shooting operation when displaying a shooting preview interface;
[0009] Alternatively, the electronic device receives a voice shooting instruction while displaying a shooting preview interface;
[0010] Alternatively, the electronic device receives a shooting instruction from another electronic device while displaying the shooting preview interface.
[0011] Optionally, after displaying the shooting preview interface, the electronic device may perform target recognition on the scene to be shot in the shooting preview frame in real time to determine whether the scene to be shot is the target scene.
[0012] Optionally, after receiving the shooting instruction, the electronic device may perform target recognition on the scene to be shot in the shooting preview frame to determine whether the scene to be shot is the target scene.
[0013] For example, the target scene may refer to a high-brightness scene that changes dynamically at night. For example, the target scene may include a fireworks scene or a scene of a light source (such as a car light or a rolling light sign) that moves rapidly at night.
[0014] It should be noted that better photosensitivity and worse photosensitivity are relative terms. Specifically, a camera's better photosensitivity may refer to the overall higher sensitivity of the image sensor in the camera to light. A camera's worse photosensitivity may refer to the overall lower sensitivity of the image sensor in the camera to light. For example, taking multiple cameras including a main camera and a telephoto camera as an example, since the overall sensitivity of the image sensor in the main camera to light is higher than the sensitivity of the image sensor in the telephoto camera to light, the main camera is the camera with better photosensitivity and the telephoto camera is the camera with worse photosensitivity. In addition, smaller exposure and larger exposure are also relative terms. For example, assuming that each group of images includes a first exposure image and a second exposure image, and the exposure of the first exposure image is less than that of the second exposure image, then the first exposure image in each group of images is the image with the smaller exposure, and the second exposure image is the image with the larger exposure.
[0015] In specific applications, the field of view angles of the multiple different cameras used to capture each set of images may be the same, partially the same, or different. In some embodiments, the multiple different cameras may include a main camera and a telephoto camera. In other embodiments, the multiple different cameras may include a main camera and a wide-angle camera. In still other embodiments, the multiple different cameras may include two main cameras with different performance. In still other embodiments, the multiple different cameras may include a telephoto camera, a main camera, and a wide-angle camera.
[0016] Since the exposure of an image is determined by the aperture value, exposure time, and sensitivity of the camera that captures the image, the electronic device can control the exposure of each image in each group of images to be different by the following methods, while ensuring that the exposure time of multiple images in each group of images is the same:
[0017] For an electronic device with a fixed aperture value and an adjustable ISO value of a camera, the electronic device can control the ISO value of a camera with better photosensitivity to be smaller than the ISO value of a camera with worse photosensitivity.
[0018] For an electronic device with an adjustable aperture value of a camera and a fixed ISO value, the electronic device can control the aperture value of a camera with better photosensitivity to be larger than the aperture value of a camera with worse photosensitivity.
[0019] For an electronic device in which both the aperture value and ISO value of a camera are adjustable, the electronic device can control the ISO value of a camera with better photosensitivity to be smaller than the ISO value of a camera with worse photosensitivity, and / or control the aperture value of a camera with better photosensitivity to be larger than the aperture value of a camera with worse photosensitivity.
[0020] The field of view angle corresponding to the image may refer to the field of view angle of the camera that captured the image. The field of view angle of the camera may refer to the maximum angle range that the camera can capture.
[0021] The dimensions of an image may refer to the actual size of the image in physical space.
[0022] The resolution of an image can be used to indicate how the pixels in the image are distributed across the width and height of the image.
[0023] According to the shooting method provided in the embodiment of the present application, since the same camera is only used to shoot images at the same exposure level, rather than for shooting multiple images at different exposure levels, the coherence of multiple frames of images at the same exposure level shot by the same camera can be improved, and the inter-frame differences of multiple frames of images at the same exposure level can be reduced, thereby reducing the difficulty of subsequent image registration. At the same time, since the images at different exposure levels are shot by different cameras, and the shooting time and exposure duration of the images at different exposure levels are the same, it can be ensured that the content of the images at different exposure levels is highly consistent, thereby further reducing the difficulty of subsequent image registration and improving the quality of the final fused image. Furthermore, since the images at different exposure levels are shot simultaneously by multiple cameras, rather than by a single camera, the shooting efficiency of the fused image can be improved, the user's shooting waiting time can be shortened, and the user's shooting experience can be enhanced.
[0024] In addition, the embodiment of the present application uses a camera with better photosensitivity to shoot images with a smaller exposure, and uses a camera with worse photosensitivity to shoot images with a larger exposure. This can simultaneously improve the signal-to-noise ratio of images with a smaller exposure and the signal-to-noise ratio of images with a larger exposure, thereby enabling the electronic device to subsequently fuse images with a smaller exposure and images with a larger exposure to produce a higher quality fused image, further improving the overall quality of fireworks images taken at night.
[0025] In an optional implementation of the first aspect, the image denoising fusion model includes multiple denoising units and an image fusion unit; the total number of the multiple denoising units is equal to the total number of the multiple cameras, and the multiple denoising units correspond one-to-one to the multiple cameras respectively; the denoising unit is used to fuse and denoise multiple images with the same exposure value taken by the corresponding camera; correspondingly, multiple groups of images to be fused at adjacent shooting moments are input into the trained image denoising fusion model for processing to obtain a fused image, including: through each denoising unit, fusing and denoising multiple images with the same exposure value taken by the camera corresponding to the denoising unit, and outputting a denoised image to the image fusion unit respectively; through the image fusion unit, fusing multiple denoised images with different exposure values from multiple denoising units to obtain a fused image.
[0026] According to the shooting method provided in the embodiment of the present application, since the image denoising fusion model is configured with multiple denoising units corresponding one-to-one to multiple cameras, and different denoising units are used to fuse and denoise multiple frames of images with the same exposure taken by different cameras, the image denoising fusion model can achieve accurate denoising of images with different exposures, which is beneficial to improving the overall quality of the fused image output by the image denoising fusion model.
[0027] In an optional implementation of the first aspect, different denoising units are trained through different training data sets; the training data set includes multiple sample data, each sample data includes multiple noisy images with noise and a real image without noise; the multiple noisy images are used as inputs of the denoising unit during training, and the real images are used as outputs of the denoising unit during training; correspondingly, the training data sets corresponding to the multiple denoising units are generated in the following manner: obtaining multiple high-definition night scene images; for each high-definition night scene image, randomly cutting out an area from the high-definition night scene image as the image to be degraded; performing random preprocessing on the image to be degraded; the random preprocessing includes: random white balance processing, random rotation, random addition of ghosting, random addition of color blocks, and random addition of jitter blur; the image to be degraded that has undergone random preprocessing Determine it as a real image; copy and obtain multiple copy images of the real image; the total number of the multiple copy images is equal to the total number of multiple cameras; each copy image corresponds to a different camera among the multiple cameras; perform a first noise addition process on the copy image corresponding to the large field of view camera among the multiple cameras to obtain a noise image corresponding to the large field of view camera; perform a second noise addition process on the copy image corresponding to the small field of view camera among the multiple cameras to obtain a noise image corresponding to the small field of view camera; the second noise addition process is different from the first noise addition process; use the noise image corresponding to the large field of view camera and the real image as a sample data of the denoising unit corresponding to the large field of view camera; use the noise image corresponding to the small field of view camera and the real image as a sample data of the denoising unit corresponding to the small field of view camera.
[0028] The high-definition night scene image may be an image of any night scene captured by a professional shooting device (such as a SLR camera or a compact camera, etc.). Exemplarily, the high-definition night scene image may be an image in RGB format.
[0029] In an optional implementation of the first aspect, a first noise addition process is performed on a copy image corresponding to a camera with a wide field of view among multiple cameras to obtain a noise image corresponding to the camera with a wide field of view, including: performing image quality blurring processing on the copy image corresponding to the camera with a wide field of view; performing sensitivity compensation on the copy image that has undergone image quality blurring processing; adding a first noise to the copy image that has undergone sensitivity compensation based on a noise coefficient of the camera with a wide field of view; and converting the copy image to which the first noise has been added into an image in a raw format to obtain a noise image corresponding to the camera with a wide field of view.
[0030] Among them, the raw format can be understood as the original format.
[0031] Exemplarily, the first noise may include a first Gaussian white noise and a first Poisson granular noise corresponding to a noise coefficient of a camera with a large field of view angle.
[0032] Exemplarily, the electronic device may perform a random downsampling operation on the copy image corresponding to the wide field of view camera, and a random upsampling operation on the copy image that has undergone the random downsampling operation, so as to add image blur introduced by image cropping to the copy image corresponding to the wide field of view camera.
[0033] The electronic device can multiply the pixel value of each pixel in the copy image corresponding to the wide field of view camera that has undergone image blur processing by a sensitivity difference value to achieve sensitivity compensation for the copy image corresponding to the wide field of view camera.
[0034] In an optional implementation of the first aspect, a second noise addition process is performed on a copy image corresponding to a small field of view angle camera among multiple cameras to obtain a noise image corresponding to the small field of view angle camera, including: brightening the copy image corresponding to the small field of view angle camera; adding a second noise to the brightened copy image based on the noise coefficient of the small field of view angle camera; converting the copy image to which the second noise is added into an image in a raw format to obtain a noise image corresponding to the small field of view angle camera.
[0035] Exemplarily, the second noise may include a second Gaussian white noise and a second Poisson granular noise corresponding to a noise coefficient of a camera with a small field of view angle.
[0036] Since the electronic device can obtain a piece of sample data corresponding to each denoising unit from each high-definition night scene image, a training data set corresponding to each denoising unit can be obtained from multiple high-definition night scene images.
[0037] According to the shooting method provided in the embodiment of the present application, since different denoising units in the image denoising fusion model are trained by training data sets that match the denoising capabilities they need to have, different denoising units can have different denoising capabilities, so that the image denoising fusion model can perform more accurate denoising on images taken by different cameras at different exposure levels, thereby improving the overall quality of the fused image output by the image denoising fusion model.
[0038] In an optional implementation of the first aspect, the first processing includes: cropping, scaling, and upsampling; correspondingly, the first processing is performed on each group of images, including: cropping the large field of view angle image in each group of images to obtain a cropped image corresponding to the large field of view angle image; the field of view corresponding to the cropped image is the same as the field of view angle corresponding to the small field of view angle image in the same group of images; the large field of view angle image refers to an image captured by a large field of view angle camera among multiple cameras, and the small field of view angle image refers to an image captured by a small field of view angle camera among multiple cameras; scaling the cropped image corresponding to the large field of view angle image in each group of images, A scaled image corresponding to the large field of view image is obtained; the size of the scaled image is the same as the size of the small field of view image in the same group of images, and the field of view corresponding to the scaled image is the same as the field of view corresponding to the small field of view image in the same group of images; the scaled image corresponding to the large field of view image in each group of images is upsampled to obtain an upsampled image corresponding to the large field of view image; the resolution of the upsampled image is the same as the resolution of the small field of view image in the same group of images, the size of the upsampled image is the same as the size of the small field of view image in the same group of images, and the field of view corresponding to the upsampled image is the same as the field of view corresponding to the small field of view image in the same group of images.
[0039] It should be noted that the terms "large field of view angle image" and "small field of view angle image" are relative. Specifically, a large field of view angle image may refer to an image captured by a camera with a larger field of view angle among multiple cameras, and a small field of view angle image may refer to an image captured by a camera with a smaller field of view angle among multiple cameras. For example, assuming that the first exposure image in each set of images is captured by the main camera and the second exposure image is captured by the telephoto camera, since the field of view angle of the main camera is larger than that of the telephoto camera, the first exposure image in each set of images is a large field of view angle image, and the second exposure image is a small field of view angle image.
[0040] According to the shooting method provided in the embodiment of the present application, by processing the large field of view angle image in each group of images so that the field of view angle, size and resolution corresponding to the small field of view angle image are consistent, the alignment accuracy during subsequent image alignment can be improved.
[0041] In an optional implementation of the first aspect, the first processing also includes: distortion correction; correspondingly, the first processing is performed on each group of images, and also includes: based on the calibrated internal parameters of the camera corresponding to each image in each group of images, distortion correction is performed on each image respectively; the calibrated internal parameters include distortion coefficient and focal length.
[0042] The calibrated internal parameters of the camera may refer to the pre-calibrated internal parameters of the camera.
[0043] For example, the internal parameters of the camera may include the distortion coefficient and focal length of the camera. The distortion coefficient of the camera may include a radial distortion coefficient and a tangential distortion coefficient. Among them, the radial distortion coefficient can be used to describe the degree to which each pixel in the image captured by the camera bends outward or inward as it is farther away from the center pixel. The tangential distortion coefficient can be used to describe the degree of deflection of each pixel in the image caused by the camera not being flush with the image plane. The focal length of the camera may include the focal length of the camera in the horizontal direction and the focal length in the vertical direction.
[0044] According to the shooting method provided in the embodiment of the present application, by performing distortion correction on each image in each group of images, the images shot by different cameras can be corrected to the same reference plane, thereby reducing or eliminating the distortion differences between the images in each group of images and improving the registration accuracy during subsequent image registration.
[0045] In an optional implementation of the first aspect, the first processing also includes: affine transformation; correspondingly, performing the first processing on each group of images also includes: performing an affine transformation on each image based on the calibrated extrinsic parameters of the camera corresponding to each image in each group of images; the calibrated extrinsic parameters are used to describe the position of the camera relative to the spatial coordinate system.
[0046] The camera's extrinsic parameters may refer to pre-calibrated external parameters of the camera. The camera's extrinsic parameters may be used to describe the camera's position, for example, the camera's position relative to the world coordinate system.
[0047] For example, the external parameters of the camera may include: the translation amount of the camera in the three degrees of freedom (ie, x-axis, y-axis, and z-axis) of the spatial coordinate system, and the rotation angle of the camera around the three degrees of freedom.
[0048] According to the shooting method provided in the embodiment of the present application, by performing an affine transformation on each image in each group of images, all images in each group of images can be aligned to a unified reference coordinate system, thereby reducing or eliminating the physical perspective difference between each image in each group of images and improving the registration accuracy during subsequent image registration.
[0049] In an optional implementation of the first aspect, the first processing may further include: inter-group registration. Correspondingly, performing the first processing on each group of images further includes: inter-group registration of multiple images in each group of images.
[0050] For example, the electronic device may use the image with the smaller exposure in each set of images as a reference frame and align the image with the larger exposure in the same set of images to the image with the smaller exposure, thereby achieving registration between multiple images in the same set of images. Aligning the image with the larger exposure to the image with the smaller exposure may include aligning each pixel in the image with the larger exposure with a corresponding pixel (i.e., pixels with identical content) in the image with the smaller exposure.
[0051] According to the shooting method provided in the embodiment of the present application, by performing inter-group registration on each group of images, the contents of images taken by different cameras at different exposure levels can be aligned as much as possible, thereby improving the subsequent image fusion effect.
[0052] In an optional implementation of the first aspect, the first processing may further include brightness alignment. Accordingly, performing the first processing on each group of images further includes calculating the exposure difference between every two images in each group of images, and adjusting the brightness of each two images based on the exposure difference between the two images, so that all images in each group of images have the same brightness.
[0053] According to the shooting method provided in the embodiment of the present application, by aligning the brightness of each image in each group of images, the brightness of all images in each group of images can be made consistent, which is conducive to improving the subsequent image fusion effect.
[0054] In an optional implementation of the first aspect, after performing the first processing on each group of images, the method further includes:
[0055] Perform inter-frame registration on all images in all groups of images.
[0056] According to the shooting method provided in the embodiment of the present application, by registering all images in all groups of images, all images can be further aligned, thereby further improving the subsequent image fusion effect.
[0057] In a second aspect, an embodiment of the present application provides an electronic device, comprising: one or more processors, and a memory;
[0058] The memory is coupled to one or more processors, and the memory is used to store computer program code, which includes computer instructions. One or more processors call the computer instructions to enable the electronic device to execute the shooting method of any implementation of the first aspect mentioned above.
[0059] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which includes instructions. When the instructions are executed on an electronic device, the electronic device executes a shooting method as described in any implementation of the first aspect above.
[0060] In a fourth aspect, an embodiment of the present application provides a computer executable program product. When the computer executable program product runs on an electronic device, the electronic device executes the shooting method of any implementation manner of the above-mentioned first aspect.
[0061] In a fifth aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to execute a shooting method such as any implementation method of the first aspect mentioned above.
[0062] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 A schematic diagram of the field of view angles corresponding to the main camera, wide-angle camera, and telephoto camera respectively;
[0064] Figure 2 A schematic diagram of a fireworks scene shot using a single-camera multi-frame multi-exposure fusion shooting strategy;
[0065] Figure 3 A schematic diagram of a fused fireworks image obtained based on a single-camera multi-frame multi-exposure fusion shooting strategy;
[0066] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0067] Figure 5 A schematic architecture diagram of a software system of an electronic device provided in an embodiment of the present application;
[0068] Figure 6 A schematic flow chart of a photographing method provided in an embodiment of the present application;
[0069] Figure 7 A schematic diagram of a shooting preview interface provided in an embodiment of the present application;
[0070] Figure 8 A schematic diagram of multiple groups of images captured by an electronic device based on the capturing method provided in an embodiment of the present application;
[0071] Figure 9 This is a schematic diagram of a specific implementation process of S603 in a shooting method provided in an embodiment of the present application;
[0072] Figure 10A schematic diagram of a processing process involved in a first processing performed on each group of images by an electronic device provided in an embodiment of the present application;
[0073] Figure 11 This is a schematic diagram of a specific implementation flow of S603 in a shooting method provided in another embodiment of the present application;
[0074] Figure 12 A schematic diagram of the structure of an image denoising fusion model provided in an embodiment of the present application;
[0075] Figure 13 A schematic diagram of the generation process of a training data set corresponding to a denoising unit in an image denoising fusion model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0076] It should be noted that the terms used in the implementation methods section of the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two, "at least one" and "one or more" mean one, two or more than two.
[0077] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly specifying the number of the technical features indicated. Therefore, the definition of "first" and "second" features may explicitly or implicitly include one or more of the features.
[0078] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0079] First, some of the terms used in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.
[0080] 1. The aperture is a component within a camera that controls the amount of light entering the image sensor. The size of the aperture can be expressed as an aperture value. Specifically, a smaller aperture value indicates a larger aperture, allowing more light to enter the image sensor; a larger aperture value indicates a smaller aperture, allowing less light to enter the image sensor.
[0081] 2. Shutter speed: This indicates the length of time the camera's shutter is open to allow light to enter the image sensor (i.e., exposure time). A faster shutter speed means a shorter exposure time; a slower shutter speed means a longer exposure time.
[0082] 3. Sensitivity, which indicates the sensitivity of the image sensor in the camera to light. The sensitivity can be expressed by a sensitivity value. For example, the sensitivity value can be specified by the International Organization for Standardization (ISO). The sensitivity value specified by ISO can be referred to as the ISO value. Specifically, the smaller the ISO value, the lower the sensitivity of the image sensor to light, which is more suitable for shooting scenes with sufficient light; the larger the ISO value, the higher the sensitivity of the image sensor to light, which is more suitable for shooting scenes with insufficient light.
[0083] 4. Exposure: This refers to the amount of light actually received by the camera's image sensor, which directly affects image brightness. Specifically, a moderate exposure results in a balanced image. Excessive exposure (overexposure) results in an image that is often too bright. Underexposure (underexposure) results in an image that is often too dark.
[0084] The exposure can be determined by the aperture value, shutter speed, and ISO value. For example, the exposure can be represented by the exposure value (EV). Specifically, a larger EV value indicates a smaller exposure; a smaller EV value indicates a larger exposure. The relationship between the EV value and the aperture value, shutter speed, and ISO value can be expressed by the following formula (1):
[0085]
[0086] Among them, EV represents exposure value, N represents aperture value, t represents exposure time, and ISO represents ISO value.
[0087] According to the above formula (1), when the exposure time and ISO value remain unchanged, the larger the aperture value (i.e., the smaller the aperture), the larger the exposure value and the smaller the exposure amount; the smaller the aperture value (i.e., the larger the aperture), the smaller the exposure value and the larger the exposure amount. When the aperture value and ISO value remain unchanged, the shorter the exposure time, the larger the exposure value and the smaller the exposure amount; the longer the exposure time, the smaller the exposure value and the larger the exposure amount. When the aperture value and exposure time remain unchanged, the smaller the ISO value, the larger the exposure value and the smaller the exposure amount; the larger the ISO value, the smaller the exposure value and the larger the exposure amount.
[0088] 5. Field of view (FOV) is used to indicate the maximum angle range that a camera can capture. When the object to be photographed is within this angle range, the object to be photographed can be captured by the camera; when the object to be photographed is outside this angle range, the object to be photographed will not be captured by the camera. Generally, the larger the field of view of a camera, the larger the shooting range and the shorter the focal length; the smaller the field of view of a camera, the smaller the shooting range and the longer the focal length. Therefore, according to the different field of view angles, cameras can be divided into main cameras, wide-angle cameras, and telephoto cameras.
[0089] For example, see Figure 1 , a schematic diagram showing the corresponding field of view angles for the main camera, wide-angle camera, and telephoto camera, respectively. As can be seen, the wide-angle camera's field of view is larger than that of the main camera. Therefore, the wide-angle camera is more suitable for capturing close-up shots than the main camera. The telephoto camera's field of view is smaller than that of the main camera. Therefore, the telephoto camera is more suitable for capturing distant shots than the main camera. The main camera is more suitable for capturing scenes that are moderately close or far away.
[0090] The above is a brief introduction to some of the terms involved in the embodiments of this application, which will not be repeated below.
[0091] In daily life, people sometimes celebrate holidays by setting off fireworks and use electronic devices such as mobile phones to capture the beautiful moment of fireworks blooming in the night sky. Because the blooming process of fireworks is fleeting, and the brightness contrast between the bright sparks produced by the fireworks and the dark night sky is quite obvious, in order to ensure that the fireworks images captured by electronic devices at night can preserve the details of both the brightest and darkest parts, thereby expanding the dynamic range of the fireworks images and improving the overall quality of the fireworks images, electronic devices usually use a single-camera multi-frame multi-exposure fusion shooting strategy when capturing dynamic, high-brightness scenes like fireworks at night.
[0092] The specific shooting strategy of single-camera multi-frame multi-exposure fusion is as follows: for the same scene to be shot, multiple frames of scene images of the scene to be shot at different exposure values are sequentially shot with the same camera, and the multiple frames of scene images at different exposure values are aligned and fused, and the fused image is output. Among them, different exposure values are achieved by controlling the exposure time to be different and keeping the values of other parameters affecting the exposure value (such as aperture value and ISO value) the same. For example, please refer to Figure 2 , is a schematic diagram of shooting fireworks scenes based on a single camera multi-frame multi-exposure fusion shooting strategy. Figure 2 As shown, when capturing fireworks images, the electronic device can use a telephoto camera to sequentially capture fireworks images of a fireworks scene at a first exposure level and a second exposure level, and then register and fuse the multiple frames of fireworks images captured at the first and second exposure levels to generate a fused fireworks image. The first exposure level is greater than the second exposure level, i.e., the first exposure level is the larger exposure level and the second exposure level is the smaller exposure level. The specific values of the first and second exposure levels can be set based on actual conditions.
[0093] However, when electronic devices capture fireworks images at night based on the aforementioned shooting strategy, the image sensor in the night camera receives insufficient light. Therefore, to ensure that the exposure of each frame of the fireworks image meets the required level, each frame requires a relatively long exposure time. The sparks produced by fireworks move quickly, and a long exposure time results in large differences in the spark content between two adjacent frames of fireworks images (i.e., large inter-frame differences). Consequently, when multiple frames of fireworks images shot sequentially at different exposure levels are fused together, the resulting fused fireworks image will have large fusion holes and a poor high dynamic range (HDR) effect, resulting in poor overall quality of the final fused fireworks image.
[0094] For example, see Figure 3 , is a schematic diagram of a fused fireworks image obtained based on a single camera multi-frame multi-exposure fusion shooting strategy. It should be noted that, Figure 3 The fireworks image under the first exposure and the fireworks image under the second exposure are two adjacent frames of images taken in sequence by the same camera for the same fireworks scene. Figure 3It can be seen that the spark content of the fireworks image under the first exposure amount and the fireworks image under the second exposure amount is quite different. Specifically, taking the overexposed area 30 in the fireworks image under the first exposure amount as an example, the area corresponding to the overexposed area 30 in the fireworks image under the second exposure amount should be bright, but the corresponding area in the fireworks image under the second exposure amount actually taken is dark. This will cause the overexposed area 30 in the fireworks image under the first exposure amount to be unable to be aligned with the bright area in the fireworks image under the second exposure amount, that is, it is impossible to find a bright area in the fireworks image under the second exposure amount that can be matched with the overexposed area 30 to enrich the details of the overexposed area 30, resulting in an overexposure anomaly similar to a hole in the fused fireworks image after fusion (i.e., a fusion hole). In addition, in some scenes, pseudo-color problems may also occur in the fused fireworks image. For example, if Figure 3 As shown, the pseudo color problem refers to that the color of the sparks in the fused fireworks image changes compared to the color of the sparks in the original image, resulting in the color of the sparks in the fused fireworks image being inconsistent with the color of the sparks in the actual scene.
[0095] Based on this, related technologies typically address issues such as fusion holes and false colors in fused fireworks images by reducing the exposure time of fireworks images. Specifically, they use the shortest possible exposure time to ensure consistency between the content of each adjacent fireworks frame. However, due to the rapid movement of sparks produced by fireworks, simply reducing the exposure time cannot fundamentally address the large differences between fireworks image frames. Furthermore, shorter exposure times result in a lower signal-to-noise ratio in the captured fireworks images, leading to poor overall quality of the fused fireworks images.
[0096] In view of this, an embodiment of the present application provides a shooting method, comprising: upon receiving a shooting instruction and the scene to be shot in the shooting preview box is a target scene, shooting at least one group of images through multiple cameras; the target scene includes a fireworks scene; each group of images includes: multiple images with the same exposure time and different exposure amounts shot simultaneously by multiple cameras for the scene to be shot; in each group of images, images with a smaller exposure amount are shot by a camera with better photosensitivity among the multiple cameras, and images with a larger exposure amount are shot by a camera with a worse photosensitivity among the multiple cameras; different groups of images have different shooting times; each group of images is subjected to a first processing so that the corresponding field of view angles, sizes, and resolutions of all images in each group of images are the same; each group of images after the first processing is a group of images to be fused; multiple groups of images to be fused with adjacent shooting times are input into a trained image denoising and fusion model for processing to obtain a fused image.
[0097] In the embodiment of the present application, since the same camera is only used to capture images at the same exposure level, rather than for capturing images at multiple different exposure levels, the coherence of multiple frames of images at the same exposure level captured by the same camera can be improved, and the inter-frame differences of multiple frames of images at the same exposure level can be reduced, thereby reducing the difficulty of subsequent image registration. At the same time, since the images at different exposure levels are captured by different cameras, and the shooting time and exposure duration of the images at different exposure levels are the same, it is possible to ensure that the content of the images at different exposure levels is highly consistent, thereby further reducing the difficulty of subsequent image registration and improving the quality of the final fused image. Furthermore, since the images at different exposure levels are captured simultaneously by multiple cameras, rather than by a single camera, the shooting efficiency of the fused image can be improved, the user's shooting waiting time can be shortened, and the user's shooting experience can be enhanced.
[0098] In addition, the embodiment of the present application uses a camera with better photosensitivity to shoot images with a smaller exposure, and uses a camera with worse photosensitivity to shoot images with a larger exposure. This can simultaneously improve the signal-to-noise ratio of images with a smaller exposure and the signal-to-noise ratio of images with a larger exposure, thereby enabling the electronic device to subsequently fuse images with a smaller exposure and images with a larger exposure to produce a higher quality fused image, further improving the overall quality of fireworks images taken at night.
[0099] The shooting method provided in the embodiments of the present application can be applied to various electronic devices. For example, the electronic devices may include single-lens reflex cameras, compact cameras and other camera devices, mobile phones, tablet computers, wearable devices, augmented reality (AR) devices, virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not limit the specific types of electronic devices.
[0100] The structure of the electronic device is described below by taking a mobile phone as an example.
[0101] See also Figure 4 , is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. Figure 4As shown, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180I, a touch sensor 180J, a bone conduction sensor 180K, an ambient light sensor 180L, etc.
[0102] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0103] In specific applications, the electronic device can realize the shooting function through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0104] Among them, the camera 193 can be used to capture images. Exemplarily, the camera 193 may include components such as a lens, a filter, and an image sensor. When the camera 193 captures an image, the light emitted or reflected by the object enters the lens of the camera 193, passes through the filter, and finally converges on the image sensor. The image sensor is mainly used to converge and image the light emitted or reflected by all objects within the field of view of the camera 193 (also referred to as the scene to be captured). The filter is mainly used to filter out unnecessary light waves in the light (for example, light waves other than visible light, such as infrared light waves).
[0105] Specifically, the image sensor can be used to convert the received light into an electrical signal and transmit the electrical signal to the ISP. The ISP can be used to convert the electrical signal into a digital image signal and transmit the digital image signal to the DSP. The DSP can be used to convert the digital image signal into an image signal in a standard three primary colors (red, green, blue, RGB) or YUV format, where Y represents brightness and U and V represent chrominance. In some embodiments, the ISP can be set in the camera 193.
[0106] For example, the image sensor may be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor, etc.
[0107] In some embodiments, the camera 193 can be disposed on the front or back of the electronic device. The specific number and arrangement of the cameras 193 are not particularly limited in this embodiment of the application.
[0108] For example, the electronic device may include a front camera and a rear camera. Each of the front camera and the rear camera may include one or more cameras. For example, if the electronic device includes three rear cameras, where the three rear cameras are a main camera, a telephoto camera, and a wide-angle camera, the electronic device may employ the shooting method provided in the embodiments of the present application when simultaneously activating at least two rear cameras or simultaneously activating at least two front cameras for shooting.
[0109] It is understandable that Figure 4 The illustrated structure does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0110] The software system of the electronic device may adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture, etc. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the software system of the electronic device.
[0111] See also Figure 5 , which is a schematic architecture diagram of a software system of an electronic device provided in an embodiment of the present application.
[0112] like Figure 5 As shown in Figure 1, a layered architecture divides software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. For example, a layered architecture can divide the Android system into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0113] The application layer can include a series of application packages, such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0114] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer may include some predefined functions.
[0115] Exemplarily, the application framework layer may include a window manager, a content provider, a telephony manager, a resource manager, a notification manager, a view system, and the like.
[0116] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0117] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0118] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0119] The system library can include multiple functional modules, such as a surface manager, a 3D graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), and media libraries.
[0120] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, sensor driver, etc.
[0121] It should be noted that Figure 5 Only modules related to the embodiments of the present application are shown. In other embodiments, each layer may also include any other possible modules, and each module may also include one or more sub-modules, which are not limited in this application.
[0122] The following is a detailed introduction to the shooting method provided in the embodiments of the present application in conjunction with the accompanying drawings.
[0123] See also Figure 6 , is a schematic flow chart of a shooting method provided in an embodiment of the present application. The shooting method can be applied to Figure 4 The hardware structure shown and Figure 5 In the electronic device with the software architecture shown in FIG, taking the electronic device as a mobile phone as an example, as shown in FIG. Figure 6 As shown, the shooting method provided in the embodiment of the present application may include S601 to S604, which are detailed as follows:
[0124] S601: The electronic device displays a shooting preview interface.
[0125] The electronic device may have a camera application installed. The electronic device may display a shooting preview interface upon receiving a camera startup instruction. The camera startup instruction may be used to instruct the camera application to be started.
[0126] For example, the electronic device may receive a camera activation instruction in the following situations, but is not limited to:
[0127] Scenario 1.1: The electronic device detects a first operation on an icon of a camera application.
[0128] Exemplarily, the first operation may be a click operation or a long press operation.
[0129] In specific applications, the camera application icon can be included in multiple display interfaces of the electronic device. The multiple display interfaces may include, for example, the desktop, the negative first screen, the control center, the lock screen, or the interface of other applications. Other applications may refer to any application other than the camera application in the electronic device, such as a social application or an image processing application.
[0130] Based on this, for example, the electronic device may determine that a camera startup instruction has been received when it detects that the icon of the camera application on the desktop has been clicked. For another example, the electronic device may determine that a camera startup instruction has been received when it detects that the icon of the camera application on the negative one screen has been clicked. For another example, the electronic device may determine that a camera startup instruction has been received when it detects that the icon of the camera application in the control center has been clicked. For another example, the electronic device may determine that a camera startup instruction has been received when it detects that the icon of the camera application on the lock screen interface has been clicked or long pressed. For another example, the electronic device may determine that a camera startup instruction has been received when it detects that the icon of the camera application in another application interface has been clicked.
[0131] Scenario 1.2: The electronic device detects a camera quick launch operation.
[0132] For example, the camera quick launch operation may include but is not limited to: sliding upward from the lower left corner or lower right corner of the lock screen interface; or pressing the same volume button twice in a row in the lock screen state, etc.
[0133] Based on this, for example, the electronic device can determine that a camera activation instruction has been received when it detects that a user swipes up from the lower left corner or lower right corner of the lock screen. For another example, the electronic device can determine that a camera activation instruction has been received when it detects that the same volume button is pressed twice in succession while the screen is locked.
[0134] Scenario 1.3: The electronic device receives a voice command for launching a camera application.
[0135] For example, the voice command for starting the camera application may be “Xiaoyi, open the camera.” Based on this, the electronic device may determine that a camera start instruction has been received when receiving the voice command “Xiaoyi, open the camera.”
[0136] For example, see Figure 7 , is a schematic diagram of a shooting preview interface provided by an embodiment of the present application. Figure 7 As shown in (a) of FIG, the shooting preview interface may include a shooting preview box 701 and a shooting control 702. Figure 7 As shown in (b) of FIG. 7 , the shooting preview box 701 can be used to display a preview image 7011. The preview image 7011 can include a scene to be shot captured by the camera of the electronic device, such as a fireworks scene. The shooting control 702 can be used to control the electronic device to start or stop shooting, etc.
[0137] S602: When the electronic device receives a shooting instruction and the scene to be shot in the shooting preview box is a target scene, the electronic device shoots at least one set of images through multiple cameras.
[0138] For example, the electronic device may receive a shooting instruction in the following situations, but is not limited to:
[0139] Scenario 2.1: The electronic device detects a second operation on the shooting control in the shooting preview interface.
[0140] For example, the second operation may be a click operation or a continuous press operation, etc. Based on this, the electronic device may determine that a shooting instruction is received when detecting that the shooting control in the shooting preview frame is clicked or continuously pressed.
[0141] Scenario 2.2: When the electronic device displays a shooting preview interface, it detects a quick shooting operation.
[0142] For example, quick shooting operations may include, but are not limited to: pressing a preset volume button, performing a preset air gesture, or clicking anywhere in the shooting preview box. The preset volume button may be a volume down button or a volume up button. Preset air gesture operations may include an air grab gesture or an air V-sign gesture.
[0143] In specific applications, the quick shooting operation can be set by the user according to his or her own needs.
[0144] Based on this, for example, if a user sets the action of pressing the volume down key as a quick capture key, the electronic device may determine that a capture instruction has been received when the volume down key is detected to be pressed while the capture preview interface is displayed. For another example, if a user sets the action of pressing the volume up key as a quick capture key, the electronic device may determine that a capture instruction has been received when the volume up key is detected to be pressed while the capture preview interface is displayed. For another example, if a user sets the air grab gesture as a quick capture action, the electronic device may determine that a capture instruction has been received when the air grab gesture is detected while the capture preview interface is displayed. For another example, if a user sets the air V-sign gesture as a quick capture action, the electronic device may determine that a capture instruction has been received when the air V-sign gesture is detected while the capture preview interface is displayed. For another example, if a user sets the action of clicking anywhere in the capture preview box as a quick capture action, the electronic device may determine that a capture instruction has been received when the capture preview interface is displayed and a click is detected anywhere in the capture preview box.
[0145] Scenario 2.3: The electronic device receives a voice shooting command while displaying a shooting preview interface.
[0146] For example, the voice shooting command may be “take a photo” or “take a video”, etc.
[0147] Based on this, the electronic device can determine that a shooting instruction has been received when it receives a voice shooting instruction "take a photo" or "take a video" while displaying the shooting preview interface.
[0148] Scenario 2.4: When the electronic device displays a shooting preview interface, it receives a shooting instruction from another electronic device.
[0149] The other electronic device may be any electronic device that has established a communication connection with the electronic device. For example, the other electronic device may include a smartwatch or smart bracelet. The communication connection may include a wired connection or a wireless connection. For example, a wired connection may include a connection via a USB cable. For example, a wireless connection may include a Bluetooth connection or a Wi-Fi connection.
[0150] For example, other electronic devices may be provided with a shooting button. Based on this, when the shooting button on the other electronic device is pressed, the other electronic device may send a shooting instruction to the electronic device.
[0151] In some embodiments, after displaying the shooting preview interface, the electronic device may perform target recognition on the scene to be shot in the shooting preview frame in real time to determine whether the scene to be shot is the target scene.
[0152] In other embodiments, after receiving the shooting instruction, the electronic device may perform target recognition on the scene to be shot in the shooting preview frame to determine whether the scene to be shot is the target scene.
[0153] For example, the electronic device may use a target recognition algorithm to identify whether the scene to be photographed is the target scene. It should be noted that the specific recognition process of the target recognition algorithm can be referred to the description in the relevant technology and will not be described in detail here.
[0154] For example, the target scene may refer to a high-brightness scene that changes dynamically at night. For example, the target scene may include a fireworks scene or a scene of a light source (such as a car light or a rolling light sign) that moves rapidly at night.
[0155] Taking the target scene as a fireworks scene as an example, the electronic device can capture at least one set of images and execute subsequent S603 to S604 when receiving a shooting instruction and recognizing that the scene to be captured in the shooting preview box is a fireworks scene.
[0156] Each set of images may include multiple images. The multiple images in each set of images may be images captured simultaneously by an electronic device using multiple different cameras, each capturing a scene with the same exposure time and different exposure levels. In other words, the multiple images in each set of images may be captured at the same time, with the same exposure time and different exposure levels.
[0157] For example, when capturing each set of images, the electronic device may use a camera with better light sensitivity among the multiple cameras to capture images with smaller exposures, and use a camera with worse light sensitivity among the multiple cameras to capture images with larger exposures. That is, within each set of images, the images with smaller exposures may be captured by the camera with better light sensitivity, and the images with larger exposures may be captured by the camera with worse light sensitivity.
[0158] It should be noted that better photosensitivity and worse photosensitivity are relative terms. Specifically, a camera's better photosensitivity may refer to the overall higher sensitivity of the image sensor in the camera to light. A camera's worse photosensitivity may refer to the overall lower sensitivity of the image sensor in the camera to light. For example, taking multiple cameras including a main camera and a telephoto camera as an example, since the overall sensitivity of the image sensor in the main camera to light is higher than the sensitivity of the image sensor in the telephoto camera to light, the main camera is the camera with better photosensitivity and the telephoto camera is the camera with worse photosensitivity. In addition, smaller exposure and larger exposure are also relative terms. For example, assuming that each group of images includes a first exposure image and a second exposure image, and the exposure of the first exposure image is less than that of the second exposure image, then the first exposure image in each group of images is the image with the smaller exposure, and the second exposure image is the image with the larger exposure.
[0159] It should also be noted that different groups of images are captured at different times. For example, a preset time interval may be provided between the capture times of two adjacent groups of images. That is, the electronic device may capture a group of images at intervals of a preset time interval. The preset time interval may be set based on actual needs. For example, the preset time interval may be 0.1 seconds.
[0160] In specific applications, the field of view angles of the multiple different cameras used to capture each set of images may be the same, partially the same, or different. For example, in some embodiments, the multiple different cameras may include a main camera and a telephoto camera. In other embodiments, the multiple different cameras may include a main camera and a wide-angle camera. In still other embodiments, the multiple different cameras may include two main cameras with different performance. In still other embodiments, the multiple different cameras may include a telephoto camera, a main camera, and a wide-angle camera.
[0161] For example, see Figure 8 , is a schematic diagram of multiple groups of images captured by an electronic device based on the shooting method provided in an embodiment of the present application. Figure 8As shown, taking a case where multiple different cameras include a main camera and a telephoto camera, the electronic device shoots a first exposure amount image through the main camera and shoots a second exposure amount image through the telephoto camera, and the exposure amount of the first exposure amount image is less than the exposure amount of the second exposure amount image as an example, when the electronic device receives a shooting instruction and recognizes that the scene to be shot in the shooting preview box is the target scene, it can simultaneously shoot a group of images (i.e., Figure 8 The first exposure image and the second exposure image in each dotted box are a group of images), thereby obtaining at least one group of images.
[0162] It is understandable that since the exposure of an image is determined by the aperture value, exposure time, and sensitivity of the camera that captures the image, the electronic device can control the exposure of each image in each group of images to be different by the following methods, while ensuring that the exposure time of multiple images in each group of images is the same:
[0163] Method 1: For an electronic device with a fixed aperture value and an adjustable ISO value of a camera, the electronic device can control the ISO value of a camera with better photosensitivity to be smaller than the ISO value of a camera with worse photosensitivity.
[0164] For example, taking multiple cameras including a main camera and a telephoto camera, the main camera is used to shoot an image with a first exposure amount, the telephoto camera is used to shoot an image with a second exposure amount, and the exposure amount of the first exposure amount image is less than the exposure amount of the second exposure amount image, the electronic device can control the exposure amount of the first exposure amount image to be less than the exposure amount of the second exposure amount image by controlling the ISO value of the main camera to be less than the ISO value of the telephoto camera.
[0165] Method 2: For an electronic device with an adjustable camera aperture value and a fixed ISO value, the electronic device can control the aperture value of a camera with better photosensitivity to be larger than the aperture value of a camera with worse photosensitivity.
[0166] For example, taking multiple cameras including a main camera and a telephoto camera, the main camera is used to shoot a first exposure image, the telephoto camera is used to shoot a second exposure image, and the exposure of the first exposure image is less than the exposure of the second exposure image, the electronic device can control the exposure of the first exposure image to be less than the exposure of the second exposure image by controlling the aperture value of the main camera to be greater than the aperture value of the telephoto camera.
[0167] Method 3: For an electronic device in which both the aperture value and ISO value of a camera are adjustable, the electronic device can control the ISO value of a camera with better photosensitivity to be smaller than the ISO value of a camera with worse photosensitivity, and / or control the aperture value of a camera with better photosensitivity to be larger than the aperture value of a camera with worse photosensitivity.
[0168] For example, taking multiple cameras including a main camera and a telephoto camera, the main camera is used to shoot an image with a first exposure amount, the telephoto camera is used to shoot an image with a second exposure amount, and the exposure amount of the first exposure amount image is less than the exposure amount of the second exposure amount image, the electronic device can control the exposure amount of the first exposure amount image to be less than the exposure amount of the second exposure amount image by controlling the ISO value of the main camera to be less than the ISO value of the telephoto camera, and / or controlling the aperture value of the main camera to be greater than the aperture value of the telephoto camera.
[0169] It is understandable that since the telephoto camera is more prone to insufficient light input than the main camera, the embodiment of the present application uses the main camera with better photosensitivity to shoot a first exposure image with a smaller exposure, and uses the telephoto camera with worse photosensitivity to shoot a second exposure image with a larger exposure. This can simultaneously improve the signal-to-noise ratio of the first exposure image and the signal-to-noise ratio of the second exposure image, thereby enabling the electronic device to subsequently fuse the first exposure image and the second exposure image to produce a higher quality fused image.
[0170] In specific applications, when an electronic device uses multiple different cameras to capture each group of images, in order to ensure that the shooting times of multiple images in each group of images are the same, the electronic device can use software synchronization control or hardware synchronization control to control the multiple cameras to shoot simultaneously.
[0171] In a specific implementation, the electronic device controls multiple cameras to shoot simultaneously using software synchronization control, which may include: the electronic device sends shooting instructions to the multiple cameras simultaneously through a processor. This ensures that the multiple cameras can receive the shooting instructions at the same time and shoot simultaneously.
[0172] In another specific implementation, a hardware synchronization circuit may be provided between multiple cameras of an electronic device. The hardware synchronization circuit is configured to control the multiple cameras to simultaneously perform a capture operation upon receiving a capture instruction. Based on this, the electronic device using hardware synchronization control to control the multiple cameras to simultaneously capture may include: the electronic device, via a processor, issuing a capture instruction to the hardware synchronization circuit, thereby triggering the hardware synchronization circuit to control the multiple cameras to simultaneously perform a capture operation.
[0173] S603, the electronic device performs a first processing on each group of images so that all images in each group of images have the same field of view angle, the same size, and the same resolution; each group of images after the first processing is an image group to be fused.
[0174] It can be understood that since the multiple images in each group of images may be taken by the electronic device through multiple cameras with the same field of view angle, or may be taken by the electronic device through multiple cameras with partially the same field of view angle, or may be taken by the electronic device through multiple cameras with different field of view angles, therefore, in order to ensure that the field of view angle corresponding to each image subsequently used for image fusion is the same, the size is the same, and the resolution is the same, so as to improve the quality of the fused image, in some embodiments, when the field of view angles corresponding to the multiple images in each group of images are not exactly the same, the electronic device may perform a first processing on each group of images so that the field of view angles corresponding to the multiple images in each group of images are the same, the size is the same, and the resolution is the same.
[0175] In a specific implementation, the first processing may include: cropping, scaling, and upsampling. Figure 9 As shown, the electronic device performs a first process on each group of images, which may include S6031 to S6033, as detailed below:
[0176] S6031, the electronic device crops the large field-of-view angle image in each group of images to obtain a cropped image corresponding to the large field-of-view angle image; the field-of-view angle corresponding to the cropped image is the same as the field-of-view angle corresponding to the small field-of-view angle image.
[0177] It should be noted that the terms "large field of view angle image" and "small field of view angle image" are relative. Specifically, a large field of view angle image may refer to an image captured by a camera with a larger field of view angle among multiple cameras, and a small field of view angle image may refer to an image captured by a camera with a smaller field of view angle among multiple cameras. For example, assuming that the first exposure image in each set of images is captured by the main camera and the second exposure image is captured by the telephoto camera, since the field of view angle of the main camera is larger than that of the telephoto camera, the first exposure image in each set of images is a large field of view angle image, and the second exposure image is a small field of view angle image.
[0178] For example, see Figure 10 , is a schematic diagram of the processing process involved when an electronic device according to an embodiment of the present application performs a first processing on each group of images. Figure 10As shown, assuming that each group of images includes a first exposure image and a second exposure image, the first exposure image is taken by the main camera, and the second exposure image is taken by the telephoto camera, that is, the first exposure image is a large field of view image, and the second exposure image is a small field of view image; and the shadow portion in the first exposure image is the overlapping portion of the field of view corresponding to the first exposure image and the field of view corresponding to the second exposure image, that is, the field of view corresponding to the shadow portion in the first exposure image is the same as the field of view corresponding to the second exposure image, then the electronic device can crop the first exposure image in each group of images according to the field of view corresponding to the second exposure image in each group of images to crop out the shadow portion from the first exposure image, and the shadow portion is the cropped image corresponding to the first exposure image, and the field of view corresponding to the cropped image is the same as the field of view corresponding to the second exposure image.
[0179] S6032, the electronic device scales the cropped image corresponding to the large field of view angle image in each group of images to obtain a scaled image corresponding to the large field of view angle image; the size of the scaled image is the same as the size of the small field of view angle image, and the field of view corresponding to the scaled image is the same as the field of view angle corresponding to the small field of view angle image.
[0180] The size of an image may refer to the actual size of the image in physical space. For example, the size of an image may be represented by the width and height of the image. For example, the size of an image may be 25.6×34.14 centimeters, which may be used to indicate that the width of the image is 25.6 centimeters and the height is 34.14 centimeters.
[0181] It is understandable that since the original sizes of the images in each group are the same, and the size of the cropped image obtained after the electronic device crops the large field of view image is smaller than the size of the small field of view image, to facilitate subsequent image fusion, the electronic device can scale the cropped image corresponding to the large field of view image so that the size of the scaled cropped image is the same as the size of the small field of view image in the same group. For ease of explanation, the embodiments of the present application describe the scaled cropped image as the scaled image corresponding to the large field of view image.
[0182] Specifically, the electronic device scales the cropped image corresponding to the large field of view image, which may include: enlarging the width of the cropped image to be equal to the width of the small field of view image, and enlarging the height of the cropped image to be equal to the height of the small field of view image.
[0183] For example, please refer to Figure 10Assuming that each set of images includes a first exposure image and a second exposure image, the first exposure image is captured by the main camera, and the second exposure image is captured by the telephoto camera, the electronic device may crop a cropped image corresponding to the first exposure image in each set of images and then scale the cropped image corresponding to the first exposure image to obtain a scaled image corresponding to the first exposure image. The scaled image has the same size as the second exposure image, and the scaled image has the same field of view as the second exposure image.
[0184] S6033, the electronic device upsamples the scaled image corresponding to the large field of view angle image in each group of images to obtain an upsampled image corresponding to the large field of view angle image; the resolution of the upsampled image is the same as the resolution of the small field of view angle image, the size of the upsampled image is the same as the size of the small field of view angle image, and the field of view corresponding to the upsampled image is the same as the field of view corresponding to the small field of view angle image.
[0185] The resolution of an image can be used to indicate the distribution of pixels in the image across its width and height. For example, the resolution of an image can be 1920×1080 pixels, which indicates that the image has 1920 pixels in the width direction (i.e., horizontal direction) and 1080 pixels in the height direction (i.e., vertical direction).
[0186] It is understandable that although the size of the scaled image corresponding to the large field of view image in each group of images is the same as the size of the small field of view image in the same group, since the cropped image is cropped from the large field of view image, the resolution of the cropped image will be smaller than the resolution of the small field of view image. The scaled image is simply a stretch of the width and height of the cropped image, and does not increase the resolution of the cropped image. Therefore, it will further cause the resolution of the scaled image to be smaller than the resolution of the small field of view image. Based on this, in order to facilitate subsequent image fusion, the electronic device can upsample the scaled image corresponding to the large field of view image in each group of images so that the resolution of the upsampled scaled image is the same as the resolution of the small field of view image. For ease of explanation, the embodiment of the present application describes the upsampled scaled image as the upsampled image corresponding to the large field of view image.
[0187] For example, the electronic device may use an interpolation method to upsample the scaled image. The interpolation method may include nearest neighbor interpolation, bilinear interpolation, and Lagrange interpolation. It should be noted that the specific operation process of the interpolation method can be found in the description of the relevant art and will not be described in detail here.
[0188] For example, the electronic device may perform upsampling processing on the scaled image based on super-resolution technology. It should be noted that the specific operation process of the super-resolution technology can be referred to the description in the related art and will not be described in detail here.
[0189] It is understood that for ease of explanation, Figure 8 As shown, in the embodiment of the present application, each group of images that has undergone the first processing can be described as a group of images to be fused.
[0190] It can also be understood that in specific applications, due to differences in the manufacturing process and optical characteristics of different cameras, different distortions (i.e., geometric distortion or distortion, etc.) exist between the images taken by different cameras. Therefore, in order to align the images taken by different cameras to the same plane to reduce or eliminate the distortion differences between the images taken by different cameras, in another specific implementation method, the first processing can also include: distortion correction.
[0191] Based on this, Figure 11 As shown, before S6031, S603 may further include S6034, which is described in detail as follows:
[0192] S6034: The electronic device performs distortion correction on each image based on the calibrated internal parameters of the camera corresponding to each image.
[0193] The camera corresponding to the image may refer to the camera used to capture the image.
[0194] The calibrated internal parameters of a camera may refer to the pre-calibrated internal parameters of the camera.
[0195] For example, the internal parameters of the camera may include the distortion coefficient and focal length of the camera. The distortion coefficient of the camera may include a radial distortion coefficient and a tangential distortion coefficient. Among them, the radial distortion coefficient can be used to describe the degree to which each pixel in the image captured by the camera bends outward or inward as it is farther away from the center pixel. The tangential distortion coefficient can be used to describe the degree of deflection of each pixel in the image caused by the camera not being flush with the image plane. The focal length of the camera may include the focal length of the camera in the horizontal direction and the focal length in the vertical direction.
[0196] For example, S6034 may specifically include steps 1.1 to 1.2, which are described in detail as follows:
[0197] In step 1.1, the electronic device determines the intrinsic parameter matrix of the camera corresponding to each image based on the coordinates of the central pixel point of each image and the focal length of the camera corresponding to each image.
[0198] Among them, the intrinsic parameter matrix of the camera can be used to describe the intrinsic characteristics of the camera.
[0199] For example, the intrinsic parameter matrix of the camera corresponding to each image can be expressed as:
[0200]
[0201] Among them, K is the intrinsic parameter matrix of the camera, f x Indicates the horizontal focal length of the camera, f y Indicates the focal length of the camera in the vertical direction, c x Represents the horizontal coordinate of the center pixel of the image, c y Indicates the vertical coordinate of the center pixel of the image.
[0202] In step 1.2, the electronic device performs distortion correction on each image based on the intrinsic parameter matrix and distortion coefficient of the camera corresponding to each image, thereby obtaining a distortion-corrected image corresponding to each image.
[0203] For example, the electronic device can calculate the degree of distortion (i.e., offset) of each pixel in each image based on the distortion coefficient of the camera corresponding to each image; and determine the new coordinates of each pixel in each image based on the original coordinates and the degree of distortion; and determine the actual position of each pixel in the world coordinate system based on the new coordinates of each pixel in each image and the intrinsic parameter matrix of the camera corresponding to each image, thereby obtaining the distortion-corrected image corresponding to each image.
[0204] The embodiment of the present application corrects the distortion of each image in each group of images, so that the images taken by different cameras can be corrected to the same reference plane, thereby reducing or eliminating the distortion differences between the images in each group of images and improving the registration accuracy during subsequent image registration.
[0205] It can also be understood that in specific applications, the installation positions and / or installation angles of multiple cameras of an electronic device cannot completely overlap, which will result in the images taken by different cameras being unable to be aligned to the same reference coordinate system. Therefore, in order to reduce the perspective differences between different images caused by the different installation positions and / or installation angles of different cameras, in another specific implementation method, the first processing can also include: affine transformation.
[0206] Based on this, please continue to refer to Figure 11 Before S6031, S603 may further include S6035, which is described in detail as follows:
[0207] S6035: The electronic device performs an affine transformation on each image based on the calibrated extrinsic parameters of the camera corresponding to each image.
[0208] The camera's extrinsic parameters may refer to pre-calibrated external parameters of the camera. The camera's extrinsic parameters may be used to describe the camera's position, for example, the camera's position relative to the world coordinate system.
[0209] For example, the external parameters of the camera may include: the translation amount of the camera in the three degrees of freedom (ie, x-axis, y-axis, and z-axis) of the spatial coordinate system, and the rotation angle of the camera around the three degrees of freedom.
[0210] For example, S6035 may specifically include steps 2.1 to 2.4, which are described in detail as follows:
[0211] In step 2.1, the electronic device selects a number of feature points from each image in each set of images.
[0212] For example, the aforementioned feature points may be corner points, edge points, or other significant feature points of the image.
[0213] For example, the electronic device may select a number of feature points from each image using a feature point detection algorithm, such as a scale-invariant feature transform (SIFT) algorithm or a speeded-up robust features (SURF) algorithm.
[0214] In step 2.2, the electronic device matches the feature points of different images in each group of images using a feature point matching algorithm to obtain multiple feature point pairs.
[0215] Exemplarily, the feature matching algorithm may include but is not limited to a K nearest neighbors (KNN) algorithm.
[0216] In step 2.3, the electronic device determines an affine transformation matrix based on the coordinates of the plurality of feature point pairs.
[0217] For example, the electronic device may determine an affine transformation matrix based on the coordinates of the plurality of feature point pairs using an affine transformation function. The affine transformation function may include, for example, the cv2.getAffineTransform() function in the OpenCV library. For details about the cv2.getAffineTransform() function, reference may be made to the description in the related art and will not be described in detail here.
[0218] For example, the affine transformation matrix can be expressed as:
[0219]
[0220] Among them, M is the affine transformation matrix, a 11 Indicates the scaling of the pixel in the image on the x-axis, a 22 Indicates the scaling of the pixel in the image on the y-axis, a 12 Indicates the amount of cropping of pixels in the image on the x-axis, a 21 Indicates the amount of cropping of pixels in the image on the y-axis, t x Indicates the translation of the pixel in the image on the x-axis, t y Indicates the translation of the pixel in the image on the y-axis.
[0221] In step 2.4, the electronic device performs an affine transformation on each image in each group of images based on the affine transformation matrix to obtain an affine transformed image corresponding to each image.
[0222] For example, the electronic device may process each image in each group of images using the following formula (2) based on the affine transformation matrix:
[0223]
[0224] Among them, (x i ,y i ) is the coordinate of the i-th pixel in each image, (x i ',y i ') is the coordinate of the i-th pixel in each image after affine transformation.
[0225] The embodiment of the present application performs an affine transformation on each image in each group of images, so that all images in each group of images can be aligned to a unified reference coordinate system, thereby reducing or eliminating the physical perspective difference between the images in each group of images and improving the registration accuracy during subsequent image registration.
[0226] It is also understood that in specific applications, when an electronic device uses software synchronization control to control multiple cameras to capture images simultaneously, there may still be millimeter-level errors between the capture times of different cameras. Therefore, in order to ensure that the content of images captured by different cameras in the same image set is aligned as much as possible, the first processing may also include: inter-group registration. Inter-group registration can refer to the registration of multiple frames in the same image set.
[0227] Based on this, please continue to refer to Figure 11 Before S6031, S603 may further include S6036, which is described in detail as follows:
[0228] S6036: The electronic device performs inter-group registration on multiple images in each group of images.
[0229] For example, the electronic device may use the image with the smaller exposure in each set of images as a reference frame and align the image with the larger exposure in the same set of images to the image with the smaller exposure, thereby achieving registration between multiple images in the same set of images. Aligning the image with the larger exposure to the image with the smaller exposure may include aligning each pixel in the image with the larger exposure with a corresponding pixel (i.e., pixels with identical content) in the image with the smaller exposure.
[0230] It should be noted that the above S6034, S6035 and S6036 may be located before S6031, that is, the electronic device may first execute S6034 and / or S6035 and / or S6036, and then execute S6031 to S6033.
[0231] The embodiment of the present application performs inter-group registration on each group of images, so that the contents of images taken by different cameras at different exposure levels can be aligned as much as possible, thereby improving the subsequent image fusion effect.
[0232] In yet another specific implementation, the first processing may further include: brightness alignment.
[0233] Based on this, please continue to refer to Figure 11 After S6033, S603 may further include S6037, which is described in detail as follows:
[0234] S6037, the electronic device calculates the exposure value difference between each two images in each group of images, and adjusts the brightness of each two images based on the exposure difference between each two images, so that the brightness of all images in each group of images is the same.
[0235] It should be noted that S6037 may be located after S6031 to S6033, that is, the electronic device may first execute S6031 to S6033 and then execute S6037.
[0236] The embodiment of the present application aligns the brightness of each image in each group of images, so that the brightness of all images in each group of images is consistent, which is beneficial to improving the subsequent image fusion effect.
[0237] In some other embodiments of the present application, the electronic device may further perform a second process on all the image groups. For example, the second process may include: inter-frame registration. The inter-frame registration may refer to registering all images in all the image groups.
[0238] Based on this, after S603 and before S604, the shooting method may further include step 3.1, which is described in detail as follows:
[0239] In step 3.1, the electronic device performs inter-frame registration on all images in all groups of images.
[0240] Exemplarily, the electronic device may use an image in a group of images as a reference image, calculate the registration matrix between other images except the reference image and the reference image, and align each other image based on the registration matrix between each other image and the reference image to achieve registration between all images in all groups of images.
[0241] The embodiment of the present application can further align all images by registering all images in all groups of images, thereby further improving the subsequent image fusion effect.
[0242] Based on this, in some embodiments, when the shooting method does not include the above step 3.1, the electronic device can directly determine each group of images that have undergone the first processing as a group of images to be fused.
[0243] In other embodiments, when the shooting method includes the above step 3.1, the electronic device may determine each group of images that have undergone the first processing and the second processing as a group of images to be fused.
[0244] S604: The electronic device inputs a plurality of image groups to be fused that are adjacent in shooting time into a trained image denoising and fusion model for processing to obtain a fused image.
[0245] Among them, the image denoising fusion model may include multiple denoising units and an image fusion unit. The total number of the multiple denoising units may be equal to the total number of multiple cameras used by the electronic device when shooting multiple groups of images, and the multiple denoising units may correspond one to one to the multiple cameras respectively. Exemplarily, the input of each denoising unit may be multiple images with the same exposure taken by its corresponding camera, and the output of each denoising unit may be connected to the image fusion unit. Based on this, each denoising unit can be used to fuse and denoise multiple images with the same exposure taken by its corresponding camera, and output a denoised image with the corresponding exposure to the fusion unit. The fusion unit can be used to fuse multiple denoised images with different exposures from multiple denoising units to obtain a fused image.
[0246] Based on this, S604 may specifically include steps 4.1 and 4.2, which are described in detail as follows:
[0247] In step 4.1, the electronic device uses each denoising unit in the image denoising fusion model to fuse and denoise multiple images with the same exposure value taken by the camera corresponding to each denoising unit, and outputs a denoised image with the corresponding exposure value to the image fusion unit.
[0248] In step 4.2, the electronic device fuses the denoised images with different exposure amounts from the multiple denoising units through the image fusion unit to obtain a fused image.
[0249] For ease of understanding, the following example uses each set of images including a first exposure image and a second exposure image, where the first exposure image is taken by the main camera and the second exposure image is taken by the telephoto camera, to illustrate the specific structure of the image denoising fusion model. Figure 12 , is a structural diagram of an image denoising fusion model provided in an embodiment of the present application.
[0250] like Figure 12 As shown, the image denoising and fusion model may include a first denoising unit 1201, a second denoising unit 1202, and an image fusion unit 1203. The first denoising unit 1201 may correspond to the main camera, and the second denoising unit 1202 may correspond to the telephoto camera. The first denoising unit 1201 may be specifically configured to fuse and denoise multiple first-exposure images captured by the main camera, obtaining a denoised image corresponding to the first-exposure image, and output the denoised image corresponding to the first-exposure image to the fusion unit 1203. The second denoising unit 1202 may be specifically configured to fuse and denoise multiple second-exposure images captured by the telephoto camera, obtaining a denoised image corresponding to the second-exposure image, and output the denoised image corresponding to the second-exposure image to the fusion unit 1203. The fusion unit 1203 may be specifically configured to fuse the denoised image corresponding to the first-exposure image with the denoised image corresponding to the second-exposure image to obtain a fused image.
[0251] Based on this, the electronic device can fuse and denoise multiple first exposure images taken by the main camera through the first denoising unit 1201, and output a denoised image corresponding to the first exposure image to the image fusion unit 1203; and can fuse and denoise multiple second exposure images taken by the telephoto camera through the second denoising unit 1202, and output a denoised image corresponding to the second exposure image to the image fusion unit 1203; and can fuse the denoised image corresponding to the first exposure image and the denoised image corresponding to the second exposure image through the image fusion unit 1203 to obtain a fused image.
[0252] For example, the structure of each denoising unit in the image denoising fusion model can adopt a convolutional neural network structure based on deep learning. For example, the convolutional neural network structure can include one or more convolution blocks (transformer blocks) based on the attention mechanism and one or more ordinary convolution blocks (conv).
[0253] The image fusion unit in the image denoising fusion model can adopt the HDR multi-frame fusion module in the related art. For the specific content of the HDR multi-frame fusion module, please refer to the description in the related art and will not be described in detail here.
[0254] It is understandable that, since different denoising units in the image denoising fusion model are used to denoise images taken by different cameras, and the noise coefficients of different cameras are different, the noise carried by the images taken by different cameras will be different. Therefore, different denoising units need to have different denoising capabilities. Based on this, in specific applications, when training the image denoising fusion model, different denoising units can be trained using different training data sets. For example, Figure 12 Taking the image denoising fusion model shown as an example, the first denoising unit 1201 can be trained using a first training data set, and the second denoising unit 1202 can be trained using a second training data set.
[0255] That is to say, different denoising units in the image denoising fusion model are trained using different training data sets. Exemplarily, the training data set corresponding to each denoising unit may include multiple sample data. Each sample data may include multiple noise images carrying noise and one real image without noise. The multiple noise images may be obtained by adding noise to the real image. Based on this, in a specific application, when training the image denoising fusion model, the electronic device may use the multiple noise images in each sample data corresponding to each noise unit as the input of the noise unit, and use a real image in each sample data as the output of the noise unit, and train each noise unit separately, so that each noise unit learns different denoising capabilities.
[0256] For example, see Figure 13 , is a schematic diagram of the generation process of a training data set corresponding to a denoising unit in an image denoising fusion model provided in an embodiment of the present application. Figure 13 As shown, the electronic device can generate a training data set corresponding to each denoising unit through S1301 to S1308, as detailed below:
[0257] S1301, the electronic device obtains multiple high-definition night scene images.
[0258] The high-definition night scene image may be an image of any night scene captured by a professional shooting device (such as a SLR camera or a compact camera, etc.). Exemplarily, the high-definition night scene image may be an image in RGB format.
[0259] S1302 : For each high-definition night scene image, the electronic device randomly cuts out an area from the high-definition night scene image as an image to be degraded.
[0260] S1303: The electronic device performs random preprocessing on the image to be degraded.
[0261] For example, random pre-processing may include, but is not limited to: random white balance processing, random rotation, random addition of ghosting, random addition of color blocks, and random addition of jitter blur, etc.
[0262] The purpose of the electronic device performing random white balance processing on the image to be degraded is to randomly modify the color temperature of the image to be degraded. For example, the electronic device can use any white balance coefficient to perform white balance processing on the image to be degraded.
[0263] It should be noted that operations such as rotating, sharpening, adding ghosting, adding color blocks, and adding jitter blur to images are all commonly used image processing methods in the field of image processing. For the specific processing procedures of these image processing methods, please refer to the descriptions in the relevant technologies and will not be described in detail here.
[0264] S1304: The electronic device determines the image to be degraded after random preprocessing as a real image.
[0265] S1305: The electronic device copies and obtains multiple copies of the real image.
[0266] The total number of the multiple duplicate images may be equal to the total number of multiple cameras used by the electronic device to capture multiple sets of images. The multiple duplicate images may correspond one-to-one to the multiple cameras, respectively.
[0267] S1306: The electronic device performs a first noise addition process on the copy image corresponding to the camera with a wide field of view angle to obtain a noise image corresponding to the camera with a wide field of view angle.
[0268] The term "wide field of view camera" may refer to a camera with a larger field of view among multiple cameras. It is understood that images captured by a wide field of view camera may be referred to as wide field of view images. For example, the first exposure image captured by the main camera may be referred to as a wide field of view image.
[0269] Exemplarily, the noise addition process corresponding to the first noise addition process may include: image quality blurring processing, sensitivity compensation, adding first noise corresponding to a camera with a large field of view, and format conversion, etc.
[0270] Based on this, S1306 may specifically include S13061 to S13064, which are described in detail as follows:
[0271] S13061, the electronic device performs image blurring processing on the copy image corresponding to the camera with a wide field of view.
[0272] It is understandable that since the electronic device will crop the wide field of view image in the aforementioned S6031, there will be image blur in the wide field of view image due to image cropping. Based on this, in order to more realistically simulate the image input into the noise unit corresponding to the wide field of view camera, the electronic device can perform image blur processing on the copy image corresponding to the wide field of view camera to add the image blur introduced by image cropping.
[0273] Exemplarily, the electronic device may perform a random downsampling operation on the copy image corresponding to the wide field of view camera, and a random upsampling operation on the copy image that has undergone the random downsampling operation, so as to add image blur introduced by image cropping to the copy image corresponding to the wide field of view camera.
[0274] At S13062, the electronic device performs sensitivity compensation on the copy image that has been blurred.
[0275] It is understandable that since different cameras have different sensitivities, in order to simulate the sensitivity differences between different cameras, the electronic device can perform sensitivity compensation on the blurred copy image based on the sensitivity difference between different cameras.
[0276] Exemplarily, the electronic device may multiply the pixel value of each pixel in the copy image that has undergone image blurring by a sensitivity difference value to achieve sensitivity compensation for the copy image corresponding to the camera with a wide field of view.
[0277] S13063, the electronic device adds a first noise to the sensitivity-compensated copy image based on the noise coefficient of the wide-field-of-view camera.
[0278] The first noise may include a first Gaussian white noise and a first Poisson granular noise corresponding to the noise coefficient of the camera with a large field of view angle.
[0279] S13064: The electronic device converts the copy image with the first noise added thereto into a raw format image to obtain a noise image corresponding to the camera with a wide field of view.
[0280] Among them, the raw format can also be understood as the original format.
[0281] It can be understood that since the high-definition night scene image is an image in RGB format, the real image obtained from the high-definition night scene image is also an image in RGB format, and the copy image of the real image is also an image in RGB format. In specific applications, the image input into the image noise fusion model is an image in the original format (i.e., raw format) taken by the camera. Therefore, the electronic device needs to convert the copy image corresponding to the wide field of view camera with the first noise added into a raw format image. The raw format image is the noise image corresponding to the wide field of view camera.
[0282] S1307: The electronic device performs a second noise addition process on the copy image corresponding to the camera with a small field of view angle to obtain a noise image corresponding to the camera with a small field of view angle.
[0283] The term "small field of view camera" may refer to a camera with a smaller field of view among multiple cameras. It is understood that an image captured by a small field of view camera may be referred to as a small field of view image. For example, an image with a second exposure value captured by a telephoto camera may be referred to as a small field of view image.
[0284] The noise addition process corresponding to the second noise addition process is different from the noise addition process corresponding to the first noise addition process.
[0285] Exemplarily, the noise addition process corresponding to the second noise addition processing may include: exposure brightening, adding second noise corresponding to a camera with a small field of view angle, and format conversion, etc.
[0286] Based on this, S1307 may specifically include S13071 to S13073, which are described in detail as follows:
[0287] S13071, the electronic device brightens the copy image corresponding to the small field of view camera.
[0288] Exemplarily, the electronic device may perform random brightening processing on the copy image corresponding to the small field-of-view camera to simulate overexposure in the image taken by the small field-of-view camera.
[0289] S13072: The electronic device adds a second noise to the brightened copy image based on the noise coefficient of the small field-of-view camera.
[0290] The second noise may include a second Gaussian white noise and a second Poisson granular noise corresponding to the noise coefficient of the camera with a small field of view angle.
[0291] S13073, the electronic device converts the copy image with the second noise added thereto into a raw format image to obtain a noise image corresponding to the camera with a small field of view angle.
[0292] It can be understood that since the high-definition night scene image is an image in RGB format, the real image obtained from the high-definition night scene image is also an image in RGB format, and the copy image of the real image is also an image in RGB format. In specific applications, the image input into the image noise fusion model is an image in the original format (i.e., raw format) taken by the camera. Therefore, the electronic device needs to convert the copy image corresponding to the small field of view camera with the second noise added into a raw format image. The raw format image is the noise image corresponding to the small field of view camera.
[0293] S1038 uses the noise image and the real image corresponding to the large field of view camera obtained from each high-definition night scene image as a sample data of the denoising unit corresponding to the large field of view camera; and uses the noise image and the real image corresponding to the small field of view camera obtained from each high-definition night scene image as a sample data of the denoising unit corresponding to the small field of view camera.
[0294] Since the electronic device can obtain a piece of sample data corresponding to each denoising unit from each high-definition night scene image, a training data set corresponding to each denoising unit can be obtained from multiple high-definition night scene images.
[0295] In the embodiment of the present application, since different denoising units in the image denoising fusion model are trained through training data sets that match the denoising capabilities they need to have, different denoising units can have different denoising capabilities, so that the image denoising fusion model can perform more accurate denoising on images taken by different cameras at different exposure levels, thereby improving the overall quality of the fused image output by the image denoising fusion model.
[0296] Based on the same technical concept, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer-executable program. When the computer-executable program is called by the computer, the computer executes one or more steps in any of the above method embodiments.
[0297] Based on the same technical concept, an embodiment of the present application further provides a chip system, including a processor coupled to a memory, wherein the processor executes a computer-executable program stored in the memory to implement one or more steps in any of the above method embodiments. The chip system can be a single chip or a chip module composed of multiple chips.
[0298] Based on the same technical concept, an embodiment of the present application also provides a computer executable program product. When the computer executable program product runs on an electronic device, it enables the electronic device to execute one or more steps in any of the above method embodiments.
[0299] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described or recorded in detail in a particular embodiment, please refer to the relevant descriptions of other embodiments. It should be understood that the order of the sequence numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0300] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0301] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0302] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A shooting method, characterized in that: include: When a shooting instruction is received and the scene to be shot in the shooting preview box is a target scene, at least one set of images is shot by multiple cameras; The target scene includes a fireworks scene; each group of images includes: multiple images simultaneously captured by the multiple cameras of the scene to be captured with the same exposure time and different exposure amounts; in each group of images, the image with the smaller exposure amount is captured by a camera with better photosensitivity among the multiple cameras, and the image with the larger exposure amount is captured by a camera with worse photosensitivity among the multiple cameras; different groups of images are captured at different times; performing a first processing on each group of images so that all images in each group have the same field of view angle, the same size, and the same resolution; each group of images after the first processing constitutes an image group to be fused; A plurality of image groups to be fused that are adjacent in shooting time are input into a trained image denoising and fusion model for processing to obtain a fused image.
2. The shooting method according to claim 1, wherein: The image denoising and fusion model includes a plurality of denoising units and an image fusion unit; the total number of the plurality of denoising units is equal to the total number of the plurality of cameras, and the plurality of denoising units correspond one-to-one to the plurality of cameras respectively; the denoising unit is used to fuse and denoise a plurality of images with the same exposure value taken by the corresponding camera; Correspondingly, a plurality of image groups to be fused that are taken at adjacent times are input into a trained image denoising and fusion model for processing to obtain a fused image, including: By means of each of the denoising units, a plurality of images with the same exposure value captured by the camera corresponding to the denoising unit are fused and denoised, and a denoised image is output to the image fusion unit respectively; The image fusion unit fuses the multiple denoised images with different exposure amounts from the multiple denoising units to obtain the fused image.
3. The shooting method according to claim 2, characterized in that: Different denoising units are trained by different training data sets; the training data sets include multiple sample data, each of which includes multiple noisy images with noise and one real image without noise; The multiple noisy images are used as inputs of the denoising unit during training, and the real images are used as outputs of the denoising unit during training; Correspondingly, the training data sets corresponding to the multiple denoising units are generated in the following manner: Acquire multiple high-definition night scene images; For each high-definition night scene image, randomly cropping an area from the high-definition night scene image as the image to be degraded; Performing random preprocessing on the image to be degraded; The random pre-processing includes: random white balance processing, random rotation, random addition of ghosting, random addition of color blocks and random addition of jitter blur; Determining the image to be degraded after the random preprocessing as a real image; Copying the real image to obtain multiple copy images; the total number of the multiple copy images is equal to the total number of the multiple cameras; the multiple copy images respectively correspond to the multiple cameras one by one; performing a first noise addition process on a copy image corresponding to a camera with a large field of view angle among the multiple cameras to obtain a noise image corresponding to the camera with a large field of view angle; performing a second noise addition process on a copy image corresponding to a camera with a small field of view angle among the multiple cameras to obtain a noise image corresponding to the camera with a small field of view angle; wherein the second noise addition process is different from the first noise addition process; The noise image corresponding to the large field of view camera and the real image are used as a sample data of the denoising unit corresponding to the large field of view camera; the noise image corresponding to the small field of view camera and the real image are used as a sample data of the denoising unit corresponding to the small field of view camera.
4. The shooting method according to claim 3, characterized in that: Performing a first noise addition process on a copy image corresponding to a camera with a wide field of view among the multiple cameras to obtain a noise image corresponding to the camera with a wide field of view, comprising: Performing image blurring processing on the copy image corresponding to the camera with a wide field of view; Performing sensitivity compensation on the blurred copy image; adding a first noise to the sensitivity-compensated copy image based on a noise coefficient of the wide-field-of-view camera; The copy image with the first noise added thereto is converted into an image in a raw format to obtain a noise image corresponding to the camera with a large field of view.
5. The shooting method according to claim 3, wherein: Performing a second noise addition process on the copy image corresponding to the camera with a small field of view among the multiple cameras to obtain a noise image corresponding to the camera with a small field of view, comprising: Brightening the copy image corresponding to the small field of view camera; adding a second noise to the brightened copy image based on a noise coefficient of the small field of view camera; The copy image with the second noise added thereto is converted into an image in a raw format to obtain a noise image corresponding to the camera with a small field of view angle.
6. The shooting method according to any one of claims 1 to 5, characterized in that: The first processing includes: cropping, scaling, and upsampling; correspondingly, performing the first processing on each group of the images includes: cropping the large field-of-view image in each group of images to obtain a cropped image corresponding to the large field-of-view image; the field-of-view corresponding to the cropped image is the same as the field-of-view corresponding to the small field-of-view image in the same group of images; the large field-of-view image refers to an image captured by a large field-of-view camera among the multiple cameras, and the small field-of-view image refers to an image captured by a small field-of-view camera among the multiple cameras; scaling the cropped image corresponding to the large field-of-view image in each group of images to obtain a scaled image corresponding to the large field-of-view image; wherein the scaled image has the same size as the small field-of-view image in the same group of images, and the scaled image has the same field-of-view angle as the small field-of-view image in the same group of images; The scaled image corresponding to the large field of view angle image in each group of images is upsampled to obtain an upsampled image corresponding to the large field of view angle image; the resolution of the upsampled image is the same as the resolution of the small field of view angle image in the same group of images, the size of the upsampled image is the same as the size of the small field of view angle image in the same group of images, and the field of view corresponding to the upsampled image is the same as the field of view angle corresponding to the small field of view angle image in the same group of images.
7. The shooting method according to any one of claims 1 to 5, characterized in that: The first processing further includes: distortion correction; correspondingly, performing the first processing on each group of the images further includes: Based on the calibrated internal parameters of the camera corresponding to each image in each group of images, distortion correction is performed on each image respectively; the calibrated internal parameters include a distortion coefficient and a focal length.
8. The shooting method according to any one of claims 1 to 5, characterized in that: The first processing further includes: affine transformation; correspondingly, performing the first processing on each group of the images further includes: Based on the calibrated extrinsic parameters of the camera corresponding to each image in each group of images, an affine transformation is performed on each image; the calibrated extrinsic parameters are used to describe the position of the camera relative to the spatial coordinate system.
9. An electronic device, characterized in that: include: one or more processors, and memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises instructions, which, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 8.
11. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Multi-camera denoising photographing method and photographing device, mobile terminal and readable memory medium
CN108307122A
Image noise reduction method and device, electronic equipment and storage medium
CN110290289A
Image recognition method and device, equipment and medium
CN114022662A
Noise reduction method and device, electronic equipment and medium
CN114331902A
Video processing method, electronic equipment and chip system
CN117479008A