Photographing method and apparatus, storage medium, terminal, computer program product

By simultaneously capturing images from multiple cameras and fusing depth map information, and dynamically adjusting the fusion weights, the problem of blurred moving targets was solved, and image quality was improved, especially the clarity of the moving target area.

CN119485046BActive Publication Date: 2025-10-24SPREADTRUM COMMUNICATION (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411591879.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-10-24
Estimated Expiration
2044-11-07

AI Technical Summary

Technical Problem

When shooting scenes containing moving objects, existing technologies struggle to effectively address the blurring or motion blur issues of moving objects, leading to a decrease in image quality, especially with increased noise in low-light scenes.

Method used

Multiple cameras are used to simultaneously capture multiple frames of the same target scene. By combining image fusion technology with depth map information, the fusion weight of each frame image is dynamically adjusted to achieve image compensation and information complementarity among multiple cameras.

Benefits of technology

It improves the overall quality of captured images, especially the clarity of moving target areas, reduces blurring issues, and enhances image fusion effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119485046B_ABST
    Figure CN119485046B_ABST
Patent Text Reader

Abstract

A photographing method and device, a storage medium, a terminal and a computer program product, the method comprising: synchronously performing multi-frame photographing on a same target scene by using multiple cameras to obtain multiple preliminary photographing images corresponding to each camera, and performing fusion to obtain a first fusion image corresponding to the camera, the multiple cameras being divided into at least one group of cameras; determining a depth map corresponding to each group of cameras, and determining a difference in depth values between each pixel position and surrounding pixel positions of the depth map; determining, based on at least the difference in depth values, a first fusion weight of each pixel belonging to the pixel position in each frame of the first fusion image corresponding to the group of cameras, and then performing fusion on each frame of the first fusion image based on the first fusion weight to obtain a second fusion image; and performing image fusion on multiple frames of the second fusion image corresponding to each group of cameras to obtain a final photographing image. The above scheme can effectively improve the blurring or trailing problem of a moving target region in an image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of photographing, in particular to a photographing method and device, a storage medium, a terminal and a computer program product. BACKGROUND

[0002] In actual photographing process, the picture content can be preliminarily divided into static target and moving target, and the moving target can be further divided into global moving target (both the photographing device and the photographing target are in motion) and local moving target (the photographing device is fixed and the photographing target is in motion). Because the exposure time is set in a certain range during the photographing process, the moving speed of the moving target is inevitably too fast to cause the blurring or trailing of the photographing content, which affects the image quality of the photographing.

[0003] In the prior art, to solve the blurring or trailing problem caused by the target motion, a common method is to reduce the exposure time, however, this will cause the obvious increase of the image dark area noise, especially in the light source dark scene (for example, the night scene), which causes the poor image quality of the photographing. Another method is a single camera multi-frame fusion scheme, specifically, the image fusion is performed on the multiple frames of images photographed by a single camera at multiple different times, however, this scheme can effectively remove the time domain noise of the image region to which the static target belongs, but it is difficult to effectively solve the blurring problem of the fast moving target in the image, even if the noise reduction process is performed on the moving target region, on the one hand, the noise reduction process will inevitably lose the details of the moving target region, on the other hand, the effect of improving the image quality is still limited. SUMMARY

[0004] The technical problem solved by the embodiments of the present application is how to effectively improve the blurring or trailing problem of the moving target region in the image and improve the image quality when photographing the scene containing the moving target.

[0005] To solve the above technical problems, the embodiment of the present application provides a photographing method, comprising the following steps: using multiple cameras to synchronously take multiple frames for a same target scene, obtaining multiple frames of preliminary photographing images corresponding to each camera, and performing image fusion on the multiple frames of preliminary photographing images to obtain a first fused image corresponding to the camera, wherein the multiple cameras are divided into at least one group of cameras, and each group has at least two cameras; for each group of cameras, a depth map of the group of cameras relative to the target scene is determined according to multiple groups of preliminary photographing images synchronously taken by the group of cameras, each pixel position of the depth map has a respective depth value, wherein each group of preliminary photographing images contains at least two frames of preliminary photographing images synchronously taken by each camera in the group; for each pixel position of the depth map of the group of cameras, a depth value difference between the pixel position and surrounding pixel positions is determined, and at least based on the depth value difference, multiple first fusion weights corresponding to each pixel belonging to the pixel position in each frame of the first fused image corresponding to the group of cameras are determined, wherein the greater the depth value difference, the greater the difference between the multiple first fusion weights corresponding to each pixel belonging to the pixel position; for each frame of the first fused image corresponding to each group of cameras, using the multiple first fusion weights corresponding to each pixel belonging to each pixel position, weighted operation is performed on the pixel values of the pixels, thereby obtaining a second fused image corresponding to the group of cameras; and performing image fusion on the multiple frames of second fused images corresponding to each group of cameras to obtain a final photographing image.

[0006] Optionally, the multiple cameras are used to synchronously take multiple frames for the target scene using target photographing parameters; before the multiple cameras are used to synchronously take multiple frames for the target scene, the method further comprises: for at least one camera in the multiple cameras, determining multiple frames of preview images of the camera in a preview mode for the target scene, and performing motion analysis on a moving target in the multiple frames of preview images to determine motion information of the moving target; performing weighted operation on the motion information obtained by the at least one camera to obtain weighted motion information; and determining the target photographing parameters based on the weighted motion information.

[0007] Optionally, the motion information comprises one or more of a motion speed, a motion direction, and an area of an image region to which the motion target belongs, and the target shooting parameter comprises one or more of an exposure time, a camera movement direction, and a magnification; and the target shooting parameter is determined based on the weighted motion information, including: determining the exposure time corresponding to the motion speed according to the motion speed and a preset speed-exposure time mapping relationship, wherein the faster the motion speed is, the shorter the corresponding exposure time is; and / or using the motion direction as the camera movement direction; and / or determining the magnification corresponding to the area of the image region to which the motion target belongs according to the area of the image region to which the motion target belongs and a preset area-magnification mapping relationship, wherein the larger the area of the image region to which the motion target belongs is, the smaller the corresponding magnification is.

[0008] Optionally, the at least determining, based on the depth value difference, the first fusion weights corresponding to the pixels belonging to the pixel position in the first fusion image corresponding to each group of cameras, includes: determining, for the pixels belonging to the pixel position, a candidate index difference of each pixel, the candidate index difference being selected from one or more of a luminance value difference, a contrast difference, and a texture difference; performing weighted operation on the depth value difference and each candidate index difference to obtain a weighted difference, wherein the weight of the depth value difference in the weighted operation process is the largest; and determining the first fusion weights corresponding to the pixels belonging to the pixel position based on the weighted difference, wherein the larger the weighted difference is, the larger the difference between the first fusion weights is.

[0009] Optionally, in the case that there is a difference between the first fusion weights corresponding to the pixels belonging to the same pixel position in the first fusion image corresponding to each group of cameras, the first fusion weights satisfy: the higher the image quality of the first fusion image to which the pixel belongs is, the larger the first fusion weight corresponding to the pixel is.

[0010] Optionally, the determining, for each pixel position of the depth map of the group of cameras, the depth value difference between the pixel position and surrounding pixel positions, includes: respectively determining, for each pixel position of the depth map of the group of cameras, a depth difference between the depth value of the pixel position and the depth values of a plurality of pixel positions within a preset area range around the pixel position, thereby obtaining a plurality of depth differences corresponding to the pixel position; and calculating a mean square error of the plurality of depth differences, and taking the mean square error as the depth value difference.

[0011] Optionally, the method further comprises: determining, from the plurality of preliminary photographed images, a preliminary photographed image with the highest image quality as a reference frame image, and determining preliminary photographed images other than the reference frame image as non-reference frame images; for each preliminary photographed image, determining an image region to which a moving target in the preliminary photographed image belongs, and performing feature point extraction on an image region other than the image region to obtain a feature point set of the preliminary photographed image; and for each non-reference frame image, determining an alignment matrix based on the feature point set of the reference frame image and the feature point set of the reference frame image, and performing alignment processing on the reference frame image and the reference frame image by using the alignment matrix.

[0012] Optionally, in each round of fusion, the fusion result of the last round is recorded as a first to-be-fused frame, and one of the preliminary photographed images that has not participated in image fusion is recorded as a second to-be-fused frame; and the image fusion of the fusion result of the last round and one of the preliminary photographed images that has not participated in image fusion in each round of fusion to obtain the fusion result of the current round of fusion comprises: determining a pixel value difference of each pair of pixels at the same position in the first to-be-fused frame and the second to-be-fused frame; determining a second fusion weight corresponding to the pair of pixels based on the pixel value difference, and performing weighted operation on the pixel values of the pair of pixels by using the second fusion weight to obtain a fusion pixel value, wherein the greater the pixel value difference, the greater the difference between the second fusion weights corresponding to the pair of pixels; and using the fusion pixel values of each pair of pixels at the same position in the first to-be-fused frame and the second to-be-fused frame as the fusion result of the current round of fusion.

[0013] Optionally, before the image fusion of the plurality of preliminary photographed images, the method further comprises: determining, from the plurality of preliminary photographed images, a preliminary photographed image with the highest image quality as a reference frame image, and determining preliminary photographed images other than the reference frame image as non-reference frame images; for each preliminary photographed image, determining an image region to which a moving target in the preliminary photographed image belongs, and performing feature point extraction on an image region other than the image region to obtain a feature point set of the preliminary photographed image; and for each non-reference frame image, determining an alignment matrix based on the feature point set of the reference frame image and the feature point set of the reference frame image, and performing alignment processing on the reference frame image and the reference frame image by using the alignment matrix.

[0014] Optionally, for each group of cameras, a depth map of the group of cameras relative to the target scene is determined according to a plurality of groups of preliminary photograph images synchronously photographed by the group of cameras, comprising: for each group of cameras, selecting a group of preliminary photograph images with the best image quality from the plurality of groups of preliminary photograph images synchronously photographed by the group of cameras; and determining the depth map of the group of cameras relative to the target scene according to the group of preliminary photograph images with the best image quality.

[0015] The embodiment of the present application further provides a photographing device, comprising: a first fusion module, configured to perform multi-frame photographing on a same target scene by using a plurality of cameras to obtain a plurality of preliminary photograph images corresponding to each camera, and perform image fusion on the plurality of preliminary photograph images to obtain a first fusion image corresponding to the camera, wherein the plurality of cameras are divided into at least one group of cameras, and each group has at least two cameras; a depth map determination module, configured to determine, for each group of cameras, a depth map of the group of cameras relative to the target scene according to a plurality of groups of preliminary photograph images synchronously photographed by the group of cameras, each pixel position of the depth map having a respective depth value, wherein each group of preliminary photograph images comprises at least two preliminary photograph images synchronously photographed by each camera in the group; a fusion weight determination module, configured to determine, for each pixel position of the depth map of the group of cameras, a depth value difference between the pixel position and surrounding pixel positions, and determine, based at least on the depth value difference, a plurality of first fusion weights corresponding to each pixel belonging to the pixel position in each first fusion image corresponding to the group of cameras, wherein the greater the depth value difference, the greater the difference between the plurality of first fusion weights corresponding to each pixel belonging to the pixel position; a multi-camera fusion module, configured to perform weighted operation on pixel values of each pixel by using the plurality of first fusion weights corresponding to each pixel belonging to each pixel position for each first fusion image corresponding to each group of cameras, thereby obtaining a second fusion image corresponding to the group of cameras; and a second fusion module, configured to perform image fusion on a plurality of second fusion images corresponding to each group of cameras to obtain a final photograph image.

[0016] The embodiment of the present application further provides a storage medium having a computer program stored thereon, wherein the computer program is run by a processor to perform the steps of the photographing method.

[0017] The embodiment of the present application further provides a terminal comprising a memory and a processor, wherein the memory has a computer program stored thereon, the computer program being capable of being run on the processor, and the processor performs the steps of the photographing method when the computer program is run.

[0018] The embodiment of the present application further provides a computer program product comprising a computer program, wherein the computer program is run by a processor to perform the steps of the photographing method.

[0019] Compared with the prior art, the technical scheme of the embodiment of the present application has the following beneficial effects:

[0020] In the embodiment of the present application, the single-camera multi-frame fusion scheme and the multi-camera multi-frame re-fusion scheme are combined to obtain a higher quality of the photographed image. Specifically, a plurality of cameras are used to synchronously perform multi-frame photographing on the same target scene, to obtain a plurality of preliminary photographed images corresponding to each camera, and image fusion is performed on the plurality of preliminary photographed images to obtain a first fused image corresponding to the camera. The plurality of cameras are divided into at least one group of cameras, and each group has at least two cameras. Then, for each group of cameras, a depth map of the group of cameras relative to the target scene is determined according to a plurality of groups of preliminary photographed images synchronously photographed by the group of cameras, and the fusion weight (i.e., the first fusion weight) of the pixels at the same pixel position in each frame of the first fused image is adaptively assigned based on the depth map information of the group of cameras. Then, multi-camera multi-frame fusion (i.e., re-fusion of the first fused images corresponding to the plurality of cameras respectively) is performed based on the first fusion weight, to obtain a second fused image corresponding to each group of cameras. Finally, a plurality of frames of the second fused images corresponding to the groups of cameras are fused to obtain a photographed image.

[0021] From the above, on the one hand, image fusion is performed on the plurality of preliminary photographed images photographed by the same camera, which can preliminarily improve the overall quality of the image. On the other hand, by re-fusing the first fused images corresponding to the plurality of cameras respectively, depth information can be established from different perspectives, and the weight of multi-camera multi-frame fusion can be adjusted using the depth information, which can avoid blurring in the time domain. Further, the images synchronously photographed between the plurality of cameras can compensate for the missing information of different perspectives, and can also effectively reduce the blurring problem of the motion region, further improving the quality of the fused image. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a flowchart of a photographing method in the embodiment of the present application;

[0023] Figure 2 is Figure 1 is a flowchart of a specific implementation of step S11 in

[0024] Figure 3 is Figure 1 is a flowchart of a specific implementation of step S13 in

[0025] Figure 4 is a partial flowchart of another photographing method in the embodiment of the present application;

[0026] Figure 5 is an architectural schematic diagram of a photographing system for implementing the photographing method described in the embodiment of the present application;

[0027] Figure 6 is a structural schematic diagram of a photographing device in an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the above-mentioned objectives, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application are described in detail below with reference to the drawings.

[0029] Reference Figure 1 , Figure 1 is a flowchart of a photographing method in an embodiment of the present application. The method can be applied to various terminal devices with multi-camera synchronous photographing function, such as a mobile phone, a tablet computer, a computer, a smart wearable device (for example, a smart watch), an autonomous driving vehicle, etc. configured with multiple cameras.

[0030] The method can include steps S11 to S14:

[0031] Step S11: multiple cameras are used to synchronously take multiple frames for a same target scene, to obtain multiple frames of preliminary photographing images corresponding to each camera, and image fusion is performed on the multiple frames of preliminary photographing images to obtain a first fused image corresponding to the camera, wherein the multiple cameras are divided into at least one group of cameras, and each group has at least two cameras;

[0032] Step S12: for each group of cameras, a depth map of the group of cameras relative to the target scene is determined according to multiple groups of preliminary photographing images synchronously taken by the group of cameras, each pixel position of the depth map has a respective depth value, wherein each group of preliminary photographing images contains at least two frames of preliminary photographing images synchronously taken by each camera of the group;

[0033] Step S13: for each pixel position of the depth map of the group of cameras, a depth value difference between the pixel position and surrounding pixel positions is determined, and at least based on the depth value difference, multiple first fusion weights corresponding to each pixel belonging to the pixel position in each frame of first fused image corresponding to the group of cameras are determined, wherein the greater the depth value difference, the greater the difference between the multiple first fusion weights corresponding to each pixel belonging to the pixel position;

[0034] Step S14: for each frame of first fused image corresponding to each group of cameras, multiple first fusion weights corresponding to each pixel belonging to each pixel position are used to perform weighted operation on the pixel values of each pixel, thereby obtaining a second fused image corresponding to the group of cameras;

[0035] Step S15: multiple frames of second fused images corresponding to each group of cameras are subjected to image fusion to obtain a final photographing image.

[0036] It can be understood that in specific implementation, the method can be implemented in the form of a software program running in a processor integrated inside a chip or a chip module (for example, an Image Signal Processor (ISP) chip or a chip module); or the method can be implemented in the form of hardware or a combination of software and hardware.

[0037] In specific implementation of step S11, taking a smartphone as an example, the plurality of cameras can be cameras arranged on the back of the smartphone, and the plurality of cameras have an overlapping area between a plurality of preliminary shooting images obtained by shooting the target scene at the same time (i.e., one-time synchronous shooting). The target scene can include stationary targets and / or moving targets. The moving targets can be, for example, a person in a running or walking state, a vehicle in a driving state, an animal in a moving state, and the like.

[0038] The synchronous multi-frame shooting of the plurality of cameras on the same target scene can specifically refer to: synchronous multi-frame shooting of the plurality of cameras on the target scene multiple times, each camera shooting one preliminary shooting image each time, and multiple synchronous shooting times corresponding to a plurality of preliminary shooting images obtained by each camera.

[0039] As a non-linear example, three cameras (denoted as camera A, camera B, and camera C) are used to perform synchronous multi-frame shooting on the same target scene, and the plurality of preliminary shooting images obtained by each camera have a shooting time. For example, camera A, B, and C perform one-time synchronous shooting at t1, t2, …, t n , respectively, to obtain preliminary shooting images a1, b1, and c1 obtained by camera A, B, and C at t1, preliminary shooting images a2, b2, and c2 obtained by camera A, B, and C at t2, …, and preliminary shooting images a n , b n , and c n , obtained by camera A, B, and C at t n .

[0040] For camera A, the corresponding plurality of preliminary shooting images are a1, a2, …, a n ; for camera B, the corresponding plurality of preliminary shooting images are b1, b2, …, b n ; and for camera C, the corresponding plurality of preliminary shooting images are c1, c2, …, c n .

[0041] It can be understood that in actual application, in order to improve the blurring or smear problem of the moving target as much as possible, the time t1, t2, …, t nThe time interval between each two times of shooting is as small as possible, for example, the multiple times of shooting are realized by using the way of continuous multiple times of snapshot.

[0042] In a specific implementation, the grouping manner of the multiple cameras can be selected from, but not limited to, the following manners.

[0043] Manner one: each N cameras are grouped as a group (for example, the grouping can be random or all possible combinations are exhausted), and at least one group of cameras is obtained, wherein N is greater than or equal to 2 and less than or equal to the total number of the multiple cameras, and N is a positive integer (when N is 2, each group has two cameras; when N is the total number of the multiple cameras, the multiple cameras are divided into one group of cameras).

[0044] Manner two: the camera with the best shooting quality is selected from the multiple cameras as a reference camera, and the remaining cameras are grouped as a group of M cameras (for example, the grouping can be random or all possible combinations are exhausted), and at least one group of initial cameras is obtained; then each group of initial cameras and the reference camera are grouped as a group, and at least one group of cameras is obtained, wherein M is greater than or equal to 1 and less than or equal to the total number of the multiple cameras minus 1, and M is a positive integer (when M is 1, each group has two cameras, one of which is the reference camera; when M is the total number of the multiple cameras minus 1, the multiple cameras are divided into one group of cameras).

[0045] Further, in a specific implementation, the image fusion scheme for the multiple frames of preliminary shooting images corresponding to each camera (for example, the multiple frames of preliminary shooting images a1, a2, …, aN corresponding to camera A, and the multiple frames of preliminary shooting images b1, b2, …, bN corresponding to camera B) can adopt the image fusion scheme shown in FIG. 5. n ) are fused to obtain the first fused image corresponding to the camera. Figure 2

[0046] Referring to FIG. 6, Figure 2 , Figure 2 is a flow chart of a specific implementation of step S11 in Figure 1 , the step S11 being to fuse the multiple frames of preliminary shooting images to obtain the first fused image corresponding to the camera, which can specifically include steps S111 to S112.

[0047] In step S111, the preliminary shooting image with the highest image quality is determined from the multiple frames of preliminary shooting images, and is recorded as a reference frame image.

[0048] In a specific implementation, the method for determining the preliminary shooting image with the highest image quality can be a manual image quality evaluation method or a machine learning method.

[0049] ​In step S112, one or more rounds of fusion operations are performed based on the reference frame image and the remaining preliminary captured images, and the fusion result of the last round is taken as the first fusion image corresponding to the camera.

[0050] In each round of fusion, the fusion result of the last round and one of the remaining preliminary captured images that does not participate in image fusion are fused to obtain the fusion result of the current round; and the fusion result of the first round is obtained by fusing the reference frame image and one of the preliminary captured images other than the reference frame image.

[0051] Further, in each round of fusion, the fusion result of the last round is denoted as a first to-be-fused frame, and one of the remaining preliminary captured images that does not participate in image fusion is denoted as a second to-be-fused frame; and the fusion of the fusion result of the last round and one of the remaining preliminary captured images that does not participate in image fusion in each round of fusion to obtain the fusion result of the current round includes: determining, for each pair of pixels in the same position in the first to-be-fused frame and the second to-be-fused frame, a pixel value difference of the pair of pixels; determining a second fusion weight corresponding to the pair of pixels based on the pixel value difference, and performing a weighted operation on the pixel values of the pair of pixels by using the second fusion weight to obtain a fusion pixel value, wherein the greater the pixel value difference, the greater the difference between the two second fusion weights corresponding to the pair of pixels; and using the fusion pixel values of each pair of pixels in the same position in the first to-be-fused frame and the second to-be-fused frame as the fusion result of the current round.

[0052] Specifically, if the pixel difference of the pair of pixels in the same position is less than a first preset threshold (for example, which can be set to 0 or a smaller value close to 0), the two second fusion weights corresponding to the pair of pixels can be equal (for example, each can be 0.5); if the pixel difference of the pair of pixels in the same position is greater than or equal to the first preset threshold, the two second fusion weights corresponding to the pair of pixels are different, and the greater the pixel value difference, the greater the difference between the two second fusion weights corresponding to the pair of pixels.

[0053] Further, for each pair of pixels in the same position, in the case where the two second fusion weights corresponding to the pair of pixels are different, the second fusion weight corresponding to the pixel belonging to the to-be-fused frame with higher image quality is greater than the second fusion weight corresponding to the pixel belonging to the to-be-fused frame with lower image quality.

[0054] In specific implementation, the sum of the two second fusion weights corresponding to each pair of pixels in the same position in the first to-be-fused frame and the second to-be-fused frame can be 1, that is, 1 can be weightedly allocated according to the pixel difference of the pair of pixels to determine the two second fusion weights corresponding to the pair of pixels.

[0055] As a non-limiting example, in the case that the sum of the two second fusion weights corresponding to the pixel pair is 1 and the two second fusion weights are different, the second fusion weight corresponding to the pixel belonging to the frame with higher image quality is, for example, 0.6 (or 0.7), and the second fusion weight corresponding to the pixel belonging to the frame with lower image quality is, for example, 0.4 (or 0.3).

[0056] As to the specific manner of weight distribution, the difference between the two second fusion weights corresponding to each pixel pair can be positively correlated with the pixel value difference of the pixel pair, and the actual application scene requirements can be combined for appropriate distribution, which is not specifically limited in the embodiments of the present application.

[0057] It can be understood that, for each pixel pair with the same position in the first frame and the second frame, the greater the pixel value difference of the pixel pair, the greater the possibility that the target belonging to the pixel pair is a moving target, and the greater the pixel value difference often means the faster the moving speed; the smaller the pixel value difference of the pixel pair, the greater the possibility that the target belonging to the pixel pair is a stationary target. Therefore, in the embodiments of the present application, for a moving target, the two second fusion weights corresponding to the pixel pair are set to be different, that is, the pixel in one of the two frames is taken as the main factor affecting the weighted pixel value, and the pixel in the other frame is taken as the secondary factor affecting the weighted pixel value; and for a stationary target, the two second fusion weights corresponding to the pixel pair are set to be as consistent as possible, that is, the two pixels at the same position in the two frames have close effects on the weighted fusion. Thus, it is helpful to improve the image fusion effect, and especially helpful to improve the clarity of the moving target region after fusion.

[0058] Further, for each pixel pair with the same position, in the case that the two second fusion weights corresponding to the pixel pair are different (that is, it means that the pixel pair is more likely to belong to a moving target), the second fusion weight corresponding to the pixel belonging to the frame with higher image quality is set to be greater than the second fusion weight corresponding to the pixel belonging to the frame with lower image quality. Thus, for a moving target, the pixel characteristics of the frame with higher image quality can be mainly retained, and the image fusion quality is further improved.

[0059] It should be noted that, in specific implementation, in addition to the above scheme, other appropriate methods can also be used for image fusion on the plurality of preliminary shooting images corresponding to each camera, for example, the following steps can be used for image fusion:

[0060] Step (1): determining the preliminary shooting image with the highest image quality from the plurality of preliminary shooting images, and recording it as a reference frame image, and recording the remaining preliminary shooting images as non-reference frame images;

[0061] Step (2): for each non-reference frame image, determining pixel value difference of each pair of pixels belonging to the same position in the non-reference frame image and the reference frame image, and determining third fusion weight corresponding to the pair of pixels based on the pixel value difference, wherein the greater the pixel value difference, the greater the difference between the two third fusion weights corresponding to the pair of pixels;

[0062] Step (3): normalizing each third fusion weight corresponding to each pair of pixels belonging to the same position in each non-reference frame image and the reference frame image to obtain target weight corresponding to each pixel belonging to the same position in each non-reference frame image and the reference frame image;

[0063] Step (4): for all pixels belonging to the same position in each non-reference frame image and the reference frame image, performing weighted operation on pixel values of all pixels in the same position by using the target weight corresponding to each pixel to obtain weighted pixel value corresponding to the position, thereby obtaining the first fusion image corresponding to the camera.

[0064] Further, before image fusion is performed on the plurality of preliminary shooting images, the method further comprises: determining a preliminary shooting image with the highest image quality from the plurality of preliminary shooting images, and marking the preliminary shooting image as a reference frame image; marking preliminary shooting images other than the reference frame image as non-reference frame images; for each preliminary shooting image, determining an image region to which a moving target in the preliminary shooting image belongs, and performing feature point extraction on an image region other than the image region to obtain a feature point set of the preliminary shooting image; for each non-reference frame image, determining an alignment matrix based on the feature point set of the reference frame image and the feature point set of the reference frame image, and performing alignment processing on the reference frame image and the reference frame image by using the alignment matrix.

[0065] In the embodiment of the application, the image alignment operation is performed based on the feature points of the image region other than the image region to which the moving target belongs, which helps to reduce the influence of the moving region which is blurred itself on the alignment result, improves the alignment accuracy, and thus improves the image fusion effect of the plurality of preliminary shooting images shot by each camera.

[0066] Continuing to refer to Figure 1 In the implementation of step S12, for each group of cameras, each set of preliminary shooting images synchronously shot by the group of cameras comprises at least two preliminary shooting images synchronously shot by each camera in the group.

[0067] Continuing to combine the foregoing description of step S11, for the example of using three cameras A, B and C to perform multiple synchronous shooting, it is assumed that the three cameras are divided into a group, and the group of cameras synchronously shoot at t1, t2, t nIf the multiple groups of preliminary photograph images synchronously photographed by the multiple groups of cameras are respectively photographed at different time points, the multiple groups of preliminary photograph images synchronously photographed by the multiple groups of cameras are as follows:

[0068] The preliminary photograph images a1, b1 and c1 synchronously photographed by the cameras A, B and C at the time point t1; the preliminary photograph images a2, b2 and c2 synchronously photographed by the cameras A, B and C at the time point t2; and the preliminary photograph images a3, b3 and c3 synchronously photographed by the cameras A, B and C at the time point t3. n The preliminary photograph image a synchronously photographed by the cameras A and B at the time point t1 n , b n , c n .

[0069] Further, in the step S12, the method for determining the depth map of each group of cameras relative to the target scene can comprise: for each group of cameras, selecting a group of preliminary photograph images with the best image quality from the multiple groups of preliminary photograph images synchronously photographed by the group of cameras; and determining the depth map of the group of cameras relative to the target scene according to the group of preliminary photograph images with the best image quality. In this way, a more accurate depth map can be obtained, thereby improving the accuracy of the subsequent determination of the fusion weight of the pixel based on the depth map.

[0070] In the implementation of the step S13, the depth value difference of each pixel position of the depth map and the surrounding pixel positions can be determined in the following manner: for each pixel position of the depth map of the group of cameras, the depth difference between the depth value of the pixel position and the depth values of multiple pixel positions within a preset region around the pixel position is determined respectively, thereby obtaining multiple depth differences corresponding to the pixel position; and the mean square deviation of the multiple depth differences is calculated, and the mean square deviation is taken as the depth value difference.

[0071] In the implementation, in addition to taking the mean square deviation of the depth difference between each pixel position and the surrounding pixel positions as the depth difference, other appropriate indicators can also be taken as the depth difference.

[0072] For example, for each pixel position of the depth map, the absolute values of the differences between the depth values of the pixel position and the surrounding multiple pixel positions are calculated respectively; and the sum of the multiple absolute values is taken as the depth value difference.

[0073] For another example, for each pixel position of the depth map, the average values of the depth values of the surrounding multiple pixel positions of the pixel position are calculated respectively; and the difference between the depth value of the pixel position and the average value is taken as the depth value difference.

[0074] In one specific implementation, the multiple first fusion weights corresponding to the pixels belonging to the same pixel position in each frame of the first fusion image corresponding to the group of cameras can be determined directly based on the single difference item of the depth value difference.

[0075] In another specific implementation, in addition to considering the depth value difference, other suitable candidate indicator differences can also be introduced, and based on the weighted results of the depth value difference and each candidate indicator difference, the first fusion weights corresponding to each pixel belonging to the same pixel position in each frame of the first fusion image corresponding to the group of cameras are determined. In this way, it is helpful to improve the accuracy of each first fusion weight and improve the fusion effect, especially for the motion target region, which is more helpful to improve the blur problem and improve the clarity of the corresponding image region after fusion. For the another specific implementation, please refer to the description of Figure 3 .

[0076] Referring to Figure 3 , Figure 3 is Figure 1 a flow chart of one specific implementation of step S13, and in this implementation, the step S13 can specifically include steps S131 to S133.

[0077] In step S131, for each pixel belonging to the same pixel position, the candidate indicator difference of each pixel is determined, and the candidate indicator difference is selected from one or more of the following: brightness value difference, contrast difference, and texture difference.

[0078] In a specific implementation, the candidate indicator difference of each pixel belonging to the same pixel position can be determined in the following manner: taking the brightness value difference as an example, for each pixel belonging to the same pixel position, the initial brightness difference between the pixel and the surrounding pixels is determined (for example, the mean square error of the difference between the brightness values of the pixel and the surrounding pixels can be used); the initial brightness difference corresponding to each pixel belonging to the same pixel position is averaged to obtain the brightness value difference of each pixel belonging to the same pixel position. Alternatively, other suitable methods can be used to determine the candidate indicator difference of each pixel belonging to the same pixel position, as long as the candidate indicator difference can be used to represent the possibility that the pixels of the pixel position and the surrounding pixel positions belong to the same target.

[0079] In step S132, the depth value difference and each candidate indicator difference are weighted to obtain a weighted difference, wherein the weight of the depth value difference in the weighted operation process is the largest.

[0080] In step S133, based on the weighted difference, the first fusion weights corresponding to each pixel belonging to the same pixel position are determined, and the larger the weighted difference is, the larger the difference between the first fusion weights is.

[0081] It can be understood that, for each pixel position of each frame of the first fusion image corresponding to each group of cameras, the greater the difference (or weighted difference) in depth values between the pixel position and surrounding pixel positions, the less likely the pixel position and the surrounding pixel positions belong to the same target (i.e., the greater the probability that the pixel position is a boundary pixel between two different targets), and accordingly, the greater the difference between the plurality of first fusion weights corresponding to the pixels belonging to the pixel position in the first fusion image. That is, the pixel of one of the frames of the first fusion image is taken as the main factor affecting the weighted pixel value, and the pixels of the remaining first fusion images are taken as the secondary factor affecting the weighted pixel value.

[0082] For the pixel position with a smaller difference in depth values from surrounding pixel positions (meaning the greater the likelihood that the pixel position and the surrounding pixel positions belong to the same target), the plurality of first fusion weights corresponding to the pixels belonging to the pixel position in the first fusion image are made as consistent as possible, that is, the influence of the pixels at the same position in the first fusion image on the weighted fusion is made close to consistent. Thus, it helps to improve the noise removal effect on the moving target region and improve the clarity of the fused image.

[0083] Specifically, if the weighted difference is less than a second preset threshold (which can be set to 0 or a smaller value close to 0), the plurality of first fusion weights corresponding to the pixels belonging to the pixel position can be equal; if the weighted difference is greater than or equal to the second preset threshold, the plurality of first fusion weights corresponding to the pixels belonging to the pixel position are different, and the greater the weighted difference, the greater the difference between the plurality of first fusion weights corresponding to the pixels belonging to the pixel position.

[0084] Further, in the case that there is a difference between the plurality of first fusion weights corresponding to the pixels belonging to the same pixel position in each frame of the first fusion image corresponding to each group of cameras, the plurality of first fusion weights satisfy: the higher the image quality of the first fusion image to which the pixel belongs, the greater the first fusion weight corresponding to the pixel. In this way, for pixels with a low probability of belonging to the same target (i.e., boundary pixels between different targets), the pixel features of the first fusion image with higher image quality can be mainly retained, further improving the image fusion quality.

[0085] In specific implementation, the sum of the plurality of first fusion weights corresponding to the pixels belonging to the same pixel position can be 1, that is, the weighted allocation of 1 can be made according to the weighted difference to determine the plurality of first fusion weights corresponding to the pixels belonging to the same pixel position.

[0086] In a non-limiting embodiment, each group of cameras has 2 cameras, the sum of the two first fusion weights corresponding to two pixels belonging to the same pixel position in the two first fusion images corresponding to the group of cameras is 1, and in the case of a difference between the two first fusion weights, the first fusion weight corresponding to the pixel belonging to the first fusion image with higher image quality is, for example, 0.6 (or 0.7), and the first fusion weight corresponding to the pixel belonging to the first fusion image with lower image quality is, for example, 0.4 (or 0.3).

[0087] In another non-limiting embodiment, each group of cameras has 3 cameras, the sum of the three first fusion weights corresponding to three pixels belonging to the same pixel position in the three first fusion images corresponding to the group of cameras is 1, and in the case of a difference between the three first fusion weights, the first fusion weight corresponding to the pixel belonging to the first fusion image with the highest image quality is, for example, 0.5, the first fusion weight corresponding to the pixel belonging to the first fusion image with the second highest image quality is, for example, 0.3, and the first fusion weight corresponding to the pixel belonging to the first fusion image with the lowest image quality is, for example, 0.2. Alternatively, the first fusion weights corresponding to the pixels belonging to the first fusion images with the second highest and the lowest image qualities can also be the same, for example, both 0.25.

[0088] As for the specific manner of weight distribution, the difference between the multiple first fusion weights corresponding to each pixel can be positively correlated with the weighted difference, and appropriate distribution can be made in combination with the actual application scene requirements, which is not specifically limited in the embodiments of the present application.

[0089] Continuing to refer to Figure 1 In the specific implementation of step S14, for each first fusion image corresponding to each group of cameras, the multiple first fusion weights corresponding to each pixel belonging to each pixel position are used to perform weighted operation on the pixel values of each pixel, thereby obtaining a second fusion image corresponding to the group of cameras.

[0090] In the specific implementation of step S15, the multiple second fusion images corresponding to each group of cameras are fused to obtain a final shooting image. As for the detailed implementation scheme of fusing the multiple second fusion images, reference can be made to the scheme of fusing the multiple preliminary shooting images corresponding to each camera in step S11, which will not be described here again.

[0091] Referring to Figure 4 , Figure 4 is a partial flowchart of another photographing method in the embodiments of the present application, in which the multiple cameras are used to synchronously shoot multiple frames of images by using the target shooting parameters for the target scene; the another photographing method can include Figure 1The steps S11 to S15 shown can further include steps S41 to S43, which can be executed before step S11.

[0092] In step S41, for at least one camera of the plurality of cameras, a plurality of preview images of the camera in a preview mode for the target scene are determined, and motion analysis is performed on a moving target in the plurality of preview images to determine motion information of the moving target.

[0093] The at least one camera can be one or more cameras of the plurality of cameras with higher shooting quality.

[0094] The preview mode (or preview state) can refer to a mode or state before a photographing function key is triggered. In the preview mode, the terminal device usually collects and stores preview frames collected by each camera for the target scene in real time. For each camera, the plurality of preview images of the camera in the preview mode for the target scene can be selected from the preview frames collected by the camera for the target scene. For example, a plurality of preview frames collected within a preset time period before the photographing function key is triggered can be selected.

[0095] The method of performing motion analysis on the moving target in the plurality of preview images can use existing conventional methods, such as a motion analysis algorithm based on optical flow or a motion analysis method based on a fast corner detection and binary basic feature extraction algorithm (Oriented FAST and Rotated BRIEF, ORB).

[0096] In step S42, a weighted operation is performed on the motion information obtained by the at least one camera to obtain weighted motion information.

[0097] In step S43, the target shooting parameter is determined based on the weighted motion information.

[0098] Compared to motion analysis based on a plurality of preview images collected by a single camera to obtain motion information and determine a target shooting parameter, the weighted motion information determined by the plurality of cameras in the present embodiment can obtain a more accurate and reliable target shooting parameter.

[0099] Further, the motion information can include one or more of a motion speed, a motion direction, and an area of an image region to which the moving target belongs, and the target shooting parameter can include one or more of an exposure time, a camera movement direction, and a magnification. In step S43, the target shooting parameter is determined based on the weighted motion information, which can specifically include:

[0100] According to the motion speed and a preset speed-exposure time mapping relationship, a corresponding exposure time of the motion speed is determined, wherein the faster the motion speed is, the shorter the corresponding exposure time is; and / or the motion direction is used as the camera moving direction; and / or according to an area of the image region to which the motion target belongs and a preset area-magnification mapping relationship, a corresponding magnification (i.e., zoom multiple) of the area of the image region to which the motion target belongs is determined, wherein the larger the area of the image region to which the motion target belongs is, the smaller the corresponding magnification is.

[0101] In a specific implementation, the speed-exposure time mapping relationship and the area-magnification mapping relationship can be empirical data determined according to a pre-experiment, and can be in the form of a mapping relationship table or a fitted mapping relationship function.

[0102] In a specific implementation, after the target shooting parameter is determined, the target shooting parameter can be displayed on a display screen (or a user interaction screen) of the terminal device, and the display form includes but is not limited to text and arrow graphics. The user can select whether to apply the target shooting parameter to the multiple cameras of the terminal device according to actual needs, so as to synchronously perform multi-frame shooting on the target scene.

[0103] In another specific implementation, after the target shooting parameter is determined, the target shooting parameter can be automatically applied to the multiple cameras of the terminal device, so as to synchronously perform multi-frame shooting on the target scene.

[0104] In the embodiment of the application, on the basis of obtaining the first fused image corresponding to each camera by using the single-camera multi-frame (i.e., multiple preliminary shooting images shot by a single camera) fusion scheme and re-fusing the first fused image corresponding to each camera by using the multi-camera multi-frame fusion scheme, the motion of the motion target is additionally analyzed according to the preview image of each camera for the target scene, and the target shooting parameter (including but not limited to exposure time, camera moving direction, and magnification) suitable for shooting the target scene is determined according to the analysis result. Therefore, the preliminary shooting image (which can be referred to as preliminary data) with better quality can be obtained, and especially for the scene containing a fast motion target, the blurring and smear problems of the image can be improved as much as possible in the preliminary data acquisition stage, and the quality of the shooting image obtained by final fusion is further improved.

[0105] Reference Figure 5 , Figure 5 is a schematic architecture diagram of a photographing system for implementing the photographing method in the embodiment of the application.

[0106] The photographing system can include two cameras (wherein the number of cameras is only a non-limiting example), a first multi-frame snapshot module 52, a second multi-frame snapshot module 53, a preview image acquisition module 54, a motion analysis module 55, a strategy analysis module 56, a user interaction module 57, a depth information calculation module 58, a first multi-frame alignment fusion module 59, a second multi-frame alignment fusion module 510, and a multi-camera compensation fusion module 511.

[0107] Specifically, the first multi-frame snapshot module 52 and the second multi-frame snapshot module 53 are configured to control the main camera 50 and the auxiliary camera 51 to synchronously perform multi-frame shooting on the same target scene by using the target shooting parameter, so as to obtain a plurality of preliminary shooting images corresponding to the main camera 50 and a plurality of preliminary shooting images corresponding to the auxiliary camera 51.

[0108] It should be noted that the photographing system includes a plurality of cameras, and the number of cameras is not limited to two.

[0109] The preview image acquisition module 54 is configured to acquire a plurality of preview images of the target scene in the preview mode by using the main camera 50.

[0110] The motion analysis module 55 is configured to perform motion analysis on the motion target in the plurality of preview images, and determine the motion information of the motion target.

[0111] The strategy analysis module 56 is configured to determine the target shooting parameter based on the motion information.

[0112] The user interaction module 57 is configured to display the target shooting parameter on the display screen (or user interaction screen) of the terminal device, so that the user can select whether to apply the target shooting parameter to the main camera 50 and the auxiliary camera 51 according to actual needs.

[0113] The depth information calculation module 58 is configured to determine a depth map of the target scene with respect to the main camera 50 and the auxiliary camera 51 according to a plurality of groups of preliminary shooting images synchronously shot by the main camera 50 and the auxiliary camera 51, each pixel position of the depth map has a respective depth value, wherein each group of preliminary shooting images includes two preliminary shooting images synchronously shot by the main camera 50 and the auxiliary camera 51.

[0114] The first multi-frame alignment fusion module 59 is configured to perform alignment processing on the multi-frame preliminary shooting images corresponding to the main camera 50, and then perform image fusion processing on the multi-frame preliminary shooting images after the alignment processing, to obtain a first fusion image corresponding to the main camera 50; further, the first alignment information provided by the motion analysis module 55 can be used for the alignment processing, and the first alignment information specifically refers to image region information (for example, the position, area, etc. of the image region) of a motion target determined by the motion analysis module 55 by performing motion analysis on each preliminary shooting image from the first multi-frame snapshot module 52.

[0115] The second multi-frame alignment module 510 is configured to perform alignment processing on the multi-frame preliminary shooting images corresponding to the auxiliary camera 51, and then perform image fusion processing on the multi-frame preliminary shooting images after the alignment processing, to obtain a first fusion image corresponding to the auxiliary camera 51; further, the second alignment information provided by the motion analysis module 55 can be used for the alignment processing, and the second alignment information specifically refers to image region information (for example, the position, area, etc. of the image region) of a motion target determined by the motion analysis module 55 by performing motion analysis on each preliminary shooting image from the second multi-frame snapshot module 53.

[0116] For the specific implementation of the alignment processing using the first alignment information and the second alignment information, refer to the description of the related scheme in the foregoing for the feature point extraction on the image region other than the image region to which the motion target belongs, and the alignment processing based on the extracted feature point set, which will not be repeated here.

[0117] The multi-camera compensation fusion module 511 performs multi-camera multi-frame fusion processing on the two first fusion images output by the first multi-frame alignment fusion module 59 and the second multi-frame alignment module 510 based on the depth map determined by the depth information calculation module 58, and outputs the fusion result (i.e., the final shooting image).

[0118] For more implementation contents, principles, technical effects, etc. of the modules of the photographing system, refer to the related description of any embodiment in the foregoing and Figures 1 to 4 The related description of any embodiment will not be repeated here.

[0119] Refer to Figure 6 , Figure 6 is a structural schematic diagram of a photographing device in an embodiment of the present application. The photographing device can include:

[0120] The first fusion module 61 is configured to perform multi-frame shooting on a same target scene by using a plurality of cameras, to obtain a plurality of preliminary shooting images corresponding to each camera, and perform image fusion on the plurality of preliminary shooting images, to obtain a first fusion image corresponding to the camera, wherein the plurality of cameras are divided into at least one group of cameras, and each group has at least two cameras.

[0121] a depth map determination module 62 configured to determine, for each group of cameras, a depth map of the group of cameras relative to the target scene according to a plurality of preliminary captured images synchronously captured by the group of cameras, each pixel position of the depth map having a respective depth value, wherein each group of preliminary captured images comprises at least two frames of preliminary captured images synchronously captured by the respective cameras of the group of cameras;

[0122] a fusion weight determination module 63 configured to determine, for each pixel position of the depth map of the group of cameras, a depth value difference between the pixel position and surrounding pixel positions, and determine, based at least on the depth value difference, a plurality of first fusion weights corresponding to the respective pixels belonging to the pixel position in the respective first fused images corresponding to the group of cameras, wherein the greater the depth value difference, the greater the difference between the plurality of first fusion weights corresponding to the respective pixels belonging to the pixel position;

[0123] a multi-camera fusion module 64 configured to perform, for each group of cameras, a weighted operation on pixel values of the respective pixels of the respective first fused images corresponding to the group of cameras using the plurality of first fusion weights corresponding to the respective pixels belonging to each pixel position, to obtain a second fused image corresponding to the group of cameras;

[0124] a second fusion module 65 configured to perform image fusion on the plurality of second fused images corresponding to the respective groups of cameras to obtain a final captured image.

[0125] For the principle, specific implementation and beneficial effects of the photographing device, please refer to the foregoing and Figures 1 to 4 For the related description of the photographing method shown in any of the embodiments, please refer to the foregoing.

[0126] The embodiment of the present application further provides a storage medium, for example, a computer readable storage medium, which stores a computer program, and the computer program is run by a processor to execute the photographing method shown in any of the embodiments. Figures 1 to 4 The steps of the photographing method shown in any of the embodiments. The computer readable storage medium can include a non-volatile memory or a non-transitory memory, and can further include an optical disc, a mechanical hard disk, a solid state disk, etc.

[0127] Specifically, in the embodiments of the present application, the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0128] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).

[0129] The embodiments of the present application also provide a terminal, comprising a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor executes the computer program to perform the above-mentioned Figures 1 to 4The steps of the photographing method shown in any of the embodiments.

[0130] The embodiment of the present application further provides a computer program product, comprising a computer program, which is executed by a processor to perform the above-mentioned Figures 1 to 4 The steps of the photographing method shown in any of the embodiments.

[0131] It should be understood that the term "and / or" in the present document is only used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present document means that the front and rear associated objects are in an "or" relationship.

[0132] The "multiple" appearing in the embodiments of the present application means two or more.

[0133] The first, second, and the like appearing in the embodiments of the present application are only used for description and distinction of the description objects, and there is no order, nor does it represent a special limitation on the number of devices in the embodiments of the present application, which cannot constitute any limitation on the embodiments of the present application.

[0134] It should be noted that the serial numbers of the steps in the embodiments do not represent the limitation on the execution order of the steps.

[0135] Although the present application is disclosed as above, the present application is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application, and therefore the protection scope of the present application should be subject to the scope defined by the claims.

Claims

1. A photographing method, characterized by, The method comprises: Synchronously capturing multiple frames of images of a same target scene by using multiple cameras to obtain multiple preliminary captured images corresponding to each camera, and performing image fusion on the multiple preliminary captured images to obtain a first fused image corresponding to the camera, wherein the multiple cameras are divided into at least one group of cameras, and each group has at least two cameras; For each group of cameras, determining a depth map of the group of cameras relative to the target scene according to multiple groups of preliminary captured images synchronously captured by the group of cameras, each pixel position of the depth map having a respective depth value, wherein each group of preliminary captured images comprises at least two frames of preliminary captured images synchronously captured by each camera in the group; For each pixel position of the depth map of the group of cameras, determining a difference in depth values between the pixel position and surrounding pixel positions, and determining, based at least on the difference in depth values, multiple first fusion weights corresponding to each pixel belonging to the pixel position in each frame of first fused image corresponding to the group of cameras, wherein the greater the difference in depth values, the greater the difference between the multiple first fusion weights corresponding to each pixel belonging to the pixel position; For each frame of first fused image corresponding to each group of cameras, performing weighted operation on pixel values of each pixel using the multiple first fusion weights corresponding to each pixel belonging to each pixel position, thereby obtaining a second fused image corresponding to the group of cameras; Performing image fusion on multiple frames of second fused images corresponding to each group of cameras to obtain a final captured image.

2. The method of claim 1, wherein, The multiple cameras are used to synchronously capture multiple frames of images of the target scene using target shooting parameters; Before synchronously capturing multiple frames of images of the target scene using the multiple cameras, the method further comprises: For at least one camera of the multiple cameras, determining multiple frames of preview images of the target scene in a preview mode by the camera, and performing motion analysis on a moving target in the multiple frames of preview images to determine motion information of the moving target; Performing weighted operation on the motion information obtained by the at least one camera to obtain weighted motion information; Determining the target shooting parameters based on the weighted motion information.

3. The method of claim 2, wherein, The motion information comprises one or more of motion speed, motion direction, and area of an image region to which the moving target belongs, and the target shooting parameters comprise one or more of exposure time, camera movement direction, and magnification; Determining the target shooting parameters based on the weighted motion information comprises: Determining an exposure time corresponding to the motion speed according to the motion speed and a preset speed-exposure time mapping relationship, wherein the faster the motion speed, the shorter the corresponding exposure time; and / or, Using the motion direction as the camera movement direction; and / or, Determining a magnification corresponding to the area of the image region to which the moving target belongs according to the area of the image region to which the moving target belongs and a preset area-magnification mapping relationship, wherein the larger the area of the image region to which the moving target belongs, the smaller the corresponding magnification.

4. The method of claim 1, wherein, The at least one depth value difference is used to determine a plurality of first fusion weights corresponding to each pixel at the pixel position in each frame of the first fusion image corresponding to the group of cameras, including: For each pixel at the pixel position, a candidate index difference of the pixel is determined, and the candidate index difference is selected from one or more of the following: a luminance value difference, a contrast difference, and a texture difference; The depth value difference and each candidate index difference are subjected to a weighted operation to obtain a weighted difference, wherein the depth value difference has the largest weight in the weighted operation; The plurality of first fusion weights corresponding to each pixel at the pixel position are determined based on the weighted difference, and the larger the weighted difference is, the larger the difference between the plurality of first fusion weights is.

5. The method according to claim 1 or 4, characterized in that, In the case that there is a difference between the plurality of first fusion weights corresponding to each pixel at the same pixel position in each frame of the first fusion image corresponding to each group of cameras, the plurality of first fusion weights satisfy: the higher the image quality of the first fusion image to which the pixel belongs, the larger the first fusion weight corresponding to the pixel is.

6. The method of claim 1, wherein, The depth value difference between each pixel position and a surrounding pixel position of each pixel position of the depth map of the group of cameras is determined, including: For each pixel position of the depth map of the group of cameras, a plurality of depth differences between the depth value of the pixel position and the depth values of a plurality of pixel positions within a preset region range around the pixel position are determined respectively, thereby obtaining a plurality of depth differences corresponding to the pixel position; A mean square error of the plurality of depth differences is calculated, and the mean square error is taken as the depth value difference.

7. The method of claim 1, wherein, Image fusion is performed on the plurality of preliminary captured images to obtain a first fusion image corresponding to the camera, including: A preliminary captured image with the highest image quality is determined from the plurality of preliminary captured images, and is recorded as a reference frame image; One or more rounds of fusion operations are performed based on the reference frame image and the remaining preliminary captured images, and the fusion result of the last round is taken as the first fusion image corresponding to the camera; In each round of fusion, image fusion is performed on the fusion result of the previous round and one of the preliminary captured images that has not participated in image fusion to obtain the fusion result of the current round; The fusion result of the first round is obtained by performing image fusion on the reference frame image and one preliminary captured image other than the reference frame image.

8. The method of claim 7, wherein, In each round of fusion, the fusion result of the previous round is recorded as a first to-be-fused frame, and one of the preliminary captured images that has not participated in image fusion is recorded as a second to-be-fused frame; In each round of fusion, image fusion is performed on the fusion result of the previous round and one of the preliminary captured images that has not participated in image fusion to obtain the fusion result of the current round, including: For each pair of pixels at the same position in the first to-be-fused frame and the second to-be-fused frame, a pixel value difference of the pair of pixels is determined; A second fusion weight corresponding to the pair of pixels is determined based on the pixel value difference, and a weighted operation is performed on the pixel values of the pair of pixels by using the second fusion weight to obtain a fusion pixel value, wherein the larger the pixel value difference is, the larger the difference between the second fusion weights corresponding to the pair of pixels is. The fusion pixel value of each pair of pixels in the first to-be-fused frame and the second to-be-fused frame is used as the fusion result of the current round.

9. The method of claim 1, wherein, Before image fusion is performed on the multiple preliminary shooting images, the method further includes: An image with the highest image quality is determined from the multiple preliminary shooting images, and is recorded as a reference frame image; and preliminary shooting images other than the reference frame image are recorded as non-reference frame images. For each preliminary shooting image, an image region to which a moving target in the preliminary shooting image belongs is determined, and feature point extraction is performed on an image region other than the image region, to obtain a feature point set of the preliminary shooting image. For each non-reference frame image, an alignment matrix is determined based on the feature point set of the non-reference frame image and the feature point set of the reference frame image, and the non-reference frame image and the reference frame image are subjected to alignment processing by using the alignment matrix.

10. The method of claim 1, wherein, For each group of cameras, a depth map of the group of cameras relative to the target scene is determined according to multiple groups of preliminary shooting images synchronously shot by the group of cameras, including: For each group of cameras, a group of preliminary shooting images with the best image quality is selected from the multiple groups of preliminary shooting images synchronously shot by the group of cameras. The depth map of the group of cameras relative to the target scene is determined according to the group of preliminary shooting images with the best image quality.

11. A photographing apparatus, characterized by comprising: It includes: A first fusion module is configured to perform multiple shooting on a same target scene synchronously by using multiple cameras to obtain multiple preliminary shooting images corresponding to each camera, and perform image fusion on the multiple preliminary shooting images to obtain a first fusion image corresponding to the camera, wherein the multiple cameras are divided into at least one group of cameras, and each group has at least two cameras. A depth map determination module is configured to determine, for each group of cameras, a depth map of the group of cameras relative to the target scene according to multiple groups of preliminary shooting images synchronously shot by the group of cameras, each pixel position of the depth map having a respective depth value, wherein each group of preliminary shooting images includes at least two preliminary shooting images synchronously shot by cameras in the group. A fusion weight determination module is configured to determine, for each pixel position of the depth map of the group of cameras, a depth value difference between the pixel position and surrounding pixel positions, and determine, based on at least the depth value difference, multiple first fusion weights corresponding to each pixel belonging to the pixel position in each first fusion image corresponding to the group of cameras, wherein the greater the depth value difference, the greater the difference between the multiple first fusion weights corresponding to each pixel belonging to the pixel position. A multi-camera fusion module is configured to perform weighted operation on pixel values of each pixel by using multiple first fusion weights corresponding to each pixel belonging to each pixel position for each first fusion image corresponding to each group of cameras, to obtain a second fusion image corresponding to the group of cameras. A second fusion module is configured to perform image fusion on multiple second fusion images corresponding to each group of cameras to obtain a final shooting image.

12. A storage medium having stored thereon a computer program, characterized in that The computer program is run by a processor to execute the steps of the photographing method in any one of claims 1 to 10.

13. A terminal comprising a memory and a processor, said memory having stored thereon a computer program capable of running on said processor, characterized in that, The processor executes the steps of the photographing method of any one of claims 1 to 10 when the computer program is run by the processor.

14. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, performs the steps of the photographing method of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image fusion method and terminal, and computer readable storage medium

    CN107507160A

  • Image fusion method, electronic equipment, storage medium and computer program product

    CN113888452A