Electronic device and control method thereof

By integrating multiple cameras and neural network models in electronic devices, segmenting and adjusting image frame areas and optimizing camera parameters, the problem of poor image quality in multiple camera devices under low light conditions is solved, and high-quality image frame generation is achieved.

CN115868170BActive Publication Date: 2025-08-01SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180045892.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-03
Filing Date
2021-06-28
Publication Date
2025-08-01
Estimated Expiration
2041-06-28

AI Technical Summary

Technical Problem

The prior art is difficult to generate high-quality image frames in electronic devices using multiple cameras, especially in low-light conditions, with poor image quality and ghosting and noise problems.

Method used

By integrating multiple cameras and processors in electronic devices, using neural network models to segment image frames into multiple regions, adjust camera parameters according to area characteristics, and merge image frames to generate high-quality image frames. The method includes using the first neural network model to obtain a set of camera parameter setting values, the second neural network model optimizes image indexes, and performs image merging and processing through the third neural network model.

Benefits of technology

Improve the quality of image frames, especially in low-light conditions, reduce ghosting and noise, and generate image frames with clear edges and uniform brightness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115868170B_ABST
    Figure CN115868170B_ABST
Patent Text Reader

Abstract

An electronic device includes a plurality of cameras and at least one processor connected to the plurality of cameras. The at least one processor is configured to, based on a first user command to obtain a live view image, segment an image frame obtained via a camera among the plurality of cameras into a plurality of regions based on the luminance and objects of pixels included in the image frame; obtain a plurality of sets of camera parameter setting values, each set including a plurality of parameter values regarding the plurality of regions; based on a second user command to capture the live view image, obtain a plurality of image frames using the plurality of sets of camera parameter setting values and at least one camera among the plurality of cameras; and obtain an image frame by merging the plurality of obtained image frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an electronic device and a method for controlling the electronic device, and more particularly, to an electronic device that obtains an image frame based on a user command for capturing and a method for controlling the electronic device.

[0002] Cross - Reference to Related Applications

[0003] This application is based on and claims priority to Korean Patent Application No. 10 - 2020 - 0167701, filed with the Korean Intellectual Property Office on December 3, 2020, and Indian Patent Application No. 202041027623, filed with the Indian Patent Office on June 29, 2020, the disclosures of which are incorporated herein by reference in their entirety. Background Art

[0004] In recent years, with the development of electronic technology, electronic devices including multiple cameras have been variously developed.

[0005] When generating an image frame based on a capture command of a user using an electronic device including multiple cameras, it is necessary to provide an image frame of higher quality. Summary of the Invention

[0006] Technical Problem

[0007] Provided is an electronic device that obtains a plurality of parameter values using a live view image and obtains an image frame corresponding to a command using the plurality of parameter values, and a method for controlling the electronic device.

[0008] Technical Solution

[0009] According to an aspect of the present disclosure, there is provided an electronic device including a plurality of cameras and at least one processor communicatively connected to the plurality of cameras, wherein the at least one processor is configured to, based on a first user command for obtaining a live view image, divide an image frame into a plurality of regions based on the luminance and objects of pixels included in the image frame obtained via a camera among the plurality of cameras; obtain a plurality of sets of camera parameter setting values, each set of camera parameter setting values including a plurality of parameter values respectively regarding the plurality of regions; based on a second user command for capturing the live view image, obtain a plurality of image frames by applying the plurality of sets of camera parameter setting values to at least one of the plurality of cameras; and obtain an image frame corresponding to the second user command by merging the plurality of obtained image frames.

[0010] The at least one processor may be configured to obtain a plurality of first sets of camera parameter settings by inputting a plurality of regions into a first neural network model, and to obtain a plurality of sets of camera parameter settings by inputting the plurality of first sets of camera parameter settings, an image frame obtained via the camera, and a plurality of parameter values of the camera into a second neural network model.

[0011] The first neural network model may be a neural network model trained based on regions obtained from the image frame, a plurality of parameter values of the camera that has obtained the image frame, and edge metrics, contrast metrics, and background metrics of the regions.

[0012] Based on the input of the set of camera parameter settings, the image frame obtained via the camera, and the plurality of parameter values of the camera, the second neural network model may be configured to output information about the edge metrics, contrast metrics, and background metrics of the input image frame corresponding to the input set of camera parameter settings, and the at least one processor may be configured to obtain, based on the information about the edge metrics, contrast metrics, and background metrics obtained from the second neural network model, the set of camera parameter settings corresponding to the maximum edge metric, the set of camera parameter settings corresponding to the maximum contrast metric, and the set of camera parameter settings corresponding to the maximum background metric in the plurality of first sets of camera parameter settings.

[0013] The at least one processor may be configured to identify pixel values of pixels included in each set of camera parameter settings of the plurality of regions, and to obtain the plurality of sets of camera parameter settings based on the identified pixel values and predefined rules.

[0014] The at least one processor may be configured to, based on a second user command to capture a live view image, obtain an image frame by applying one set of camera parameter settings from the plurality of sets of camera parameter settings to a first camera among the plurality of cameras, and obtain at least two image frames by applying at least two sets of camera parameter settings from the plurality of sets of camera parameter settings to a second camera among the plurality of cameras.

[0015] The at least one processor may be configured to obtain an image frame corresponding to the second user command by inputting the merged image frame into a third neural network model.

[0016] Each of the plurality of obtained image frames may be a Bayer raw image.

[0017] The image frame obtained from the third neural network model may be an image frame that has undergone at least one of black level adjustment, color correction, gamma correction, or edge enhancement.

[0018] According to another aspect of the present disclosure, there is provided a method for controlling an electronic device including a plurality of cameras, the method including: based on a first user command to obtain a live view image, segmenting an image frame obtained by a camera among the plurality of cameras into a plurality of regions based on the brightness and objects of pixels included in the image frame; obtaining a plurality of sets of camera parameter setting values, each set of camera parameter setting values including a plurality of parameter values respectively regarding the plurality of regions; based on a second user command to capture the live view image, obtaining a plurality of image frames by applying the plurality of sets of camera parameter setting values to the cameras among the plurality of cameras; and obtaining an image frame corresponding to the second user command by merging the plurality of obtained image frames.

[0019] Obtaining a plurality of sets of camera parameter setting values may include: obtaining a plurality of first sets of camera parameter setting values by inputting the plurality of regions into a first neural network model, and obtaining a plurality of sets of camera parameter setting values by inputting the plurality of first sets of camera parameter setting values, an image frame obtained via a camera, and a plurality of parameter values of the camera into a second neural network model.

[0020] The first neural network model may be a neural network model trained based on regions obtained from an image frame, a plurality of parameter values of a camera that has obtained the image frame, and edge metrics, contrast metrics, and background metrics of the regions.

[0021] Based on the input sets of camera parameter setting values, an image frame obtained by a camera, and a plurality of parameter values of the camera, the second neural network model may be configured to output information on edge metrics, contrast metrics, and background metrics of the input image frame corresponding to the input set of camera parameter setting values, and obtaining a plurality of sets of camera parameter setting values may include: based on the information on edge metrics, contrast metrics, and background metrics obtained from the second neural network model, obtaining a set of camera parameter setting values corresponding to the maximum edge metric, a set of camera parameter setting values corresponding to the maximum contrast metric, and a set of camera parameter setting values corresponding to the maximum background metric in the plurality of first sets of camera parameter setting values.

[0022] Obtaining a plurality of sets of camera parameter setting values may include: identifying pixel values of pixels included in each of the plurality of regions, and obtaining a plurality of sets of camera parameter setting values based on the identified pixel values and predefined rules.

[0023] Obtaining a plurality of image frames may include: based on a second user command to capture the live view image, obtaining an image frame by applying one set of camera parameter setting values among the plurality of sets of camera parameter setting values to a first camera among the plurality of cameras, and obtaining at least two image frames by applying at least two sets of camera parameter setting values among the plurality of sets of camera parameter setting values to a second camera among the plurality of cameras.

[0024] Obtaining an image frame may include obtaining an image frame corresponding to a second user command by inputting the merged image frame into a third neural network model.

[0025] Each of the plurality of obtained image frames may be a Bayer raw image.

[0026] The image frame obtained from the third neural network model may be an image frame that has undergone at least one of black level adjustment, color correction, gamma correction, or edge enhancement.

[0027] The plurality of parameter values included in each set of camera parameter setting values may include an International Organization for Standardization (ISO) value, a shutter speed, and an aperture, and the aperture of at least one camera to which the plurality of sets of camera parameter setting values are applied may have a size equal to the aperture included in the plurality of sets of camera parameter setting values.

[0028] According to another aspect of an exemplary embodiment of the present disclosure, there is provided a non-transitory computer-readable storage medium storing program code that can be executed by at least one processor to cause the at least one processor to control an electronic device including a plurality of cameras by performing the foregoing method. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The above and other aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0030] Figure 1 is a diagram schematically showing an electronic device according to an embodiment;

[0031] Figure 2 is a flowchart showing a method for controlling an electronic device according to an embodiment;

[0032] Figure 3A is a diagram showing a method for dividing an image frame into a plurality of regions according to an embodiment;

[0033] Figure 3B is a diagram showing a method for dividing an image frame into a plurality of regions according to an embodiment;

[0034] Figure 3C is a diagram showing a method for dividing an image frame into a plurality of regions according to an embodiment;

[0035] Figure 3D is a diagram showing a method for dividing an image frame into a plurality of regions according to an embodiment;

[0036] Figure 4 is a flowchart showing a method for obtaining a plurality of sets of camera parameter setting values according to an embodiment;

[0037] Figure 5 is a diagram showing a process of training a first neural network model according to an embodiment;

[0038] Figure 6 is a diagram showing an example of a method of obtaining a plurality of first camera parameter setting value sets using a first neural network model according to an embodiment;

[0039] Figure 7 is a diagram showing an example of a method of obtaining a plurality of camera parameter setting value sets using a second neural network model according to an embodiment;

[0040] Figure 8 is a diagram showing a plurality of image frames obtained using a plurality of cameras according to an embodiment;

[0041] Figure 9 is a flowchart showing a method of obtaining an image frame by merging a plurality of image frames according to an embodiment;

[0042] Figure 10 is a diagram showing a plurality of Gaussian pyramids according to an embodiment;

[0043] Figure 11 is a diagram showing a method for merging a plurality of image frames according to an embodiment;

[0044] Figure 12 is a diagram showing a method of obtaining an image frame using a third neural network model according to an embodiment;

[0045] Figure 13 is a diagram showing an example of a method of obtaining a plurality of camera parameter setting value sets using a first neural network model according to an embodiment;

[0046] Figure 14 is a flowchart showing an example of a method of obtaining a plurality of camera parameter setting value sets using rules according to an embodiment;

[0047] Figure 15 is a flowchart showing an example of a method of obtaining a plurality of camera parameter setting value sets using rules according to an embodiment;

[0048] Figure 16 is a block diagram showing a configuration of an electronic device according to an embodiment;

[0049] Figure 17 is a block diagram showing more specifically a configuration of an electronic device according to an embodiment;

[0050] Figure 18A is a diagram showing an example of a method for calculating an image metric according to an embodiment;

[0051] Figure 18B is a diagram showing an example of a method for calculating an image metric according to an embodiment;

[0052] Figure 18C is a diagram showing an example of a method for calculating an image metric according to an embodiment;

[0053] Figure 19 is a diagram showing the structure of a first neural network model according to an embodiment;

[0054] Figure 20 is a diagram showing the structure of a first neural network model according to an embodiment;

[0055] Figure 21 is a diagram showing the training of a first neural network model according to an embodiment;

[0056] Figure 22 is a diagram showing the structure of a second neural network model according to an embodiment;

[0057] Figure 23A is a diagram showing the structure of a third neural network model according to an embodiment; and

[0058] Figure 23B is a diagram showing the structure of a third neural network model according to an embodiment. Detailed Description of the Invention

[0059] Hereinafter, embodiments will be described in more detail with reference to the accompanying drawings. These embodiments can be variously changed and may include other embodiments. Specific embodiments will be shown in the drawings and described in detail in the specification. It should be noted that these embodiments are not intended to limit the scope of the present disclosure, but should be construed as including all modifications, equivalents, and / or alternatives of the exemplary embodiments of the present disclosure. Regarding the interpretation of the drawings, similar reference numerals may be used for similar elements.

[0060] When describing the present disclosure, the detailed description of related technologies or configurations may be omitted when it is determined that the detailed description may unnecessarily obscure the gist of the present disclosure.

[0061] In addition, the embodiments described herein can be changed in various forms, and the scope of the present disclosure is not limited to these embodiments. The embodiments are provided to make the present disclosure complete and to fully inform those of ordinary skill in the art of the scope of the present disclosure.

[0062] The terms used in the present disclosure are only for describing specific embodiments and do not limit the scope of the present disclosure. Unless specifically defined, singular expressions may include plural expressions.

[0063] In the present disclosure, terms such as "comprising", "may comprise", "consisting of", or "may consist of" are used herein to indicate the presence of corresponding features (e.g., constituent elements such as quantity, function, operation, or components), and do not exclude the presence of additional features.

[0064] In the present disclosure, expressions such as "A or B", "at least one of A [and / or] B", or "one or more of A [and / or] B" include all possible combinations of the listed items. For example, "A or B", "at least one of A and B", or "at least one of A or B" includes (1) at least one A, (2) at least one B, or (3) any one of at least one A and at least one B.

[0065] In the present disclosure, expressions such as "first", "second", etc. may represent various elements regardless of order and / or importance, and may be used to distinguish one element from another, and do not limit these elements.

[0066] If a specific element (e.g., the first element) is described as "operatively or communicatively coupled to" or "connected to" another element (e.g., the second element), it should be understood that the specific element may be directly or through yet another element (e.g., the third element) connected to the other element.

[0067] On the other hand, if a specific element (e.g., the first element) is described as "directly coupled to" or "directly connected to" another element (e.g., the second element), it can be understood that there is no element (e.g., the third element) between the specific element and the other element.

[0068] Furthermore, depending on the embodiment, the expression "configured to" used in the present disclosure may be interchangeably used with other expressions such as "suitable for", "capable of...", "designed to", "adapted to", "manufactured to", and "able to". The expression "configured to" does not necessarily refer to a device "specially designed" in terms of hardware.

[0069] On the contrary, in some cases, the expression "the device is configured to" may mean that the device "is able to" perform an operation together with another device or component. For example, the phrase "a unit or processor configured (or set) to perform an operation" may refer to, for example but not limited to, a dedicated processor (e.g., an embedded processor), a general-purpose processor (e.g., a central processing unit (CPU) or an application processor), etc. that is able to perform the corresponding operation by executing one or more software programs stored in a memory device.

[0070] In an embodiment, terms such as "module" or "unit" can perform at least one function or operation and can be implemented as hardware, software, or a combination of hardware and software. In addition, components can be integrated into at least one module and implemented in at least one processor, except when each of a plurality of "modules", "units", etc. needs to be implemented in specific hardware.

[0071] The various elements and regions in the drawings are schematically illustrated. Accordingly, the scope of the present disclosure is not limited by the relative sizes or intervals shown in the drawings.

[0072] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art to which the present disclosure pertains can easily practice the present disclosure.

[0073] Figure 1 is a diagram schematically showing an electronic device according to an embodiment.

[0074] As Figure 1 shown, the electronic device 100 according to an embodiment of the present disclosure may include a plurality of cameras 111, 112, and 113.

[0075] For example, the electronic device 100 may be implemented as a Figure 1 smartphone as shown. However, the electronic device 100 according to the present disclosure is not limited to a specific type of device and may be implemented as various types of electronic devices, such as a tablet personal computer (PC), a personal digital assistant (PDA), a smartwatch, a laptop computer, a virtual reality (VR) device, an Internet of Things (IoT) device, and a digital camera.

[0076] Each of the plurality of cameras 111, 112, and 113 according to the present disclosure may include an image sensor and a lens.

[0077] Here, the field of view (FOV) of the lenses may be different from each other. For example, as Figure 1 shown, the plurality of cameras 111, 112, and 113 may include a telephoto lens, a wide-angle lens, and an ultra-wide-angle lens disposed on the rear surface of the electronic device 100. However, the number and type of the lenses according to the present disclosure are not particularly limited.

[0078] Figure 2 is a flowchart showing a method for controlling an electronic device according to an embodiment.

[0079] As Figure 2 shown, the electronic device 100 according to an embodiment of the present disclosure may receive a first user command for obtaining a live view image (S210).

[0080] The live view image in this document may refer to an image displayed on the display of the electronic device 100, which is obtained by converting the light incident through the lens of the camera into an electrical image signal via the image sensor of the camera.

[0081] The user command may be, for example, a user command for driving a camera application stored in the electronic device 100. Such a user command may be received based on a user touch input via the display of the electronic device 100, a user voice received via the microphone of the electronic device 100, an input via a physical button provided on the electronic device 100, a control signal transmitted by a remote control device for controlling the electronic device 100, and the like.

[0082] Based on the received user command for obtaining a live view image, the electronic device 100 may use the camera to obtain a plurality of image frames (S220).

[0083] Hereinafter, a description will be made based on the assumption that the electronic device 100 obtains a plurality of image frames by driving the camera 111 among the plurality of cameras 111, 112, and 113. However, this is only an example for describing the operation of the electronic device 100 according to the present disclosure, and does not limit the operation of the electronic device 100. For example, the electronic device 100 may obtain a plurality of image frames by driving one of the cameras 112 and 113.

[0084] When the camera 111 is used to obtain a plurality of image frames, the camera 111 may obtain a plurality of image frames based on a plurality of parameter values of the camera 111.

[0085] Hereinafter, the plurality of parameters may include the International Organization for Standardization (ISO), shutter speed, and aperture. However, this is only an example, and the plurality of parameters are not limited thereto.

[0086] ISO may be a parameter for controlling the responsiveness (or sensitivity) of the image sensor to light. Although the same amount of light is incident on the image sensor of the camera, the brightness of the image frame varies depending on the ISO value set for the camera. For example, although the same amount of light is incident on the image sensor of the camera, as the ISO value becomes larger, the image frame becomes brighter, and as the ISO value becomes smaller, the image frame becomes darker.

[0087] The shutter speed may be a parameter for controlling the length of the exposure time of the image sensor. According to the shutter speed value set for the camera, the length of time for which the image sensor is exposed to obtain an image frame can be determined. For example, when the shutter speed is high, the exposure time of the image sensor is reduced, and when the shutter speed is low, the exposure time of the image sensor is extended.

[0088] The aperture can be a parameter related to the area where light enters the image sensor via the aperture (i.e., the hole of the aperture), and can be represented by an f-number, which is the ratio of the focal length of the lens to the aperture diameter. The amount of light incident on the image sensor can vary depending on the size of the area corresponding to the aperture. For example, when the aperture value (e.g., f-number) is small, the size of the area is large, and a larger amount of light can be incident on the image sensor.

[0089] All of the multiple cameras can have a fixed aperture value. In other words, the size of the hole of the aperture can be unchanged, and a fixed size can be provided for each camera.

[0090] Therefore, camera 111 can obtain multiple image frames based on the ISO value, shutter speed value, and aperture value of camera 111.

[0091] In other words, the electronic device 100 can use the ISO value and shutter speed value determined to correspond to camera 111 to set the ISO and shutter speed of camera 111, and obtain multiple image frames by driving camera 111 based on the set ISO value and shutter speed value. The hole of the aperture of camera 111 can have a size corresponding to the aperture value.

[0092] In the present disclosure, it is assumed that the ISO values and shutter speed values of multiple cameras 111, 112, and 113 are normalized to values between 0 and 1 (e.g., 0.1, 0.2, 0.3,....... 0.8, 0.9, 1.0). In addition, in the present disclosure, it is assumed that the aperture value of camera 111 is 1, the aperture value of camera 112 is 2, and the aperture value of camera 113 is 3. Here, the aperture values of 1, 2, or 3 do not mean that the actual aperture values of cameras 111, 112, and 113 are 1, 2, or 3, but are set to distinguish cameras 111, 112, and 113 according to the aperture value. In other words, the aperture value of each of cameras 111, 112, and 113 can be, for example, one of various aperture values such as F / 1.8, F / 2.0, F / 2.2, F / 2.4, and F / 2.8.

[0093] In addition, the electronic device 100 can use the multiple image frames to display a live view image on the display of the electronic device 100.

[0094] Specifically, the electronic device 100 can temporarily store the multiple image frames obtained by camera 111 in a volatile memory (such as the frame buffer of the electronic device 100), and use the multiple image frames to display a live view image on the display of the electronic device 100.

[0095] The electronic device 100 may obtain a plurality of regions by dividing an image frame (S230). For example, the electronic device 100 may obtain an image frame through the camera 111 and obtain a plurality of regions by dividing the image frame.

[0096] Specifically, the electronic device 100 may obtain a plurality of regions by dividing the image frame based on the luminance of the pixels included in the image frame and an object.

[0097] First, the electronic device 100 may obtain a plurality of regions by dividing the image frame based on the luminance of the pixels included in the image frame.

[0098] The luminance of the pixels in this document may refer to the pixel value of the pixel (e.g., the pixel value of the pixel in a grayscale image).

[0099] In addition, the division based on the luminance of the pixels may mean that the division is based on the light intensity. This is because the luminance of the pixels corresponding to the subject in the image frame obtained through the camera varies depending on the light intensity applied to the subject.

[0100] In this case, the electronic device 100 may obtain information about the plurality of regions divided from the image frame by using a neural network model. The neural network model in this document may refer to a model trained by image frames, each of which is divided into a plurality of regions based on the luminance of the pixels.

[0101] In other words, the electronic device 100 may input the image frame into the neural network model and obtain information about the plurality of regions divided from the image frame based on the luminance of the pixels from the neural network model.

[0102] In this document, inputting the image frame into the neural network model may refer to inputting the pixel values of the pixels included in the image frame (e.g., R, G, and B pixel values). In addition, the information about the plurality of regions may include the position information of the plurality of regions. In this case, the position information may include the coordinate values of the pixels included in each region.

[0103] For example, as Figure 3A shown, the electronic device 100 may obtain a plurality of regions 311, 312, 313, and 314 divided from the image frame 310 based on the luminance of the pixels. In this case, regions with a luminance equal to or greater than a threshold luminance are determined from the image frame 310 based on the luminance, and the regions with a luminance equal to or greater than the threshold luminance may be divided into a plurality of regions 311, 312, 313, and 314.

[0104] In addition, the electronic device 100 may obtain a plurality of regions by dividing the image frame based on the object included in the image frame.

[0105] The objects in this document can include various types of objects with various structures (or shapes), such as people, animals, plants, buildings, objects, landforms, etc.

[0106] The electronic device 100 can use a neural network model to obtain information about multiple regions segmented from an image frame based on the objects included in the image frame.

[0107] The neural network model in this document can refer to a model trained by image frames, and each image frame is divided into multiple regions based on objects. For example, the neural network model can refer to an object segmentation model. The object segmentation model can identify the objects included in the input image frame, the structures of the objects, etc., and output information about the object regions corresponding to the objects.

[0108] In other words, the electronic device 100 can input the image frame into the neural network model and obtain information about multiple regions segmented from the image frame based on the objects from the neural network model.

[0109] In this document, inputting the image frame into the neural network model can refer to inputting the pixel values of the pixels included in the image frame (e.g., R, G, and B pixel values). In addition, the information about the multiple regions can include the position information of each of the multiple regions. In this case, the position information can include the coordinate values of the pixels included in each region.

[0110] For example, as Figure 3B shown, the electronic device 100 can obtain multiple regions 321, 322, 323, and 324 segmented from the image frame 310 based on the objects.

[0111] Then, the electronic device 100 can segment the image frame 310 into multiple regions according to the multiple regions segmented based on the pixel brightness and the multiple regions segmented based on the objects included in the image frame 310.

[0112] The electronic device 100 can use k - means clustering to segment the image frame into multiple regions. The k - means clustering in this document can refer to an algorithm that divides data into k clusters based on the distance between the data and the centroid of each cluster.

[0113] Specifically, the electronic device 100 can identify multiple clusters according to the multiple regions segmented from the image frame based on the pixel brightness and the multiple regions segmented from the image frame based on the objects.

[0114] The electronic device 100 can use k - means clustering to segment the image frame for each of the identified multiple clusters.

[0115] For example, if as Figure 3AThe pixel-based luminance shown divides the image frame 310 into four regions 311, 312, 313, and 314, and as Figure 3B shown, the object-based division of the image frame 310 into four regions 321, 322, 323, and 324, the electronic device 100 may assume that there are a total of eight clusters (i.e., the four regions based on pixel-based luminance and the four regions based on objects).

[0116] As described above, if there are eight clusters, the electronic device 100 may use k-means clustering (k = 1, 2,......, 8) to divide the image frame into one region, two regions,......, and eight regions based on the pixel values of the pixels.

[0117] Then, the electronic device 100 may calculate the entropy of each divided image frame and identify the image frame with the maximum calculated entropy.

[0118] As in the above example, if the image frame is divided into one region, two regions,......, or eight regions, there may be an image frame divided into one region, an image frame divided into two regions,......, and an image frame divided into eight regions.

[0119] In this case, the electronic device 100 may calculate the entropy E(i) of each divided image frame based on the following mathematical expression 1.

[0120]

Mathematical Expression 1

[0121]

[0122] Here, satisfying n represents the number of pixels in each region of the divided image frame, and N represents the total number of pixels in each divided image frame.

[0123] For example, as Figure 3C shown, assume that the image frame 330 configured with a total of 16 pixels is divided into two regions 331 and 332 using k-means clustering (k = 2).

[0124] If both regions 331 and 332 consist of eight pixels, then P1 for region 331 is 8 / 16, and P2 for region 332 is 8 / 16. Therefore, the entropy E of the image frame can be calculated as:

[0125]

[0126] Through the above method, the electronic device 100 can calculate the entropy of the image frame divided into one region, the entropy of the image frame divided into two regions,......, and the entropy of the image frame divided into eight regions.

[0127] In addition, the electronic device 100 may divide an image frame into a plurality of regions based on the image frame having the maximum computational entropy. In other words, the electronic device 100 may identify the plurality of regions divided from the image frame having the maximum computational entropy as a plurality of regions divided from the image frame based on the luminance of pixels and objects.

[0128] In the above example, if the entropy of the image frame divided into four regions has the maximum value among the calculated entropies, the electronic device 100 may divide the image frame 310 into four regions 341, 342, 343, and 344, as Figure 3D shown.

[0129] Therefore, through the above method, the electronic device 100 may divide an image frame into a plurality of regions considering the luminance of pixels and objects.

[0130] The electronic device 100 may obtain a plurality of sets of camera parameter setting values, each set including a plurality of parameter values based on the plurality of regions (S240).

[0131] The set of camera parameter setting values in this document may be configured with a plurality of parameter values. For example, the set of setting values may include an ISO value, a shutter speed value, and an aperture value.

[0132] The electronic device 100 may obtain a plurality of sets of camera parameter setting values using a neural network model.

[0133] Figure 4 is a flowchart illustrating a method for obtaining a plurality of sets of camera parameter setting values according to an embodiment. Referring to Figure 4 , the electronic device 100 may obtain a plurality of first sets of camera parameter setting values by inputting the plurality of regions into a first neural network model (S410).

[0134] In this document, inputting a region into the first neural network model may refer to inputting pixel values (e.g., R, G, and B pixel values) of pixels included in the region.

[0135] The first neural network model may output a plurality of sets of camera parameter setting values respectively corresponding to the plurality of input regions.

[0136] The first neural network model may output sets of camera parameter setting values respectively corresponding to the plurality of input regions.

[0137] The set of camera parameter setting values in this document may include an ISO value, a shutter speed value, and an aperture value.

[0138] The expression "corresponding to..." may mean that the set of camera parameter setting values output from the first neural network model includes multiple desired (or optimal) parameter values for the input region. "Desired" in this context may mean that if the input region is obtained by a camera using the set of camera parameter setting values output from the first neural network model, the corresponding region will be further enhanced in terms of an image index.

[0139] To this end, the first neural network model can be trained based on the region obtained from the image frame, the multiple parameter values of the camera that has obtained the image frame, and the image index of the region.

[0140] The image index in this context may include an edge index, a contrast index, and a background index (or background noise index).

[0141] Hereinafter, reference will be made to Figure 5 describe the process of training the first neural network model in more detail.

[0142] Reference Figure 5 , the image frame 510 can be divided into multiple regions 521, 522, and 523.

[0143] For example, based on the information corresponding to the position of the object obtained by the object detection model, the image frame 510 can be divided into multiple regions 521, 522, and 523. The object detection model can output information about the region of the object corresponding to the shape of the object included in the input image frame 510.

[0144] Then, the edge index, contrast index, and background index of each of the multiple regions 521, 522, and 523 can be calculated.

[0145] The edge index can be an index for indicating the sharpness of the image frame (or a region of the image frame). A large edge index value indicates a clear image frame.

[0146] The contrast index can be an index for indicating the brightness of the image frame (or a region of the image frame). A large contrast index indicates a bright image frame.

[0147] The background index can be an index for indicating the degree of noise in the background region of the image frame (or a region of the image frame). A large background index indicates that the image frame includes a background region with less noise.

[0148] The enhancement ratio of each of the multiple regions 521, 522, and 523 can be calculated using the edge index, contrast index, and background index calculated for each of the multiple regions 521, 522, and 523.

[0149] Specifically, when the edge index, contrast index, and background index calculated for a given segmented region are defined as Edgeindex, Contrastindex, and Backgroundindex respectively, the enhancement ratio for the corresponding region can be calculated as In other words, the enhancement ratio can be the average of the edge index, contrast index, and background index.

[0150] Multiple regions 521, 522, and 523 can be converted into separate image frames 531, 532, and 533.

[0151] Converting multiple regions 521, 522, and 523 into separate image frames 531, 532, and 533 can mean obtaining the pixel values (e.g., R, G, and B pixel values) of the pixels included in each of the regions 521, 522, and 523, and generating the image frames 531, 532, and 533 by using the obtained pixel values.

[0152] In this case, each of the image frames 531, 532, and 533 can be the input data of the first neural network model. In other words, the input data can include the pixel values of the pixels included in each of the regions 521, 522, and 523.

[0153] Furthermore, the output data of the first neural network model corresponding to each of the multiple image frames (e.g., input data) 531, 532, and 533 can be obtained by multiplying the calculated enhancement ratios of each of the multiple regions 521, 522, and 523 by the multiple parameter values of the camera used to obtain the image frame 510.

[0154] For example, referring to Figure 5 , it is assumed that the image frame 510 is obtained by a camera (not shown) with ISO value, shutter speed value, and aperture value of 0.5, 0.3, and 1 respectively, and the enhancement ratios of the regions 521, 522, and 523 are calculated as 0.53, 0.31, and 0.16 respectively.

[0155] In this case, the output data corresponding to the image frame 531 can be obtained as 0.53 x (^0.5, 0.3, 1^) = (^0.265, 0.159, 0.53^). Furthermore, the output data corresponding to the image frame 532 can be obtained as 0.31 x (^0.5, 0.3, 1^) = (^0.155, 0.093, 0.31^), and the output data corresponding to the image frame 533 can be obtained as 0.16 x (^0.5, 0.3, 1^) = (^0.08, 0.048, 0.16^).

[0156] In the following, a first neural network model can be trained based on input / output data (e.g., input data and output data pairs). In this case, the first neural network model can be trained based on the input / output data obtained from a large number of image frames by referring to the method described in Figure 5 The method described above is used to train the first neural network model based on the input / output data obtained from a large number of image frames.

[0157] Specifically, a first neural network model that outputs an ISO value, a shutter speed value, and an aperture value with respect to the input data can be generated. For ease of description, each of the ISO value and the shutter speed value in this document can be a value between 0 and 1, and the aperture value can be values of 1, 2, and 3.

[0158] In the following, the input data can be input into the first neural network model, and the first neural network model can be trained to minimize the loss (e.g., loss function) between the ISO value, the shutter speed value, and the aperture value output from the first neural network model and the output data.

[0159] As described above, in the present disclosure, the enhancement ratio is calculated based on an image metric indicating the quality of a region. Therefore, the enhancement ratio can be an indicator for evaluating whether multiple parameter values of a camera used to obtain a region are appropriate values in terms of the image metric for the corresponding region. Therefore, if the enhancement ratio is used as a weight factor for multiple parameter values of a camera used to obtain a region during the training of the first neural network model, the first neural network model can be trained to output multiple parameter values for enhancing the image metric of the input image.

[0160] As described above, the electronic device 100 can obtain multiple sets of first camera parameter setting values by inputting multiple regions into the first neural network model.

[0161] Figure 6 FIG. is an example diagram showing a method of obtaining multiple sets of first camera parameter setting values using the first neural network model according to an embodiment. Referring to Figure 6 , the electronic device 100 can obtain each of multiple sets of first camera parameter setting values 621, 622, 623, and 624 by inputting each of multiple regions 611, 612, 613, and 614 into the first neural network model 121.

[0162] A set of camera parameter setting values can include (ISO value, shutter speed value, aperture value). For example, the set of camera parameter setting values corresponding to region 611 can be (0.2, 0.1, 1) 621, the set of camera parameter setting values corresponding to region 612 can be (0.1, 0.3, 0.1) 622, the set of camera parameter setting values corresponding to region 613 can be (0.4, 0.6, 2) 623, and the set of camera parameter setting values corresponding to region 614 can be (0.6, 0.7, 3) 625.

[0163] The electronic device 100 may obtain a plurality of sets of camera parameter setting values (S420) by inputting a plurality of sets of first camera parameter setting values, image frames obtained by the camera 111, and a plurality of parameter values of the camera 111 into a second neural network model.

[0164] In this document, inputting an image frame into the second neural network model may refer to inputting the pixel values of the pixels included in the image frame (e.g., R, G, and B pixel values).

[0165] Based on the input sets of camera parameter setting values, image frames, and the plurality of parameter values of the camera used to obtain the image frames, the second neural network model may output information about the edge metric, contrast metric, and background metric of the image frame corresponding to the input set of camera parameter setting values.

[0166] Specifically, when the camera obtains an input image frame using the input set of camera parameter setting values, the second neural network model may estimate the edge metric, contrast metric, and background metric of the image frame based on the input image frame and the plurality of parameter values used to obtain the input image frame, and output information about the estimated edge metric, contrast metric, and background metric.

[0167] In this case, the second neural network model may be, for example, a parameter estimation network. The second neural network model may be trained based on image frames obtained by capturing images of an object using various parameter values of the camera, the parameter values used to obtain the image frames, the image frame with the highest edge metric among the obtained image frames, the image frame with the highest contrast metric, and the image frame with the highest background metric. However, this is only an example, and the second neural network model may be trained by various methods.

[0168] Therefore, by inputting each of the plurality of sets of first camera parameter setting values obtained from the first neural network model together with the image frames and the plurality of parameter values of the camera 111 used to obtain the image frames into the second neural network model, the electronic device 100 may obtain information about the edge metric, contrast metric, and background metric corresponding to each of the plurality of sets of first camera parameter setting values.

[0169] The electronic device 100 may select a plurality of sets of camera parameter setting values from the plurality of sets of first camera parameter setting values based on the information about the edge metric, contrast metric, and background metric obtained from the second neural network model.

[0170] Specifically, the electronic device 100 may obtain a set of camera parameter settings with the maximum edge metric, a set of camera parameter settings with the maximum contrast metric, and a set of camera parameter settings with the maximum background metric among multiple first sets of camera parameter settings based on the information about the edge metric, contrast metric, and background metric obtained from the second neural network model.

[0171] Figure 7 FIG. is an example diagram showing a method of obtaining multiple sets of camera parameter settings using a second neural network model according to an embodiment. Refer to Figure 7 , the electronic device 100 may select the set of camera parameter settings 621 among the multiple first sets of camera parameter settings 621, 622, 623, and 624, and input the ISO value, shutter speed value, and aperture value included in the set of camera parameter settings 621 (i.e., (ISO, shutter speed, aperture) = (0.2, 0.1, 1)), the image frame 310, and the ISO value, shutter speed value, and aperture value of the camera 111 for obtaining the image frame 310 (i.e., (ISO, shutter speed, aperture) = (0.2, 0.5, 3)) 710 to the second neural network model 122 to obtain information 720 about the edge metric, contrast metric, and background metric corresponding to the set of camera parameter settings 621.

[0172] The electronic device 100 may obtain information about the edge metric, contrast metric, and background metric corresponding to each of the sets of camera parameter settings 622, 623, and 624 by performing the above process on the sets of camera parameter settings 622, 623, and 624.

[0173] When assuming that the edge metric corresponding to the set of camera parameter settings 622 is the highest among the obtained edge metrics, the contrast metric corresponding to the set of camera parameter settings 623 is the highest among the obtained contrast metrics, and the background metric corresponding to the set of camera parameter settings 624 is the highest among the obtained background metrics, the electronic device 100 may obtain the sets of camera parameter settings 622, 623, and 624 among the multiple first sets of camera parameter settings 621, 622, 623, and 624.

[0174] Therefore, the electronic device 100 may obtain multiple sets of camera parameter settings through the above method.

[0175] Returning to reference Figure 2 , the electronic device 100 may receive a second user command (S250) for capturing a live view image.

[0176] The capture of a live view image may refer to obtaining an image frame corresponding to the timing at which a user command for capture is input, so as to store the image frame in the electronic device 100.

[0177] Storing the image frame in this case may be different from storing the image frame of the live view image in a volatile memory (such as a frame buffer), and may mean that the image frame is stored in a non-volatile memory (such as the flash memory of the electronic device 100). If the gallery application is driven according to a user command, the image frame stored as described above may be displayed on the display of the electronic device 100 through the gallery application.

[0178] The second user command may be, for example, a user command for selecting an image capture button included in a user interface provided by the camera application. Without limiting this example, the second user command may be received in various ways, such as based on a user touch input via the display of the electronic device 100, a user voice received via the microphone of the electronic device 100, an input via a physical button provided on the electronic device 100, a control signal transmitted by a remote control device for controlling the electronic device 100, etc.

[0179] Based on receiving the second user command (S250 - Yes) for capturing a live view image, the electronic device 100 may obtain a plurality of image frames (S260) using a plurality of sets of camera parameter setting values and a camera among the plurality of cameras 111, 112, and 113.

[0180] To this end, the electronic device 100 may identify the set of camera parameter setting values corresponding to the camera among the plurality of sets of camera parameter setting values.

[0181] Specifically, the electronic device 100 may identify the set of camera parameter setting values including the same aperture value as the aperture value of the camera as the set of camera parameter setting values corresponding to the camera based on the aperture value included in the plurality of sets of camera parameter setting values and the aperture values of the plurality of cameras 111, 112, and 113.

[0182] For example, assume that the plurality of sets of camera parameter setting values are (ISO, shutter speed, aperture) = (0.1, 0.3, 1), (0.4, 0.6, 2), and (0.6, 0.7, 3), the aperture value of camera 111 is 1, the aperture value of camera 112 is 2, and the aperture value of camera 113 is 3.

[0183] In this case, the electronic device 100 can recognize that (0.1, 0.3, 1) is a set of camera parameter setting values corresponding to the camera 111, recognize that (0.4, 0.6, 2) is a set of camera parameter setting values corresponding to the camera 112, and recognize that (0.6, 0.7, 3) is a set of camera parameter setting values corresponding to the camera 113.

[0184] The electronic device 100 can use (0.1, 0.3) recognized as corresponding to the camera 111 to set the ISO and shutter speed of the camera 111, and obtain an image frame through the camera 111 based on the set ISO value and shutter speed value of the camera 111. In addition, the electronic device 100 can use (0.4, 0.6) recognized as corresponding to the camera 112 to set the ISO and shutter speed of the camera 112, and obtain an image frame through the camera 112 based on the set ISO value and shutter speed value of the camera 112. In addition, the electronic device 100 can use (0.6, 0.7) recognized as corresponding to the camera 113 to set the ISO and shutter speed of the camera 113, and obtain an image frame through the camera 113 based on the set ISO value and shutter speed value of the camera 113.

[0185] In this case, the electronic device 100 can obtain multiple image frames at the same timing through the multiple cameras 111, 112, and 113. The same timing in this article can include not only exactly the same timing but also the timing within a threshold time range.

[0186] As described above, the electronic device 100 can use multiple sets of camera parameter setting values and multiple cameras 111, 112, and 113 to obtain multiple image frames.

[0187] In another example, the electronic device 100 can use the first camera among the multiple cameras 111, 112, and 113 and apply one set of camera parameter setting values among the multiple sets of camera parameter setting values to the first camera to obtain an image frame, and can use the second camera among the multiple cameras 111, 112, and 113 and apply at least two sets of camera parameter setting values among the multiple sets of camera parameter setting values to the second camera to obtain at least two image frames.

[0188] In other words, if the aperture values in at least two sets of camera parameter setting values included in the multiple sets of camera parameter setting values are the same, the electronic device 100 can recognize that at least two sets of camera parameter setting values with the same aperture value correspond to one camera, and obtain at least two image frames through this one camera using the at least two sets of camera parameter setting values.

[0189] For example, assume that multiple sets of camera parameter setting values are (ISO, shutter speed, aperture) = (0.1, 0.3, 1), (0.9, 0.7, 1), and (0.8, 0.4, 2), the aperture value of camera 111 is 1, the aperture value of camera 112 is 2, and the aperture value of camera 113 is 3.

[0190] In this case, the electronic device 100 can recognize that (0.1, 0.3, 1) and (0.9, 0.7, 1) are the sets of camera parameter setting values of camera 111, and recognize that (0.8, 0.4, 2) is the set of camera parameter setting values of camera 112.

[0191] The electronic device 100 can use (0.1, 0.3) corresponding to camera 111 to set the ISO and shutter speed of camera 111, and obtain an image frame through camera 111. In addition, the electronic device 100 can use (0.9, 0.7) corresponding to camera 111 to set the ISO and shutter speed of camera 111, and obtain an image frame through camera 111.

[0192] In addition, the electronic device 100 can use (0.8, 0.4) corresponding to camera 112 to set the ISO and shutter speed of camera 112, and obtain an image frame through camera 112.

[0193] In this case, at least one of the at least two image frames obtained through camera 111 and the image frame obtained through camera 112 can be obtained at the same timing. The same timing in this article can include not only exactly the same timing, but also the timing within a threshold time range.

[0194] As described above, the electronic device 100 can use multiple sets of camera parameter setting values and two cameras among multiple cameras 111, 112, and 113 to obtain multiple image frames.

[0195] In another example, the electronic device 100 can use a camera having the same aperture value as the aperture value included in the multiple sets of camera parameter setting values to obtain multiple image frames. In other words, if all the aperture values included in the multiple sets of camera parameter setting values are the same, the electronic device 100 can use one camera having the same aperture value as the aperture value included in the multiple sets of camera parameter setting values to obtain multiple image frames.

[0196] For example, assume that multiple sets of camera parameter setting values are (0.1, 0.3, 1), (0.9, 0.7, 1), and (0.8, 0.4, 1), the aperture value of camera 111 is 1, the aperture value of camera 112 is 2, and the aperture value of camera 113 is 3.

[0197] In this case, the electronic device 100 can recognize that (0.1, 0.3, 1), (0.9, 0.7, 1), and (0.8, 0.4, 1) are all sets of camera parameter setting values of the camera 111.

[0198] The electronic device 100 can use (0.1, 0.3) to set the ISO and shutter speed of the camera 111, and obtain an image frame through the camera 111. In addition, the electronic device 100 can use (0.9, 0.7) to set the ISO and shutter speed of the camera 111, and obtain an image frame through the camera 111. In addition, the electronic device 100 can use (0.8, 0.4) to set the ISO and shutter speed of the camera 111, and obtain an image frame through the camera 111.

[0199] Therefore, the electronic device 100 can use multiple sets of camera parameter setting values and one of the multiple cameras 111, 112, and 113 to obtain multiple image frames.

[0200] In the above embodiment, the multiple image frames obtained by the electronic device 100 can be Bayer raw data (or Bayer pattern data). In other words, in order to generate an image frame (for example, an image frame in which each of multiple pixels has R, G, and B pixel values (an image frame with an RGB format)), the data obtained from the image sensor has undergone processing such as interpolation, but the Bayer raw data can be data that has not undergone such processing.

[0201] Specifically, the pixels (or units) of the image sensor included in the camera are monochromatic pixels. Therefore, only a color filter that transmits a specific color can be located on the pixels of the image sensor. The color filter can include an R filter that only transmits the R color, a G filter that only transmits the G color, or a B filter that only transmits the B color, and the color filter located on the pixels of the image sensor can refer to a color filter array. In the color filter array, the R, G, and B color filters are set in a specific pattern. In an example, in a Bayer filter, the percentage of the G filter is 50%, the percentage of each of the R filter and the B filter is 25%, and a pattern in which two G filters and the R and B filters cross each other can be provided.

[0202] As described above, one of the R, G, and B filters can be combined with one pixel of the image sensor, and one pixel can only detect one of R, G, and B. In this case, the data generated by the pixels of the image sensor using the Bayer filter can refer to Bayer raw data (hereinafter, referred to as a Bayer raw image). However, this example is not limited, and the data generated by the pixels of the image sensor is not limited to Bayer raw data, and various types of color filter arrays can be used.

[0203] Figure 8This is a diagram showing multiple image frames obtained using multiple cameras according to an embodiment. Refer to Figure 8 , based on a second user command for capturing a live view image, the electronic device 100 may use multiple cameras 111, 112, and 113 respectively to obtain Bayer raw images 821, 822, and 823.

[0204] Therefore, the electronic device 100 may obtain multiple image frames based on the second user command for capturing through the above method.

[0205] Return to reference Figure 2 , the electronic device 100 may obtain an image frame corresponding to the second user command based on multiple image frames (S270).

[0206] In this case, the image frame may be stored in a non-volatile memory (such as a flash memory) of the electronic device 100 and then displayed on the display of the electronic device 100 through a gallery application driven by a user command.

[0207] Each of the multiple image frames may be a Bayer raw image as described above.

[0208] Figure 9 This is a flowchart showing a method of obtaining an image frame by merging multiple image frames according to an embodiment. Refer to Figure 9 , the electronic device 100 may select a reference image frame from multiple image frames (S910).

[0209] Specifically, the electronic device 100 may determine one of the multiple image frames as a reference image frame based on sharpness. In an example, the electronic device 100 may determine the image frame with the highest edge metric as the reference image frame.

[0210] The electronic device 100 may generate multiple image frames corresponding to multiple levels for each image frame by scaling each of the multiple image frames to different sizes from each other (S920).

[0211] Generating multiple image frames corresponding to multiple levels by scaling the image frames to different sizes from each other may refer to generating an image pyramid. An image pyramid may refer to an assembly of images obtained by scaling a base image to different sizes from each other.

[0212] A Gaussian pyramid serves as an example of an image pyramid. A Gaussian pyramid may include an assembly of images with different sizes generated by repeatedly performing blurring and subsampling (or downsampling) on an image. In this case, a Gaussian filter may be used for blurring. Additionally, subsampling may be performed by a method for removing pixels included in even columns and even rows of an image to reduce the resolution of the image.

[0213] For example, an image at level 1 can be generated by performing blurring and subsampling on an image at level 0 (e.g., a base image). In this case, the image at level 1 can have a resolution that is 1 / 2 of the base image. Additionally, an image at level 2 can be generated by performing blurring and subsampling on the image at level 1. In this case, the image at level 2 can have a resolution that is 1 / 4 of the base image. This method can be performed in a step-by-step manner to generate a Gaussian pyramid including images corresponding to multiple levels.

[0214] In other words, the electronic device 100 can set each of the multiple image frames as a base image and generate multiple Gaussian pyramids of the multiple image frames. In this case, the Gaussian pyramid can include image frames corresponding to multiple levels.

[0215] The electronic device 100 can select one image frame at each level by comparing the image frames at each level (S930).

[0216] Specifically, the electronic device 100 can select, among the multiple image frames included at each level, the image frame having the smallest difference between the image frames. In other words, the electronic device 100 can calculate the difference in pixel values between an image frame and other image frames included at the same level as the image frame, and select the image frame having the smallest calculated difference at each level.

[0217] Figure 10 is a diagram showing multiple Gaussian pyramids according to an embodiment. As Figure 10 shown, multiple Gaussian pyramids 1010, 1020, and 1030 can be generated for multiple image frames 1011, 1021, and 1031. Figure 10 shows the generation of a Gaussian pyramid having four levels, but the present disclosure is not limited to this example, and the Gaussian pyramid can include images of different levels.

[0218] In this case, the electronic device 100 can identify the image frame 1011 having the smallest pixel value difference among the multiple image frames 1011, 1021, and 1031 at level 0. Additionally, the electronic device 100 can identify the image frame 1022 having the smallest pixel value difference among the multiple image frames 1012, 1022, and 1032 at level 1. The electronic device 100 can perform such a process at levels 2 and 3 to identify the image frame 1033 having the smallest pixel value difference at level 2, and identify the image frame 1014 having the smallest pixel value difference at level 3.

[0219] As described above, the electronic device 100 can select the multiple image frames 1011, 1022, 1033, and 1014 at multiple levels.

[0220] Then, the electronic device 100 may scale at least one of the multiple selected image frames such that the multiple selected image frames have the same size (S940).

[0221] In this case, the electronic device 100 may perform upsampling on the remaining image frames except for the image frame used as the base image to have the same size as the image frame used as the base image when generating the Gaussian pyramid.

[0222] The electronic device 100 may identify at least one tile in each of the multiple image frames having the same size that is to be merged with the reference image frame (S950).

[0223] To this end, the electronic device 100 may perform tiling on the reference image frame and each of the multiple image frames having the same size. Tiling means that an image frame is divided into multiple regions. In this case, each region may be referred to as a tile, and the size of the tile may be smaller than the size of the image frame.

[0224] The electronic device 100 may identify at least one tile in the tiles of each of the multiple image frames that is to be merged with the tile of the reference image frame.

[0225] Specifically, the electronic device 100 may identify at least one tile in the tiles of any given image frame whose difference from the tile of the reference image frame is equal to or less than a preset threshold. In this case, the electronic device 100 may apply a Laplacian filter to the reference image frame and the image frame, identify the difference between the tile of the reference image frame to which the Laplacian filter is applied and the tile of the image frame to which the Laplacian filter is applied, and identify at least one tile in the tiles of each image frame having a difference equal to or less than the preset threshold. The preset threshold herein may be, for example, 5%.

[0226] Figure 11 is a diagram illustrating a method for merging multiple image frames according to an embodiment. Refer to Figure 11 , image frames 1022-1, 1033-1, and 1014-1 having the same size as image frame 1011 may be generated by scaling the selected image frames 1022, 1033, and 1014 among the multiple image frames at multiple levels.

[0227] The electronic device 100 may calculate the difference between tile 1 of the image frame 1014-1 to which the Laplacian filter is applied and tile 1 of the reference image frame to which the Laplacian filter is applied, and if the calculated difference is equal to or less than a preset threshold, the electronic device 100 may identify tile 1 of the image frame 1014-1 as an image frame to be merged with tile 1 of the reference image frame. In addition, the electronic device 100 may calculate the difference between tile 2 of the image frame 1014-1 to which the Laplacian filter is applied and tile 2 of the reference image frame to which the Laplacian filter is applied, and if the calculated difference is equal to or less than a preset threshold, the electronic device 100 may identify tile 2 of the image frame 1014-1 as an image frame to be merged with tile 2 of the reference image frame. The electronic device 100 may identify the image frames to be merged with tiles 1 and 2 of the reference image frame by performing the above process on the image frames 1011, 1022-1, and 1033-1. Figure 11 It is shown that the image frame is divided into two tiles. However, this example is not limited thereto, and the image frame may be divided into three or more tiles.

[0228] The electronic device 100 may merge the reference image frame with the image frame identified as being to be merged with the reference image frame, and generate a merged image frame (S960).

[0229] Specifically, the electronic device 100 may use the discrete Fourier transform (DFT) to transform the reference image frame (or a tile of the reference image frame) and the image frame to be merged (or a tile of the image frame to be merged) into the frequency domain, and merge the image frames (or tiles of the image frames) in the frequency domain.

[0230] If the image frame I(x, y) (where x and y are coordinate values of pixels) is represented by T(wx, wy) in the frequency domain, the electronic device 100 may use the following Mathematical Expression 2 to merge the image frames in the frequency domain.

[0231]

Mathematical Expression 2

[0232]

[0233] Here, To(w x , wy) is the merged image frame to which the discrete Fourier transform is applied, Tr(w x , w y ) is the reference image frame to which the discrete Fourier transform is applied, Tz(w x , wy) is the image frame to be merged to which the discrete Fourier transform is applied, and Kz corresponds to a scale factor determined based on the difference in average pixel values between the reference image frame and the image frame to be merged with the reference image frame.

[0234] In addition, the electronic device 100 may use an inverse Fourier transform to convert the merged frame image into the image domain. In other words, the electronic device 100 may use an inverse Fourier transform to convert To(w x , wy) into the image domain and generate a merged image frame Io(x, y).

[0235] According to an embodiment, the resolutions of at least two of the plurality of cameras may be different from each other. Accordingly, the sizes of at least two of the plurality of image frames obtained by the electronic device 100 according to a second user command for capturing a live view image may be different from each other.

[0236] In this case, before generating a plurality of Gaussian pyramids, the electronic device 100 may crop at least one of the plurality of image frames so that the plurality of image frames have the same size. Accordingly, the plurality of image frames used as reference images for generating a plurality of Gaussian pyramids may have the same size.

[0237] The above-described merging process may be performed on a Bayer raw image, and thus, the merged image frame in the image domain may also be a Bayer raw image.

[0238] In this case, the electronic device 100 may generate an image frame (e.g., an image frame in which each of a plurality of pixels has R, G, and B pixel values (an image frame in RGB format)) based on the merged image frame.

[0239] Herein, the electronic device 100 may use a neural network model to obtain an image frame. The third neural network model may be a conditional generative adversarial network (cGAN). The conditional GAN (cGAN) may include a generator and a discriminator. In this case, the generator and the discriminator may be trained in an adversarial manner.

[0240] The generator may generate and output an image frame in which each of a plurality of pixels has R, G, and B pixel values corresponding to the input Bayer raw image.

[0241] The input data x for training the generator of the cGAN may be a Bayer raw image obtained by an image sensor of a camera when a user captures an image using various parameter values, and the output data y of the generator may be an image frame (e.g., an image frame in JPEG format) obtained by performing a process such as interpolation on the Bayer raw image.

[0242] In this case, the output G(x) of the generator may be obtained by inputting the input data x and noise data to the generator.

[0243] In addition, pairs of (input data x, output data y) (hereinafter referred to as real pair data) and (input data x, G(x)) (hereinafter referred to as fake pair data) can be input into the discriminator. The discriminator can output probability values indicating whether the real pair data and the fake pair data are real or fake, respectively.

[0244] If the output of the discriminator for the input real pair data is defined as D(x, y), and the output of the discriminator for the input fake pair data is defined as D(x, G(x)), then the loss of the generator can be represented by the following Mathematical Expression 3, and the loss of the discriminator can be represented by the following Mathematical Expression 4.

[0245]

Mathematical Expression 3

[0246]

[0247]

Mathematical Expression 4

[0248]

[0249] Here, λ1, λ2, and λ3 are hyperparameters.

[0250] L1[G(x), y] is the loss of the sharpness between the output G(x) of the generator and the output data y, and can indicate that the difference in pixel values is |G(x) - y|. In addition, L2[G(x), y] is the loss of the pixel values between the output G(x) of the generator and the output data y, and can be expressed as (where i = 1, 2, 3,......, n (n is the total number of pixels)).

[0251] In addition, α is the number of downsampling layers of the generator, and T batch corresponds to the average processing time per batch. As described above, the loss is calculated based on α because the processing performance of the electronic device using the model is useful for determining the appropriate number of layers of the model. In this case, when the processing performance of the electronic device is sufficiently high (e.g., high-end electronic device), the calculation speed is high, and thus, a large number of layers can be used. When the processing performance of the electronic device is not high enough (e.g., low-end electronic device), a small number of layers can be used.

[0252] In this case, the discriminator can be trained to maximize the loss of the discriminator and the generator can be trained to minimize the loss of the generator minimize.

[0253] Accordingly, the electronic device 100 can obtain an image frame by inputting the merged image frame into the third neural network model. In addition, the electronic device 100 can store the obtained image frame in a non-volatile memory (such as a flash memory) of the electronic device 100.

[0254] Here, the image frame can be an image frame in which each of a plurality of pixels has R, G, and B pixel values. In this case, the image frame can have the same size as the merged image frame.

[0255] Figure 12 FIG. is a diagram illustrating a method of obtaining an image frame using a third neural network model according to an embodiment. Refer to Figure 12 , the electronic device 100 can obtain an image frame 1220 by inputting the merged image frame 1210 into the third neural network model 123. In this case, the electronic device 100 can obtain the image frame 1220 by inputting only the merged image frame 1210 into the third neural network model 123 without inputting separate noise data into the third neural network model 123. This is because, in the present disclosure, a merged image frame is generated by merging a plurality of image frames having different attributes, and in this process, it is considered that the merged image frame guarantees a certain degree of randomness. However, this is only an example, and the electronic device 100 can input noise data and the merged image frame 1210 into the third neural network model 123.

[0256] In the present disclosure, the image frame obtained from the third neural network model 1 can be an image frame that has undergone at least one of black level adjustment, color correction, gamma correction, or edge enhancement.

[0257] Here, black level adjustment can refer to image processing for adjusting the black level of pixels included in a Bayer raw image. When generating an image frame by performing interpolation on a Bayer raw image, distortion may occur in color representation, and color correction can refer to image processing for correcting such distortion. Gamma correction can refer to image processing for non-linearly transforming the intensity signal of light using a non-linear transfer function for an image. Edge enhancement can refer to image processing for highlighting image edges.

[0258] To this end, the generator can have a U-net (U-shaped network) structure. U-net can refer to a neural network capable of extracting features (such as edges, etc.) of an image frame using both low-order information and high-order information. In the present disclosure, when the image frame generated by the generator is recognized as a real image frame based on the output of the discriminator, the generator can perform a filtering operation through each layer of the U-net and perform at least one of black level adjustment, color correction, gamma correction, or edge enhancement of the image frame.

[0259] As described above, according to an embodiment of the present disclosure, the electronic device 100 may obtain an image frame based on a second user command for capturing a live view image.

[0260] According to the above disclosure, a plurality of sets of camera parameter setting values corresponding to a plurality of regions obtained from the image frame may be obtained, a set of camera parameter setting values selected from the plurality of sets of camera parameter setting values based on an image metric may be used to obtain a plurality of image frames, and the obtained image frames may be combined to obtain an image frame corresponding to a user command for capturing. Accordingly, a high-quality image frame may be generated, an image frame with relatively sharp edges of regions having different brightnesses from each other in the image frame may be obtained, a patch phenomenon (e.g., a horizontal line of a sky background) may be hardly observed, and a ghosting phenomenon may also be reduced.

[0261] In the above embodiment, the neural network model may refer to an artificial intelligence model including a neural network and may be trained through deep learning. For example, the neural network model may include at least one artificial neural network model among a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), and a generative adversarial network (GAN). However, the neural network model according to the present disclosure is not limited to the above examples.

[0262] In addition, according to an embodiment of the present disclosure, the electronic device 100 obtaining an image frame through the above method may be performed in a low light mode (e.g., low light shooting).

[0263] Specifically, if the light intensity is less than a preset threshold, the electronic device 100 may activate the low light mode and obtain an image frame in the low light mode through the above method. Here, the light intensity may be calculated based on the brightness of pixels included in an image frame obtained through a camera for a live view image. However, the present disclosure is not limited to this example, and the electronic device 100 may measure the ambient brightness using an illuminance sensor of the electronic device 100, and if the measured ambient brightness is less than a preset threshold, the electronic device may activate the low light mode. Due to blurring, noise, and uneven contrast, the quality of an image frame obtained in the low light mode is poor. Accordingly, in the present disclosure, an enhanced image frame may be obtained in the low light mode through the above method. However, the present disclosure is not limited to this example, and even if the low light mode is not activated, the electronic device 100 may obtain an image frame through the above method.

[0264] In addition, in the above-described embodiments, it is described that the first neural network model 121 outputs a set of camera parameter setting values corresponding to each of a plurality of regions. However, the present disclosure is not limited to this example. In another example, the first neural network model 121 may additionally generate a plurality of sets of camera parameter setting values by combining sets of camera parameter setting values corresponding to at least two regions, and output sets of camera parameter setting values corresponding to the plurality of regions and the plurality of sets of camera parameter setting values generated by the combination.

[0265] The combination of sets of camera parameter setting values may refer to changing the parameter values of the same parameters included in the sets of camera parameter setting values.

[0266] Specifically, the first neural network model 121 may identify sets of camera parameter setting values having the same aperture value among the plurality of sets of camera parameter setting values corresponding to the plurality of regions, change the ISO value included in one of the identified sets of camera parameter setting values to the ISO value included in another of the identified sets of camera parameter setting values, or change the shutter speed value included in one of the identified sets of camera parameter setting values to the shutter speed value included in another of the identified sets of camera parameter setting values, and generate a plurality of sets of camera parameter setting values based on the changed sets of camera parameter setting values.

[0267] Figure 13 is a diagram showing an example of a method of obtaining a plurality of sets of camera parameter setting values using a first neural network model according to an embodiment. Refer to Figure 13 , as described in the example of Figure 6 , the first neural network model 121 may generate sets of camera parameter setting values 621, 622, 623, and 624 based on each of the plurality of regions 611, 612, 613, and 614.

[0268] Here, the set of camera parameter setting values 621 corresponding to the region 611 and the set of camera parameter setting values 622 corresponding to the region 612 have the same aperture value (i.e., aperture value 1).

[0269] In this case, the first neural network model 121 may generate sets of camera parameter setting values 625 and 626 by changing the shutter speed values of the set of camera parameter setting values 621 and the set of camera parameter setting values 622. Specifically, the first neural network model 121 may generate the set of camera parameter setting values 625 by changing the shutter speed value included in the set of camera parameter setting values 621 to the shutter speed value included in the set of camera parameter setting values 622, and generate the set of camera parameter setting values 626 by changing the shutter speed value included in the set of camera parameter setting values 622 to the shutter speed value included in the set of camera parameter setting values 621.

[0270] Therefore, the first neural network model 121 can output multiple sets of camera parameter settings 621, 622, 623, 624, 625, and 626. In this case, the electronic device 100 can obtain multiple sets of camera parameter settings by inputting the multiple sets of camera parameter settings obtained from the first neural network model 121 into the second neural network model 122. The subsequent operations are the same as those in the above embodiment, so the specific redundant descriptions will not be repeated.

[0271] In addition, in the above embodiment, the electronic device 100 uses the second neural network model 122 to obtain multiple sets of camera parameter settings for capturing a live view image. However, the present disclosure is not limited to this example, and the electronic device 100 can generate image frames corresponding to a second user command for capturing a live view image by using the multiple sets of camera parameter settings obtained from the first neural network model 121. Specifically, the electronic device 100 can identify the set of camera parameter settings corresponding to a camera based on the aperture value included in each of the multiple sets of camera parameter settings obtained from the neural network model 121 and the aperture values of the multiple cameras 111, 112, and 113, and obtain multiple image frames by using the multiple sets of camera parameter settings and at least one camera.

[0272] In addition, in the above embodiment, it is described that the electronic device 100 uses neural network models (e.g., the first neural network model 121 and the second neural network model 122) to obtain multiple sets of camera parameter settings. However, it is not limited thereto, and the electronic device 100 can also obtain multiple sets of camera parameter settings based on rules.

[0273] Specifically, the electronic device 100 can identify the pixel values of the pixels included in each of the multiple regions, and obtain multiple sets of camera parameter settings based on the identified pixel values and predefined rules.

[0274] In this case, the electronic device 100 can use a rule-based engine. The rule-based engine can analyze the input data based on predefined rules and output the results.

[0275] In the present disclosure, obtaining multiple sets of camera parameter settings based on rules can include obtaining multiple sets of camera parameter settings for capturing image frames based on rules.

[0276] Figure 14 is a flowchart showing an example of a method for obtaining multiple sets of camera parameter settings using rules according to an embodiment. First, as Figure 14As shown, the electronic device 100 may identify the pixel values of the pixels included in each of the multiple regions (S1410), and generate input data for the rule-based engine based on the identified pixel values (S1420).

[0277] The input data here may include at least one of the pixel values of the pixels included in each region, the average of the pixel values, the deviation, or the variance.

[0278] The electronic device 100 may obtain multiple sets of camera parameter setting values by inputting the generated input data into the rule-based engine (S1430).

[0279] Specifically, the electronic device 100 may input the generated input data for each region into the rule-based engine. In the present disclosure, if the input data of the generated region satisfies a predefined first rule, the rule-based engine may determine a set of camera parameter setting values corresponding to the region according to the first rule, and if the multiple sets of camera parameter setting values determined for the multiple regions satisfy a predefined second rule, the rule-based engine may determine and output, according to the second rule, a set of camera parameter setting values for capturing an image frame among the multiple sets of camera parameter setting values.

[0280] In addition, in the present disclosure, obtaining multiple sets of camera parameter setting values based on rules may include obtaining a set of camera parameter setting values corresponding to each region based on rules.

[0281] Figure 15 is a flowchart showing an example of a method for obtaining multiple sets of camera parameter setting values using rules according to an embodiment. First, as Figure 15 shown, the electronic device 100 may identify the pixel values of the pixels included in each of the multiple regions (S1510), and generate input data for the rule-based engine based on the identified pixel values (S1520).

[0282] Here, the input data may include at least one of the pixel values of the pixels included in each region, the average of the pixel values, the deviation, or the variance.

[0283] The electronic device 100 may obtain multiple sets of second camera parameter setting values by inputting the generated input data into the rule-based engine (S1530).

[0284] Specifically, the electronic device 100 may input the generated input data for each region into a rule-based engine. In the present disclosure, if the input data for the generated region satisfies a predefined rule, the rule-based engine may determine a set of camera parameter setting values corresponding to the region according to the rule, and output multiple sets of second camera parameter setting values determined for multiple regions.

[0285] The electronic device 100 may obtain a set of camera parameter setting values for capturing an image frame by inputting multiple sets of second camera parameter setting values into the second neural network model 122. Specifically, the electronic device 100 may obtain a set of camera parameter setting values for capturing an image frame by inputting multiple sets of second camera parameter setting values, the image frame obtained through the camera 111, and multiple parameter values set on the camera 111 into the second neural network model 122 (S1540).

[0286] In addition, in the above embodiment, it is described that the image metrics include an edge metric, a contrast metric, and a background metric. However, the present disclosure is not limited to this example, and the image metrics may include various metrics related to the quality of the image frame.

[0287] In addition, the image metrics may include at least one of an edge metric, a contrast metric, or a background metric. In this case, even during the training process of the first neural network model 121, an enhancement ratio may be calculated based on at least one metric of the region, and the first neural network model 121 may be trained by using the enhancement ratio. In addition, for the input set of camera parameter setting values, the second neural network model 122 may output information about at least one metric of the image frame corresponding to the input set of camera parameter setting values.

[0288] The electronic device 100 may select at least one set of camera parameter setting values from multiple sets of first camera parameter setting values based on the information obtained from the second neural network model 122. Based on receiving a second user command for capturing a live view image, the electronic device 100 may use the selected at least one set of camera parameter setting values and at least one of the multiple cameras 111, 112, and 113 to obtain an image frame.

[0289] In addition, in the above embodiment, it is described that the electronic device 100 includes three cameras 111, 112, and 113. However, it is not limited thereto, and the electronic device 100 may include at least one camera.

[0290] In addition, in the above embodiment, it is described that the number of the multiple cameras 111, 112, and 113 of the electronic device 100 is the same as the number of the multiple sets of camera parameter setting values for capturing a live view image. However, it is not limited thereto.

[0291] According to an embodiment, the number of cameras of the electronic device 100 may be less than the number of sets of multiple camera parameter setting values.

[0292] In this case, the electronic device 100 may obtain multiple image frames by applying at least two sets of camera parameter setting values from the multiple sets of camera parameter setting values to one camera several times. For example, when using one camera to capture a live view image, the electronic device 100 may use multiple camera parameter setting values on one camera to obtain multiple image frames. In another example, when using two cameras to capture a live view image, the electronic device 100 may use at least two sets of camera parameter setting values from the multiple sets of camera parameter setting values on one camera to obtain multiple image frames, and use the remaining one set of camera parameter setting values on the other camera to obtain one image frame.

[0293] In these cases, the camera used to obtain multiple image frames may obtain an image frame using multiple parameter values included in one set of camera parameter setting values, and then obtain an image frame using multiple parameter values included in another set of camera parameter setting values.

[0294] According to an embodiment, the number of cameras of the electronic device 100 may be greater than the number of sets of multiple camera parameter setting values. In this case, the electronic device 100 may use at least one camera from the multiple sets of camera parameter setting values and the multiple cameras to obtain multiple image frames.

[0295] In addition, in the above embodiment, the electronic device 100 obtains an image frame by inputting the merged image frames into the third neural network model 123. However, this is not limited thereto, and the electronic device 100 may obtain an image frame (i.e., an image frame in which each of the multiple pixels has R, G, and B pixel values (an image frame having an RGB format)) by performing interpolation, black level adjustment, color correction, gamma correction, edge enhancement, etc. on the merged image frames without using the third neural network model 123.

[0296] In addition, in the above embodiment, the electronic device 100 obtains multiple Bayer raw images and merges the multiple Bayer raw images based on a second user command for capturing a live view image. However, this is not limited thereto, and the electronic device 100 may obtain multiple image frames (i.e., an image frame in which each of the multiple pixels has R, G, and B pixel values (an image frame having an RGB format)) and obtain an image frame corresponding to the second user command for obtaining a live view image by merging the obtained image frames.

[0297] In addition, in the above-described embodiment, the electronic device 100 combines multiple image frames to generate a combined image frame. In this case, the electronic device 100 may remove noise from the multiple image frames by using a time-domain filter and combine the multiple noise-removed image frames.

[0298] Figure 16 is a block diagram showing a configuration of an electronic device according to an embodiment.

[0299] Referring Figure 16 , the electronic device 100 may include multiple cameras 111, 112, and 113, a memory 120, and a processor 130.

[0300] Each of the multiple cameras 111, 112, and 113 may obtain at least one image frame.

[0301] Specifically, each of the multiple cameras 111, 112, and 113 may include an image sensor and a lens.

[0302] In this case, the multiple lenses of the multiple cameras 111, 112, and 113 may have different viewing angles. For example, the multiple lenses may include a telephoto lens, a wide-angle lens, and an ultra-wide-angle lens. However, according to the present disclosure, there is no particular limitation on the number and type of the lenses, and in the example, the camera may include a standard lens.

[0303] Here, the viewing angle of the telephoto lens is wider than that of the extreme telephoto lens, the viewing angle of the standard lens is wider than that of the telephoto lens, the viewing angle of the wide-angle lens is wider than that of the standard lens, and the viewing angle of the ultra-wide-angle lens is wider than that of the wide-angle lens. For example, the viewing angle of the extreme telephoto lens is 3 degrees to 6 degrees, the viewing angle of the telephoto lens is 8 degrees to 28 degrees, the viewing angle of the standard lens is 47 degrees, the viewing angle of the wide-angle lens is 63 degrees to 84 degrees, and the viewing angle of the ultra-wide-angle lens is 94 degrees to 114 degrees.

[0304] The memory 120 may store at least one instruction regarding the electronic device 100. The memory 120 may store an operating system (O / S) for driving the electronic device 100. In addition, the memory 120 may store various software programs or applications for driving the electronic device 100 according to various embodiments of the present disclosure. The memory 120 may include a volatile memory such as a frame buffer, a semiconductor memory such as a flash memory, or a magnetic storage medium such as a hard disk drive.

[0305] Specifically, the memory 120 may store various software modules for driving the electronic device 100 according to various embodiments of the present disclosure, and the processor 130 may control the operation of the electronic device 100 by executing the various software modules stored in the memory 120. In other words, the processor 130 may access the memory 120, and the processor 130 may read, record, edit, delete, and / or update data in the memory 120.

[0306] The term “memory” in the present disclosure (e.g., memory 120) may include a read-only memory (ROM) (not shown) and a random access memory (RAM) (not shown) in the processor 130, or a memory card (not shown) (e.g., a micro secure digital (SD) card or a memory stick) mounted on the electronic device 100.

[0307] The processor 130 may control the general operation of the electronic device 100. Specifically, the processor 130 may be connected to elements of the electronic device 100, including the plurality of cameras 111, 112, and 113 and the memory 120 as described above, and may generally control the operation of the electronic device 100 by executing at least one instruction stored in the memory 120 as described above.

[0308] Specifically, the processor 130 may include at least one processor 130. For example, the processor 130 may include an image signal processor, an image processor, and an application processor (AP). Depending on the embodiment, at least one of the image signal processor or the image processor may be included in the AP.

[0309] In addition, in various embodiments according to the present disclosure, based on receiving a first user command for obtaining a live view image, the processor 130 may obtain a plurality of regions by segmenting an image frame obtained by camera 111 among the multiple cameras 111, 112, and 113 based on brightness and objects of pixels included in the image frame, obtain a plurality of camera parameter setting value sets including a plurality of parameter values regarding the plurality of segmented regions, and based on receiving a second user command for capturing a live view image, the processor 130 may obtain a plurality of image frames using the plurality of camera parameter setting value sets and the cameras among the multiple cameras 111, 112, and 113, and obtain an image frame corresponding to the second user command by merging the plurality of image frames.

[0310] Specifically, refer to Figure 16 , the processor 130 can load and use the first neural network model 121, the second neural network model 122 and the third neural network model 123 stored in the memory 120, and can include a segmentation module 131, a parameter determination module 132, a camera control module 133 and an alignment / merging module 134.

[0311] As described above with reference to Figures 1 to 15 various embodiments of the present disclosure based on the control of the processor 130 have been described in detail, and thus redundant descriptions will not be repeated.

[0312] The segmentation module 131 can obtain multiple regions by segmenting an image frame obtained by one of the multiple cameras based on the luminance of the pixels included in the image frame and the object.

[0313] The parameter determination module 132 can obtain multiple sets of camera parameter setting values based on the multiple regions, and each set of camera parameter setting values includes multiple parameter values.

[0314] Specifically, the parameter determination module 132 can obtain multiple first sets of camera parameter setting values by inputting the multiple regions into the first neural network model 121, and obtain multiple sets of camera parameter setting values by inputting the multiple first sets of camera parameter setting values, the image frame obtained by the camera 111, and the multiple parameter values of the camera 111 into the second neural network model 122.

[0315] The first neural network model 121 can be a neural network model trained based on the regions obtained from the image frame; the multiple parameter values of the camera that has obtained the image frame; the edge index, contrast index, and background index of the region.

[0316] In addition, based on the input of the set of camera parameter setting values, the image frame obtained by the camera 111, and the multiple parameter values of the camera 111, the second neural network model 122 can output information about the edge index, contrast index, and background index of the input image frame corresponding to the input set of camera parameter setting values.

[0317] The first neural network model 121 can refer to a camera calibration network, and the second neural network model 122 can refer to an ensemble parameter network.

[0318] The parameter determination module 132 can obtain the set of camera parameter setting values with the maximum edge index, the set of camera parameter setting values with the maximum contrast index, and the set of camera parameter setting values with the maximum background index in the multiple first sets of camera parameter setting values based on the information about the edge index, contrast index, and background index obtained from the second neural network model 122.

[0319] In the above example, obtaining multiple sets of camera parameter setting values using a neural network model is described, but it is not limited thereto. The parameter determination module 132 can identify the pixel values of the pixels included in each of the multiple regions, and obtain multiple sets of camera parameter setting values based on the identified pixel values and predefined rules.

[0320] Based on receiving a second user command for capturing a live view image, the camera control module 133 can obtain a plurality of image frames by using a plurality of sets of camera parameter setting values obtained by the parameter determination module 132 (e.g., a set of camera parameter setting values having a maximum edge metric, a set of camera parameter setting values having a maximum contrast metric, and a set of camera parameter setting values having a maximum background metric) and at least two of the plurality of cameras 111, 112, and 113.

[0321] In an example, based on receiving a second user command for capturing a live view image, the camera control module 133 can use a plurality of sets of camera parameter setting values and the plurality of cameras 111, 112, and 113 to obtain a plurality of image frames.

[0322] In another example, based on receiving a second user command for capturing a live view image, the camera control module 133 can obtain a plurality of image frames by obtaining an image frame by using one set of camera parameter setting values from the plurality of sets of camera parameter setting values and a first camera among the plurality of cameras 111, 112, and 113, and obtaining at least two image frames by using at least two sets of camera parameter setting values from the plurality of sets of camera parameter setting values and a second camera among the plurality of cameras 111, 112, and 113.

[0323] In yet another example, based on receiving a second user command for capturing a live view image, the camera control module 133 can use a plurality of sets of camera parameter setting values and one of the plurality of cameras 111, 112, and 113 to obtain a plurality of image frames.

[0324] As described above, the camera control module 133 can obtain a plurality of image frames. Each of the plurality of image frames can be a Bayer raw image.

[0325] The alignment / merging module 134 can obtain an image frame corresponding to the second user command by merging the plurality of obtained image frames and inputting the merged image frames into the third neural network model 123.

[0326] In this case, the image frame obtained from the third neural network model 123 can be an image frame that has undergone at least one of black level adjustment, color correction, gamma correction, or edge enhancement.

[0327] As described above, according to various embodiments of the present disclosure, the processor 130 can obtain an image frame. In this case, the processor 130 can store the obtained image frame in the memory 120 (specifically, a flash memory).

[0328] According to an embodiment of the present disclosure, operations of the parameter determination module 132 for obtaining a plurality of sets of camera parameter setting values may be performed by an image processor, and operations of the alignment / merging module 134 for obtaining an image frame corresponding to a user command for capture by using a plurality of image frames may be performed by an image signal processor.

[0329] Figure 17 is a block diagram more specifically showing a hardware configuration of an electronic device according to an embodiment.

[0330] Reference Figure 17 , in addition to the plurality of cameras 111, 112, and 113, the memory 120, and the processor 130, the electronic device 100 according to the present disclosure may further include a communication interface 140, an input interface 150, and an output interface 160. However, this configuration is merely an example, and additional elements may be added to the above configuration, or some elements in the above configuration may be omitted.

[0331] The communication interface 140 may include circuitry and perform communication with an external device. Specifically, the processor 130 may receive various data or information from an external device connected via the communication interface 140 and send various data or information to the external device.

[0332] The communication interface 140 may include at least one of a Wi-Fi module, a Bluetooth module, a wireless communication module, and a near field communication (NFC) module. Specifically, the Wi-Fi module and the Bluetooth module may perform communication by using Wi-Fi TM and Bluetooth TM protocols, respectively. In the case of using the Wi-Fi module and the Bluetooth module, various connection information such as a service set identifier (SSID) may be first sent or received to allow communication connection via Wi-Fi TM and Bluetooth TM , and then various information may be sent and received.

[0333] The wireless communication module may perform communication according to various communication standards such as IEEE, Zigbce TM , third generation (3G), third generation partnership project (3GPP), long term evolution (LTE), and fifth generation (5G). The NFC module may perform communication by using a near field communication (NFC) method in a 13.56 MHz frequency band among various radio frequency identification (RFID) frequency bands such as 135 kHz, 13.56 MHz, 433 MHz, 860 MHz to 960 MHz, 2.45 GHz, etc.

[0334] Specifically, according to an embodiment of the present disclosure, when the electronic device 100 downloads a neural network model from an external device, the communication interface 140 may receive the first neural network model 121 , the second neural network model 122 , and the third neural network model 123 from the external device.

[0335] The input interface 150 may include a circuit, and the processor 130 may receive a user command for controlling the operation of the electronic device 100 via the input interface 150. Specifically, the input interface 150 may be configured with a configuration such as a microphone and a remote signal receiver (not shown), and implemented to be included in the display 161 as a touch screen.

[0336] Specifically, in various embodiments according to the present disclosure, the input interface 150 may receive a first user command for obtaining a live view image and a second user command for capturing a live view image.

[0337] The output interface 160 may include a circuit, and the processor 130 may output various functions that can be performed by the electronic device 100 via the output interface 160. The output interface 160 may include a display 161 and a speaker 162. For example, the display 161 may be implemented as a liquid crystal display (LCD) panel or an organic light emitting diode (OLED), and the display 161 may also be implemented as a flexible display or a transparent display. However, the display 161 according to the present disclosure is not limited to a specific type. Specifically, in various embodiments according to the present disclosure, the display 161 may display a live view image. In addition, the display 161 may display an image frame stored in the memory 120 according to a user command.

[0338] The speaker 162 may output sounds. Specifically, the processor 130 may output various alarms or voice guidance messages related to the operation of the electronic device 100 via the speaker 162 .

[0339] Figure 18A 、 Figure 18B and Figure 18C is a diagram illustrating an example of a method for calculating an image index according to an embodiment.

[0340] Figure 18A is a diagram illustrating a method for calculating edge indices.

[0341] Specifically, a filter may be applied to an image frame, and after applying the filter, the width (or width), height (or length) of the image frame, and the pixel value of each pixel included in the image frame may be obtained. An edge indicator may be determined based on the obtained width, height, and pixel value of the image.

[0342] For example, reference Figure 18A, a Laplacian filter with a mask 1812 is applied to the image frame 1811, and an image frame 1813 with enhanced edges can be obtained. The edge index can be calculated by applying the width, height, and pixel values of the image frame 1813 to the following mathematical expression 5.

[0343]

Mathematical Expression 5

[0344]

[0345] Here, m represents the width of the image frame (i.e., the number of pixels in the width direction), n represents the height of the image frame (i.e., the number of pixels in the height direction), and F(x, y) represents the pixel value of the pixel with coordinate values (x, y) in the image frame.

[0346] Figure 18B is a diagram showing a method for calculating a contrast index.

[0347] For example, referring to Figure 18B , a histogram 1822 of the image frame 1821 can be generated. Here, the histogram 1822 shows the pixels classified by pixel value included in the image frame 1821, and can represent the number of pixels having corresponding pixel values for each pixel value.

[0348] In addition, an equalized histogram 1823 can be generated by performing equalization on the histogram 1822.

[0349] Referring to Mathematical Expression 6, the deviation between the histogram 1822 and the equalized histogram 1823 can be calculated to calculate the contrast index.

[0350]

Mathematical Expression 6

[0351]

[0352] Here, m represents the width of the image frame (i.e., the number of pixels in the width direction), n represents the height of the image frame (i.e., the number of pixels in the height direction), H(x) represents the number of pixels with pixel value x calculated from the histogram, H(y) represents the number of pixels with pixel value y calculated from the equalized histogram, and N represents the number of pixels included in the image frame.

[0353] Figure 18C is a diagram showing a method for calculating a background index.

[0354] Regarding the background index, a filter can be applied to the image frame, the width and height of the image frame can be obtained after applying the filter, and the background index can be determined based on the obtained width and height.

[0355] For example, referring to Figure 18C , the image frame 1831 can be divided into a plurality of regions as shown in 1832, and one region F(x, y) 1833 (for example, the region corresponding to the background) can be selected from these regions. A random pixel can be selected from the selected region 1833, the pixel value of the selected pixel can be randomly set to 0 or 255, and salt-and-pepper noise can be added to the selected region 1833. Then, a median filter with a kernel size of 5 can be applied to the region 1834 to which the salt-and-pepper noise has been applied, and a region G(x, y) 1835 to which the median filter has been applied can be generated.

[0356] The background metric can be calculated based on the following mathematical expression 7.

[0357]

Mathematical Expression 7

[0358]

[0359] Here, F(x, y) represents the pixel value of the pixel with the coordinate value (x, y) in the region 1833, G(x, y) represents the pixel value of the pixel with the coordinate value (x, y) in the region 1835, m represents the width of the image frame (i.e., the number of pixels in the width direction), and n represents the height of the image frame (i.e., the number of pixels in the height direction).

[0360] Figures 18A to 18C This is only an example of the method for calculating the image metric and is not limited thereto, and the image metric can be calculated by various known methods.

[0361] In addition, referring to Figures 18A to 18C describes the calculation of the image metric of the image frame, but the image metric of the region of the image frame can also be calculated by the above method.

[0362] Figure 19 is a diagram showing the structure of the first neural network model according to an embodiment.

[0363] Referring to Figure 19 , the first neural network model 121 can receive a region as input and output a set of camera parameter setting values corresponding to the region. In this case, the set of camera parameter setting values can include an ISO value, a shutter speed value, and an aperture value.

[0364] In the case of the first neural network model 121, there is a common part (for example, a residual block (Res block)). As described above, the first neural network model 121 is designed such that the features used in one model can be used for the shared feature concept of another model, and thus, the computational amount can be relatively reduced.

[0365] In addition, the final layer of the first neural network model 121 is designed considering the final estimated parameters (i.e., ISO value, shutter speed value, and aperture value).

[0366] Specifically, in the case of ISO, the effect caused by the color of the image is more significant than the shape, so a convolutional block is used. In addition, in the case of the aperture, the shape and edges are considered important. Therefore, two convolutional blocks are used for the aperture. In the case of the shutter speed, it is necessary to distinguish edges and features at a high level, so a residual block is used. Residual blocks are effective for maintaining most of the input image when learning new features, while convolutional blocks are effective for learning filters to reduce input calculations.

[0367] In the present disclosure, the residual block and the convolutional block can have Figure 20 the relationships (a) and (b).

[0368] Hereinafter, with reference to Figure 21 , the training of the first neural network model 121 will be briefly described.

[0369] With reference to Figure 21 (a), the forward propagation learning of the first neural network model 121 is as follows.

[0370] The input data x is input into the first neural network model 121 to obtain the x iso value, the x shutter_speed direct sum, and the x aperture value from the first neural network model 121.

[0371] Then, the loss between the data obtained from the first neural network model 121 and the output data y can be calculated. Specifically, the loss L iso between (x iso , y iso ), the loss L shutter_speed between (x shutter_speed ), and the loss L shutter_speed between (x aperture , y aperture ) can be calculated based on Mathematical Expression 8. The output data y can satisfy y aperture = r × Y iso , y captured_iso = r × Y shutter_speed , and y captured_shutter_speed = r × Y aperture . Here, Y captured_aperture , Y captured_iso , and Y captured_shutter_speed , and Y captured_opertureThey are the ISO value, shutter speed value, and aperture value for obtaining the input data x, respectively. In addition, r = (ei + bi + ci) / 3. Here, ei represents the edge index of the input data x, bi represents the background index of the input data x, and ci represents the contrast index of the input data x.

[0372]

Mathematical Expression 8

[0373]

[0374] Reference Figure 21 of (b), the backpropagation learning of the first neural network model 121 is performed as follows.

[0375] Specifically, the weights of each layer can be updated to using the losses Liso, Lshutter_speed, and Laperture by gradient descent. The combined gradient can be calculated using the average of the gradients before backpropagation through the common part, and backpropagation is performed based on this to update the weights.

[0376] The structure and training method of the first neural network model 121 are only examples, and the first neural network model 121 with various structures can be trained by various methods.

[0377] Figure 22 is a diagram showing the structure of the second neural network model according to the embodiment.

[0378] Based on the input of the camera parameter setting value set (x, y, z), the image frame, and the multiple parameter values (x0, y0, z0) of the camera that has obtained the input image frame, the second neural network model 122 can output information on the edge index Ei, contrast index Ci, and background index Bi of the image frame corresponding to the input camera parameter setting value set. In this case, the second neural network model 122 can use the dense module of the parameter estimation network model to estimate the edge index, contrast index, and background index, and output these indices.

[0379] Figure 22 Each block of the second neural network model 122 shown in (a) can have Figure 22 the configuration shown in (b).

[0380] The above structure of the second neural network model 122 is only an example, and the second neural network model 122 can have various structures.

[0381] Figure 23A and Figure 23B is a diagram showing the structure of the third neural network model according to the embodiment.

[0382] First,Figure 23A The structure of the generator of the third neural network model 123 is shown.

[0383] Reference Figure 23A , based on the input of the merged image frames, the generator can output image frames (i.e., image frames in which each of a plurality of pixels has R, G, and B pixel values). In this case, adaptive GaN can be used to counteract the real features of the image frames ensured by the adversarial loss term.

[0384] In addition, the generator can have a U-net structure to compensate for the information loss that occurs during downsampling of multiple layers. The number α of downsampling layers is only an example and can be 3. In this case, the generator can perform a filtering operation through each layer having a U-net structure and perform at least one of black level adjustment, color correction, gamma correction, or edge enhancement of the image frames.

[0385] Figure 23B The structure of the discriminator of the third neural network model 123 is shown.

[0386] Reference Figure 23B , the discriminator can receive input data 2310 and 2320 and output a probability value indicating whether the input data is real or fake. The input data 2310 and 2320 can be real pair data or fake pair data. In other words, the discriminator can be used as an image classifier for classifying the input data as real data or fake data.

[0387] For this purpose, the residual blocks of the discriminator can convert the Bayer raw image and the RGB format image frames into a common format with the same number of channels. The data in the same format can be concatenated, and the residual blocks and convolutional blocks in the final layer can obtain significant features for the final labeling.

[0388] The structure of the above-mentioned third neural network model 123 is only an example, and the third neural network model 123 can have various structures.

[0389] The functions related to the above neural network model can be executed by a memory and a processor. The processor can include one or more processors. The one or more processors can be a general-purpose processor such as a CPU or an AP, a graphics dedicated processor such as a GPU or a vision processing unit (VPU), or an artificial intelligence dedicated processor such as a neural processing unit (NPU), etc. The one or more processors can execute control to process the input data according to predefined action rules or artificial intelligence models stored in the non-volatile memory and the volatile memory. The predefined action rules or artificial intelligence models are formed through training.

[0390] In this document, forming through training can, for example, mean forming a predefined action rule or an artificial intelligence model with desired characteristics by applying a learning algorithm to a plurality of learning data. Such training can be performed in a device exhibiting artificial intelligence according to the present disclosure, or by a separate server and / or system.

[0391] The artificial intelligence model can include multiple neural network layers. Each layer has multiple weight values, and the processing of the layer is performed through the processing results of the previous layer and the processing between multiple weights. Examples of neural networks include convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks. Unless otherwise stated, the neural networks of the present disclosure are not limited to the above examples.

[0392] The learning algorithm can be a method of training a predetermined target machine (e.g., a robot) using a plurality of learning data so that the predetermined target machine can make determinations or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. However, unless otherwise stated, the learning algorithms in the present disclosure are not limited to these examples.

[0393] The machine-readable storage medium can be provided in the form of a non-transitory storage medium. Here, the "non-transitory" storage medium is tangible and may include not only signals, and it does not distinguish whether the data is stored in the storage medium semi-permanently or temporarily. For example, the "non-transitory storage medium" can include a buffer for temporarily storing data.

[0394] According to an embodiment, a method according to various embodiments disclosed in the present disclosure can be provided in a computer program product. The computer program product can be exchanged between a seller and a buyer as a commercially available product. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or distributed online (e.g., downloaded or uploaded) through an app store (e.g., PlayStoreTM), or directly distributed between two user devices (e.g., smartphones). In the case of online distribution, at least a part of the computer program product can be at least temporarily stored or temporarily generated in a machine-readable storage medium (such as the memory of a manufacturer's server, an app store's server, or a relay server).

[0395] Each of the elements (e.g., modules or programs) according to the various embodiments disclosed above may include a single entity or multiple entities, and some of the above sub-elements may be omitted, or other sub-elements may be further included in the various embodiments. Alternatively or additionally, some elements (e.g., modules or programs) may be integrated into one entity to perform the same or similar functions as performed by each corresponding element before integration.

[0396] According to various embodiments, the operations performed by a module, program, or other element may be sequentially executed in a parallel, repetitive, or heuristic manner, or at least some operations may be executed in a different order, omitted, or different operations may be added.

[0397] In the present disclosure, the terms "unit" or "module" may include units implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logical blocks, components, or circuits. A "unit" or "module" may be an integrally formed component or the smallest unit or a part of a component that performs one or more functions. For example, a module may be implemented as an application specific integrated circuit (ASIC).

[0398] The various embodiments of the present disclosure may be implemented as software, including instructions stored in a machine (e.g., computer) readable storage medium. A machine is a device that invokes instructions stored in the storage medium and operates according to the invoked instructions, and may include an electronic device (e.g., electronic device 100) according to the disclosed embodiments.

[0399] In the case where the instructions are executed by a processor, the processor may directly or under the control of the processor use other elements to perform functions corresponding to the instructions. The instructions may include code generated by a compiler or code executed by an interpreter.

[0400] The electronic device and the method for controlling the electronic device according to the present disclosure may obtain a high-quality image frame. In particular, even in low grayscale, a high-quality image frame with minimized noise influence may be obtained.

[0401] The electronic device and the method for controlling the electronic device according to the present disclosure may provide a high-quality image frame to a user according to a user command for capturing by real-time identifying a set of optimal parameter setting values via a live view image.

[0402] The electronic device and the method for controlling the electronic device according to the present disclosure may generate a high-quality image frame by combining multiple images obtained by multiple cameras to which different sets of setting values are applied.

[0403] In addition, in the related art, moiré phenomena (e.g., horizontal lines in a sky background) are likely to occur in high-resolution images. However, in the electronic device and the method for controlling the electronic device according to the present disclosure, moiré phenomena are hardly observable, and the effect of reducing ghosting phenomena is reduced because a live view image is segmented into a plurality of regions according to brightness and objects included in an image frame, a plurality of image frames are obtained using a set of camera parameter setting values corresponding to the plurality of regions, and an image frame corresponding to a user command for capturing is generated using the plurality of image frames.

[0404] Although the exemplary embodiments have been described in detail above, it should be understood that various modifications and variations can be made without departing from the spirit and scope of the present disclosure defined by the appended claims and their equivalents.

Claims

1. An electronic device, comprising: a plurality of cameras; and at least one processor communicatively connected to the plurality of cameras, wherein the processor is configured to: Based on a first user command to obtain a live view image, divide an image frame obtained via a camera among the plurality of cameras into a plurality of regions based on the luminance and objects of the pixels included in the image frame; Obtain a plurality of sets of camera parameter setting values, each set of camera parameter setting values among the plurality of sets of camera parameter setting values including a plurality of parameter values corresponding to a respective region among the plurality of regions; Based on a second user command to capture a live view image, obtain a plurality of image frames by applying the plurality of sets of camera parameter setting values to at least one camera among the plurality of cameras; and Obtain an image frame corresponding to the second user command by merging the plurality of obtained image frames, wherein the at least one processor is further configured to: Obtain a plurality of first sets of camera parameter setting values by inputting the plurality of regions into a first neural network model, and Obtain the plurality of sets of camera parameter setting values by inputting the plurality of first sets of camera parameter setting values, the image frames obtained via the at least one camera, and the plurality of parameter values of the camera into a second neural network model.

2. The electronic device according to claim 1, wherein The first neural network model is a neural network model trained based on regions obtained from training image frames, a plurality of parameter values of the cameras that have obtained the training image frames, and edge metrics, contrast metrics, and background metrics of the regions, wherein the edge metrics include metrics for indicating the sharpness of the image frame, the contrast metrics include metrics for indicating the luminance of the image frame, and the background metrics include metrics for indicating the noise level of the background region of the image frame.

3. The electronic device according to claim 1, wherein, The second neural network model is configured to output information on edge metrics, contrast metrics, and background metrics of an input image frame corresponding to the input set of camera parameter setting values based on the input set of camera parameter setting values, the image frames obtained by the camera, and the plurality of parameter values of the camera, wherein the edge metrics include metrics for indicating the sharpness of the image frame, the contrast metrics include metrics for indicating the luminance of the image frame, and the background metrics include metrics for indicating the noise level of the background region of the image frame, and wherein the at least one processor is further configured to obtain, based on the information on edge metrics, contrast metrics, and background metrics obtained from the second neural network model, a set of camera parameter setting values corresponding to the maximum edge metric, a set of camera parameter setting values corresponding to the maximum contrast metric, and a set of camera parameter setting values corresponding to the maximum background metric among the plurality of first sets of camera parameter setting values.

4. The electronic device according to claim 1, wherein, The at least one processor is further configured to identify the pixel values of the pixels included in each of the plurality of regions, and obtain the plurality of sets of camera parameter setting values based on the identified pixel values and predefined rules.

5. The electronic device according to claim 1, wherein, The at least one processor is further configured to obtain an image frame by applying a set of camera parameter settings from among a plurality of sets of camera parameter settings to a first camera of the plurality of cameras based on a second user command to capture a live view image, and obtain at least two image frames by applying at least two sets of camera parameter settings from among the plurality of sets of camera parameter settings to a second camera of the plurality of cameras.

6. The electronic device according to claim 1, wherein, The at least one processor is further configured to obtain an image frame corresponding to the second user command by inputting the merged image frame into a third neural network model.

7. The electronic device according to claim 1, wherein, Each of the plurality of obtained image frames is a Bayer raw image.

8. The electronic device according to claim 6, wherein, The image frame obtained from the third neural network model has undergone at least one of black level adjustment, color correction, gamma correction, or edge enhancement.

9. A method for controlling an electronic device, the electronic device including a plurality of cameras, the method including: Based on a first user command to obtain a live view image, segmenting an image frame obtained via a camera of the plurality of cameras into a plurality of regions based on the luminance and objects of pixels included in the image frame; Obtaining a plurality of sets of camera parameter settings, each set of camera parameter settings among the plurality of sets of camera parameter settings including a plurality of parameter values corresponding to a respective one of the plurality of regions; Based on a second user command to capture a live view image, obtaining a plurality of image frames by applying the plurality of sets of camera parameter settings to at least one of the plurality of cameras; And Obtaining an image frame corresponding to the second user command by merging the plurality of obtained image frames, wherein obtaining the plurality of sets of camera parameter settings includes: Obtaining a plurality of first sets of camera parameter settings by inputting the plurality of regions into a first neural network model, and Obtaining the plurality of sets of camera parameter settings by inputting the plurality of first sets of camera parameter settings, the image frame obtained via the camera, and the plurality of parameter values of the camera into a second neural network model.

10. The method according to claim 9, wherein, The first neural network model is a neural network model trained based on regions obtained from training image frames, a plurality of parameter values of a camera from which the training image frames have been obtained, and edge metrics, contrast metrics, and background metrics of the regions, wherein the edge metrics include a metric for indicating the sharpness of the image frame, the contrast metrics include a metric for indicating the luminance of the image frame, and the background metrics include a metric for indicating the degree of noise in the background region of the image frame.

11. The method according to claim 9, wherein, The second neural network model is configured to output information on edge metrics, contrast metrics, and background metrics of an input image frame corresponding to an input set of camera parameter settings based on the input set of camera parameter settings, the image frame obtained via the camera, and the plurality of parameter values of the camera, wherein the edge metrics include a metric for indicating the sharpness of the image frame, the contrast metrics include a metric for indicating the luminance of the image frame, and the background metrics include a metric for indicating the degree of noise in the background region of the image frame, and Among them, obtaining a plurality of sets of camera parameter setting values includes: based on the information about the edge index, contrast index, and background index obtained from the second neural network model, obtaining the set of camera parameter setting values corresponding to the maximum edge index, the set of camera parameter setting values corresponding to the maximum contrast index, and the set of camera parameter setting values corresponding to the maximum background index in the plurality of first sets of camera parameter setting values.

12. The method according to claim 9, wherein, Obtaining a plurality of sets of camera parameter setting values includes: identifying the pixel values of the pixels included in each of the plurality of regions, and obtaining a plurality of sets of camera parameter setting values based on the identified pixel values and predefined rules.

13. A non-transitory computer-readable storage medium storing program code that can be executed by at least one processor to cause the at least one processor to control an electronic device including a plurality of cameras by performing the following steps: Based on a first user command to obtain a live view image, segmenting an image frame obtained by a camera among the plurality of cameras into a plurality of regions based on the brightness and objects of the pixels included in the image frame; Obtaining a plurality of sets of camera parameter setting values, each set of camera parameter setting values among the plurality of sets of camera parameter setting values including a plurality of parameter values corresponding to a corresponding region among the plurality of regions; Based on a second user command to capture a live view image, obtaining a plurality of image frames by applying the plurality of sets of camera parameter setting values to at least one camera among the plurality of cameras; And Obtaining an image frame corresponding to the second user command by combining the plurality of obtained image frames, wherein obtaining the plurality of sets of camera parameter setting values includes: Obtaining a plurality of first sets of camera parameter setting values by inputting the plurality of regions into a first neural network model, and Obtaining the plurality of sets of camera parameter setting values by inputting the plurality of first sets of camera parameter setting values, the image frames obtained via the at least one camera, and the plurality of parameter values of the camera into a second neural network model.

Citation Information

Patent Citations

  • Picture focusing method and device, terminal and corresponding storage medium

    CN109561257A

  • Image processing method and device, storage medium and electronic equipment

    CN110493538A