Electronic device, method, and storage medium for switching camera

A depth estimation model in electronic devices with multiple cameras addresses the challenge of FoV switching and depth estimation, ensuring seamless image transitions and accurate distance information.

WO2026014722A1PCT designated stage Publication Date: 2026-01-15SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/007658
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-24
Filing Date
2025-06-04
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing electronic devices with multiple cameras struggle to efficiently switch between different field of views (FoVs) and accurately estimate depth information for seamless image transitions, leading to suboptimal user experience and functionality.

Method used

The implementation of a depth estimation model that learns to estimate depth information using a converted image based on a wider second FoV, allowing for seamless switching between cameras and displaying preview images adjusted according to reference values, enhancing depth information accuracy.

Benefits of technology

Enables smooth transitions between camera views with improved depth estimation, providing enhanced user experience and accurate distance information for various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025007658_15012026_PF_FP_ABST
    Figure KR2025007658_15012026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise: a memory for storing instructions; at least one processor; a first camera for providing a first field of view (FoV); a second camera for providing a second FoV; and a display. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to: acquire first depth information about a first preview image while displaying the first preview image acquired via the first camera; display a second preview image acquired via the second camera according to the first depth information that is not greater than a reference value; acquire second depth information about the second preview image on the basis of a depth estimation model by using data pertaining to the second preview image while displaying the second preview image; and display a third preview image acquired via the first camera according to the second depth information that is greater than the reference value.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and storage medium for switching a camera

[0001] The descriptions below relate to electronic devices, methods, and storage media for switching cameras.

[0002] An electronic device may include multiple cameras. For example, the electronic device may acquire an image of the external environment using at least one of the multiple cameras. For example, the electronic device may use the image to identify the distance from the camera (or the electronic device) to an external object within the external environment.

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.

[0004] An electronic device according to embodiments of the present disclosure is provided. The electronic device may include a memory that stores instructions and includes one or more storage media. The electronic device may include at least one processor including a processing circuit. The electronic device may include a first camera that provides a first field of view (FoV). The electronic device may include a second camera that provides a second FoV wider than the first FoV. The electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire first depth information for a region of interest (ROI) within the first preview image based on the first camera while displaying a first preview image acquired through the first camera on the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, on the display, a second preview image, which is changed from the first preview image and acquired through the second camera, according to the first depth information that is less than or equal to a reference value. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain second depth information for an ROI within the second preview image based on a depth estimation model learned to estimate depth information for the converted image using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera using data about the second preview image while displaying the second preview image on the display.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, on the display, a third preview image, which is changed from the second preview image and acquired through the first camera, according to the second depth information exceeding the reference value.

[0005] A method performed by an electronic device according to embodiments of the present disclosure is provided. The method performed by the electronic device may include an operation of acquiring first depth information for a region of interest (ROI) within a first preview image based on a first camera, wherein the first preview image is acquired through a first camera providing a first field of view (FoV) of the electronic device, while displaying the first preview image. The method may include an operation of displaying a second preview image, wherein the second preview image is acquired through a second camera, wherein the second camera provides a second FoV wider than the first FoV of the electronic device and is changed from the first preview image, based on the first depth information being less than or equal to a reference value. The method may include an operation of acquiring second depth information for the ROI within the second preview image based on a depth estimation model learned to estimate depth information for the converted image using a converted image based on the second FoV, generated from an image representing the first FoV acquired through the first camera, using data regarding the second preview image, while displaying the second preview image. The method may include an operation of displaying a third preview image obtained through the first camera and changed from the second preview image according to the second depth information exceeding the reference value.

[0006] A non-transitory computer-readable storage medium according to embodiments of the present disclosure is provided. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by at least one processor of an electronic device including a first camera providing a first field of view (FoV), a second camera providing a second FoV wider than the first FoV, and a display, cause the electronic device to acquire first depth information for a region of interest (ROI) within the first preview image based on the first camera while displaying a first preview image acquired through the first camera on the display. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to display, on the display, a second preview image, which is changed from the first preview image and acquired through the second camera, according to the first depth information that is less than or equal to a reference value. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to obtain second depth information for an ROI within the second preview image based on a depth estimation model learned to estimate depth information for the converted image using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera using data about the second preview image while displaying the second preview image on the display.The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to display a third preview image, which is changed from the second preview image and acquired through the first camera, on the display according to the second depth information exceeding the reference value.

[0007] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.

[0008] FIG. 2 is a block diagram illustrating a camera module according to various embodiments.

[0009] Figure 3 illustrates an exemplary block diagram of an electronic device.

[0010] Figures 4a and 4b illustrate examples of attributes of multiple cameras of an electronic device.

[0011] FIGS. 5A to 5C illustrate examples of a method for performing switching between multiple cameras of an electronic device using depth information acquired based on a depth estimation model.

[0012] FIG. 6A illustrates an example of an operational flow for a method in which an electronic device learns a depth estimation model and acquires depth information based on the depth estimation model.

[0013] FIG. 6b illustrates an example of how an electronic device learns a depth estimation model and acquires depth information based on the depth estimation model.

[0014] Figure 6c illustrates an example of a method for identifying distance information using depth information.

[0015] FIG. 7 illustrates an example of an operational flow for a method in which an electronic device learns a depth estimation model using a transformed image based on a second FoV generated from an image representing a first FoV (field of view).

[0016] Figure 8 illustrates an example of an operational flow for a method of performing switching between multiple cameras using a learned depth estimation model.

[0017] The terms used in this disclosure are merely used to describe specific embodiments and may not be intended to limit the scope of the disclosed embodiments. Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this disclosure. Terms defined in general dictionaries among the terms used in this disclosure may be interpreted as having the same or similar meaning as the meaning they have in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude embodiments of the present disclosure.

[0018] The various embodiments of the present disclosure described below illustrate a hardware-based approach as an example. However, since the various embodiments of the present disclosure include techniques utilizing both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.

[0019] In addition, in the present disclosure, expressions such as "more than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled, but this is merely a description for expressing an example and does not exclude descriptions such as "more than" or "less than." A condition described as "more than" may be replaced with "more than," a condition described as "less than" may be replaced with "less than," and a condition described as "more than and less than" may be replaced with "more than and less than." In addition, hereinafter, "A" to "B" mean at least one of the elements from A (including A) to B (including B). hereinafter, "C" and / or "D" mean at least one of "C" or "D," that is, including {"C", "D", "C" and "D"}. hereinafter, the meaning of "about E" may be replaced with a value within a margin of error of ±5% or ±10% based on E.

[0020] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.

[0021] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0022] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0023] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0024] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0025] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0026] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0027] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0028] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0029] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0030] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0031] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0032] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0033] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0034] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0035] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0036] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0037] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0038] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0039] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0040] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0041] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0042] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0043] FIG. 2 is a block diagram illustrating a camera module according to various embodiments.

[0044] Referring to FIG. 2, the camera module (180) may include a lens assembly (210), a flash (220), an image sensor (230), an image stabilizer (240), a memory (250) (e.g., a buffer memory), or an image signal processor (260). The lens assembly (210) may collect light emitted from a subject that is a target of image capturing. The lens assembly (210) may include one or more lenses. According to one embodiment, the camera module (180) may include a plurality of lens assemblies (210). In this case, the camera module (180) may form, for example, a dual camera, a 360-degree camera, or a spherical camera. Some of the plurality of lens assemblies (210) may have the same lens properties (e.g., angle of view, focal length, auto focus, f-number, or optical zoom), or at least one lens assembly may have one or more lens properties that are different from the lens properties of the other lens assemblies. The lens assembly (210) may include, for example, a wide-angle lens, an ultra wide-angle lens, or a telephoto lens.

[0045] The flash (220) can emit light used to enhance light emitted or reflected from a subject. According to one embodiment, the flash (220) can include one or more light-emitting diodes (e.g., a red-green-blue (RGB) light-emitting diode (LED), a white LED, an infrared LED, or an ultraviolet LED), or a xenon lamp. The image sensor (230) can acquire an image corresponding to the subject by converting light emitted or reflected from the subject and transmitted through the lens assembly (210) into an electrical signal. According to one embodiment, the image sensor (230) can include one image sensor selected from among image sensors having different properties, such as an RGB sensor, a black and white (BW) sensor, an IR sensor, or a UV sensor, a plurality of image sensors having the same property, or a plurality of image sensors having different properties. Each image sensor included in the image sensor (230) can be implemented using, for example, a CCD (charged coupled device) sensor or a CMOS (complementary metal oxide semiconductor) sensor.

[0046] The image stabilizer (240) can move at least one lens or image sensor (230) included in the lens assembly (210) in a specific direction or control the operating characteristics of the image sensor (230) (e.g., adjusting the read-out timing, etc.) in response to the movement of the camera module (180) or the electronic device (101) including the same. This allows compensating for at least some of the negative effects of the movement on the captured image. In one embodiment, the image stabilizer (240) can detect such movement of the camera module (180) or the electronic device (101) using a gyro sensor (not shown) or an acceleration sensor (not shown) disposed inside or outside the camera module (180). In one embodiment, the image stabilizer (240) can be implemented as, for example, an optical image stabilizer. The memory (250) can temporarily store at least a portion of the image acquired through the image sensor (230) for the next image processing task. For example, when image acquisition is delayed due to the shutter, or when multiple images are acquired at high speed, the acquired original image (e.g., a Bayer-patterned image or a high-resolution image) is stored in the memory (250), and a corresponding copy image (e.g., a low-resolution image) can be previewed through the display module (160). Thereafter, when a specified condition is satisfied (e.g., a user input or a system command), at least a portion of the original image stored in the memory (250) can be acquired and processed, for example, by the image signal processor (260). According to one embodiment, the memory (250) can be configured as at least a portion of the memory (130) or as a separate memory that operates independently therefrom.

[0047] The image signal processor (260) can perform one or more image processing operations on an image acquired through an image sensor (230) or an image stored in a memory (250). The one or more image processing operations may include, for example, depth map generation, 3D modeling, panorama generation, feature point extraction, image synthesis, or image compensation (e.g., noise reduction, resolution adjustment, brightness adjustment, blurring, sharpening, or softening). Additionally or alternatively, the image signal processor (260) may perform control (e.g., exposure time control, read-out timing control, etc.) on at least one of the components included in the camera module (180) (e.g., image sensor (230)). An image processed by the image signal processor (260) may be stored back in the memory (250) for further processing or provided to an external component of the camera module (180) (e.g., memory (130), display module (160), electronic device (102), electronic device (104), or server (108)). In one embodiment, the image signal processor (260) may be at least a part of the processor (120). It may be configured as a separate processor that is configured or operates independently from the processor (120). If the image signal processor (260) is configured as a separate processor from the processor (120), at least one image processed by the image signal processor (260) may be displayed through the display module (160) as is or after undergoing additional image processing by the processor (120).

[0048] According to one embodiment, the electronic device (101) may include a plurality of camera modules (180), each having different properties or functions. In this case, for example, at least one of the plurality of camera modules (180) may be a wide-angle camera, and at least another may be a telephoto camera. For example, the plurality of camera modules (180) may further include an ultra-wide-angle camera. Similarly, at least one of the plurality of camera modules (180) may be a front camera, and at least another may be a rear camera.

[0049] Figure 3 illustrates an exemplary block diagram of an electronic device.

[0050] Fig. 3 illustrates an exemplary block diagram of an electronic device (101). The electronic device (101) of Fig. 3 may be an example of the electronic device (101) of Fig. 1.

[0051] Referring to FIG. 3, according to one embodiment, an electronic device (101) may include a processor (310), a display (320), a camera assembly (330), and a memory (340). However, the embodiments of the present disclosure are not limited thereto. For example, the processor (310), the display (320), the camera assembly (330), and the memory (340) may be electrically and / or operably coupled with each other by a communication bus. Hereinafter, operably coupled hardware components may mean that a direct connection or an indirect connection is established between the hardware components, either wired or wireless, such that a second hardware component is controlled by a first hardware component among the hardware components. Although illustrated based on different blocks, the embodiment is not limited thereto, and some of the hardware components illustrated in FIG. 3 (e.g., at least a portion of the processor (310) and the memory (340)) may be included in a single integrated circuit such as a system on a chip (SoC) or a system in package (SIP). The type and / or number of hardware components included in the electronic device (101) is not limited to those illustrated in FIG. 3. For example, the electronic device (101) may include only some of the hardware components illustrated in FIG. 3.

[0052] According to one embodiment, the processor (310) of the electronic device (101) may include a hardware component for processing data based on one or more instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), and a field programmable gate array (FPGA). As an example, the hardware component for processing data may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a microcontroller (MCU), and / or a neural processing unit (NPU). The number of processors (310) may be one or more. For example, the processor (310) may have a multi-core processor structure such as a dual core, a quad core, or a hexa core. The processor (310) of FIG. 3 may be substantially identically applied to the processor (120) of FIG. 1.

[0053] For example, the processor (310) may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include various processing circuits including at least one processor, one or more of which may be configured to individually and / or collectively perform the various functions described below in a distributed manner. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms encompass, for example, and without limitation, situations where one processor performs some of the recited functions and other processor(s) perform other parts of the recited functions, and also situations where one processor may perform all of the recited functions. Additionally, the at least one processor may include a combination of processors that perform the various functions enumerated / disclosed, for example, in a distributed manner. At least one processor may execute program instructions to achieve or perform the various functions.

[0054] According to one embodiment, a display (320) of an electronic device (101) can output visualized information to a user of the electronic device (101). For example, the display (320) can be controlled by a processor (310) including circuits such as a CPU, a GPU (graphic processing unit), and / or a DPU (display processing unit) to output visualized information to the user. The display (320) can include a flexible display, a flat panel display (FPD), and / or electronic paper. The display (320) can include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LEDs can include organic LEDs (OLEDs). Embodiments are not limited thereto, and for example, if the electronic device (101) includes a lens for transmitting external light (or ambient light), the display (320) may include a projector (or projection assembly) for projecting light onto the lens. In one embodiment, the display (320) may be referred to as a display panel and / or a display module. For example, the display (320) may be an example of the display module (160) of FIG. 1.

[0055] According to one embodiment, the camera assembly (330) of the electronic device (101) may be used to acquire an image of the external environment of the electronic device (101). For example, the external environment may be referred to as the surrounding environment of the electronic device (101). For example, the camera assembly (330) may include a plurality of cameras. For example, the plurality of cameras may include a first camera (331) and a second camera (332). For example, the first camera (331) may be referred to as a wide-angle camera, or a W (wide) camera. For example, the second camera (332) may be referred to as an ultra-wide-angle camera, or a UW (ultra-wide) camera. However, the present disclosure is not limited thereto. For example, the camera assembly (330) may further include a telephoto camera. For example, the first camera (331) and the second camera (332) can be distinguished based on the properties of each camera (or lens). For example, the properties can include at least one of a field of view (FoV) (e.g., a diagonal field of view), a focal length, or a depth of field (DoF). In one example, the first camera (331) can have a first FoV, and the second camera (332) can have a second FoV that is higher (or wider) than the first FoV. As a non-limiting example, the first FoV can be about 85°, and the second FoV can be about 120°. For specific details on the first camera (331) and the second camera (332) distinguished based on the properties, reference may be made to FIGS. 4A and 4B below. The camera assembly (330) of FIG. 3 may include at least a portion of the camera module (180) of FIGS. 1 and 2.

[0056] According to one embodiment, the memory (340) of the electronic device (101) may include a hardware component for storing data and / or instructions input to and / or output from the processor (310). The memory may include, for example, a volatile memory such as a random-access memory (RAM), and / or a non-volatile memory such as a read-only memory (ROM). The volatile memory may include, for example, at least one of a dynamic RAM (DRAM), a static RAM (SRAM), a cache RAM, and a pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a flash memory, a hard disk, a compact disc, and an embedded multimedia card (eMMC). The specific details of the memory (211) of Fig. 2 can be applied substantially identically to the details of the memory (130) of Fig. 1.

[0057] According to one embodiment, one or more instructions (or commands) representing operations and / or actions to be performed on data by the processor (310) of the electronic device (101) may be stored in the memory (340) of the electronic device (101). A set of one or more instructions may be referred to as a program, firmware, an operating system, a process, a routine, a sub-routine, and / or an application. Hereinafter, when an application is installed in the electronic device (e.g., the electronic device (101)), it may mean that one or more instructions provided in the form of an application are stored in the memory (340), and that the one or more applications are stored in a format executable by the processor of the electronic device (e.g., a file having an extension designated by the operating system of the electronic device (101)). According to one embodiment, the electronic device (101) may execute one or more instructions stored in the memory (340) to perform the operations of FIGS. 6A, 7, and 8. For example, the one or more instructions, when executed by the processor (310), may cause the electronic device (101) to perform at least some of the operations of FIGS. 6A, 7, and 8.

[0058] According to one embodiment, the memory (340) may include a depth estimation model (350). The electronic device (101) may obtain distance information (or distance) for the external environment by identifying depth information for the external environment using an image (or a red-green-blue image (RGB)) acquired through one camera of the camera assembly (330). For example, the depth information may be referenced as a depth map. For example, the depth information may be used to identify distance information (or distance, distance value) for the external environment. The depth information may include a depth for each pixel of the image. For example, a monocular depth estimation (MDE) may be used to identify the depth information using an image acquired through one camera. For example, the MDE may represent a technique for identifying the depth information based on a depth estimation model (350) using an image acquired through the one camera. For example, the depth estimation model (350) may be referred to as an MDE model, an artificial intelligence model, a deep-learning model, a neural network, an AI engine, an AI module, or a machine-learning model. In the present disclosure, the depth estimation model (350) may be referred to as a distance estimation model. In other words, the MDE may output depth information corresponding to an image by using an image acquired through a single camera as an input without using information from an additional device (e.g., a sensor). For example, the depth information may be used to identify the distance information (or distance, distance value).

[0059] In one example, the memory (340) may store data (or learning data) for learning the depth estimation model (350). For example, the learning data may include an image acquired through a camera of the camera assembly (330) and depth information corresponding to the image. For example, the image may be an image representing the second FoV (or based on the second FoV). For example, the depth information may be depth information representing the second FoV (or based on the second FoV).

[0060] For example, the image of the learning data may be generated using an image acquired immediately before switching from the first camera (331) to the second camera (332) when switching from the first camera (331) to the second camera (332) of the camera assembly (330) is performed. The image may include a transformation image generated from a last image representing the first FoV among one or more images acquired through the first camera (331). For example, the transformation image may be generated based on (or according to) the second FoV. The transformation image may include information within a range (or region) that matches (corresponds to) the last image representing the first FoV acquired through the first camera (331) to represent the second FoV. For example, the final image representing the first FoV may be included in the converted image by adjusting the view point of the final image and the size and position of the contents within the image to represent the second FoV. For specific details related thereto, reference may be made to FIG. 6B. Alternatively, for example, the image may include an image acquired immediately after switching from the first camera (331) to the second camera (332) when switching from the first camera (331) to the second camera (332) of the camera assembly (330) is performed. The image may include a first image representing the second FoV among one or more images acquired through the second camera (332). For example, the depth information may include converted depth information based on the second FoV generated from depth information corresponding to the final image among one or more images acquired through the first camera (331) and representing the first FoV. For example, the transformed depth information may be generated based on (or according to) the second FoV.The above-described converted depth information may include information within a range (or region) that matches (or corresponds to) the depth information representing the first FoV in order to represent the second FoV. For example, the depth information representing the first FoV may be included in the converted depth information by adjusting the reference time point (corresponding to the shooting time point) of the depth information and the size and position of contents within the depth information so as to represent the second FoV. For specific details related thereto, reference may be made to FIG. 6B. Learning the training data including images before or after switching is performed by the depth estimation model (350) may include performing fine tuning for a specific distance (e.g., a reference distance for the switching) (or a depth corresponding to the specific distance). Accordingly, the accuracy of the estimation of the depth estimation model (350) for the specific distance may increase. For specific details on the learning data used for learning the depth estimation model (350) and specific details on the method by which the depth estimation model (350) learns the learning data, reference may be made to FIG. 7 below.

[0061] Figures 4a and 4b illustrate examples of attributes of multiple cameras of an electronic device.

[0062] FIG. 4A illustrates an example (410) of the FoV and the focal length of the lens of the first camera (331) and an example (420) of the FoV and the focal length of the lens of the second camera (332). In FIG. 4A, for convenience of explanation, it is assumed that the electronic device (101) acquires an image of an external object (405) located at the same distance from the lens (411 or 421) using each of the first camera (331) and the second camera (332).

[0063] Referring to example (410), the first camera (331) may include an image sensor (230) and a lens (411). For example, the electronic device (101) may obtain an image of an external object (405) through the first camera (331) having a first focal length (419) of the lens (411). For example, the first focal length (419) of the lens (411) may be referred to as a distance between the image sensor (230) and the lens (411). For example, an image of an external object (405) obtained through the first camera (331) having the first focal length (419) may represent a first FoV (415). For example, an image of an external object (405) representing a first FoV (415) may include an area of ​​an external environment including the external object (405) that can be represented by an image of the external object (405) acquired through the first camera (331) corresponding to the first FoV (415). For example, the first FoV (415) may include a first angular FoV (AFoV) (417). For example, the first AFoV (417) may be determined based on light that passes through the center of the lens (411) and is received at a periphery portion of the image sensor (230). For example, the first camera (331) may provide (or support) the first FoV (415).

[0064] Referring to example (420), the second camera (332) may include an image sensor (230) and a lens (421). For example, the electronic device (101) may acquire an image of an external object (405) through the second camera (332) having a second focal length (429) of the lens (421). For example, the second focal length (429) of the lens (421) may be referred to as a distance between the image sensor (230) and the lens (421). For example, the second focal length (429) may be shorter than the first focal length (419). For example, an image of the external object (405) acquired through the second camera (332) having the second focal length (429) may represent a second FoV (425). For example, the second FoV (425) may be higher (or wider) than the first FoV (415). For example, the image for the external object (405) representing the second FoV (425) may include that an area of ​​the external environment including the external object (405) that can be represented by the image for the external object (405) acquired through the second camera (332) corresponds to the second FoV (425). For example, the second FoV (425) may include a second AFoV (427). For example, the second AFoV (427) may be determined based on light that passes through the center of the lens (421) and is received at an edge portion of the image sensor (230). For example, the second camera (332) may provide (or support) the second FoV (425).

[0065] Referring to FIG. 4A, the electronic device (101) can obtain an image of an external object (405) through the first camera (331) or the second camera (332). For example, the image can be obtained using a plurality of partial images. For example, each of the plurality of partial images can represent image information obtained by each of the plurality of photo diodes (PDs) included in the image sensor (230). For example, the image sensor (230) can include a plurality of PD sets (431, 432, 433) and microlenses (441, 442, 443) corresponding to each of the plurality of PD sets (431, 432, 433). For example, the first microlens (441) can correspond to the first PD set (431). For example, the second microlens (442) can correspond to the second PD set (432). For example, the third microlens (443) may correspond to the third PD set (433). For example, each PD set may include multiple PDs. For example, the first PD set (431) may include PD#1 (431a) and PD#2 (431b). For example, the second PD set (432) may include PD#3 and PD#4. For example, the third PD set (433) may include PD#N-1 and PD#N. In FIG. 4A, an example of an image sensor (230) according to a pixel type of 2PD is illustrated, but the present disclosure is not limited thereto. For example, the image sensor (230) may be formed with another pixel type (e.g., 4PD (or 2x2 OCL), or 2x1 OCL (on-chip lens)). For example, PD#1 (431a) may acquire a first partial image. For example, PD#2 (431b) can acquire the second partial image. For example, PD#N-1 can acquire the first partial image. PD#N can acquire the second partial image.For example, the image can be obtained using the partial images obtained using the PDs of the image sensor (230).

[0066] In addition, in one example, depth information about an external environment including an external object (405) represented by the image may be obtained using the partial images. For example, the depth information may be obtained based on a phase difference between partial images obtained from each of the PDs of the PD sets. For example, the phase difference may represent a phase difference between a plurality of PDs. The phase difference between the PDs of each of the plurality of PD sets may be defined as a disparity map. Obtaining the depth information using the phase difference between the PDs as described above may be referred to as phase detection. For example, the depth information may be obtained based on a phase difference (e.g., number of pixels) corresponding to an area according to the PDs using a first partial image obtained using PD#1 (431a) of the first PD set (431) and a second partial image obtained using a PD of the second PD set (432) (e.g., PD#3 of the second PD set (432)). In the above example, an example is described in which the depth information is acquired by utilizing the phase difference between one first partial image and one second partial image, but the present disclosure is not limited thereto. For example, the depth information can be acquired by utilizing the phase difference identified by utilizing a plurality of first partial images and a plurality of second partial images acquired from a plurality of PD sets (431, 432, 433) of the image sensor (230).

[0067] In FIG. 4A, both the first camera (331) and the second camera (332) are illustrated as including the same image sensor (230), but the present disclosure is not limited thereto. For example, the image sensor (230) of the first camera (331) may be different from the image sensor (230) of the second camera (331).

[0068] FIG. 4B illustrates an example (460) of the depth of field (DoF) of the first camera (331) and an example (470) of the DoF of the second camera (332). In FIG. 4B, for convenience of explanation, it is assumed that the electronic device (101) acquires an image of an external object (455) located at the same distance from each of the first camera (331) and the second camera (332). For example, the depth (459) from the electronic device (101) to the external object (455) in example (460) may be the same as the depth (459) from the electronic device (101) to the external object (455) in example (470).

[0069] Referring to example (460), a user (450) can obtain an image of an external object (455) through a first camera (331) of an electronic device (101). For example, a depth (459) from the electronic device (101) to the external object (455) can be referred to as a distance between the first camera (331) and the external object (455). For example, the first camera (331) can provide a first DoF (465).

[0070] Referring to example (470), the user (450) can obtain an image of an external object (455) through the second camera (332) of the electronic device (101). For example, the depth (459) from the electronic device (101) to the external object (455) can be referred to as the distance between the second camera (332) and the external object (455). For example, the second camera (332) can provide a second DoF (475).

[0071] For example, the DoF may represent a range for depth with a certain level of clarity or higher in an area including an external object (455). For example, the first DoF (465) may be narrower (or shallower) than the second DoF (475). In other words, the first range for depth with a certain level of clarity or higher in an image including an external object (455) acquired through the first camera (331) may be narrower (or shallower) than the second range for depth with a certain level of clarity or higher in an image including an external object (455) acquired through the second camera (332).

[0072] Referring to FIGS. 4A and 4B , the first camera (331) and the second camera (332) may have different properties. For example, the first FoV (415) of the first camera (331) may be narrower (or lower) than the second FoV (425) of the second camera (332). For example, the first focal length (419) of the first camera (331) may be longer than the second focal length (429) of the second camera (332). For example, the first DoF (465) of the first camera (331) may be narrower than the second DoF (475) of the second camera (332).

[0073] In order to train a distance estimation model for the above MDE, an image and actual distance information (or actual distance information, actual distance, actual distance) corresponding to the image may be required. For example, depth information for identifying the distance information may be acquired based on various methods. For example, the depth information may be acquired using a sensor for ToF (time of flight) or LiDAR (light detection and ranging). Or, for example, the depth information may be acquired through additional operations on a set of images continuously acquired (or photographed) by a single camera or multiple cameras (e.g., a stereo camera). Or, for example, the depth information may be acquired using images captured from various angles for the same external object to construct a 3D space from a 2D (dimensional) image based on SfM (structure from motion).

[0074] When using a stereo camera, two cameras are placed side by side, and by calculating disparity using images captured from one side (e.g., left) and another side (e.g., right) of an external object, distance information to the external object can be calculated using the focal length and baseline of the stereo camera. For example, the baseline can represent the distance between the center points (or reference points) of each of the two cameras of the stereo camera.

[0075] When using a set of images continuously acquired through a single camera, the distance to the external object can be calculated by applying SfM to the set of images captured from various viewpoints (or angles) of the single camera. For example, by extracting feature points of each image in the image set using a model such as SIFT (scale invariant feature transform), and matching (or mapping) the locations of similar feature points between the feature points for each image in the image set, 3D coordinates can be estimated, and a 3D space can be constructed.

[0076] The depth information (or distance information) obtained in the above examples and the image corresponding to the depth information can be used as training data for the distance estimation model. For example, the distance estimation model can be trained by performing regression or classification on each pixel value of the depth information included in the training data. Alternatively, for example, the distance estimation model can be trained through classification-regression performed by binning the distance and calculating the probability distribution for each interval of each pixel.

[0077] In one example, the distance estimation model for the MDE can be trained according to the purpose of use and learning conditions. For example, to measure (or identify) distance information within a wide range, the distance estimation model can be trained using relative depth information based on a large-scale data set. Thereafter, the distance estimation model can output depth information with high accuracy by performing fine-tuning. The fine-tuning may represent additional learning for adjusting a pre-trained model to fit a specific domain. In other words, the fine-tuning may be used to adjust the accuracy and directionality of the output result by further training the distance estimation model using training data similar to the target information (or target data). For example, the fine-tuning may be performed using a pair of an image to be trained and a ground truth (GT) corresponding to the image. The GT may be referred to as target data to be acquired (or predicted) through the distance estimation model.

[0078] For example, for the learned distance estimation model, depth information can be output by inputting an image. By comparing the output depth information with the depth information used as GT, the accuracy of the distance estimation model can be calculated. The distance estimation model can be evaluated according to the accuracy. For example, the accuracy can be evaluated by δ (threshold accuracy), REL (absolute relative error), RMSE (root mean squared error), and average log 10 It can be calculated based on indicators such as (error).

[0079] When utilizing the above distance estimation model, distance information to external objects included in the image can be identified by utilizing an image acquired through a single camera without using additional devices (e.g., sensors). For example, distance information to a region of interest (ROI) within the image can be identified. The ROI may represent an area within the image corresponding to the external object in the external environment.

[0080] As described above, when acquiring depth information using a sensor, high accuracy and stable measurement of the depth information are possible, but additional sensors and corresponding modules are required, which increases the production cost of the electronic device, requires space for placement within the electronic device, and increases the power consumption of the electronic device. Alternatively, even when a stereo camera is used but no sensor is used, the production cost and placement space for arranging and using multiple cameras are high, and additional calculations may be required for images acquired through the stereo camera. Alternatively, when a distance estimation model learned using random data is used, the depth information for a specific range of distances (or distance information) of the distance estimation model may have low accuracy. In one example, additional learning of the distance estimation model to increase the accuracy of the depth information may cause a disadvantage of a method of acquiring depth information using a sensor in that it utilizes depth information acquired through an additional sensor.

[0081] Hereinafter, the present disclosure may utilize MDE for switching between a plurality of cameras (e.g., a first camera (331) and a second camera (332)) of an electronic device (101). For example, the electronic device (101) may obtain depth information to an external object located within a reference distance based on a depth estimation model (350) for the MDE using an image obtained from a camera. As a non-limiting example, the electronic device (101) may perform a comparison between the depth information and a reference value indicating a reference distance (or a reference depth). In one example, among the plurality of cameras, the first camera (331) (or a wide-angle camera) may have a longer focal length than the second camera (332) (or an ultra-wide-angle camera), and thus the magnification of the image may be high. Since the disparity (or phase difference) between the PDs of the first camera (331) is accurate, it may be easier to obtain distance information to the external object located within the reference distance using the depth information obtained through the first camera (331) than to use the image obtained through the second camera (332). However, in order to obtain an image of the external object located within the reference distance, using the second camera (332) having a wider FoV than the first camera (331) may provide a relatively more convenient user experience to the user. In other words, from the user's perspective, the first camera (331) may be used to obtain an image of the external object located outside the reference distance, and the second camera (332) may be used to obtain an image of the external object located within the reference distance.At this time, while the second camera (332) is used to acquire an image of an external object located within the reference distance, if the distance between the external object and the electronic device (101) changes outside the reference distance or the electronic device (101) moves, MDE may be used to switch between the first camera (331) and the second camera (332) according to the reference distance. At this time, in order to perform the switching more naturally (seamlessly) and accurately, learning data for the depth estimation model (350) of the MDE may be required. For example, phase detection may be used to acquire the learning data (or depth information used as the correct answer (or GT) among the learning data) without using a device (e.g., a sensor) for measuring distance information.

[0082] For example, the present disclosure can use an image acquired through the first camera (331) and depth information about the image immediately before switching from the first camera (331) to the second camera (332) as learning data for learning the depth estimation model (350) of the MDE. The present disclosure can perform learning of the depth estimation model (350) by using an image acquired from one camera and depth information about the image without using an additional device. The present disclosure can acquire depth information for performing switching from the second camera (332) to the first camera (331) by using the learned depth estimation model (350). Accordingly, the production cost of an electronic device (101) using the depth estimation model (350) according to the present disclosure can be reduced, the deployable space within the electronic device (101) can be expanded, and the power consumption of the electronic device (101) can be reduced. In addition, by utilizing the image and corresponding depth information immediately before switching from the first camera (331) to the second camera (332) for learning the depth estimation model (350) according to the present disclosure, the accuracy of the depth estimation model (350) can be increased. Accordingly, as natural switching between cameras is performed while the electronic device (101) acquires an image of an external object (or performs shooting to acquire the image), the user's usability of the electronic device (101) can be improved.

[0083] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure pertains.

[0084] For specific details on a method for performing switching from the first camera (331) to the second camera (332) and switching from the second camera (332) to the first camera (331) based on the depth estimation model (350), reference may be made to FIGS. 5A to 5C below.

[0085] FIGS. 5A to 5C illustrate examples of a method for performing switching between multiple cameras of an electronic device using depth information acquired based on a depth estimation model.

[0086] FIGS. 5A to 5C illustrate examples (501, 502, 503) of a method in which an electronic device (101) performs switching between multiple cameras by using distance information to an external object (507) obtained based on a depth estimation model (350). The electronic device (101) of FIGS. 5A to 5C may be an example of the electronic device (101) of FIG. 3.

[0087] In FIGS. 5A to 5C , for convenience of explanation, it is assumed that a user (505) acquires an image of an external environment including an external object (507) using an electronic device (101). Examples (501 and 502) may represent a case where the user (505) moves closer to the external object (507), and examples (502 and 503) may represent a case where the user (505) moves away from the external object (507). In addition, in FIGS. 5A to 5C , for convenience of explanation, it is assumed that the electronic device (101) executes an application (hereinafter, referred to as a camera application) for acquiring an image through a camera (e.g., a first camera (331) or a second camera (332)) and displays a preview image corresponding to the acquired image within a UI (user interface) of the camera application on a display (320). However, the present disclosure is not limited thereto. For example, the present disclosure may also be applied when an electronic device (101) executes an application (or the camera application) for acquiring a video through a camera (e.g., a first camera (331) or a second camera (332)). In FIGS. 5A to 5C , the preview image displayed on the display (320) may be understood to be substantially identical to an image acquired by the electronic device (101) through a camera (e.g., a first camera (331) or a second camera (332)).

[0088] Referring to example (501) of FIG. 5A, the electronic device (101) can obtain an image of an external object (507) through the first camera (331). In example (501), the camera for obtaining the preview image may be the first camera (331). In this case, the camera for obtaining the preview image may be referred to as a main camera or a shooting camera. For example, the electronic device (101) can obtain a first preview image (513) obtained through the first camera (331) and display it on the display (320). For example, the first preview image (513) may represent at least a portion of an external environment including the external object (507). For example, the first preview image (513) may represent a first FoV (e.g., the first FoV (415) of FIG. 4A).

[0089] For example, the electronic device (101) may obtain first depth information to an external object (507) based on the first camera (331). For example, the electronic device (101) may obtain the first depth information for a ROI (517) within a first preview image (513). For example, the ROI (517) may include an area within the first preview image (513) corresponding to the external object (507) or a visual object.

[0090] For example, ROI (517) may include a pre-designated object or an area including the center coordinates of the preview image. For example, the pre-designated object may include a human face in the preview image. For example, the pre-designated object may be determined according to an auto focus function. For example, when the preview image includes multiple objects, ROI (517) may be determined as an area including the center coordinates of the preview image. Or, for example, ROI (517) may be determined according to an input received from the user (505) regarding the UI of the camera application.

[0091] For example, the first depth information may indicate a first distance (511) from the electronic device (101) (or the first camera (331)) to the external object (507). For example, the first depth information may include the first distance (511) or a phase difference (or phase value) corresponding to the first distance (511). For example, the electronic device (101) may obtain the first depth information associated with the ROI (517) by calculating a phase difference for each pixel of the ROI (517) of the first preview image (513). For example, the first depth information may be used to identify first distance information indicating the first distance (511).

[0092] For example, the electronic device (101) may compare the first depth information with a reference value. For example, the reference value may represent a reference depth (or reference distance) for switching the capturing camera for obtaining a preview image from the first camera (331) to the second camera (332). For example, the electronic device (101) may determine (or identify) whether the first depth information exceeds the reference value. For example, if the first depth information exceeds the reference value, the electronic device (101) may identify (or maintain) the capturing camera as the first camera (331). Alternatively, if the first depth information is less than or equal to the reference value, the electronic device (101) may change the capturing camera from the first camera (331) to the second camera (332). In example (501), the electronic device (101) can acquire the first depth information while the user (505) approaches the external object (507) in a first direction (519). Accordingly, the electronic device (101) can identify the photographing camera by comparing the acquired first depth information with the reference value. For example, the first depth information can be acquired periodically or aperiodically.

[0093] As the first depth information changes below the reference value according to movement (519) in example (501), a method for the electronic device (101) to acquire an image through the second camera (332) may be referred to example (502).

[0094] Referring to example (502) of FIG. 5B, the electronic device (101) can acquire an image of an external object (507) through the second camera (332). In example (502), the capturing camera for acquiring the preview image may be the second camera (332). For example, the electronic device (101) can acquire a second preview image (523) acquired through the second camera (332) and display it on the display (320). For example, the second preview image (523) can represent at least a portion of an external environment including the external object (507).

[0095] For example, the second preview image (523) may represent the same FoV as the first preview image (513) (e.g., the first FoV (415) of FIG. 4A). For example, the second preview image (523) may represent a first FoV that is the same as the first FoV of the first preview image (513). Accordingly, the user (505) may not recognize that the capturing camera is switched from the first camera (331) to the second camera (332). In the example, the electronic device (101) may modify (or crop) at least a portion of the image representing the second FoV when generating the second preview image (523) from the image representing the second FoV obtained through the second camera (332). As described above, generating the modified second preview image (523) may be performed by taking into account occlusion. For example, the occlusion may be caused by different capturing regions between cameras (e.g., the first camera (331) and the second camera (332)). For example, there may be a region that is not visible from the capturing region of the second camera (332) among the capturing region (or viewing region) of the first camera (331). For example, the electronic device (101) may generate the modified second preview image (523) by taking into account both occlusion by the second camera (332) with respect to the first camera (331) and / or occlusion by the first camera (331) with respect to the second camera (332).

[0096] However, the present disclosure is not limited thereto. For example, the electronic device (101) may acquire a second preview image (523a) obtained through the second camera (332) and display it on the display (320). For example, the second preview image (523a) may represent at least a portion of an external environment including an external object (507). For example, the second preview image (523a) may represent the second FoV.

[0097] For example, the electronic device (101) can obtain second depth information about an external object (507) based on the depth estimation model (350) by using data about the second preview image (523). For example, the electronic device (101) can obtain the second depth information about the second preview image (523) based on the depth estimation model (350) by using the second preview image (523). Alternatively, for example, the electronic device (101) can obtain the second depth information about the second preview image (523a) based on the depth estimation model (350) by using the second preview image (523a).

[0098] For example, the depth estimation model (350) used to obtain the second depth information may be in a learned state. For example, the depth estimation model (350) may be learned to estimate depth information for the converted image by using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera (331). For example, the depth information for the converted image may represent the second FoV. For example, the electronic device (101) may generate the converted image based on the second FoV by using the image representing the first FoV acquired through the first camera (331). For example, the electronic device (101) may generate the converted image based on the second FoV by using depth information corresponding to the image and representing the first FoV. For example, the converted image and the converted depth information may be used to learn the depth estimation model (350). For example, the above-described converted image and the above-described converted depth information may be referred to as learning data. Alternatively, in one example, the learning data may include an image representing the second FoV acquired through the second camera (332) instead of the above-described converted image, and the above-described converted depth information.

[0099] For example, the electronic device (101) can obtain second depth information for the second preview image (523) (or the second preview image (523a)) by inputting data about the second preview image (523) (or the second preview image (523a)) into the depth estimation model (350). For example, the electronic device (101) can obtain second distance information for the ROI (527) of the second preview image (523) (or the ROI (527a) of the second preview image (523a)) using the second depth information. For example, the ROI (527) can include an area within the second preview image (523) corresponding to an external object (507) or a visual object. For example, the ROI (527a) may include an area within the second preview image (523a) corresponding to an external object (507) or a visual object.

[0100] For example, the second depth information may indicate a second distance (521) from the electronic device (101) (or the second camera (332)) to the external object (507). For example, the second depth information may include the second distance (521) or a phase difference (or phase value) corresponding to the second distance (521). For example, the electronic device (101) may obtain the second distance information indicating the second distance (521) related to the ROI (527) (or the ROI (527a)) by using the second depth information obtained based on the depth estimation model (350).

[0101] In one example, the electronic device (101) may stop acquiring images through the first camera (331) to reduce power consumption while acquiring a second preview image (523) (or a second preview image (523a)) through the second camera (332). Accordingly, the electronic device (101) may stop acquiring depth information to an external object (507) through the first camera (331).

[0102] For example, the electronic device (101) may compare the second depth information with the reference value. For example, the reference value may represent a reference depth (or reference distance) for switching the capturing camera for obtaining a preview image from the second camera (332) to the first camera (331). In the following, for convenience of explanation, the reference value for switching from the first camera (331) to the second camera (332) and the reference value for switching from the second camera (332) to the first camera (331) are illustrated as being the same, but the present disclosure is not limited thereto. For example, the first reference value for switching from the first camera (331) to the second camera (332) may be different from the second reference value for switching from the second camera (332) to the first camera (331). In this case, the second reference value may be a value within a reference range from the first reference value. At this time, the reference range may be determined based on at least one of the properties of the first camera (331) and the second camera (332). In addition, in the present disclosure, if the camera assembly (330) further includes an additional camera in addition to the first camera (331) and the second camera (332), an additional reference value (or a third reference value) may be further utilized.

[0103] For example, the electronic device (101) can determine (or identify) whether the second depth information exceeds the reference value. For example, if the second depth information exceeds the reference value, the electronic device (101) can switch the photographing camera from the second camera (332) to the first camera (331). Alternatively, if the second depth information is less than or equal to the reference value, the electronic device (101) can identify (or maintain) the photographing camera as the second camera (332). In example (502), the electronic device (101) can acquire the second depth information while the user (505) moves away from the external object (507) in accordance with movement (529) in a second direction opposite to the first direction. Accordingly, the electronic device (101) can identify the photographing camera by comparing the acquired second depth information with the reference value. For example, the second depth information may be acquired periodically or aperiodically.

[0104] In FIG. 5B, a case is exemplified where the depth estimation model (350) is trained using data regarding the second preview image (523) (or the second preview image (523a)), but the present disclosure is not limited thereto. In the above example, the preview image is an image displayed through the display (320) of the electronic device (101) before capturing an image, and the resolution of the preview image may be determined according to the resolution of the display (320). Depending on the input for capturing, the image stored in the electronic device (101) may have a resolution set for the electronic device (101). The data regarding the second preview image (523) (or the second preview image (523a)) may include an image stored according to the input for capturing, or an image having a predetermined resolution. For example, the depth estimation model (350) may estimate depth information (or distance information) using a stored image or an image with a predetermined resolution according to the input for shooting.

[0105] As the second depth information changes to exceed the reference value according to movement (529) in example (502), the method of the electronic device (101) acquiring an image through the first camera (331) may refer to example (503).

[0106] Referring to example (503) of FIG. 5c, the electronic device (101) can obtain an image of an external object (507) through the first camera (331). In example (503), the photographing camera for obtaining a preview image may be the first camera (331). For example, the electronic device (101) can obtain a third preview image (533) obtained through the first camera (331) and display it on the display (320). For example, the third preview image (533) can represent at least a portion of an external environment including the external object (507). For example, the third preview image (533) can represent the first FoV (e.g., the first FoV (415) of FIG. 4a).

[0107] For example, the electronic device (101) may obtain third depth information about an external object (507) based on the first camera (331). For example, the electronic device (101) may obtain the third depth information about an ROI (537) within a third preview image (533). For example, the ROI (537) may include an area within the third preview image (533) corresponding to the external object (507) or a visual object.

[0108] For example, the third depth information may indicate a third distance (531) from the electronic device (101) (or the first camera (331)) to the external object (507). For example, the third depth information may include the third distance (531) or a phase difference (or phase value) corresponding to the third distance (531). For example, the electronic device (101) may obtain the third depth information related to the ROI (537) by calculating the phase difference for each pixel of the ROI (537) of the third preview image (533).

[0109] For example, the electronic device (101) may compare the third depth information with a reference value. For example, the reference value may represent a reference depth (or reference distance) for switching the capturing camera for obtaining a preview image from the first camera (331) to the second camera (332). For example, the electronic device (101) may determine (or identify) whether the third depth information exceeds the reference value. For example, if the third depth information exceeds the reference value, the electronic device (101) may identify (or maintain) the capturing camera as the first camera (331). Alternatively, if the third depth information is less than or equal to the reference value, the electronic device (101) may switch the capturing camera from the first camera (331) to the second camera (332). For example, the third depth information may be obtained periodically or aperiodically.

[0110] Referring to FIGS. 5A to 5C, the electronic device (101) according to the present disclosure can acquire an image of an external object through the first camera (331) and acquire depth information based on the first camera (331) when the electronic device (101) is positioned outside the reference distance. Since the accuracy of the depth information acquired based on the first camera (331) may be relatively high, the electronic device (101) can acquire (or calculate, identify) the depth information by using phase detection based on partial images acquired through the first camera (331). Accordingly, when the position of the electronic device (101) moves from outside the reference distance to within the reference distance, the electronic device (101) can perform switching from the first camera (331) to the second camera (331).

[0111] In addition, the electronic device (101) according to the present disclosure can acquire an image of an external object through the second camera (332) when the electronic device (101) is positioned within the reference distance, and acquire depth information based on the depth estimation model (350) using the image acquired through the second camera (332). Since the accuracy of the depth information acquired based on the second camera (332) within the reference distance may be relatively low, the electronic device (101) can acquire (or estimate, predict, or identify) the depth information using the image acquired through the second camera (332) as an input of the depth estimation model (350). Accordingly, when the position of the electronic device (101) moves from within the reference distance to outside the reference distance, the electronic device (101) can perform switching from the second camera (332) to the first camera (331) using depth information having high accuracy.

[0112] More specific details on how the electronic device (101) acquires learning data for learning the depth estimation model (350), performs learning of the depth estimation model (350), and acquires depth information based on the depth estimation model (350) are illustrated and described with reference to FIGS. 6A and 6B.

[0113] Figure 6a illustrates an example of an operational flow for a method in which an electronic device learns a depth estimation model and acquires depth information based on the depth estimation model. Figure 6b illustrates an example of a method in which an electronic device learns a depth estimation model and acquires depth information based on the depth estimation model. Figure 6c illustrates an example of a method in which depth information is used to identify distance information.

[0114] At least some of the methods of FIG. 6A may be performed by the electronic device (101) of FIG. 3. For example, at least some of the methods may be controlled by the processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0115] Referring to FIG. 6A, in operation (600), the electronic device (101) may acquire a first image through the first camera (331). For example, the first image may represent a first FoV (e.g., the first FoV (415) of FIG. 4A).

[0116] For example, the electronic device (101) can obtain first depth information corresponding to the first image based on the first camera (331). For example, the first depth information can be obtained based on phase detection using partial images used to generate the first image. For example, the electronic device (101) can obtain the first depth information by using the phase difference between the first partial image and the second partial image obtained through PDs (e.g., PD#1 (331a) and PD#2 (331b) of FIG. 4A) within each PD set of the first camera (331). For example, the electronic device (101) can also obtain first distance information (or distance, distance value) for an ROI within the first image by using the first depth information.

[0117] In operation (610), the electronic device (101) may generate a converted image from the first image acquired through the first camera (331). For example, the electronic device (101) may generate the converted image by performing coordinate converting on the first image. For example, the converted image may represent a second FoV (e.g., the second FoV (425) of FIG. 4A). For example, the converted image may represent the second FoV. The converted image may include information within a range (or region) that matches (corresponds to) the first image representing the first FoV acquired through the first camera (331) to represent the second FoV. For example, the first image representing the first FoV may be included in the converted image by adjusting the view point of the first image and the size of contents within the first image to represent the second FoV.

[0118] For example, the electronic device (101) may generate transformed depth information using the first depth information corresponding to the first image and representing the first FoV. For example, the electronic device (101) may generate the transformed depth information by performing a coordinate system transformation on the first depth information. For example, the transformed depth information may represent the second FoV. The transformed depth information may include information within a range (or region) that matches (or corresponds to) the first depth information representing the first FoV in order to represent the second FoV. For example, the first depth information representing the first FoV may be included in the transformed depth information by adjusting the size of contents within the first depth information and a reference time point (corresponding to the capturing time point) of the first depth information to represent the second FoV.

[0119] For example, the converted image and the converted depth information can be used for learning the depth estimation model (350). For example, the converted image and the converted depth information can be referred to as learning data.

[0120] In operation (620), the electronic device (101) can perform learning of the depth estimation model (350). For example, the electronic device (101) can control (or cause) the depth estimation model (350) to learn the learning data. The depth estimation model (350) can learn the learning data.

[0121] For example, for specific details on how to generate the above learning data and perform learning of the depth estimation model (350), reference may be made to FIG. 6b.

[0122] FIG. 6B illustrates an example (650) of a method in which an electronic device (101) trains a depth estimation model (350) and estimates depth information based on the trained depth estimation model (350). The example (650) illustrates an example (651) of a method in which a depth estimation model (350) is trained and an example (652) of a method in which depth information is estimated based on the trained depth estimation model (350).

[0123] Referring to example (651), the electronic device (101) can obtain a first image (661) through the first camera (331). For example, the electronic device (101) can obtain the first image (661) by integrating partial images obtained using PDs of the first camera (331). For example, the electronic device (101) can obtain first depth information (663) based on the first camera (331). For example, the electronic device (101) can identify a phase difference for each pixel between the partial images obtained using the PDs of the first camera (331). For example, the electronic device (101) can obtain the first depth information (663) by using the phase difference for each pixel of the first image (661). For example, the visually expressed first depth information (663) may represent the depth for each pixel of the first image (661) using brightness (or color). For example, the depth may represent the depth from the first camera (331) to the external environment (or an external object in the external environment) included in the first image (661). For example, the brighter the brightness, the shallower the depth (or the shorter the distance from the first camera (331) to the external environment).

[0124] For example, the electronic device (101) can obtain a transformed image (671) using the first image (661). For example, the transformed image (671) can be obtained by performing a coordinate system transformation on the first image (661) representing the first FoV. For example, the transformed image (671) can represent the second FoV. For example, the electronic device (101) can obtain transformed depth information (673) using the first depth information (663). For example, the transformed depth information (673) can be obtained by performing a coordinate system transformation on the first depth information (663) representing the first FoV. For example, the transformed depth information (673) can represent the second FoV.

[0125] For example, coordinate transformation can be used to transform data representing a first FoV (e.g., a first image (661) and first depth information (663)) into data representing a second FoV. For example, a pixel (p) representing a specific coordinate within a 2D coordinate system for the first FoV x , p y ) can be changed to a specific coordinate (X, Y, Z) of the 3D coordinate system (or, world coordinate system). The X is, can be determined according to the following formula, wherein Y is It can be determined according to a formula such as . For example, the above f is the focal length of the camera for acquiring the data, and the above c x and c yThe data may represent the center coordinate of the 2D coordinate system, and the Z may represent the depth within the 3D coordinate system. As described above, the data representing the first FoV may be converted according to the 3D coordinate system. Thereafter, the data converted into the 3D coordinate system may be converted again into the 2D coordinate system for the second FoV. For example, specific coordinates (X, Y, Z) of the 3D coordinate system may be converted into pixels (u, v) representing specific coordinates within the 2D coordinate system for the second FoV. For example, the u may be, can be determined according to a formula such as: For example, the above v is It can be determined according to the formula as above f x The focal length on the x-axis related to the above second FoV can be represented. The above f y The focal length on the y-axis related to the above second FoV can be represented. At this time, the z uw Assuming that the data of the first FoV and the data of the second FoV have the same depth, it can have the same value as the depth (Z) within the 3D coordinate system.

[0126] For example, the electronic device (101) can obtain a transformed image (671) generated using the first image (661) and transformed depth information (673) generated using the first depth information (663) by using the coordinate system transformation as described above.

[0127] For example, the electronic device (101) can perform training of the depth estimation model (350) using training data including the converted image (671) and the converted depth information (673). For example, the depth estimation model (350) can generate (or estimate) estimated depth information using the converted image (671) as input. For example, the depth estimation model (350) can identify a difference between the estimated depth information and the converted depth information (673). For example, the depth estimation model (350) can be trained (or adjusted) so that the difference has a minimum value. For example, based on the training data, the weights of the depth estimation model (350) can be adjusted (or updated) so that the difference has a minimum value. For example, the estimated depth information can be an output of the depth estimation model (350). For example, the converted depth information (673) can be used as the correct answer (or GT) of the depth estimation model (350). In the above example, a transformed image (671) generated from a first image (661) acquired through a first camera (331) may be included in the learning data. In one example, the first image (661) of the learning data may include the last image among one or more images acquired through the first camera (331) before the point in time at which the first camera (331) is switched to the second camera (332). For specific details on a method for performing learning of a depth estimation model (350) using the learning data, reference may be made to FIG. 7 below.

[0128] As a non-limiting example, the learning data may include an image representing the second FoV acquired through the second camera (332) instead of the converted image (671) (e.g., the second image (681)), and converted depth information (673) generated from the first depth information (663). For example, the electronic device (101) may acquire an image representing the second FoV through the second camera (332) when switching from the first camera (331) to the second camera (332) based on a comparison between the first distance information generated using the first depth information (663) and the reference distance. For example, the image representing the second FoV may include a first image among one or more images acquired through the second camera (332) from the time point when switching from the first camera (331) to the second camera (332). For example, each of the one or more images may represent the second FoV. Referring to the above, the electronic device (101) can perform learning of a depth estimation model (350) using the image representing the second FoV acquired through the second camera (332) and the converted depth information.

[0129] As described above, by training the depth estimation model (350) using the transformed image based on the second FoV generated from the last image or the first image representing the second FoV, the accuracy of estimation for a distance substantially equal to (or similar to, or adjacent to) the reference distance for switching from the first camera (331) to the second camera (332) or switching from the second camera (332) to the first camera (331) can be increased, rather than increasing the accuracy of estimation for a range of distances.

[0130] Referring back to FIG. 6A, in operation (630), the electronic device (101) may acquire a second image through the second camera (332). For example, after the depth estimation model (350) learns the training data, the electronic device (101) may acquire the second image through the second camera (332). For example, the second image may represent the second FoV. In one example, the electronic device (101) may stop acquiring images through the first camera (331) to reduce power consumption while acquiring the second image through the second camera (332). Accordingly, the electronic device (101) may stop acquiring depth information through the first camera (331).

[0131] In operation (640), the electronic device (101) may obtain depth information based on the depth estimation model (350) using the second image. For example, the electronic device (101) may input the second image into the learned depth estimation model (350). For example, the electronic device (101) may obtain second depth information based on the depth estimation model (350). For example, the second depth information may represent depth information estimated based on the depth estimation model (350) using the second image. For example, the second depth information may represent the second FoV. For example, the electronic device (101) may also obtain distance information using the second depth information. For example, the distance information may be identified using the second depth information. In one example, the distance information may be distance information for an ROI of the second image. Alternatively, in one example, the distance information may be distance information for each pixel of the second image.

[0132] For example, for specific details on a method for obtaining the depth information based on a learned depth estimation model (350), reference may be made to FIG. 6b.

[0133] Referring to example (652) of FIG. 6B, the electronic device (101) can obtain a second image (681) through the second camera (332). For example, the second image (681) can represent the second FoV. For example, the electronic device (101) can input the second image (681) to the depth estimation model (350). For example, the electronic device (101) can obtain (or estimate) second depth information (683) using the second image (681) based on the depth estimation model (350). For example, the second depth information (683) can be output from the depth estimation model (350) using the input second image (681). For example, the electronic device (101) can identify distance information using the second depth information (683). For example, the distance information can be identified using the second depth information (683). For example, the brightness (or color) of the visually expressed second depth information (683) can represent the depth for each pixel of the second image (681). For example, the depth can represent the depth from the second camera (332) to the external environment (or an external object in the external environment) included in the second image (681). For example, the brighter the brightness, the shallower the depth (or the shorter the distance from the second camera (332) to the external environment). For example, the value corresponding to each pixel of the second depth information (683) can be used to identify the distance value. Specific details on a method for identifying the distance information from the depth information are exemplified and described with reference to FIG. 6C.

[0134] FIG. 6c illustrates an example (690) of how an electronic device (101) identifies the distance information using depth information (e.g., first depth information (663) or second depth information (683)).

[0135] For example, the electronic device (101) can generate first depth information (663) by using disparity between PDs of PD sets of the first camera (331). For example, the disparity can represent a phase difference between a first partial image acquired through a specific PD (e.g., PD#1 (431a)) and a second partial image acquired through another specific PD (e.g., PD#3) and corresponding to the first partial image. Alternatively, for example, the electronic device (101) can generate second depth information (683) based on a depth estimation model (350) learned by using a second image (681) acquired through the second camera (332). In the example (690), the object (691) can represent a subject that is a target to be photographed through the camera.

[0136] For example, the electronic device (101) can use the depth information to identify a depth (693) from a sensor plane of the image sensor (230) to an object (691). For example, the sensor plane may represent a virtual plane extending from the image sensor (230). As a non-limiting example, the image sensor (230) of the example (690) may include the center of the image sensor (230). For example, the depth (693) to the object (691) may be defined from a focal position of a lens included in the image sensor (230), a camera, or any location inside the electronic device (101).

[0137] For example, the electronic device (101) can identify an angle between a virtual line representing a distance (695) between an object (691) and an image sensor (230) and a virtual reference line extending from the image sensor (230) (or the center of the image sensor (230)). For example, the electronic device (101) can identify a distance (695) using the angle and the depth (693) to the object (691). For example, the identified distance (695) can be included in the distance information.

[0138] Although not illustrated in FIG. 6A, the electronic device (101) may use the acquired distance information to compare it with a reference distance. For example, the reference distance may include a reference value for switching the photographing camera from the first camera (331) to the second camera (332) or from the second camera (332) to the first camera (331).

[0139] FIG. 7 illustrates an example of an operational flow for a method in which an electronic device learns a depth estimation model using a transformed image based on a second FoV generated from an image representing a first FoV (field of view).

[0140] At least some of the methods of FIG. 7 may be performed by the electronic device (101) of FIG. 3. For example, at least some of the methods may be controlled by the processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0141] In FIG. 7, for convenience of explanation, it is assumed that the electronic device (101) switches the capturing camera to be used to acquire an image from the first camera (331) to the second camera (332). For example, the electronic device (101) may switch the capturing camera from the first camera (331) to the second camera (332) as it moves closer to an external object in the external environment corresponding to the ROI in the image.

[0142] Referring to FIG. 7, in operation (700), the electronic device (101) may acquire one or more images representing a first FoV. For example, the electronic device (101) may acquire the one or more images through the first camera (331). For example, each of the one or more images may represent the first FoV.

[0143] For example, the one or more images may be acquired during a time period extending from a time point at which the photographing camera of the electronic device (101) is switched from the first camera (331) to the second camera (332). For example, the electronic device (101), after acquiring the one or more images through the first camera (331), may switch the photographing camera from the first camera (331) to the second camera (332) at the time point. For example, the time period may represent a time between the time point and a time point in the past. The electronic device (101), after switching the photographing camera from the first camera (331) to the second camera (332) at the time point, may acquire one or more other images through the second camera (332) during another time period extending from the time point. For example, the other time period may represent a time between the time point and a time point in the future. For example, each of the one or more other images may represent a second FoV. In the above example, based on the point in time, the time period may represent a time period before the point in time, and the other time period may represent a time period after the point in time.

[0144] In operation (710), the electronic device (101) may identify a last image among the one or more images. For example, the electronic device (101) may identify the last image that was most recently acquired among the one or more images acquired through the first camera (331). For example, the last image may be an image acquired last through the first camera (331) within the time period. In other words, the last image may be an image acquired immediately before switching from the first camera (331) to the second camera (332).

[0145] In operation (720), the electronic device (101) may identify depth information for the final image. For example, the depth information for the final image may represent the first FoV. For example, the depth information may be identified based on partial images used to acquire the final image.

[0146] In operation (730), the electronic device (101) may generate and store a transformation image and transformation depth information. For example, the electronic device (101) may generate the transformation image using the final image. For example, the transformation image may be generated by performing a coordinate system transformation on the final image. For example, the transformation image may represent the second FoV. For example, the electronic device (101) may generate the transformation depth information using the depth information on the final image. For example, the transformation depth information may be generated by performing a coordinate system transformation on the depth information. For example, the transformation depth information may represent the second FoV. For example, the electronic device (101) may store the generated transformation image and the transformation depth information within the electronic device (101).

[0147] In operation (740), the electronic device (101) may perform learning on a depth estimation model (350) using the converted image and the converted depth information. For example, the electronic device (101) may store the converted image and the converted depth information, and perform learning on a depth estimation model (350) using the converted image and the converted depth information.

[0148] Alternatively, for example, the electronic device (101) may perform learning on the depth estimation model (350) using the converted image and the converted depth information based on identifying an event after storing the converted image and the converted depth information. For example, the event may include that the number of stored converted images and converted depth information exceeds a specified number. Or, for example, the event may include that the remaining battery amount of the electronic device (101) exceeds a reference battery amount. For example, the event may include that the resource usage of the electronic device (101) is less than or equal to a reference usage amount. However, the present disclosure is not limited thereto.

[0149] In the above example, the case where the transformed image and the transformed depth information are used as learning data to perform learning for the depth estimation model (350) is described, but the present disclosure is not limited thereto. For example, the learning data may use the initial image and the transformed depth information among the one or more other images. For example, the electronic device (101) may identify the earliest acquired initial image among the one or more other images acquired through the second camera (332). For example, the initial image may be an image acquired for the first time through the second camera (332) within the different time period. In other words, the initial image may be an image acquired immediately after switching from the first camera (331) to the second camera (332). The initial image may represent the second FoV. For example, the learning data may include the initial image among the one or more other images and the transformed depth information.

[0150] Figure 8 illustrates an example of an operational flow for a method of performing switching between multiple cameras using a learned depth estimation model.

[0151] At least some of the methods of FIG. 8 may be performed by the electronic device (101) of FIG. 3. For example, at least some of the methods may be controlled by the processor (310) of the electronic device (101). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0152] In FIG. 8, for convenience of explanation, it is assumed that the electronic device (101) acquires an image of an external environment including an external object. In addition, for convenience of explanation, it is assumed that the electronic device (101) executes an application (hereinafter, referred to as a camera application) for acquiring an image through a camera (e.g., a first camera (331) or a second camera (332)) and displays a preview image corresponding to the acquired image on the display (320) within the UI (user interface) of the camera application. However, the present disclosure is not limited thereto. For example, the present disclosure may also be applied to a case where the electronic device (101) executes an application (or the camera application) for acquiring a video through a camera (e.g., a first camera (331) or a second camera (332)). The preview image displayed on the display (320) can be understood to be substantially identical to an image acquired by the electronic device (101) through a camera (e.g., the first camera (331) or the second camera (332)).

[0153] Referring to FIG. 8, in operation (800), the electronic device (101) may obtain first depth information for an ROI within a first preview image based on the first camera (331). For example, the electronic device (101) may obtain the first preview image through the first camera (331), which is a photographing camera, and display the first preview image on the display (320).

[0154] For example, the ROI within the first preview image may include a pre-designated object or an area including the center coordinates of the preview image. For example, the pre-designated object may include a human face within the preview image. For example, the pre-designated object may be determined according to an auto focus function. For example, when the preview image includes multiple objects, the ROI may be determined as an area including the center coordinates of the preview image. Or, for example, the ROI may be determined according to an input received with respect to the UI of the camera application.

[0155] For example, the electronic device (101) can obtain the first depth information based on the first camera (331). For example, the first depth information can be identified by calculating a phase difference for each pixel of the ROI of the first preview image. For example, the first distance information for the ROI in the first preview image can be calculated from the depth of the first depth information for the ROI. For example, the first distance information can include a first distance to the external object corresponding to the ROI.

[0156] In operation (810), the electronic device (101) can display a second preview image acquired through the second camera (332) according to the first depth information that is less than or equal to a reference value.

[0157] For example, the electronic device (101) may compare the first depth information with the reference value. For example, the reference value may represent a reference depth (or reference distance) for switching the capturing camera for obtaining a preview image from the first camera (331) to the second camera (332). For example, the electronic device (101) may determine (or identify) whether the first depth information exceeds the reference value. For example, if the first depth information exceeds the reference value, the electronic device (101) may identify (or maintain) the capturing camera as the first camera (331). Alternatively, if the first depth information is less than or equal to the reference value, the electronic device (101) may change the capturing camera from the first camera (331) to the second camera (332).

[0158] For example, the electronic device (101) can change the shooting camera from the first camera (331) to the second camera (332) according to the first depth information that is less than or equal to the reference value, and then display the second preview image acquired through the second camera (332).

[0159] In operation (820), the electronic device (101) may obtain second depth information for an ROI within the second preview image based on a depth estimation model (350) using data regarding the second preview image. As a non-limiting example, the ROI within the second preview image may be identical to the ROI within the first preview image. However, the present disclosure is not limited thereto. For example, the ROI within the second preview image may change according to a user input or a change in an external object of the external environment displayed within the second preview image.

[0160] For example, the second preview image may represent the same FoV as the first preview image (e.g., the first FoV (415) of FIG. 4A). For example, the second preview image may represent a first FoV that is the same as the first FoV of the first preview image. Accordingly, the user of the electronic device (101) may not recognize that the capturing camera has switched from the first camera (331) to the second camera (332). In the example, the electronic device (101) may modify (or crop) at least a portion of the image representing the second FoV when generating the second preview image from the image representing the second FoV obtained through the second camera (332). As described above, generating the modified second preview image may be referred to as occlusion. However, the present disclosure is not limited thereto. For example, the second preview image may represent the second FoV.

[0161] For example, the depth estimation model (350) may be trained according to the method of FIG. 7. For example, the depth estimation model (350) may be trained to estimate depth information for a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera (331). For example, the depth information for the converted image may represent the second FoV. For example, the electronic device (101) may generate the converted image based on the second FoV using the image representing the first FoV acquired through the first camera (331). For example, the electronic device (101) may generate the converted image based on the second FoV using depth information corresponding to the image and representing the first FoV. For example, the converted image and the converted depth information may be used to train the depth estimation model (350). For example, the above-described converted image and the above-described converted depth information may be referred to as learning data. Alternatively, in one example, the learning data may include an image representing the second FoV acquired through the second camera (332) instead of the above-described converted image, and the above-described converted depth information.

[0162] For example, the electronic device (101) can obtain second depth information for the ROI within the second preview image by inputting the data regarding the second preview image into the depth estimation model (350). For example, the second depth information may be an output of the depth estimation model (350). For example, the second distance information for the ROI within the second preview image may be calculated from the depth of the second depth information for the ROI. For example, the second distance information may include a second distance to an external object corresponding to the ROI within the second preview image.

[0163] In one example, the electronic device (101) may stop acquiring images through the first camera (331) to reduce power consumption while acquiring the second preview image through the second camera (332). Accordingly, the electronic device (101) may stop acquiring depth information to an external object corresponding to the ROI through the first camera (331).

[0164] For example, the electronic device (101) may compare the second depth information with the reference value. For example, the reference value may represent a reference depth (or reference distance) for switching the capturing camera for obtaining a preview image from the second camera (332) to the first camera (331). In the above example, for convenience of explanation, the reference value for switching from the first camera (331) to the second camera (332) and the reference value for switching from the second camera (332) to the first camera (331) are illustrated as being the same, but the present disclosure is not limited thereto. For example, the first reference value for switching from the first camera (331) to the second camera (332) may be different from the second reference value for switching from the second camera (332) to the first camera (331). In this case, the second reference value may be a value within a reference range from the first reference value. At this time, the reference range may be determined based on at least one of the properties of the first camera (331) and the second camera (332). In addition, in the present disclosure, if the camera assembly (330) further includes an additional camera in addition to the first camera (331) and the second camera (332), an additional reference value (or a third reference value) may be further utilized.

[0165] In operation (830), the electronic device (101) can display a third preview image acquired through the first camera (331) according to the second depth information exceeding the reference value.

[0166] For example, the electronic device (101) can determine (or identify) whether the second depth information exceeds the reference value. For example, if the second depth information exceeds the reference value, the electronic device (101) can switch the photographing camera from the second camera (332) to the first camera (331). Alternatively, if the second depth information is less than or equal to the reference value, the electronic device (101) can identify (or maintain) the photographing camera as the second camera (332).

[0167] For example, the electronic device (101) may change the capturing camera from the second camera (332) to the first camera (331) according to the second depth information exceeding the reference value, and then display the third preview image acquired through the first camera (331) on the display (320). For example, the third preview image may represent the first FoV.

[0168] Referring to FIGS. 1 to 8, the present disclosure can use an image acquired through the first camera (331) and depth information about the image immediately before switching from the first camera (331) to the second camera (332) as learning data for learning the depth estimation model (350) of the MDE. In other words, the present disclosure can perform learning of the depth estimation model (350) by using an image acquired from one camera and depth information about the image without using an additional device. The present disclosure can acquire depth information for performing switching from the second camera (332) to the first camera (331) by using the learned depth estimation model (350). Accordingly, the production cost of an electronic device (101) using the depth estimation model (350) according to the present disclosure can be reduced, the deployable space within the electronic device (101) can be expanded, and the power consumption of the electronic device (101) can be reduced. In addition, by utilizing the image and corresponding depth information immediately before switching from the first camera (331) to the second camera (332) for learning the depth estimation model (350) according to the present disclosure, the accuracy of the depth estimation model (350) can be increased. Accordingly, the usability of the user's electronic device (101) can be improved as natural switching between cameras is performed while acquiring an image of an external object (or while performing a photographing operation to acquire the image).

[0169] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0170] As described above, the electronic device (101) may include a memory (340) that stores instructions and includes one or more storage media. The electronic device (101) may include at least one processor (310) that includes a processing circuit. The electronic device (101) may include a first camera (331) that provides a first field of view (FoV). The electronic device (101) may include a second camera (332) that provides a second FoV that is wider than the first FoV. The electronic device (101) may include a display (320). The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to acquire first depth information for the first preview image based on the first camera (331) while displaying the first preview image on the display (320), obtained through the first camera (331). The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), a second preview image acquired through the second camera (332) in place of the first preview image, based on the first depth information being less than or equal to a reference value. The above instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to obtain second depth information for the second preview image based on a depth estimation model using data about the second preview image while displaying the second preview image on the display (320).The depth estimation model may be trained to estimate depth information for the converted image using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera (331). The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display a third preview image, which is changed from the second preview image and acquired through the first camera (331), on the display (320) according to the second depth information exceeding the reference value.

[0171] In one embodiment, each of the first preview image and the third preview image may represent the first FoV. The second preview image may be acquired through the second camera (332) and modified from an image representing the second FoV to represent the first FoV.

[0172] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to stop acquiring the second depth information for the second preview image through the first camera (331) while displaying the second preview image on the display (320).

[0173] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to determine whether the first depth information is greater than the reference value. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to continue displaying the first preview image on the display (320) upon determining that the first depth information is greater than the reference value. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display the second preview image, which is changed from the first preview image, on the display (320) upon determining that the first depth information is less than or equal to the reference value.

[0174] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to determine whether the second depth information is greater than the reference value. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to display, on the display (320), the third preview image modified from the second preview image upon determining that the second depth information is greater than the reference value. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to continue displaying, on the display (320), the second preview image upon determining that the second depth information is less than or equal to the reference value.

[0175] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to acquire one or more images through the first camera (331) within a time period while displaying the first preview image on the display (320). Each of the one or more images may represent the first FoV. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to, in response to displaying the second preview image modified from the first preview image, store a last image of the one or more images acquired within the time period and depth information for the last image, within the memory (340). The time period may extend from a time point at which the second preview image modified from the first preview image begins to be displayed.

[0176] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate a transformed image based on the second FoV using the final image representing the first FoV. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate transformed depth information based on the second FoV using the depth information for the final image representing the first FoV. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to perform learning for the depth estimation model using the transformed image generated using the final image and the transformed depth information.

[0177] According to one embodiment, the transformed image generated using the final image may be generated based on coordinate conversion for the final image. The transformed depth information may be generated based on the coordinate conversion for the depth information for the final image.

[0178] In one embodiment, the instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to acquire, through the second camera (332), one or more other images within a different time period from the time point at which the second preview image, which is modified from the first preview image, begins to be displayed. Each of the one or more other images may represent the second FoV. The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to store a first image of the one or more other images in the memory (340). The instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to generate transformed depth information based on the second FoV using the depth information for the last image representing the first FoV. The above instructions, when individually or collectively executed by the at least one processor (310), may cause the electronic device (101) to perform learning for the depth estimation model using the original image representing the second FoV and the transformed depth information based on the second FoV.

[0179] According to one embodiment, the first depth information within the first preview image may indicate a distance from the first camera (331) to an external object corresponding to a region of interest (ROI) within the first preview image. The first depth information for the first preview image may be obtained by utilizing a phase difference between a first partial image obtained through the first camera (331) and used to generate the first preview image and a second partial image obtained through the first camera (331) and used to generate the first preview image.

[0180] According to one embodiment, the first depth information for the first preview image may include depth information for a region of interest (ROI) within the first preview image. The second depth information for the second preview image may include depth information for an ROI within the second preview image. Each of the ROI within the first preview image and the ROI within the second preview image may include at least one of a region including a pre-specified object or the center coordinates of the preview image.

[0181] According to one embodiment, the first camera (331) can provide a first depth of field (DoF). The second camera (332) can provide a second DoF that is narrower than the first DoF.

[0182] According to one embodiment, the first camera (331) may have a first focal length. The second camera (332) may have a second focal length that is shorter than the first focal length.

[0183] The method performed by the electronic device (101) as described above may include an operation of acquiring first depth information for the first preview image based on the first camera (331) while displaying the first preview image acquired through the first camera (331) providing the first FoV (field of view) of the electronic device (101). The method may include an operation of displaying a second preview image acquired through the second camera (332) providing a second FoV higher than the first FoV of the electronic device (101) by replacing the first preview image according to the first depth information being less than or equal to a reference value. The method may include an operation of acquiring second depth information for the second preview image based on a depth estimation model by using data regarding the second preview image while displaying the second preview image. The depth estimation model may be trained to estimate depth information for the converted image using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera (331). The method may include an operation of displaying a third preview image acquired through the first camera (331) and changed from the second preview image according to the second depth information exceeding the reference value.

[0184] In one embodiment, the method may include an operation of stopping acquisition of the second depth information for the second preview image through the first camera (331) while displaying the second preview image.

[0185] In one embodiment, the method may include an operation of acquiring one or more images through the first camera (331) within a time period while displaying the first preview image. Each of the one or more images may represent the first FoV. The method may include an operation of storing a last image among the one or more images acquired within the time period and depth information for the last image in response to displaying the second preview image changed from the first preview image. The time period may extend from a time point at which the second preview image changed from the first preview image begins to be displayed.

[0186] In one embodiment, the method may include an operation of generating a transformed image based on the second FoV using the final image representing the first FoV. The method may include an operation of generating transformed depth information based on the second FoV using the depth information for the final image representing the first FoV. The method may include an operation of performing learning on the depth estimation model using the transformed image generated using the final image and the transformed depth information.

[0187] According to one embodiment, the transformed image generated using the final image may be generated based on coordinate conversion for the final image. The transformed depth information may be generated based on the coordinate conversion for the depth information for the final image.

[0188] In one embodiment, the method may include an operation of acquiring, through the second camera (332), one or more other images within a different time period from the point in time at which the second preview image changed from the first preview image starts to be displayed. Each of the one or more other images may represent the second FoV. The method may include an operation of storing a first image of the one or more other images. The method may include an operation of generating transformed depth information based on the second FoV using the depth information for the last image representing the first FoV. The method may include an operation of performing learning on the depth estimation model using the first image representing the second FoV and the transformed depth information based on the second FoV.

[0189] The non-transitory computer-readable storage medium as described above can store one or more programs including instructions that, when individually or collectively executed by at least one processor (310) of an electronic device (101) including a first camera (331) providing a first field of view (FoV), a second camera (332) providing a second FoV wider than the first FoV, and a display (320), cause the electronic device (101) to acquire first depth information within a first preview image based on the first camera (331) while displaying the first preview image acquired through the first camera (331) on the display (320). The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to display, on the display (320), a second preview image acquired through the second camera (332) in place of the first preview image, according to the first depth information that is less than or equal to a reference value. The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to acquire second depth information for the second preview image based on a depth estimation model using data about the second preview image while displaying the second preview image on the display (320).The depth estimation model may be trained to estimate depth information for the converted image using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera (331). The non-transitory computer-readable storage medium may store one or more programs including instructions that, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to display, on the display (320), a third preview image that is changed from the second preview image and acquired through the first camera (331) according to the second depth information exceeding the reference value.

[0190] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0191] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0192] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0193] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0194] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0195] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device (101), A memory (340) storing instructions and including one or more storage media; At least one processor (310) comprising a processing circuit; A first camera (331) providing a first FoV (field of view); A second camera (332) providing a second FoV wider than the first FoV; and Includes a display (320), The above instructions, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to: While displaying a first preview image acquired through the first camera (331) on the display (320), first depth information for the first preview image is acquired based on the first camera (331); According to the first depth information that is less than or equal to the reference value, the second preview image obtained through the second camera (332) is displayed on the display (320) in place of the first preview image; While displaying the second preview image on the display (320), second depth information for the second preview image is acquired based on a depth estimation model using data about the second preview image, and the depth estimation model is trained to estimate depth information for the converted image using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera (331); and According to the second depth information exceeding the reference value, a third preview image acquired through the first camera (331) is displayed on the display (320) by replacing the second preview image. Electronic device (101).

2. In claim 1, Each of the first preview image and the third preview image represents the first FoV, and The second preview image is obtained through the second camera (332) and is modified to represent the first FoV from an image representing the second FoV. Electronic device (101).

3. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to: While displaying the second preview image on the display (320), causing the acquisition of the second depth information for the second preview image through the first camera (331) to be stopped. Electronic device (101).

4. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to: Determine whether the first depth information exceeds the reference value; and Upon determining that the first depth information exceeds the reference value, maintaining the display of the first preview image on the display (320); and When the first depth information is determined to be less than or equal to the reference value, the second preview image, which is changed from the first preview image, is caused to be displayed on the display (320). Electronic device (101).

5. In claim 4, The above instructions, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to: Determine whether the second depth information exceeds the reference value; and Upon determining that the second depth information exceeds the reference value, displaying the third preview image changed from the second preview image on the display (320); and When the second depth information is determined to be less than or equal to the reference value, causing the second preview image to be displayed on the display (320). Electronic device (101).

6. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to: While displaying the first preview image on the display (320), one or more images are acquired through the first camera (331) within a time period, each of the one or more images representing the first FoV; and In response to displaying the second preview image changed from the first preview image, cause the last image among the one or more images acquired within the time period and depth information for the last image to be stored in the memory (340); The above time period extends from the time at which the second preview image changed from the first preview image begins to be displayed. Electronic device (101).

7. In claim 6, The above instructions, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to: Generating a transformed image based on the second FoV using the final image representing the first FoV; Generating transformed depth information based on the second FoV using the depth information for the final image representing the first FoV; and Using the transformed image generated using the final image and the transformed depth information, learning for the depth estimation model is performed, Electronic device (101).

8. In claim 7, The transformed image generated using the final image is generated based on coordinate converting for the final image, and The above transformation depth information is generated based on the coordinate system transformation for the depth information for the final image. Electronic device (101).

9. In claim 6, The above instructions, when individually or collectively executed by the at least one processor (310), cause the electronic device (101) to: Acquiring one or more other images through the second camera (332) within a different time period from the point in time at which the second preview image changed from the first preview image begins to be displayed, each of the one or more other images representing the second FoV; Store the first image among the one or more other images in the memory (340); Generating transformed depth information based on the second FoV using the depth information for the final image representing the first FoV; and causing learning of the depth estimation model to be performed using the first image representing the second FoV and the transformed depth information based on the second FoV, Electronic device (101).

10. In claim 1, The first depth information for the first preview image indicates a distance from the first camera (331) to an external object corresponding to a region of interest (ROI) within the first preview image, and The first depth information for the first preview image is obtained by using a phase difference between a first partial image obtained through the first camera (331) and used to generate the first preview image and a second partial image obtained through the first camera (331) and used to generate the first preview image. Electronic device (101).

11. In claim 1, The first depth information for the first preview image includes depth information for a region of interest (ROI) within the first preview image, The second depth information for the second preview image includes depth information for an ROI within the second preview image, and Each of the ROI in the first preview image and the ROI in the second preview image includes at least one of an area including the center coordinates of a pre-specified object or a preview image. Electronic device (101).

12. In claim 1, The first camera (331) provides a first DoF (depth of field), and The second camera (332) provides a second DoF narrower than the first DoF. Electronic device (101).

13. In claim 1, The above first camera (331) has a first focal length, and The second camera (332) has a second focal length shorter than the first focal length. Electronic device (101).

14. In a method performed by an electronic device (101), An operation of acquiring first depth information for the first preview image based on the first camera (331) while displaying a first preview image acquired through a first camera (331) providing a first FoV (field of view) of the electronic device (101); An operation of displaying a second preview image acquired through a second camera (332) that provides a second FoV wider than the first FoV of the electronic device (101) by replacing the first preview image according to the first depth information that is less than or equal to a reference value; An operation of obtaining second depth information for the second preview image based on a depth estimation model using data about the second preview image while displaying the second preview image, wherein the depth estimation model is learned to estimate depth information for the converted image using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera (331); and An operation of displaying a third preview image acquired through the first camera (331) by replacing the second preview image according to the second depth information exceeding the reference value, method.

15. In a non-transitory computer-readable storage medium, when individually or collectively executed by at least one processor (310) of an electronic device (101) including a first camera (331) providing a first field of view (FoV), a second camera (332) providing a second FoV wider than the first FoV, and a display (320), the electronic device (101) : While displaying a first preview image acquired through the first camera (331) on the display (320), first depth information for the first preview image is acquired based on the first camera (331); According to the first distance information that is less than or equal to the reference value, the second preview image obtained through the second camera (332) is displayed on the display (320) in place of the first preview image; While displaying the second preview image on the display (320), second depth information for the second preview image is acquired based on a depth estimation model using data about the second preview image, and the depth estimation model is trained to estimate depth information for the converted image using a converted image based on the second FoV generated from an image representing the first FoV acquired through the first camera (331); and storing one or more programs including instructions that cause a third preview image acquired through the first camera (331) to be displayed on the display (320) by replacing the second preview image according to the second depth information exceeding the reference value; Non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Fingerprint certification smart card and activation method of thereof

    KR1020210156379A

  • LED lighting apparatus

    KR102474185B1

  • Method for controlling a camera and electronic device thereof

    KR102593824B1

  • Camera switchover control techniques for multiple-camera systems

    KR102669853B1

  • Dynamic camera selection

    US20230401732A1