Method for generating images and electronic device for performing same method
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-08-13
Smart Images

Figure KR2026000525_13082026_PF_FP_ABST
Abstract
Description
Method for generating images and electronic device for performing the same
[0001] One embodiment disclosed in this document relates to a method for generating an image, an electronic device for performing the method, and a storage medium. Below, a technology for generating an image using an image object received from an external electronic device is disclosed.
[0002] Extended reality technologies, such as virtual reality, augmented reality, and mixed reality, which utilize computer graphics technology, are being developed. Virtual reality technology can make a virtual space constructed by a computer that does not exist in the real world feel like reality. Augmented reality or mixed reality technology can achieve the integration of the real world and the virtual world and enable real-time interaction with the user by overlaying computer-generated information onto the real world.
[0003] According to the VST (video see-through) method, after an image of the physical environment is captured using a camera, the captured image can be displayed on a display screen. At this time, digital information can be provided superimposed on the captured image. Meanwhile, according to the OST (optical see-through) method, while the user directly views the physical environment using a transparent display or lens, digital information can be provided superimposed on the physical environment.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] According to one embodiment, the electronic device may include a display. The electronic device may include at least one processor including processing circuitry. The electronic device may include a memory including one or more storage media for storing instructions. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of receiving a first image from an external electronic device that has established a communication connection with the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of identifying objects of the first image and objects of a second image acquired using the camera of the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of determining a first tag for the objects of the first image and a second tag for the objects of the second image. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be enabled to perform an operation of determining the object type of the objects of the first image based on at least the first tag. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be enabled to perform an operation of determining the second object of the second image associated with the first object corresponding to the first object type among the objects of the first image based on at least the second tag.When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of generating a composite image in which the second object in the second image is replaced by the first object. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of displaying the composite image through the display.
[0006] According to one embodiment, the electronic device may include a display. The electronic device may include at least one processor including processing circuitry. The electronic device may include a memory including one or more storage media for storing instructions. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of receiving a first image and object region information designated for the first image from an external electronic device that has established a communication connection with the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of identifying objects of the first image and objects of a second image acquired using the camera of the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of determining a first tag for objects of the first image and a second tag for objects of the second image. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to perform an operation of determining the object type of the objects in the first image based on the first tag and the object region information of the first image. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to perform an operation of determining the second object of the second image associated with the first object corresponding to the first object type among the objects of the first image based on at least the second tag.When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of generating a composite image in which the second object in the second image is replaced by the first object. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of displaying the composite image through the display.
[0007] A method performed by an electronic device according to one embodiment may include receiving a first image from an external electronic device that has established a communication connection with the electronic device. The method may include identifying objects of the first image and objects of a second image acquired using a camera of the electronic device. The method may include determining a first tag for objects of the first image and a second tag for objects of the second image. The method may include determining the object type of the objects of the first image based at least on the first tag. The method may include determining a second object of the second image associated with a first object corresponding to the first object type among the objects of the first image based at least on the second tag. The method may include generating a composite image in which the second object in the second image is replaced by the first object. The method may include displaying the composite image through a display of the electronic device.
[0008] A method performed by an electronic device according to one embodiment may include receiving a first image and object region information designated for the first image from an external electronic device that has established a communication connection with the electronic device. The method may include identifying objects of the first image and objects of a second image acquired using a camera of the electronic device. The method may include determining a first tag for objects of the first image and a second tag for objects of the second image. The method may include determining an object type of objects of the first image based on the first tag and the object region information of the first image. The method may include determining a second object of the second image associated with a first object corresponding to a first object type among the objects of the first image based on at least the second tag. The method may include generating a composite image in which the second object in the second image is replaced by the first object. The method may include displaying the composite image through a display of the electronic device.
[0009] According to one embodiment, a non-transient computer-readable recording medium may store one or more programs including instructions. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be enabled to perform the operation of receiving a first image from an external electronic device that has established a communication connection with the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be enabled to perform the operation of identifying objects of the first image and objects of a second image acquired using the camera of the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be enabled to perform the operation of determining a first tag for objects of the first image and a second tag for objects of the second image. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be enabled to perform the operation of determining the object type of objects of the first image based on at least the first tag. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to perform an operation of determining a second object in the second image associated with a first object corresponding to a first object type among the objects of the first image based on at least the second tag. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to perform an operation of generating a composite image in which the second object in the second image is replaced by the first object.When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of displaying the synthetic image through the display.
[0010] According to one embodiment, a non-transient computer-readable recording medium may store one or more programs including instructions. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to perform the operation of receiving a first image and object region information designated for the first image from an external electronic device that has established a communication connection with the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to perform the operation of identifying objects of the first image and objects of a second image acquired using the camera of the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to perform the operation of determining a first tag for objects of the first image and a second tag for objects of the second image. When the above commands are executed individually or collectively by the at least one processor, the electronic device may be configured to perform an operation of determining the object types of the objects in the first image based on the first tag and the object region information of the first image. When the above commands are executed individually or collectively by the at least one processor, the electronic device may be configured to perform an operation of determining the second object of the second image associated with the first object corresponding to the first object type among the objects of the first image based on at least the second tag. When the above commands are executed individually or collectively by the at least one processor, the electronic device may be configured to perform an operation of generating a composite image in which the second object in the second image is replaced by the first object.When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be made to perform the operation of displaying the synthetic image through the display.
[0011] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0012] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.
[0013] FIG. 2 illustrates examples of optical see-through devices according to various embodiments.
[0014] FIGS. 3a and 3b are drawings showing examples of the front and rear of an electronic device according to various embodiments.
[0015] FIG. 4 is a drawing illustrating an artificial intelligence system according to one embodiment.
[0016] FIG. 5 is a schematic diagram of an imaging system according to one embodiment.
[0017] FIG. 6 is a flowchart of a method for generating an image according to one embodiment.
[0018] FIG. 7 is a flowchart of a method for determining an object type according to one embodiment.
[0019] Figure 8 is a diagram illustrating a method for determining an object type according to one example.
[0020] FIG. 9 is a flowchart of a method for determining an object type according to one embodiment.
[0021] FIG. 10 is a flowchart of a method for determining an object type according to one embodiment.
[0022] FIG. 11 is a diagram illustrating a method for determining an object type according to one example.
[0023] FIG. 12 is a flowchart of a method for matching objects between images according to one embodiment.
[0024] FIG. 13 is a flowchart of a method for matching objects between images according to one embodiment.
[0025] FIG. 14 is a flowchart of a method for matching objects between images according to one embodiment.
[0026] FIG. 15 is a diagram illustrating a method of matching objects between images according to one example.
[0027] FIG. 16 is a flowchart of a method for generating an image according to one embodiment.
[0028] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0029] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.
[0030] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0031] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0032] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence is performed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0033] The number of processors (120) may be one or more. For example, the processor (120) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.
[0034] The processor (120) can control the operations of the electronic device (101) by executing instructions stored in memory (130). For example, the processor (120) may correspond to a plurality of processors that divide and collectively perform a plurality of operations among the processors.
[0035] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, software (e.g., program (140)) and input data or output data for related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0036] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0037] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0038] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0039] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0040] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0041] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0042] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0043] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0044] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0045] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0046] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0047] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0048] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0049] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0050] An antenna module (197) can transmit a signal or power to an external source (e.g., an external electronic device) or receive it from an external source. According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0051] According to one embodiment, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0052] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0053] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0054] FIG. 2 illustrates examples of optical see-through devices according to various embodiments.
[0055] An electronic device (201) (e.g., the electronic device (101) of FIG. 1) may include at least one of a display (e.g., the display module (160) of FIG. 1), a vision sensor, a light source (230a, 230b), an optical element, or a substrate. According to one embodiment, the display of the electronic device (201) is transparent and may provide an image through the transparent display. According to one embodiment, the electronic device (201) may include a transparent member and a display connected to the transparent member. A user may look at an object placed in physical space (e.g., an object in the real world) through the transparent member. An electronic device (201) that allows light reflected from an object placed in physical space to pass through a transparent configuration (e.g., a transparent display, a transparent member separate from the display) and provides an image through the display may be referred to as an optical see-through device (OST device).
[0056] The display may include, for example, a liquid crystal display (LCD), a digital mirror device (DMD), a liquid crystal on silicon (LCoS), an organic light emitting diode (OLED), or a micro light emitting diode (micro LED).
[0057] In one embodiment, where the display is composed of a liquid crystal display, a digital mirror display, or a silicon liquid crystal display, the electronic device (201) may include a light source (230a, 230b) that irradiates light onto a screen output area of the display (e.g., a screen display portion (215a, 215b)). In another embodiment, where the display can generate light on its own, for example, where it is composed of an organic light-emitting diode or a micro LED, the electronic device (201) may provide a virtual image of good quality to the user without including a separate light source (230a, 230b). In one embodiment, if the display is implemented as an organic light-emitting diode or a micro LED, the light source (230a, 230b) is unnecessary, so the electronic device (201) can be made lighter.
[0058] Referring to FIG. 2, the electronic device (201) may include a display, a first transparent member (225a) and / or a second transparent member (225b), and the user may use the electronic device (201) while wearing it on their face. The first transparent member (225a) and / or the second transparent member (225b) may be formed from a glass plate, a plastic plate, or a polymer, and may be made transparent or translucent. According to one embodiment, the first transparent member (225a) may be positioned facing the user's right eye, and the second transparent member (225b) may be positioned facing the user's left eye. The display may include a first display (205) that outputs a first image (e.g., right image) corresponding to the first transparent member (225a) and a second display (210) that outputs a second image (e.g., left image) corresponding to the second transparent member (225b). According to one embodiment, when each display is transparent, each display and the transparent member may be positioned to face the user's eyes to form a screen display unit (215a, 215b).
[0059] In one embodiment, light emitted from a display (205, 210) may be guided along a light path to a waveguide through an input optical member (220a, 220b). Light traveling within the waveguide may be guided toward the user's eye through an output optical member (e.g., an output grating region). Screen display units (215a, 215b) may be determined based on the light emitted toward the user's eye.
[0060] For example, light emitted from the display (205, 210) can be reflected by the grating region of the waveguide formed in the input optical member (220a, 220b) and the screen display portions (215a, 215b) and transmitted to the user's eye.
[0061] The optical element may include at least one of a lens or an optical waveguide.
[0062] The lens can adjust the focus so that the screen output to the display can be seen by the user's eyes. The lens may include, for example, at least one of a Fresnel lens, a pancake lens, or a multichannel lens.
[0063] An optical waveguide can transmit an image ray generated from a display to the user's eye. For example, the image ray may represent a ray of light emitted by a light source (230a, 230b) that has passed through the screen output area of the display. The optical waveguide may be made of glass, plastic, or polymer. The optical waveguide may include a nano-pattern formed on some internal or external surface, for example, a polygonal or curved grating structure.
[0064] The vision sensor may include at least one of a camera sensor or a depth sensor.
[0065] The first camera (265a, 265b) is a recognition camera and may be used for 3DoF and 6DoF head tracking, hand detection, hand tracking, and spatial recognition. The first camera (265a, 265b) may primarily include a GS (global shutter) camera. Since stereo cameras are required for head tracking and spatial recognition, the first camera (265a, 265b) may include two or more GS cameras. A GS camera may have superior performance compared to a RS (rolling shutter) camera in terms of detecting fast hand movements and fine movements such as fingers, and tracking movements. For example, a GS camera may have low image blur. The first camera (265a, 265b) can capture image data used for 6DoF spatial recognition and SLAM functions through depth capture. In addition, a user gesture recognition function can be performed based on image data captured by the first camera (265a, 265b).
[0066] The second camera (270a, 270b) is an ET (eye tracking) camera and can be used to capture image data for detecting and tracking the user's pupils. The second camera (270a, 270b) can track the user's eyes, that is, the user's gaze, using light (e.g., infrared light) output from a display. The second camera (270a, 270b) may be an eye-tracking camera that collects information to position the center of a virtual image projected onto the electronic device (201) according to the direction in which the wearer's pupils of the electronic device (201) gaze. The second camera (270a, 270b) may also include a GS camera to detect the pupils and track rapid pupil movements. The ET camera may also be installed for the left eye and the right eye, respectively, and the same camera performance and specifications may be used for each. The second camera (270a, 270b) may include a gaze tracking sensor. The eye tracking sensor may be included inside the second camera (270a, 270b). Infrared light output from the display (205, 210) may be transmitted to the user's eye as infrared reflected light by a half mirror. The eye tracking sensor may detect infrared transmitted light reflected from the user's eye. The second camera (270a, 270b) may track the user's eye, that is, the user's gaze, based on the detection result of the eye tracking sensor.
[0067] The third camera (245) may be a camera for shooting. The third camera (245) may include a high-resolution camera for capturing images of HR (high resolution) or PV (photo video). The third camera (245) may include a color camera equipped with functions for obtaining high-quality images, such as AF function and optical image stabilization (OIS). The third camera (245) may be a GS camera or an RS camera.
[0068] The fourth camera (e.g., the face recognition camera (325, 326) of FIGS. 3a and 3b below) is a face recognition camera, and the FT (face tracking) camera can be used to detect and track the user's facial expressions.
[0069] A depth sensor (not shown) may represent a sensor that senses information for determining the distance to an object, such as Time of Flight (TOF). TOF is a technology that measures the distance to an object using a signal (e.g., near-infrared, ultrasound, or laser). A depth sensor based on TOF technology emits a signal from a transmitter and measures the signal at a receiver, and can measure the flight time of the signal.
[0070] A light source (230a, 230b) (e.g., an illumination module) may include a device (e.g., a light emitting diode) that emits light of various wavelengths. The illumination module may be attached in various locations depending on the application. In one use case, a first illumination module (e.g., an LED device) attached around the frame of an augmented reality glasses device may emit light to assist in gaze detection when tracking eye movements with an ET camera. The first illumination module may, for example, include an IR LED of infrared wavelength. In another use case, a second illumination module (e.g., an LED device) may be attached adjacent to a camera mounted around a hinge (240a, 240b) connecting the frame and the temple, or around a bridge connecting the frame. The second illumination module may emit light to supplement ambient brightness when the camera is taking pictures. If subject detection is not easy in a dark environment, the second illumination module may emit light.
[0071] A substrate (235a, 235b) (e.g., a printed circuit board (PCB)) can support the aforementioned components.
[0072] A printed circuit board (PCB) may be placed on the temple of the glasses. The FPCB may transmit electrical signals to each module (e.g., camera, display, audio module, sensor) and other printed circuit boards. According to one embodiment, at least one printed circuit board may be in the form of a first board, a second board, and an interposer disposed between the first board and the second board. Electrical signals may be transmitted to each module and other printed circuit boards.
[0073] Other components may include at least one of a plurality of microphones (e.g., a first microphone (250a), a second microphone (250b), a third microphone (250c)), a plurality of speakers (e.g., a first speaker (255a), a second speaker (255b)), a battery (260), an antenna, or a sensor (e.g., an accelerometer, a gyroscope, or a touch sensor).
[0074] FIGS. 3a and 3b are drawings showing examples of the front and rear of an electronic device according to various embodiments.
[0075] FIG. 3a is an external view of the electronic device (301) viewed from a first direction (①), and FIG. 3b is an external view of the electronic device (301) viewed from a second direction (②). When a user wears the electronic device (301), the external view seen by the user's eyes may be FIG. 3b.
[0076] Referring to FIG. 3a, according to various embodiments, an electronic device (301) (e.g., the electronic device (101) of FIG. 1 or the electronic device (201) of FIG. 2) may provide a service that provides an extended reality (XR) experience to a user. For example, XR or XR service may be defined as a service that collectively refers to virtual reality (VR), augmented reality (AR), and / or mixed reality (MR).
[0077] According to one embodiment, the electronic device (301) may refer to a head-mounted device or a head-mounted display worn on the head of a user, but may be configured in the form of at least one of glasses, goggles, a helmet, or a hat. The electronic device (301) may include an OST (optical see-through) type configured to allow external light to reach the user's eyes through the glass when worn, or a VST (video see-through) type configured to block external light so that light emitted from the display reaches the user's eyes when worn, but external light does not reach the user's eyes.
[0078] According to one embodiment, an electronic device (301) may be worn on the head of a user to provide the user with images related to an extended reality (XR) service. For example, the electronic device (301) may provide XR content (hereinafter referred to as XR content images) that outputs at least one virtual object superimposed on an area determined to be a display area or the user's field of view (FoV). According to one embodiment, XR content may refer to images related to real space acquired through a camera (e.g., a camera for shooting) or images or videos that appear to have at least one virtual object superimposed on a virtual space. According to one embodiment, the electronic device (301) may provide XR content based on a function being performed on the electronic device (301) and / or a function being performed on one or more external electronic devices (e.g., the electronic devices (102, 104) of FIG. 1, the server (108) of FIG. 1).
[0079] According to one embodiment, the electronic device (301) is at least partially controlled by an external electronic device (e.g., the electronic device (102, 104) of FIG. 1), and at least one function may be performed under the control of the external electronic device, but at least one function may also be performed independently.
[0080] Referring to FIG. 3a, a vision sensor may be disposed on a first surface of the housing of the main body (310) of the electronic device (301). The vision sensor may include cameras (e.g., cameras for a second function (311, 312), cameras for a first function (315)) and / or a depth sensor (317) for acquiring information related to the surrounding environment of the electronic device (301).
[0081] In one embodiment, the second function cameras (311, 312) can acquire images related to the surrounding environment of the electronic device (301). The first function cameras (315) can acquire images while the wearable electronic device is worn by a user. The first function cameras (315) can be used for hand detection, tracking, and user gesture (e.g., hand movements) recognition. The first function cameras (315) can be used for 3DoF, 6DoF head tracking, location (space, environment) recognition, and / or movement recognition. In one embodiment, the second function cameras (311, 312) may be used for hand detection and tracking, and user gestures.
[0082] In one embodiment, the depth sensor (317) may be configured to transmit a signal and receive a signal reflected from the subject, and may be used for determining the distance to the object, such as time of flight (TOF). Instead of or in addition to the depth sensor (317), cameras (311, 312, 315, 316) may determine the distance to the object.
[0083] Referring to FIG. 3b, a face recognition camera (325, 326) and / or a display (321) (and / or a lens) may be disposed on the second surface (320) of the main body (310) housing.
[0084] In one embodiment, a face recognition camera (325, 326) adjacent to the display may be used to recognize the user's face or to recognize and / or track both of the user's eyes.
[0085] In one embodiment, the display (321) (and / or lens) may be disposed on a second surface (320) of the electronic device (301). In one embodiment, the electronic device (301) may not include some of the plurality of cameras (315). Although not illustrated in FIGS. 3a and 3b, the electronic device (301) may further include at least one of the configurations illustrated in FIG. 2.
[0086] According to one embodiment, the electronic device (301) may include a main body (310) that implements at least some of the components of FIG. 1, a display (321) (e.g., the display module (160) of FIG. 1) disposed in a first direction (①) of the main body (310), a first function camera (e.g., a recognition camera) (315) disposed in a second direction (②) of the main body (310), a second function camera (e.g., a shooting camera) (311, 312) disposed in a second direction (②), a third function camera (e.g., a gaze tracking camera) (328) disposed in a first direction (①), a fourth function camera (e.g., a face recognition camera) (325, 326) disposed in a first direction (①), a depth sensor (317) disposed in a second direction (②), and a touch sensor (313) disposed in a second direction (②). Although not shown in the drawing, the main body (310) may include a memory (e.g., memory (130) of FIG. 1) and a processor (e.g., processor (120) of FIG. 1) inside, and may further include other components shown in FIG. 1.
[0087] According to one embodiment, the display (321) may include a liquid crystal display (LCD), a digital mirror device (DMD), a liquid crystal on silicon (LCoS), an organic light emitting diode (OLED), or a micro light emitting diode (micro LED).
[0088] In one embodiment, if the display (321) is one of a liquid crystal display, a digital mirror display, or a silicon liquid crystal display, the electronic device (301) may include a light source that irradiates light onto the screen output area of the display (321). In another embodiment, if the display (321) can generate light itself, for example, if the electronic device (301) is one of an organic light-emitting diode or a micro LED, the electronic device (301) can provide a user with high-quality XR content images without including a separate light source. In one embodiment, if the display (321) is implemented as an organic light-emitting diode or a micro LED, a light source is unnecessary, so the electronic device (301) can be made lighter.
[0089] According to one embodiment, the electronic device (301) may include a plurality of cameras. For example, the cameras may include a first functional camera (e.g., a recognition camera) (315) positioned in the second direction (②) of the main body (310), a second functional camera (e.g., a shooting camera) (311, 312) positioned in the second direction (②), a third functional camera (e.g., a gaze tracking camera) (328) positioned in the first direction (①) and / or a fourth functional camera (e.g., a face recognition camera) (325, 326) positioned in the first direction (①), but may further include cameras of other functions not illustrated.
[0090] The first functional camera (e.g., recognition camera) (315) may be used for detecting user movement or user gesture recognition functions. The first functional camera (315) may support at least one of head tracking, hand detection and hand tracking, and spatial recognition. For example, the first functional camera (315) may primarily use a GS (global shutter) camera, which has superior performance compared to an RS (rolling shutter) camera, to detect hand movements and fine finger movements and to track movements, and may be composed of a stereo camera including two or more GS cameras for head tracking and spatial recognition. The first functional camera (315) may perform SLAM (simultaneous localization and mapping) functions to recognize information related to the surrounding space (e.g., location and / or orientation) through spatial recognition for 6DoF and depth capture.
[0091] A second-function camera (e.g., a camera for shooting) (311, 312) can be used to capture the outside and generate an image or video corresponding to the outside and transmit it to a processor (e.g., the processor (120) of FIG. 1). The processor can display the image received from the second-function camera (311, 312) on a display (321). The second-function camera (311, 312) may be referred to as HR (high resolution) or PV (photo video) and may include a high-resolution camera. For example, the second-function camera (311, 312) may include a color camera equipped with functions for obtaining high-quality images, such as AF (auto focus) and shake correction (OIS (optical image stabilizer)), but is not limited thereto, and the second-function camera (311, 312) may also include a GS camera or an RS camera.
[0092] A third-function camera (e.g., eye-tracking camera) (328) may be placed in the display (321) (or inside the main body) such that the camera lens faces the user's eyes when the user is equipped with the electronic device (301). The third-function camera (328) may be used for detecting and tracking the pupils (ET: eye tracking). A processor may determine the direction of gaze by tracking the movement of the user's left and right eyes in the image received from the third-function camera (328). By tracking the position of the pupils in the image, the processor may position the center of the XR content image displayed in the display area according to the direction the pupils are gazing. As an example, a GS camera may be used for the third-function camera (328) to detect the pupils and track pupil movements. The third-function camera (328) may be installed for the left and right eyes respectively, and the same camera performance and specifications may be used for each.
[0093] The fourth functional camera (e.g., face recognition camera) (325, 326) can be used to detect and track (FT: face tracking) the user's facial expression when the user is wearing the electronic device (301).
[0094] According to one embodiment, the electronic device (301) may include a lighting unit (e.g., LED) (not shown) as an auxiliary means for the cameras. For example, the third function camera (325) may use lighting included in the display so that emitted light (e.g., IR LED of infrared wavelength) is directed toward both eyes of the user as an auxiliary means to facilitate gaze detection when tracking eye movements. As another example, the second function cameras (311, 312) may further include a lighting unit (e.g., flash) as an auxiliary means to supplement ambient brightness when shooting outdoors.
[0095] According to one embodiment, a depth sensor (or depth camera) (317) may be used for determining the distance to an object (e.g., object) such as time of flight (TOF). Time of flight (TOF) is a technique for measuring the distance to an object using a signal (e.g., near-infrared, ultrasound, or laser), in which a transmitter transmits a signal, a receiver measures the signal, and the distance to the object can be measured based on the flight time of the signal.
[0096] According to one embodiment, the touch sensor (313) may be positioned in the second direction (②) of the main body (310). For example, when a user wears the electronic device (301), the user's eyes may look toward the first direction (①) of the main body. The touch sensor (313) may be implemented as a single type or a type separated into left and right sides depending on the shape of the main body (310), but is not limited thereto. For example, if the touch sensor (313) is implemented as a type separated into left and right sides as shown in FIG. 3a, when a user wears the electronic device (301), the first touch sensor (313a) may be positioned at the user's left eye position as in the fourth direction (④), and the second touch sensor (313b) may be positioned at the user's right eye position as in the third direction (③).
[0097] The touch sensor (313) can recognize touch input in at least one of, for example, capacitive, pressure-sensitive, infrared, or ultrasonic methods. For example, the capacitive touch sensor (313) may be capable of recognizing physical touch (or contact) input or hovering input (or proximity) of an external object. According to some embodiments, the electronic device (301) may use a proximity sensor (not shown) to enable proximity recognition of an external object.
[0098] According to one embodiment, the touch sensor (313) has a two-dimensional surface and can transmit touch data (e.g., touch coordinates) of an external object (e.g., user finger) that contacts the touch sensor (313) to a processor (e.g., processor (120) of FIG. 1). The touch sensor (313) can detect a hovering input for an external object (e.g., user finger) that approaches within a first distance from the touch sensor (313), or detect a touch input that touches the touch sensor (313).
[0099] According to one embodiment, when an external object touches the touch sensor (313), the touch sensor (313) may provide two-dimensional information about the contact point to the processor (120) as "touch data." The touch data may be described as "touch mode." When an external object is located within a first distance from the touch sensor (313) (or is in close proximity, hovering above the touch sensor), the touch sensor (313) may provide hovering data to the processor (120) regarding the time or location of hovering around the touch sensor (313). The hovering data may be described as "hovering mode / proximity mode."
[0100] According to one embodiment, the electronic device (301) can acquire hovering data using at least one of the touch sensor (313), a proximity sensor (not shown) and / or a depth sensor (317) to generate information regarding the distance, location, or time between the touch sensor (313) and an external object.
[0101] According to one embodiment, the interior of the main body (310) may include a processor (e.g., the processor (120) of FIG. 1) and a memory (e.g., the memory (130) of FIG. 1).
[0102] Memory can store various instructions that can be executed by the processor. Instructions may include arithmetic and logical operations, data movement, or control instructions such as input / output that can be recognized by the processor. Memory may include volatile memory (e.g., volatile memory (132) of FIG. 1) and non-volatile memory (e.g., non-volatile memory (134) of FIG. 1) and may store various data temporarily or permanently.
[0103] The processor may be configured to be operatively, functionally, and / or electrically connected to each component of the electronic device (301) and capable of performing operations or data processing regarding the control and / or communication of each component. The operations performed by the processor may be executed by instructions that are stored in memory and, at execution, cause the processor to operate.
[0104] Hereinafter, although there are no limitations on the computation and data processing functions that the processor can implement on the electronic device (301), a series of operations related to XR content service functions will be described. The operations of the processor described below can be performed by executing instructions stored in memory.
[0105] According to one embodiment, the processor can create a virtual object based on virtual information based on image information. The processor can output a virtual object related to an XR service along with background space information through a display (321). For example, the processor can acquire image information by capturing an image related to a real space corresponding to the field of view of a user wearing an electronic device (301) through a second function camera (311, 312), or can create a virtual space for a virtual environment. For example, the processor can control the display (321) to display XR content (hereinafter referred to as the XR content screen) such that at least one virtual object is superimposed on an area determined to be a field of view or a user's field of view (FoV).
[0106] According to one embodiment, the electronic device (301) may have a form factor for being worn on a user's head. The electronic device (301) may further include a strap and / or a wearing member for being secured on a part of the user's body. The electronic device (301) may provide a user experience based on augmented reality, virtual reality, and / or mixed reality while being worn on the user's head.
[0107] FIG. 4 is a drawing illustrating an artificial intelligence system according to one embodiment.
[0108] The artificial intelligence system (400) may include a user query / response interface (410), an AI framework (420), a knowledge component (430), an application / service component (440), and / or a generative model (450).
[0109] In an artificial intelligence system (hereinafter, system) (400), a user query / response interface (410) may receive input. The input may include user input and / or data obtained or generated by an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, or electronic device (301) of FIG. 3a and FIG. 3b). The above data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (120)) (e.g., illumination data around the electronic device obtained from a sensor or sensor hub (e.g., auxiliary processor (123), posture data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., display module (160)) or temperature of at least one processor (120)), size information of the display area of the display module (160), and / or images obtained through an image sensor of the electronic device (e.g., included in the camera module (180)). For example, user input may be of the type of input such as natural language, touch data obtained through a touch circuit included in the display module (160) (e.g., used to identify input from a finger and / or stylus), images, audio, and / or video. Additionally, when user input is transmitted, context information may also be transmitted together. The context information may include various side information related to the time when the user input is entered into the system (400). For example, there may be application information currently being used by the user or location information of the user.Additionally, user input may be of a mixed type of input including the aforementioned natural language, images, audio, video, and / or context information. Furthermore, user input may include non-natural language input, such as selecting a menu.
[0110] The user query / response interface (410) may provide the user with the output of a generative artificial intelligence system. The output may include results (or result information) generated or obtained by the system (400) based on at least part of the input. The output may include natural language-based responses and / or specific content. The output may also include actions requested by the user. For example, the output may have a format according to the user settings of the electronic device.
[0111] The AI framework (420) can receive user input. Based on the user input (e.g., user query), the AI framework (420) can coordinate and control one or more components necessary to perform an action corresponding to the user's intent.
[0112] User input received from the user query / response interface (410) can be transmitted to a prompt design component (421). The prompt design component (421) can be used to generate a prompt suitable as input to a generative model (e.g., LLM (large language model), LVM (large vision model), and / or LMM (large multimodal model)) based on the user input.
[0113] The prompt design component (421) may be an AI component that uses a machine learning algorithm or a neural network. The prompt design component (421) may generate improved prompts through learning over time. The prompt design component (421) may access a knowledge component (430) to generate prompts based on user input. The knowledge component (430) may contain user preference data, a prompt library, and / or prompt examples. The prompt design component (423) may provide the generated prompts to a generative model (e.g., LLM, LVM, and / or LMM).
[0114] The APIs / Plugins management component (423) can communicate with an external information source based on a request for additional information when user input is transmitted to the generative model (450).
[0115] The APIs / Plugins management component (423) can establish a communication channel for communication with the outside of the system (400) via the API. The APIs / Plugins management component (423) can enable access to various data sources through the communication channel. For example, the APIs / Plugins management component (423) can be used to request other components (e.g., application / service component (440)) that perform feedback (or response) according to the prompt. The acquired information can be used to generate a prompt by the prompt design component (421) together with user input, or can be used as input to the generative model (450).
[0116] The APIs / Plugins management component (423) can request the final action via API when the final action corresponding to user input, rather than the intermediate action, needs to be performed by the application or service.
[0117] The refiner component (425) can at least partially tune (or adjust or change) the result (e.g., content) obtained (or output) from the generative model (450). For example, the refiner component (425) can determine the relevance (e.g., score) between the output (e.g., content) of the generative model (450) and the user input. For example, the refiner component (425) can determine whether the output contains biased information (e.g., selective information). For example, the refiner component (425) can determine whether the output contains harmful information (e.g., violent content or profanity).
[0118] The refinement component (425) can determine the degree of matching (e.g., score) between the output of the generative model (450) and the user input (e.g., intent of the user input). If the refinement component (425) determines that the output of the generative model (450) does not correspond to the user input, the refinement component (425) can modify the output to correspond to the user input.
[0119] The refinement component (425) can provide the user with a hint (e.g., a hint for generating a prompt) so that the user can obtain information that matches the user's intention from the generative model (450).
[0120] A generative model (450) may refer to an artificial intelligence neural network that generates new data (e.g., text, images, audio, or video) based on user input (e.g., user utterance). The generative model (450) may include an image generation model and / or a language generation model.
[0121] Image generation models may include generative adversarial networks (GANs) and / or variational autoencoders (VAEs). An example of an image generation model is a diffusion-based generative model that has the structure of a VAE and a transformer.
[0122] A language generation model (e.g., ChatGPT) may be a model trained to generate the statistically most appropriate output based on input. A language generation model may include an LLM. An LLM can identify various types of input, such as text, images, audio (e.g., speech), and / or video, and generate new data corresponding to the input.
[0123] In one embodiment, the AI framework (420) and / or generative model (450) may be included within an AI module (e.g., including a processing circuit) in the electronic device. For example, the AI module may be operatively coupled with at least one processor of the electronic device (e.g., the processor (120) of FIG. 1). For example, the AI module may be operatively coupled with a sensor hub of the electronic device for one or more sensors in the electronic device.
[0124] FIG. 5 is a schematic diagram of an imaging system according to one embodiment.
[0125] An image system (hereinafter, system) (5) according to one embodiment may include an electronic device (510) (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, or the electronic device (301) of FIG. 3a and FIG. 3b)) and an external electronic device (520) (e.g., the electronic device (102) of FIG. 1 or the electronic device (104)).
[0126] According to one embodiment, the electronic device (510) may be a device such as a mobile terminal (e.g., smartphone, tablet, or laptop) or a fixed terminal (e.g., PC (personal computer)).
[0127] According to one embodiment, the electronic device (510) may include at least some of the configurations of the electronic device (201) of FIG. 2 and / or the electronic device (301) of FIG. 3a and FIG. 3b. The electronic device (510) may be implemented in the form of a smart glass, for example, including a wearable electronic device (e.g., the electronic device (201) of FIG. 2), such as virtual reality glasses. The electronic device (510) may also be implemented in the form of a wearable electronic device (e.g., the electronic device (301) of FIG. 3a and FIG. 3b), such as a head-mounted display (HMD), including an augmented reality (AR) device, a virtual reality (VR) device, and / or a mixed reality (MR) device. The electronic device (510) may be configured to easily control user interface (UI) components provided in an AR environment, a VR environment, and / or an MR environment.
[0128] According to one embodiment, the electronic device (510) can establish a communication connection with an external electronic device (520) (e.g., the electronic device (102) or the electronic device (104) of FIG. 1). The electronic device (510) can establish a communication connection with the external electronic device (520) through a short-range wireless communication network (e.g., the first network (198) of FIG. 1) or a long-range wireless communication network (e.g., the second network (199) of FIG. 1).
[0129] According to one embodiment, the external electronic device (520) may be a device such as a mobile terminal (e.g., smartphone, tablet, or laptop) or a fixed terminal (e.g., PC).
[0130] According to one embodiment, the external electronic device (520) may be a wearable electronic device such as the electronic device (201) of FIG. 2 or the electronic device (301) of FIG. 3a and FIG. 3b.
[0131] According to one embodiment, the electronic device (510) may receive a first image (51) from an external electronic device (520) that has established a communication connection with the electronic device (510). The first image (51) may be a still image or video of the actual space (or physical environment) around the external electronic device (520).
[0132] According to one embodiment, the electronic device (510) may acquire a second image (52) using the camera of the electronic device (510) (e.g., sensor module (176) of FIG. 1, camera module (180), first camera (265a, 265b), second camera (270a, 270b), third camera (245) of FIG. 2, second functional camera (311, 312), first functional camera (315), depth sensor (317) of FIG. 3a, third functional camera (328) of FIG. 3b, or fourth functional camera (325, 326)). The second image (52) may be a still image or video of the actual space (or physical environment) around the electronic device (510).
[0133] According to one embodiment, the electronic device (510) can determine a first tag for objects in a first image (51) and a second tag for objects in a second image (52) acquired using the camera of the electronic device (510). The first tag may include one or more tags for each of the objects in the first image (51). The second tag may include one or more tags for each of the objects in the second image (52). The method of identifying objects and the tags are described in detail with reference to FIG. 6.
[0134] According to one embodiment, the electronic device (510) can determine the object type of the objects in the first image (51) based on the first tag for the objects in the first image (51).
[0135] According to one embodiment, the object type may include a background object type and a main object type.
[0136] Background object types may represent object types that reflect the indoor or outdoor environment and atmosphere of the image, such as, for example, walls, floors, sky, ground, or ceilings. The case in which an object type is determined as a background object type is described in detail with reference to FIG. 6.
[0137] The main object type may reflect the user's intention or contextual information of the image of the electronic device (510) (or external electronic device (520)), or represent an object type determined according to a set priority criterion. For example, the user's intention may be manifested by the user directly selecting or designating a specific object. The user's intention or contextual information of the image may be understood as information inferred by the user's interest or gaze regarding the image, the location where the user is situated in the image, or the activities performed. Cases in which an object type is determined as the main object type are described in detail with reference to FIGS. 7 through 11.
[0138] Referring to FIG. 5, for example, an object belonging to the region (501) of the first image (51) can be determined as a background object type. An object belonging to the region (502) of the first image (51) can be determined as a main object type.
[0139] According to one embodiment, the electronic device (510) can determine a second object of a second image (52) associated with a first object corresponding to a first object type among the objects of a first image (51) based on a second tag. According to one embodiment, the first object type may include a background object type. According to one embodiment, the first object type may include a main object type. A method for determining the second object is described in detail with reference to FIG. 6 and FIG. 12 through 14.
[0140] Referring to FIG. 5, for example, an object belonging to the area (503) of the second image (52) can be determined as a second object associated with the first object of the first image (51) corresponding to the background object type as a first object type. An object belonging to the area (504) of the second image (52) can be determined as a second object associated with the first object of the first image (51) corresponding to the main object type as a first object type.
[0141] According to one embodiment, the electronic device (510) can generate a composite image (53) in which a second object in a second image (52) is replaced with a first object. The electronic device (510) can display the composite image (53) through a display.
[0142] The composite image (53) may include an object (e.g., a first object) that reflects the environment and atmosphere of the first image (51) in the second image (52) of the actual space (or physical environment) around the electronic device (510). The composite image (53) may include an object (e.g., a first object) of the first image (51) that reflects the user's intention or contextual information of the image in the second image (52) or is determined according to a set priority criterion. Accordingly, for example, the composite image (53) generated while the electronic device (510) and the external electronic device (520) establish a call connection can enhance the spatial unity and immersion of both sides.
[0143] FIG. 6 is a flowchart of a method for generating an image according to one embodiment.
[0144] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0145] According to one embodiment, the following operations 610 to 670 may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3a and 3b, or the electronic device (510) of FIG. 5). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1) including a processing circuit. The electronic device may include a memory (e.g., the memory (130) of FIG. 1) including one or more storage media for storing instructions. The electronic device may include a display (e.g., the display module (160) of FIG. 1).
[0146] According to one embodiment, the electronic device may be a device such as a mobile terminal (e.g., a smartphone, tablet, or laptop) or a fixed terminal (e.g., a PC (personal computer)).
[0147] According to one embodiment, the electronic device may include at least some of the configurations of the electronic device (201) of FIG. 2 and / or the electronic device (301) of FIG. 3a and FIG. 3b. The electronic device may be implemented in the form of a smart glass, for example, a wearable electronic device (e.g., the electronic device (201) of FIG. 2), such as virtual reality glasses. The electronic device may also be implemented in the form of a wearable electronic device (e.g., the electronic device (301) of FIG. 3a and FIG. 3b), such as a head-mounted display (HMD), such as an augmented reality (AR) device, a virtual reality (VR) device, and / or a mixed reality (MR) device. The electronic device may be configured to easily control user interface (UI) components provided in an AR environment, a VR environment, and / or an MR environment.
[0148] According to one embodiment, the electronic device may be a VST type configured to block external light so that when worn, light emitted from a display reaches the user's eyes, but external light does not reach the user's eyes. According to one embodiment, the electronic device may include an OST type configured to allow external light to reach the user's eyes through glasses when worn.
[0149] The electronic device may include a sensor (e.g., sensor module (176) of FIG. 1, camera module (180), first camera (265a, 265b), second camera (270a, 270b), third camera (245) of FIG. 2, second functional camera (311, 312) of FIG. 3a, first functional camera (315), depth sensor (317), third functional camera (328) of FIG. 3b, and / or fourth functional camera (325, 326)). According to one embodiment, the sensor may convert the measured or detected information into an electrical signal (or sensing data) by measuring or detecting a physical quantity. For example, the sensor may include at least one camera or image sensor for capturing at least one frame of a still image or video of real space (or, physical environment). For example, the sensor may include at least one of a button for touch input, a gesture sensor, a gyroscope, a gyro sensor, a barometric pressure sensor, a magnetic sensor, a magnetometer, an accelerometer, an accelerometer, a grip sensor, a proximity sensor, an RGB sensor, a biophysical sensor, a temperature sensor, a humidity sensor, an illuminance sensor, a UV sensor, an electromyography sensor, an electroencephalography sensor, an infrared sensor, an ultrasonic sensor, an iris sensor, or a fingerprint sensor, but the present disclosure is not limited thereto.
[0150] According to one embodiment, the sensor can capture a physical environment including an object. For example, the sensor may include at least one of an image sensor, a LiDAR sensor, an RGB-D (red-green-blue depth) sensor, a depth sensor, a ToF (time of flight) sensor, an ultrasonic sensor, a radar sensor, and a stereo camera, but the present disclosure is not limited thereto.
[0151] According to one embodiment, the sensor may generate sensing data. The sensing data may be at least one still image or video of a physical environment. The sensing data may be an image (or actual space image) in which one or more objects included in the physical environment are captured. The sensing data may include depth information. For example, the sensing data may be a color image containing depth information, such as an RGB-D image.
[0152] According to one embodiment, an electronic device can acquire sensing data from a sensor. The electronic device can provide augmented reality content using the sensing data acquired from the sensor. The electronic device can generate augmented reality content (or an image of augmented reality content) by blending a physical environment and a virtual environment based on the sensing data. The augmented reality content may include one or more physical environment objects and / or virtual environment objects, such as user interface elements (e.g., avatars, control elements, interactive elements, or any graphic elements), included in a physical environment captured in real time by the electronic device.
[0153] According to one embodiment, the electronic device may establish a communication connection with an external electronic device (e.g., the electronic device (102) of FIG. 1, the electronic device (104) of FIG. 1, or the external electronic device (520) of FIG. 5). The electronic device may establish a communication connection with the external electronic device through a short-range wireless communication network (e.g., the first network (198) of FIG. 1) or a long-range wireless communication network (e.g., the second network (199) of FIG. 1).
[0154] According to one embodiment, the external electronic device may be a device such as a mobile terminal (e.g., smartphone, tablet, or laptop) or a fixed terminal (e.g., PC).
[0155] According to one embodiment, the external electronic device may be a wearable electronic device such as the electronic device (201) of FIG. 2 or the electronic device (301) of FIG. 3a and FIG. 3b.
[0156] In operation 610, the electronic device may receive a first image from an external electronic device that has established a communication connection with the electronic device. The first image may be a still image or a video of the actual space (or physical environment) around the external electronic device.
[0157] According to one embodiment, an external electronic device can transmit a first image acquired using a camera of the external electronic device to an electronic device. For example, the electronic device can receive a first image acquired using a camera of the external electronic device from an external electronic device that has established a call connection with the electronic device.
[0158] According to one embodiment, an external electronic device can transmit a first image stored in the external electronic device to an electronic device.
[0159] According to one embodiment, the electronic device may acquire a second image using the electronic device’s camera (e.g., sensor module (176) of FIG. 1, camera module (180), first camera (265a, 265b), second camera (270a, 270b), third camera (245) of FIG. 2, second functional camera (311, 312), first functional camera (315), depth sensor (317) of FIG. 3a, third functional camera (328) of FIG. 3b, or fourth functional camera (325, 326)). The second image may be a still image or video of the actual space (or physical environment) around the electronic device.
[0160] In operation 620, the electronic device can identify objects of the first image and objects of the second image acquired using the camera of the electronic device.
[0161] According to one embodiment, the electronic device can detect objects included in each of the first image and the second image. For example, the electronic device can determine the bounding boxes of the objects detected in each of the first image and the second image. The electronic device can identify each object (or object region) by performing segmentation (e.g., semantic, instance, or panoptic segmentation) on each of the first image and the second image, based on or independently of the result of object detection.
[0162] In operation 630, the electronic device can determine a first tag for objects in a first image and a second tag for objects in a second image acquired using the camera of the electronic device.
[0163] According to one embodiment, tags for an object (e.g., a first tag and a second tag) may include keywords (or a text description of the object) indicating the classification result of the object and / or characteristics regarding the object.
[0164] For example, as tags for objects, classification results may include labels (or classes) or categories. For example, keywords as tags for objects may include characteristics such as the object's shape, form, size, texture, location, whether it is a dynamic or static object, interactivity with the user, or user accessibility.
[0165] The first tag may include one or more tags for each of the objects in the first image. The second tag may include one or more tags for each of the objects in the second image.
[0166] In operation 640, the electronic device can determine the object type of the objects in the first image based on the first tag for the objects in the first image.
[0167] According to one embodiment, the object type may include a background object type and a main object type. The background object type may represent an object type that reflects the indoor or outdoor environment and atmosphere of the image, such as, for example, a wall, floor, sky, ground, or ceiling. The main object type may represent an object type that reflects the user's intent of the electronic device (or external electronic device) or contextual information of the image, or is determined according to a predetermined priority criterion.
[0168] However, not all objects in the first image may be determined to be either a background object type or a main object type. For example, object types may include background object types, main object types, and other object types (or general object types) that are not determined to be the aforementioned object types. Operation 640 may be understood as an operation of searching for whether there exists an object among the first objects that corresponds to a background object type or a main object type based on the first tag.
[0169] According to one embodiment, the electronic device can determine that an object is a background object type if a tag for any object in a first image includes a predefined classification result regarding a background object type, such as, for example, a wall, floor, sky, ground, or ceiling.
[0170] According to one embodiment, the background object type may include a first background object type and a second background object type.
[0171] The first background object type may represent an object type corresponding to a horizontal plane (or, a floor plane). According to one embodiment, the electronic device may determine that an object is a first background object type if a tag for any object in the first image includes a predefined classification result regarding the first background object type, such as, for example, floor or ground.
[0172] The second background object type may represent an object type corresponding to a vertical plane. According to one embodiment, the electronic device may determine that an object is of the second background object type if a tag for any object in the first image includes a predefined classification result regarding the second background object type, such as, for example, a wall or the sky.
[0173] According to one embodiment, an electronic device can determine the object types of objects in a first image by referring to depth information of the first image. For example, flat objects with little change in depth, such as walls and floors, may be determined as background object types. Objects located farther away than a threshold distance, such as the sky, may be determined as background object types (e.g., second background object types). Objects distributed horizontally, such as floors or ground, may be determined as first background object types. Objects distributed vertically, such as walls, may be determined as second background object types.
[0174] According to one embodiment, the electronic device can determine an object as a background object type by referring to depth information of the first image, even if the tag for any object in the first image does not include a predefined classification result regarding the background object type, if the object is farther away than a threshold distance. For example, an object farther away than a threshold distance, such as a building or a mountain, can be determined as a background object type (e.g., a second background object type).
[0175] The case in which an object type is determined as the main object type is explained in detail with reference to FIGS. 7 to 11.
[0176] According to one embodiment, an electronic device can determine the object types of objects in a first image using a trained model. The trained model may include an artificial intelligence model included (or stored) in the electronic device or a separate server (e.g., server (108) of FIG. 1). The trained model may include, for example, an artificial intelligence model based on a DNN, CNN, Transformer, or a combination of two or more of the above.
[0177] In operation 650, the electronic device can determine a second object of a second image associated with a first object corresponding to a first object type among the objects of a first image based on a second tag.
[0178] According to one embodiment, the first object type may include a background object type. A method for determining a second object in a second image associated with a first object corresponding to a background object type among the objects in a first image is described in detail with reference to FIG. 12.
[0179] According to one embodiment, the first object type may include a main object type. A method for determining a second object in a second image associated with a first object corresponding to a main object type among the objects in a first image is described in detail with reference to FIG. 13.
[0180] According to one embodiment, the first object type may include an other object type (or a general object type) that is not determined to be a background object type or a main object type. A method for determining a second object in a second image associated with a first object corresponding to an other object type among the objects in the first image is described in detail with reference to FIG. 14.
[0181] In operation 660, the electronic device can generate a composite image in which the second object in the second image is replaced with the first object.
[0182] According to one embodiment, an electronic device can generate a composite image by overlaying a first object onto a second object of a second image. The electronic device can overlay an adjusted first object onto a second object of the second image by adjusting the first object according to the viewpoint and size of the second object.
[0183] According to one embodiment, the electronic device can remove a second object from a second image and then inpaint the corresponding area with a background (e.g., an object corresponding to a background object type). Subsequently, the electronic device can generate a composite image by inserting a first object at a location corresponding to the second object in the second image.
[0184] According to one embodiment, an electronic device can generate a composite image using a generative model (e.g., the generative model (450) of FIG. 4). The electronic device can generate a composite image by using the generative model, not only by simply replacing a second object in a second image with a first object, but also by reflecting environmental attributes of the first image (e.g., whether it is indoors or outdoors, illuminance, time, or shadow) in the second image. For example, the electronic device can generate a prompt to generate a composite image by replacing a second object in a second image with a first object. The electronic device can obtain a composite image output by the generative model by inputting a prompt containing instructions (or commands) to generate the first image, the second image, and the aforementioned composite image into the generative model.
[0185] The composite image may include an object (e.g., a first object) that reflects the environment and atmosphere of the first image in a second image of the actual space (or physical environment) surrounding the electronic device. The composite image may include an object (e.g., a first object) of the first image that reflects the user's intention or contextual information of the image in the second image, or is determined according to a predetermined priority criterion. Accordingly, for example, a composite image generated while an electronic device and an external electronic device establish a call connection can enhance the spatial unity and immersion of both parties.
[0186] In operation 670, the electronic device can display the composite image through the display.
[0187] According to one embodiment, the electronic device can update the composite image at a fixed period (e.g., every n frames). For example, the electronic device can update the composite image by repeatedly performing the aforementioned operations at a fixed period. For example, if the first object corresponds to a main object type, the electronic device can identify the first object and the second object in the first image and the second image, respectively, at a fixed period, and generate a composite image in which the second object in the second image is replaced with the first object.
[0188] According to one embodiment, the electronic device may update the composite image by repeatedly performing the aforementioned operations whenever the first image and / or the second image changes by a predetermined amount (e.g., n%) compared to the corresponding frame. According to one embodiment, the electronic device may update the composite image by replacing the second object with the first object whenever characteristics such as the shape, form, texture, or location of the first object and / or the second object change by a predetermined amount (e.g., n%) compared to the corresponding frame.
[0189] According to one embodiment, the electronic device can determine whether an object corresponding to a user in a second image approaches within a threshold distance of the second object. For example, the electronic device may identify an object corresponding to a person as a user in the second image, or identify an object corresponding to a user based on a previously stored facial image of the user in the electronic device. The electronic device may track the distance between the identified object and the second object. Based on the determination that the object corresponding to the user has approached within a threshold distance of the second object, the electronic device may update the composite image so that the first object is replaced by the second object. This can be understood as updating the composite image to maintain the corresponding portion of the original second image in which the second object is not replaced by the first object. In one embodiment, if the first object does not correspond to the main object type, the electronic device may update the composite image so that the first object is replaced by the second object based on the determination that the object corresponding to the user has approached within a threshold distance of the second object. For example, when a user interacts with a second object in close proximity to it in real space, consistency between the physical environment and the virtual environment (or synthetic image) can be maintained by displaying the second object, which is a real object, instead of the first object.
[0190] According to one embodiment, an electronic device can determine environmental properties of a first image. Environmental properties may include characteristics such as whether the image is indoors or outdoors, illuminance, time (e.g., day or night), or shadows. The electronic device can generate a composite image in which a second object in a second image is replaced with a first object, and a graphic effect is applied to at least a portion of the second image based on environmental properties.
[0191] According to one embodiment, the electronic device may determine an object in the second image that satisfies a fourth defined criterion as a non-replacement object type based on a second tag for objects in the second image. The fourth defined criterion may include a tag for any object in the second image containing predefined keywords, such as being hot, having a risk of burns, being dangerous, requiring user attention, being a heavy object placed on the floor, or being an immovable object. The electronic device may determine a second object associated with a first object among the objects in the second image that are not identified as a non-replacement object type based on the second tag. Accordingly, by preventing an object corresponding to a non-replacement object type from being replaced by an object in the first image when generating a composite image, it is possible to ensure that an object that requires user attention in the actual space or is actually impossible to replace is maintained in the composite image.
[0192] According to one embodiment, an external electronic device may perform the aforementioned operations 610 to 660. For example, the external electronic device may receive a third image from an electronic device that has established a communication connection with the external electronic device. The third image may be a still image or video of the actual space (or physical environment) surrounding the electronic device. The external electronic device may determine a third tag for objects in the third image and a fourth tag for objects in the fourth image acquired using the camera of the external electronic device. The external electronic device may determine the object type of the objects in the third image based on the third tag for objects in the third image. Based on the fourth tag, the external electronic device may determine a second target object in the fourth image associated with a first target object corresponding to a first object type among the objects in the third image. The external electronic device may generate a composite image in which the second target object in the fourth image is replaced with the first target object.
[0193] According to one embodiment, a plurality of electronic devices may each perform the aforementioned operations 610 to 660. For example, a plurality of electronic devices and an external electronic device may establish a communication connection. A plurality of electronic devices may each receive a first image from an external electronic device. The first image may be a still image or video of the actual space (or physical environment) surrounding the external electronic device. A plurality of electronic devices may each determine a first tag for objects in the first image and a second tag for objects in an image acquired using the camera of each electronic device. A plurality of electronic devices may each determine the object type of the objects in the first image based on the first tag for the objects in the first image. A plurality of electronic devices may each determine a second object in an image acquired using the camera of each electronic device associated with a first object corresponding to the first object type among the objects in the first image based on the second tag. A plurality of electronic devices may each generate a composite image in which the second object in the second image is replaced by the first object.
[0194] FIG. 7 is a flowchart of a method for determining an object type according to one embodiment.
[0195] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0196] According to one embodiment, the following operations 710 and 720 may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3a and 3b, or the electronic device (510) of FIG. 5). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1) including a processing circuit. The electronic device may include a memory (e.g., the memory (130) of FIG. 1) including one or more storage media for storing instructions. The electronic device may include a display (e.g., the display module (160) of FIG. 1).
[0197] As described with reference to FIG. 6, the electronic device may receive a first image from an external electronic device that has established a communication connection with the electronic device (e.g., the electronic device (102) of FIG. 1, the electronic device (104), or the external electronic device (520) of FIG. 5). The electronic device may determine a first tag for objects in the first image and a second tag for objects in the second image acquired using the camera of the electronic device.
[0198] According to one embodiment, the operation 640 for determining the object type of the objects of the first image of FIG. 6 may include operations 710 and 720.
[0199] In operation 710, the electronic device can receive object region information specified for the first image from an external electronic device.
[0200] According to one embodiment, an external electronic device may receive user input specifying an arbitrary region or object with respect to a first image. The external electronic device may transmit object region information determined based on the user input to an electronic device.
[0201] For example, object region information may include coordinate information of a designated region within the first image. An external electronic device may transmit object region information including coordinate information of a designated region within the first image to the electronic device.
[0202] For example, object region information may include a designated object (or object region or coordinate information of the object region) and / or a tag for said object within the first image. An external electronic device may identify an object (or object region) included in a designated area within the first image based on user input. The external electronic device may determine a tag for an object identified in the first image. A tag for an object may include a keyword (or a text description of the object) indicating the classification result of the object and / or characteristics regarding the object. The external electronic device may transmit object region information including a designated object within the first image and / or a tag for said object to an electronic device.
[0203] In operation 720, the electronic device can determine the object types of the objects in the first image based on the object region information of the first tag and the first image.
[0204] As described with reference to FIG. 6, the first tag may include one or more tags for each of the objects in the first image. According to one embodiment, the electronic device may determine that an object is a background object type if the tag for any object in the first image includes a predefined classification result such as, for example, wall, floor, sky, ground, or ceiling. According to one embodiment, the electronic device may determine the object types of the objects in the first image by referring to the depth information of the first image.
[0205] According to one embodiment, the electronic device can determine an object identified based on coordinate information of a designated area within a first image included in object area information of a first image as a main object type. The electronic device can identify an object included in the coordinates of the designated area in the received first image. The electronic device can determine the identified object as a main object type that reflects the user's intention.
[0206] According to one embodiment, the electronic device can identify an object referenced by a designated object (or object region or coordinate information of the object region) and / or a tag (e.g., classification result) included in the object region information of the first image. The electronic device can determine the identified object as a main object type that reflects the user's intention.
[0207] Figure 8 is a diagram illustrating a method for determining an object type according to one example.
[0208] The screens (81, 82) of FIG. 8 are exemplary display contents of a first image obtained by an external electronic device (e.g., the electronic device (102) of FIG. 1, the electronic device (104) or the external electronic device (520) of FIG. 5) according to one embodiment.
[0209] As described with reference to FIGS. 6 and 7, an external electronic device may receive user input specifying an arbitrary region or object for a first image. The external electronic device may transmit object region information determined based on the user input to an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3a and 3b, or electronic device (510) of FIG. 5). The electronic device may receive the first image and object region information specified for the first image from an external electronic device that has established a communication connection with the electronic device.
[0210] Referring to the screen (81), the electronic device can identify objects (801, 802, 803, 804, 805, 806, 807) included in the first image.
[0211] The electronic device can determine a first tag for objects (801, 802, 803, 804, 805, 806, 807) of the first image. The first tag may include one or more tags for each of the objects (801, 802, 803, 804, 805, 806, 807) of the first image. For example, the electronic device can determine a tag such as wall, brick pattern, wallpaper, or static object for an object (801) of the first image. The electronic device can determine a tag such as drum, interactive with user, or accessible to user for an object (802) of the first image. The electronic device can determine a tag such as floor or static object for an object (804) of the first image. The electronic device can determine a tag such as person or user for an object (805) of the first image.
[0212] Referring to the screen (82), the external electronic device may receive user input specifying an arbitrary area (808) (or object) for the first image. The external electronic device may transmit object area information determined based on the user input to the electronic device.
[0213] For example, object region information may include coordinate information of a designated region (808) within the first image. An external electronic device may transmit object region information including coordinate information of a designated region (808) within the first image to the electronic device.
[0214] For example, object region information may include a designated object (802) within the first image (or, the object region (808) or coordinate information of the object region (808)) and / or a tag for said object (802). According to one embodiment, an external electronic device may determine a first tag for objects (801, 802, 803, 804, 805, 806, 807) of the first image. The object region information may include a tag determined for a designated object (802) included in the region (808) of the first image, such as drum, interactive with user, or accessible to user.
[0215] The electronic device can determine the object types of the objects (801, 802, 803, 804, 805, 806, 807) of the first image based on the object region information of the first tag and the first image. The electronic device can determine the object (802) among the objects (801, 802, 803, 804, 805, 806, 807) of the first image as the main object type that reflects the user's intention.
[0216] FIG. 9 is a flowchart of a method for determining an object type according to one embodiment.
[0217] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0218] According to one embodiment, the following operations 910 and 920 may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3a and 3b, or the electronic device (510) of FIG. 5). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1) including a processing circuit. The electronic device may include a memory (e.g., the memory (130) of FIG. 1) including one or more storage media for storing instructions. The electronic device may include a display (e.g., the display module (160) of FIG. 1).
[0219] As described with reference to FIG. 6, the electronic device may receive a first image from an external electronic device that has established a communication connection with the electronic device (e.g., the electronic device (102) of FIG. 1, the electronic device (104), or the external electronic device (520) of FIG. 5). The electronic device may determine a first tag for objects in the first image and a second tag for objects in the second image acquired using the camera of the electronic device.
[0220] According to one embodiment, the operation 640 for determining the object type of the objects of the first image of FIG. 6 may include operations 910 and 920.
[0221] In operation 910, the electronic device can determine context information regarding the first image.
[0222] Context information regarding the first image may include at least one of scene information regarding the first image determined using a trained model, topic information regarding voice acquired during a call connection with an external electronic device, eye-tracking information received from an external electronic device, or location information received from an external electronic device.
[0223] The trained model may include an artificial intelligence model contained (or stored) in an electronic device or a separate server (e.g., server (108) of FIG. 1). The trained model may include, for example, an artificial intelligence model based on a DNN, CNN, Transformer, or a combination of two or more of the above.
[0224] For example, the trained model may be an artificial intelligence model tuned to perform analytical tasks such as classification, prediction, or detection. For example, the trained model may be a generative model such as an LMM (e.g., the generative model (450) of FIG. 4).
[0225] According to one embodiment, an electronic device can determine scene information regarding a first image using a trained model. The scene information may include information describing the actions of a person included in the image, the types of objects included in the image, or the situation of the image, and / or subject information of the image determined by said information. For example, the electronic device can obtain scene information regarding a first image by inputting the first image into an artificial intelligence model tuned to determine scene information. For example, the electronic device can generate a prompt containing instructions (or commands) to output scene information regarding the first image. The electronic device can obtain scene information output by a generative model by inputting the prompt into a generative model.
[0226] According to one embodiment, an electronic device can determine topic information for voice acquired during a call connection with an external electronic device using a trained model. The electronic device can generate text by converting the voice of a conversation between users acquired during a call connection with the external electronic device into speech-to-text (STT). For example, the electronic device can acquire topic information for voice acquired during a call connection with the external electronic device by inputting text into an artificial intelligence model tuned to determine topic information. For example, the electronic device can generate a prompt containing instructions (or commands) to output topic information for the text. The electronic device can acquire topic information output by a generative model by inputting the prompt into a generative model.
[0227] According to one embodiment, the external electronic device may be a wearable electronic device such as the electronic device (201) of FIG. 2 or the electronic device (301) of FIG. 3a and FIG. 3b. The external electronic device may track the user's eyes, that is, the user's gaze, based on the detection results of the eye-tracking sensor. The external electronic device may transmit a first image or coordinates on a display regarding the user's gaze, or a gaze vector, to the electronic device. The electronic device may determine the coordinates or gaze vector received from the external electronic device as eye-tracking information.
[0228] According to one embodiment, an external electronic device can transmit location information (e.g., GPS information) of the external electronic device to an electronic device. The electronic device can receive location information from the external electronic device.
[0229] In operation 920, the electronic device can determine the object types of the objects of the first image based on context information regarding the first tag and the first image.
[0230] As described with reference to FIG. 6, the first tag may include one or more tags for each of the objects in the first image. According to one embodiment, the electronic device may determine that an object is a background object type if the tag for any object in the first image includes a predefined classification result such as, for example, wall, floor, sky, ground, or ceiling. According to one embodiment, the electronic device may determine the object types of the objects in the first image by referring to the depth information of the first image.
[0231] According to one embodiment, the electronic device may determine, based on scene information regarding the first image, at least one object with the highest importance among the objects of the first image as the main object type that reflects contextual information of the image. For example, the object with the highest importance may be defined as the object that has the highest similarity (e.g., cosine similarity) or the shortest distance (e.g., Euclidean distance) to information describing the situation of a person included in the image (e.g., the first image), or the subject information of the image.
[0232] According to one embodiment, the electronic device may determine at least one object with the highest importance among the objects of the first image as the main object type reflecting the context information of the image, based on subject information regarding voice acquired during a call connection with an external electronic device. For example, the object with the highest importance may be defined as the object that has the highest similarity (e.g., cosine similarity) or the shortest distance (e.g., Euclidean distance) to the subject information regarding voice acquired during a call connection with the external electronic device.
[0233] According to one embodiment, the electronic device may determine, based on eye tracking information, an object among the objects of the first image that corresponds to the coordinates on the first image or display or the gaze vector regarding the gaze of the user of the external electronic device as a main object type that reflects the user's intention.
[0234] According to one embodiment, the electronic device may determine, based on location information of an external electronic device, the object among the objects in the first image that has the highest similarity (e.g., cosine similarity) or the shortest distance (e.g., Euclidean distance) to the location where the actual space around the external electronic device is located as the main object type that reflects the context information of the image.
[0235] FIG. 10 is a flowchart of a method for determining an object type according to one embodiment.
[0236] According to one embodiment, the following operations 1010 and 1020 may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3a and 3b, or the electronic device (510) of FIG. 5). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1) including a processing circuit. The electronic device may include a memory (e.g., the memory (130) of FIG. 1) including one or more storage media for storing instructions. The electronic device may include a display (e.g., the display module (160) of FIG. 1).
[0237] As described with reference to FIG. 6, the electronic device may receive a first image from an external electronic device that has established a communication connection with the electronic device (e.g., the electronic device (102) of FIG. 1, the electronic device (104), or the external electronic device (520) of FIG. 5). The electronic device may determine a first tag for objects in the first image and a second tag for objects in the second image acquired using the camera of the electronic device.
[0238] According to one embodiment, the operation 640 for determining the object type of the objects of the first image of FIG. 6 may include operations 1010 and 1020.
[0239] As described with reference to FIG. 6, the first tag may include one or more tags for each of the objects in the first image. According to one embodiment, the electronic device may determine that an object is a background object type if the tag for any object in the first image includes a predefined classification result such as, for example, wall, floor, sky, ground, or ceiling. According to one embodiment, the electronic device may determine the object types of the objects in the first image by referring to the depth information of the first image.
[0240] According to one embodiment, the electronic device can determine the respective scores of the objects in the first image based on the first tags of the objects in the first image. The scores can be defined as criteria for determining any object as a main object type.
[0241] The electronic device can determine the object types of the objects in the first image according to defined priority criteria based on the individual scores of the objects in the first image.
[0242] According to one embodiment, the predetermined priority criteria may include determining an object as the main object type when the score of any object in the first image is greater than or equal to a threshold.
[0243] As described with reference to FIG. 9, the electronic device can determine context information regarding the first image. The context information regarding the first image may include at least one of scene information regarding the first image determined using a trained model, topic information regarding voice acquired during a call connection with an external electronic device, eye tracking information received from the external electronic device, or location information received from the external electronic device.
[0244] In operation 1010, the electronic device can determine individual scores for objects in the first image based on the similarity between the first tags of the objects in the first image and the context information of the first image. The electronic device can add a score for an object if the similarity between the tag of any object in the first image and the context information of the first image is greater than or equal to a threshold. For example, if the background is a park as the scene information of the first image, a score may be added to an object among the objects in the first image that has high similarity to the scene information, such as a bench or a fountain. Thus, by determining an object that reflects the environment and atmosphere of the first image as the main object type, the atmosphere of the first image can be further reflected in the second image.
[0245] According to one embodiment, the electronic device may add (or increase) a score for any object in a first image if the tag for that object includes a predefined keyword, such as, for example, dynamic object, interactive with user, or accessible to user. For example, the electronic device may add a score based on the number of tags for any object in the first image that include a predefined keyword. The electronic device may determine one or more objects among the objects in the first image whose score is above (or exceeds) a threshold as the main object type. According to one embodiment, the electronic device may determine the object with the highest score among the objects in the first image as the main object type.
[0246] In operation 1020, the electronic device can determine the object types of the objects in the first image according to a predetermined priority criterion based on the individual scores of the objects in the first image.
[0247] According to one embodiment, the predetermined priority criteria may include determining an arbitrary object in a first image as an other object type (or general object type) when the score of the arbitrary object in the first image is less than (or less than) a threshold. The electronic device may determine an object among the objects in the first image whose score is less than (or less than) a threshold as an other object type (or general object type).
[0248] According to one embodiment, the electronic device may first identify an object among the objects of a first image that corresponds to a background object type. The electronic device may determine one or more objects among the objects that do not correspond to the background object type of the first image, whose scores are greater than or equal to a threshold, as the main object type. The electronic device may determine an object among the objects that do not correspond to the background object type of the first image, whose scores are less than or equal to the threshold, as an other object type (or general object type).
[0249] FIG. 11 is a diagram illustrating a method for determining an object type according to one example.
[0250] The screen (111) of FIG. 11 is an exemplary display of a first image obtained by an external electronic device (e.g., the electronic device (102) of FIG. 1, the electronic device (104) of FIG. 1, or the external electronic device (520) of FIG. 5) according to one embodiment.
[0251] As described with reference to FIGS. 6, 9, and 10, an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3a and 3b, or electronic device (510) of FIG. 5) may receive a first image from an external electronic device. The electronic device may determine a first tag for objects in the first image.
[0252] Referring to the screen (111), the electronic device can identify objects (1101, 1102, 1103, 1104, 1105, 1106, 1107) included in the first image.
[0253] The electronic device can determine a first tag for objects (1101, 1102, 1103, 1104, 1105, 1106, 1107) of the first image. The first tag may include one or more tags for each of the objects (1101, 1102, 1103, 1104, 1105, 1106, 1107) of the first image. For example, the electronic device can determine a tag such as building or static object for an object (1101) of the first image. The electronic device can determine a tag such as bench, user accessible, or static object for an object (1102) of the first image. The electronic device can determine a tag such as sky for an object (1103) of the first image. The electronic device can determine a tag such as puppy, pet dog, animal, dynamic object, or user interactable for an object (1106) of the first image. The electronic device can determine tags such as floor, walk, or static object for the object (1107) of the first image.
[0254] According to one embodiment, the electronic device may determine an object as a background object type if a tag for any object in a first image includes a predefined classification result such as, for example, a wall, a floor, a sky, a ground, or a ceiling. For example, the electronic device may determine an object (1103, 1107) as a background object type. An object (1107) may be determined as a first background object type, and an object (1103) may be determined as a second background object type.
[0255] According to one embodiment, the electronic device can determine the object types of objects in the first image by referring to the depth information of the first image. For example, the electronic device can determine an object (1101) located farther away than a threshold distance as a background object type (e.g., a second background object type).
[0256] According to one embodiment, the electronic device can determine context information regarding the first image. Based on the first tag and the context information regarding the first image, the electronic device can determine the object types of at least some of the objects (1101, 1102, 1103, 1104, 1105, 1106, 1107) of the first image. For example, if the background is a park as the scene information of the first image, the electronic device can determine a bench, i.e., an object (1102), which has high similarity to the scene information, as the main object type.
[0257] According to one embodiment, according to one embodiment, the electronic device may add a score for any object in the first image if the tag for that object includes predefined keywords such as, for example, dynamic object, interactive with user, or accessible to user. For example, the electronic device may determine an object (1106) whose score is above (or exceeds) a threshold as the main object type.
[0258] FIG. 12 is a flowchart of a method for matching objects between images according to one embodiment.
[0259] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0260] According to one embodiment, the following operations 1210 and 1220 may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3a and 3b, or the electronic device (510) of FIG. 5). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1) including a processing circuit. The electronic device may include a memory (e.g., the memory (130) of FIG. 1) including one or more storage media for storing instructions. The electronic device may include a display (e.g., the display module (160) of FIG. 1).
[0261] As described with reference to FIG. 6, the electronic device may receive a first image from an external electronic device that has established a communication connection with the electronic device (e.g., the electronic device (102), the electronic device (104) of FIG. 1, or the external electronic device (520) of FIG. 5). The electronic device may determine a first tag for objects in the first image and a second tag for objects in the second image acquired using the camera of the electronic device. The electronic device may determine the object type of the objects in the first image based on the first tag for objects in the first image.
[0262] According to one embodiment, the operation 650 for determining a second object in a second image associated with a first object corresponding to a first object type among the objects in the first image of FIG. 6 may include operations 1210 and 1220.
[0263] In operation 1210, the electronic device can identify a first object corresponding to a background object type as a first object type among the objects of the first image.
[0264] As described with reference to FIG. 6, according to one embodiment, an electronic device may determine an object as a background object type if a tag for any object in a first image includes a predefined classification result, such as, for example, a wall, floor, sky, ground, or ceiling. The background object type may include a first background object type and a second background object type. The first background object type may represent an object type corresponding to a horizontal plane (or a floor plane). According to one embodiment, an electronic device may determine an object as a first background object type if a tag for any object in a first image includes a predefined classification result, such as, for example, a floor or ground. The second background object type may represent an object type corresponding to a vertical plane. According to one embodiment, an electronic device may determine an object as a second background object type if a tag for any object in a first image includes a predefined classification result, such as, for example, a wall or sky.
[0265] In operation 1220, the electronic device can determine a second object of a second image associated with a first object based on a second tag for objects of the second image.
[0266] According to one embodiment, the electronic device may determine, among the objects of the second image, an object in which the corresponding tag includes a predefined classification result with respect to the first background type, such as floor or ground, as a second object associated with the first object corresponding to the first background object type.
[0267] According to one embodiment, the electronic device may determine, among the objects of the second image, an object in which the corresponding tag includes a predefined classification result regarding the second background type, such as a wall or sky, as a second object associated with a first object corresponding to the second background object type.
[0268] According to one embodiment, the electronic device can determine an object as a second object associated with a first object by referring to depth information of the second image, even if the tag for any object in the second image does not include a predefined classification result regarding the background object type. For example, an object located at a distance greater than the threshold, such as a building or a mountain, can be determined as a second object associated with a first object corresponding to a background object type (e.g., a second background object type).
[0269] Operations 1210 and 1220 can be understood as operations that associate (or match) a first object corresponding to a background object type among the objects of the first image with a second object corresponding to a background object type among the objects of the second image. The electronic device can improve the spatial unity and immersion of both sides by matching the objects of the first image and the second image, respectively, corresponding to the same background object type, and by generating a composite image in which the background of the second image is replaced with the background of the first image.
[0270] According to one embodiment, if the second object is associated with a first object corresponding to a second background object type, the electronic device can determine whether a third object exists within a closed area within the region of the second object before generating a composite image. For example, if a TV or a picture frame is hanging on a wall (the second object) in the second image, the TV or the picture frame may be identified as the third object. Furniture placed on the floor in front of the wall (the second object) in the second image may not be identified as the third object because it is not included in the closed area of the wall.
[0271] If a third object is included in a closed region within the region of a second object, the electronic device can determine whether a fourth object of the first image exists, based on a comparison between a first tag for objects of the first image and a tag of the third object, such that the corresponding comparison result satisfies a first predetermined criterion.
[0272] The first defined criterion may include, for example, a comparison result between a classification result and / or keyword (e.g., shape, form, size, or location) as a tag of a third object and a tag of any object in the first image, such that a predetermined number (or exceeding) of tags are identical or have a similarity level greater than or equal to a threshold. The electronic device may determine that an object is the fourth object if, as a result of comparing a corresponding tag among the objects in the first image with the tag of the third object, there exists an object in which a predetermined number (or exceeding) of tags are identical or have a similarity level greater than or equal to a threshold.
[0273] The electronic device may determine that a third object in a second image is replaced by a fourth object if there exists a fourth object in a first image for which the corresponding comparison result satisfies a first defined criterion. For example, a TV or picture frame (third object) hanging on a wall (second object) in the second image may be replaced by a sign (fourth object) in the first image. Since the TV or picture frame in the second image and the sign in the first image are rectangular and have similar sizes, they can satisfy the first defined criterion.
[0274] The electronic device can identify a third object in the second image as part of the second object if there is no fourth object in the first image that satisfies a first defined criterion for the corresponding comparison result. For example, a TV or picture frame (third object) hanging on a wall (second object) in the second image can be identified as part of the wall (second object). If the background of the second image is indoors while the background of the first image is outdoors, a heterogeneous composite image may be generated if the third object (e.g., TV or picture frame) contained within an enclosed area of the background remains unchanged even though the background of the second image (e.g., wall) has been replaced with the outdoor background of the first image. Therefore, a natural composite image can be generated by identifying the third object in the second image as part of the second object, i.e., the background, and then replacing the entire background including the third object with the background of the first image.
[0275] FIG. 13 is a flowchart of a method for matching objects between images according to one embodiment.
[0276] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0277] According to one embodiment, the following operations 1310 and 1320 may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3a and 3b, or the electronic device (510) of FIG. 5). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1) including a processing circuit. The electronic device may include a memory (e.g., the memory (130) of FIG. 1) including one or more storage media for storing instructions. The electronic device may include a display (e.g., the display module (160) of FIG. 1).
[0278] As described with reference to FIG. 6, the electronic device may receive a first image from an external electronic device that has established a communication connection with the electronic device (e.g., the electronic device (102), the electronic device (104) of FIG. 1, or the external electronic device (520) of FIG. 5). The electronic device may determine a first tag for objects in the first image and a second tag for objects in the second image acquired using the camera of the electronic device. The electronic device may determine the object type of the objects in the first image based on the first tag for objects in the first image.
[0279] According to one embodiment, the operation 650 for determining a second object of a second image associated with a first object corresponding to a first object type among the objects of the first image of FIG. 6 may include operations 1310 and 1320.
[0280] In operation 1310, the electronic device can identify a first object corresponding to a main object type as a first object type among the objects of the first image.
[0281] As described with reference to FIGS. 7 through 11, according to one embodiment, an electronic device may receive object region information designated for a first image from an external electronic device. Based on the object region information of the first image, the electronic device may determine an identified object as a main object type. According to one embodiment, the electronic device may determine context information regarding the first image. Based on the context information regarding the first image, the electronic device may determine as a main object type an object that has the highest importance among the objects of the first image, corresponds to coordinates or gaze vectors on the first image or display regarding the gaze of the user of the external electronic device, has the highest similarity to a location where the actual space around the external electronic device is located, or is closest in distance. According to one embodiment, the electronic device may determine as a main object type at least one object to which the highest priority (or the highest weight) is assigned according to a priority criterion determined regarding the first tag for the objects of the first image.
[0282] In operation 1320, the electronic device can determine a second object of a second image whose corresponding comparison result satisfies a second determined criterion, based on a comparison between a tag of one object and a second tag for objects of a second image.
[0283] According to one embodiment, the electronic device can determine respective matching scores between the first object and the objects of the second image based on a comparison between the tag of the first object and the second tag for the objects of the second image. The result of the comparison between the tag of the first object and the second tag for the objects of the second image may include respective matching scores between the first object and the objects of the second image.
[0284] The electronic device may add a matching score between the first object and the corresponding object in the second image when the tag of the first object and the tag of any object in the second image are identical or their similarity is greater than or equal to a threshold. The electronic device may add a matching score equal to the number of tags that are identical or have a similarity greater than or equal to the threshold when two or more tags of the first object and two or more tags of any object in the second image are each identical or their similarity is greater than or equal to a threshold.
[0285] The second defined criterion may include the highest corresponding matching score resulting from a comparison between the tag of the first object and the second tag for the objects of the second image. According to one embodiment, the second defined criterion may include the highest corresponding matching score resulting from a comparison between the tag of the first object and the second tag for the objects of the second image, and being greater than or equal to a threshold. For example, even if the matching score between the first object and any object of the second image is the highest, the electronic device may not determine the object as the second object associated with the first object if the matching score is less than or equal to the threshold.
[0286] FIG. 14 is a flowchart of a method for matching objects between images according to one embodiment.
[0287] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0288] According to one embodiment, the following operations 1410 and 1420 may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3a and 3b, or the electronic device (510) of FIG. 5). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1) including a processing circuit. The electronic device may include a memory (e.g., the memory (130) of FIG. 1) including one or more storage media for storing instructions. The electronic device may include a display (e.g., the display module (160) of FIG. 1).
[0289] According to one embodiment, the operation 650 for determining a second object of a second image associated with a first object corresponding to a first object type among the objects of the first image of FIG. 6 may include operations 1410 and 1420.
[0290] As described with reference to FIG. 6, the electronic device may receive a first image from an external electronic device that has established a communication connection with the electronic device (e.g., the electronic device (102), the electronic device (104) of FIG. 1, or the external electronic device (520) of FIG. 5). The electronic device may determine a first tag for objects in the first image and a second tag for objects in the second image acquired using the camera of the electronic device. The electronic device may determine the object type of the objects in the first image based on the first tag for objects in the first image.
[0291] In operation 1410, the electronic device can identify a first object among the objects of the first image that does not correspond to a background object type.
[0292] Although the case in which an object type is determined as a main object type has been described with reference to FIGS. 7 through 11, according to one embodiment, the operations of determining an object type as a main object type may be omitted. That is, the electronic device may determine objects among the objects of the first image that do not correspond to a background object type as an other object type (or, general object type). In operation 1410, the electronic device may identify a first object among the objects of the first image that corresponds to an other object type.
[0293] In operation 1420, the electronic device can determine a second object of a second image whose corresponding comparison result satisfies a third determined criterion, based on a comparison between a tag of a first object and a second tag of objects of a second image.
[0294] According to one embodiment, the electronic device can determine respective matching scores between the first object and the objects of the second image based on a comparison between the tag of the first object and the second tag for the objects of the second image. The result of the comparison between the tag of the first object and the second tag for the objects of the second image may include respective matching scores between the first object and the objects of the second image.
[0295] The electronic device may add a matching score between the first object and the corresponding object in the second image when the tag of the first object and the tag of any object in the second image are identical or their similarity is greater than or equal to a threshold. The electronic device may add a matching score equal to the number of tags that are identical or have a similarity greater than or equal to the threshold when two or more tags of the first object and two or more tags of any object in the second image are each identical or their similarity is greater than or equal to a threshold.
[0296] As described with reference to FIG. 9, the electronic device can determine context information regarding the first image. The context information regarding the first image may include at least one of scene information regarding the first image determined using a trained model, topic information regarding voice acquired during a call connection with an external electronic device, eye tracking information received from the external electronic device, or location information received from the external electronic device.
[0297] According to one embodiment, the electronic device may apply a weight when adding matching scores based on the similarity between the tag of the first object and the context information of the first image. If the similarity between the tag of the first object and the context information of the first image is greater than or equal to a threshold, the electronic device may apply a weight when adding matching scores between the first object and the corresponding object in the second image. For example, if the background is a park as the scene information of the first image, a weight may be applied when adding matching scores if the first object in the first image is an object with high similarity to the scene information, such as a bench or a fountain. Accordingly, by displaying the first object that reflects the environment and atmosphere of the first image in the second image, the atmosphere of the first image can be further reflected in the second image.
[0298] The second defined criterion may include the highest corresponding matching score resulting from a comparison between the tag of the first object and the second tag for the objects of the second image. According to one embodiment, the second defined criterion may include the highest corresponding matching score resulting from a comparison between the tag of the first object and the second tag for the objects of the second image, and being greater than or equal to a threshold. For example, even if the matching score between the first object and any object of the second image is the highest, the electronic device may not determine the object as the second object associated with the first object if the matching score is less than or equal to the threshold.
[0299] FIG. 15 is a diagram illustrating a method of matching objects between images according to one example.
[0300] The screens (151, 152, 153) of FIG. 15 are exemplary display contents of a first image, a second image, and a composite image according to one embodiment.
[0301] As described above with reference to FIGS. 6 through 14, an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3a and 3b, or electronic device (510) of FIG. 5) can receive a first image from an external electronic device (e.g., electronic device (102) of FIG. 1, electronic device (104), or external electronic device (520) of FIG. 5) that has established a communication connection with the electronic device.
[0302] Referring to the screen (151), the electronic device can identify objects (1501, 1502, 1503, 1504, 1505, 1506, 1507) included in the first image. The first image of FIG. 15 may correspond to the first image of FIG. 11. The objects (1501, 1502, 1503, 1504, 1505, 1506, 1507) may each correspond to objects (1101, 1102, 1103, 1104, 1105, 1106, 1107).
[0303] For example, the electronic device can determine the object (1507) as a first background object type because the tag for the object (1507) includes a predefined classification result (e.g., floor or ground) regarding the first background object type. The electronic device can determine the object (1507) as a first background object type by referring to the depth information of the first image, which is flat with little depth variation and is distributed horizontally.
[0304] The electronic device can determine the object (1503) as a second background object type because the tag for the object (1503) includes a predefined classification result (e.g., sky) regarding the second background object type. The electronic device can determine the object (1503) located farther than a threshold distance as a second background object type by referring to the depth information of the first image. The electronic device can determine the object (1501) located farther than a threshold distance as a background object type (e.g., second background object type) by referring to the depth information of the first image.
[0305] As described with reference to FIG. 9, the electronic device can determine context information regarding the first image. The electronic device can determine scene information regarding the first image using a trained model. The electronic device can determine the object types of the objects in the first image based on the first tag and the context information regarding the first image. When the background of the scene information of the first image is a park, the electronic device can determine a 'bench,' i.e., an object (1502), which has high similarity to the scene information, as the main object type.
[0306] As described with reference to FIG. 10, the electronic device can determine individual scores of objects in the first image based on first tags of objects in the first image. Based on individual scores of objects in the first image, the electronic device can determine object types of objects in the first image according to a predetermined priority criterion. According to one embodiment, the predetermined priority criterion may include determining an object as a main object type when the score of any object in the first image is above (or exceeds) a threshold. The electronic device may add (or increase) the score for any object in the first image if the tag for any object in the first image includes predefined keywords such as, for example, dynamic object, interactive with user, or accessible to user. The electronic device may determine an object (1506) as a main object type when the score of an object (1506) corresponding to a 'dog' that is dynamic and interactive with user is above (or exceeds) a threshold.
[0307] Referring to the screen (152), the electronic device can identify objects (1508, 1509, 1510, 1511, 1512) included in the second image.
[0308] According to one embodiment, the electronic device can determine a second object of a second image associated with a first object based on a second tag for objects (1508, 1509, 1510, 1511, 1512) of the second image.
[0309] For example, the electronic device may determine an object (1510) in which the corresponding tag is a 'floor' which is a predefined classification result with respect to a first background type as a second object associated with an object (1507) corresponding to a first background object type. The electronic device may determine an object (1509) in which the corresponding tag is a 'wall' which is a predefined classification result with respect to a second background type as a second object associated with an object (1501, 1503) corresponding to a second background object type.
[0310] As described with reference to FIG. 13, the electronic device may determine individual matching scores between the first object and the objects (1508, 1509, 1510, 1511, 1512) of the second image based on a comparison between the tag of the first object of the first image and the second tag for the objects (1508, 1509, 1510, 1511, 1512) of the second image. The electronic device may add matching scores between the first object and the corresponding object of the second image if the tag of the first object and the tag of any object of the second image are identical or if the similarity is above (or exceeds) a threshold. For example, a matching score can be calculated between an object (1506) corresponding to a 'dog' as the first object, which is the main object type of the first image, and objects (1508, 1509, 1510, 1511, 1512) of the second image. Among the objects (1508, 1509, 1510, 1511, 1512) of the second image, the matching score between the first object and an object (1512) corresponding to a 'robot vacuum cleaner' that is dynamic, interactive with the user, or accessible to the user can be calculated to be the highest. The electronic device can determine the object (1512) with the highest matching score corresponding to the first object as the second object associated with the first object.
[0311] Referring to screen (153), in the composite image, regarding the background object type, object (1510) can be replaced with object (1507), and object (1509) can be replaced with objects (1501, 1503).
[0312] Referring to screen (153), in the composite image, regarding the main object type, object (1512) can be replaced with object (1506).
[0313] The electronic device can identify an object (1508) included in a closed area within the region of an object (1509) as a third object. Based on a comparison between a first tag for objects in a first image and a tag of an object (1508), the electronic device can determine whether there exists a fourth object in the first image for which the corresponding comparison result satisfies a first predetermined criterion. If there is no fourth object in the first image for which the corresponding comparison result satisfies the first predetermined criterion, the electronic device can identify an object (1508) that is the third object in the second image as part of the object (1509).
[0314] Accordingly, by referring to screen (153), an object (1508) is identified as an object (1509), that is, part of the background, and then a natural composite image can be produced by replacing the entire background including the object (1508) in the second image with the background of the first image (e.g., objects (1501, 1503)).
[0315] FIG. 16 is a flowchart of a method for generating an image according to one embodiment.
[0316] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0317] According to one embodiment, the following operations 1610 to 1670 may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3a and 3b, or the electronic device (510) of FIG. 5). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1) including a processing circuit. The electronic device may include a memory (e.g., the memory (130) of FIG. 1) including one or more storage media for storing instructions. The electronic device may include a display (e.g., the display module (160) of FIG. 1).
[0318] According to one embodiment, the electronic device may establish a communication connection with an external electronic device (e.g., the electronic device (102) of FIG. 1, the electronic device (104) of FIG. 1, or the external electronic device (520) of FIG. 5). The electronic device may establish a communication connection with the external electronic device through a short-range wireless communication network (e.g., the first network (198) of FIG. 1) or a long-range wireless communication network (e.g., the second network (199) of FIG. 1).
[0319] In operation 1610, the electronic device may receive a first image and object area information specified for the first image from an external electronic device that has established a communication connection with the electronic device. The first image may be a still image or video of the actual space (or physical environment) around the external electronic device.
[0320] According to one embodiment, an external electronic device can transmit a first image acquired using a camera of the external electronic device to an electronic device. For example, the electronic device can receive a first image acquired using a camera of the external electronic device from an external electronic device that has established a call connection with the electronic device.
[0321] According to one embodiment, an external electronic device can transmit a first image stored in the external electronic device to an electronic device.
[0322] According to one embodiment, an external electronic device may receive user input specifying an arbitrary region or object with respect to a first image. The external electronic device may transmit object region information determined based on the user input to an electronic device.
[0323] For example, object region information may include coordinate information of a designated region within the first image. An external electronic device may transmit object region information including coordinate information of a designated region within the first image to the electronic device.
[0324] For example, object region information may include a designated object (or object region or coordinate information of the object region) and / or a tag for said object within the first image. An external electronic device may identify an object (or object region) included in a designated area within the first image based on user input. The external electronic device may determine a tag for an object identified in the first image. A tag for an object may include a keyword (or a text description of the object) indicating the classification result of the object and / or characteristics regarding the object. The external electronic device may transmit object region information including a designated object within the first image and / or a tag for said object to an electronic device.
[0325] According to one embodiment, the electronic device may acquire a second image using the electronic device’s camera (e.g., sensor module (176) of FIG. 1, camera module (180), first camera (265a, 265b), second camera (270a, 270b), third camera (245) of FIG. 2, second functional camera (311, 312), first functional camera (315), depth sensor (317) of FIG. 3a, third functional camera (328) of FIG. 3b, or fourth functional camera (325, 326)). The second image may be a still image or video of the actual space (or physical environment) around the electronic device.
[0326] In operation 1620, the electronic device can identify objects of the first image and objects of the second image acquired using the camera of the electronic device.
[0327] According to one embodiment, the electronic device can detect objects included in each of the first image and the second image. For example, the electronic device can determine the bounding boxes of the objects detected in each of the first image and the second image. The electronic device can identify each object (or object region) by performing segmentation (e.g., semantic, instance, or panoptic segmentation) on each of the first image and the second image, based on or independently of the result of object detection.
[0328] In operation 1630, the electronic device can determine a first tag for objects in a first image and a second tag for objects in a second image acquired using the camera of the electronic device.
[0329] According to one embodiment, tags for an object (e.g., a first tag and a second tag) may include keywords (or a text description of the object) indicating the classification result of the object and / or characteristics regarding the object.
[0330] For example, as tags for objects, classification results may include labels (or classes) or categories. For example, keywords as tags for objects may include characteristics such as the object's shape, form, size, texture, location, whether it is a dynamic or static object, interactivity with the user, or user accessibility.
[0331] The first tag may include one or more tags for each of the objects in the first image. The second tag may include one or more tags for each of the objects in the second image.
[0332] In operation 1640, the electronic device can determine the object types of the objects in the first image based on the object region information of the first tag and the first image.
[0333] According to one embodiment, the electronic device can determine at least some of the objects of a first image as background object types based on a first tag. The electronic device can determine at least some of the objects of a first image as main object types based on object region information of the first image.
[0334] As described with reference to FIG. 6, the first tag may include one or more tags for each of the objects in the first image. According to one embodiment, the electronic device may determine that an object is a background object type if the tag for any object in the first image includes a predefined classification result such as, for example, wall, floor, sky, ground, or ceiling. According to one embodiment, the electronic device may determine the object types of the objects in the first image by referring to the depth information of the first image.
[0335] According to one embodiment, the electronic device can determine an object identified based on coordinate information of a designated area within a first image included in object area information of a first image as a main object type. The electronic device can identify an object included in the coordinates of the designated area in the received first image. The electronic device can determine the identified object as a main object type that reflects the user's intention.
[0336] According to one embodiment, the electronic device can identify an object referenced by a designated object (or object region or coordinate information of the object region) and / or a tag (e.g., classification result) included in the object region information of the first image. The electronic device can determine the identified object as a main object type that reflects the user's intention.
[0337] In operation 1650, the electronic device can determine a second object of a second image associated with a first object corresponding to a first object type among the objects of a first image based on a second tag.
[0338] According to one embodiment, the first object type may include a background object type. Regarding the method for determining a second object in a second image associated with a first object corresponding to a background object type among the objects in the first image, a description that overlaps with the content described above with reference to FIG. 12 is omitted.
[0339] According to one embodiment, the first object type may include an other object type (or a general object type) that is not determined to be a background object type or a main object type. Regarding the method for determining a second object in a second image associated with a first object corresponding to an other object type among the objects in the first image, a description that overlaps with the content described above with reference to FIG. 14 is omitted.
[0340] According to one embodiment, the electronic device can identify a first object corresponding to a background object type as a first object type among the objects of a first image. The background object type may include a first background object type and a second background object type. The electronic device can determine a second object of a second image associated with the first object based on a second tag for the objects of the second image.
[0341] According to one embodiment, the electronic device can identify a first object corresponding to a main object type as a first object type among the objects of a first image. The electronic device can determine an object in a second image associated with the first object as a second object based on user input of the electronic device. For example, the electronic device can display the first object determined as the main object type on a display based on object region information received from an external electronic device. The electronic device can output a guide (e.g., 'Please select an object to replace with the first object') that induces the user to select a second object to replace the first object in the second image. The electronic device can receive user input selecting any object to replace with the first object in the second image. Based on receiving user input selecting any object to replace with the first object in the second image, the electronic device can determine the selected object as the second object.
[0342] In operation 1660, the electronic device can generate a composite image in which the second object in the second image is replaced with the first object. A description of the method for generating the composite image that overlaps with the description previously given with reference to FIG. 6 is omitted.
[0343] According to one embodiment, the electronic device can identify a fifth object corresponding to a main object type as a first object type among the objects of the first image. The electronic device can determine a location to place the fifth object in the second image based on user input of the electronic device. For example, the electronic device can display the fifth object determined as the main object type on a display based on object area information received from an external electronic device. The electronic device can output a guide (e.g., 'Please specify a location to place the fifth object') that induces the user to specify a location to place the fifth object in the second image. The electronic device can receive user input specifying an arbitrary location to place the fifth object in the second image. Based on receiving user input specifying an arbitrary location to place the fifth object in the second image, the electronic device can generate a composite image in which the fifth object is placed at the specified location. According to one embodiment, the electronic device can generate a composite image in which the second object in the second image is replaced by the first object and the fifth object is placed at the specified location in the second image.
[0344] According to one embodiment, the electronic device can determine the environmental properties of a first image. The electronic device can generate a composite image in which a second object in a second image is replaced with a first object, and a graphic effect is applied to at least a part of the second image based on the environmental properties.
[0345] In operation 1670, the electronic device can display a composite image through a display.
[0346] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.
[0347] According to one embodiment, an electronic device (101, 201, 301, 510) comprises: a display (160); at least one processor (120) including processing circuitry; and a memory (130) including one or more storage media for storing instructions, and when instructions are executed individually or collectively by at least one processor (120), the electronic device (101, 201, 301, 510) causes: an operation (610) to receive a first image from an external electronic device (102, 104, 520) that has established a communication connection with the electronic device (101, 201, 301, 510); an operation (620) to identify objects of the first image and objects of a second image acquired using a camera of the electronic device; and an operation (630) to determine a first tag for objects of the first image and a second tag for objects of the second image. The operation (640) of determining the object type of the objects of the first image based on at least the first tag; the operation (650) of determining the second object of the second image associated with the first object corresponding to the first object type among the objects of the first image based on at least the second tag; the operation (660) of generating a composite image in which the second object in the second image is replaced with the first object; and the operation (670) of displaying the composite image through the display (160) may be performed.
[0348] According to one embodiment, the operation (640) of determining the object type of objects in the first image based on at least the first tag may include: the operation (710) of receiving object region information specified for the first image from an external electronic device (102, 104, 520); and the operation (720) of determining the object type of objects in the first image based on the first tag and the object region information of the first image.
[0349] According to one embodiment, the operation (640) of determining the object type of objects in the first image based on at least the first tag may include: the operation (910) of determining context information regarding the first image; and the operation (920) of determining the object type of objects in the first image based on the first tag and the context information regarding the first image.
[0350] According to one embodiment, context information regarding the first image may include at least one of scene information regarding the first image determined using a trained model, subject information regarding voice obtained during a call connection with an external electronic device (102, 104, 520), eye-tracking information received from the external electronic device (102, 104, 520), or location information received from the external electronic device (102, 104, 520).
[0351] According to one embodiment, the operation (640) of determining the object type of the objects of the first image based on at least the first tag may include: the operation (1010) of determining the individual (respective) score of the objects of the first image according to the similarity between the first tag of the objects of the first image and the context information of the first image; and the operation (1020) of determining the object type of the objects of the first image according to a defined priority criterion based on the individual score of the objects of the first image.
[0352] According to one embodiment, the operation (1010) of determining individual scores of objects in a first image based on the similarity between the first tag of objects in a first image and the context information of the first image may further include the operation of adding a score for any object based on the determination that the tag for any object in the first image includes a predefined keyword.
[0353] According to one embodiment, the operation (650) of determining a second object of a second image associated with a first object corresponding to a first object type among the objects of a first image based on at least a second tag may include: an operation (1210) of identifying a first object corresponding to a background object type as a first object type among the objects of the first image - the background object type includes a first background object type and a second background object type -; and an operation (1220) of determining a second object of a second image associated with a first object based on a second tag for the objects of the second image.
[0354] According to one embodiment, when instructions are executed individually or collectively by at least one processor (120), the electronic device (101, 201, 301, 510) may further perform: determining whether a third object is included in a closed area within the region of the second object based on determining that the second object is associated with a first object corresponding to a second background object type; determining whether a fourth object of the first image is included in a closed area within the region of the second object based on determining that a third object is included in the closed area within the region of the second object based on a comparison between a first tag for objects of the first image and a tag of the third object, such that the corresponding comparison result satisfies a first determined criterion; and determining that the third object in the second image is replaced with the fourth object based on determining that the fourth object in the first image is present, and identifying the third object in the second image as part of the second object based on determining that the fourth object in the first image is not present.
[0355] According to one embodiment, the operation (650) of determining a second object of a second image associated with a first object corresponding to a first object type among the objects of a first image based on at least a second tag may include: an operation (1310) of identifying a first object corresponding to a main object type as a first object type among the objects of the first image; and an operation (1320) of determining a second object of a second image such that a corresponding comparison result satisfies a second predetermined criterion based on a comparison between the tag of the first object and the second tag for the objects of the second image.
[0356] According to one embodiment, the operation (650) of determining a second object of a second image associated with a first object corresponding to a first object type among the objects of a first image based on at least a second tag may include: an operation (1410) of identifying a first object that does not correspond to a background object type among the objects of the first image; and an operation (1420) of determining a second object of a second image such that a corresponding comparison result satisfies a third predetermined criterion based on a comparison between the tag of the first object and the second tag for the objects of the second image.
[0357] According to one embodiment, when instructions are executed individually or collectively by at least one processor (120), the electronic device (101, 201, 301, 510) may further perform: an operation of determining whether an object corresponding to a user in a second image approaches within a threshold distance of a second object; and an operation of updating a composite image so that the first object is replaced by the second object based on the determination that the object corresponding to the user has approached within a threshold distance of a second object, and based on the determination that the first object does not correspond to the main object type.
[0358] According to one embodiment, when instructions are executed individually or collectively by at least one processor (120), the electronic device (101, 201, 301, 510) is further made to perform an operation to determine the environmental properties of a first image, and the operation (660) to generate a composite image in which a second object in a second image is replaced with a first object may include an operation to generate a composite image in which a second object in a second image is replaced with a first object and a graphic effect is applied to at least a part of the second image based on the environmental properties.
[0359] According to one embodiment, when instructions are executed individually or collectively by at least one processor (120), the electronic device (101, 201, 301, 510) further performs the operation of determining an object satisfying a fourth predetermined criterion in a second image based on a second tag as a non-replacement object type, and the operation (650) of determining a second object in a second image associated with a first object corresponding to a first object type among objects in a first image based on at least the second tag may include the operation of determining a second object associated with a first object among objects in a second image that are not identified as a non-replacement object type based on the second tag.
[0360] According to one embodiment, an electronic device (101, 201, 301, 510) comprises: a display (160); at least one processor (120) including processing circuitry; and a memory (130) including one or more storage media for storing instructions, and when instructions are executed individually or collectively by at least one processor (120), the electronic device (101, 201, 301, 510) causes: an operation (1610) to receive a first image and object region information designated for the first image from an external electronic device (102, 104, 520) that has established a communication connection with the electronic device (101, 201, 301, 510); and an operation (1620) to identify objects of the first image and objects of a second image acquired using a camera of the electronic device (101, 201, 301, 510). The operation (1630) of determining a first tag for objects in a first image and a second tag for objects in a second image; the operation (1640) of determining the object type of objects in a first image based on the first tag and object area information of the first image; the operation (1650) of determining a second object in a second image associated with a first object corresponding to the first object type among the objects in the first image based on at least the second tag; the operation (1660) of generating a composite image in which the second object in the second image is replaced with the first object; and the operation (1670) of displaying the composite image through a display (160) may be performed.
[0361] According to one embodiment, the operation (1640) of determining the object type of the objects of the first image based on the object region information of the first image and the first tag may include: the operation of determining at least some of the objects of the first image as background object types based on the first tag; and the operation of determining at least some of the objects of the first image as main object types based on the object region information of the first image.
[0362] According to one embodiment, the operation (1650) of determining a second object of a second image associated with a first object corresponding to a first object type among the objects of a first image based on at least a second tag may include: an operation of identifying a first object corresponding to a background object type as a first object type among the objects of a first image - the background object type includes a first background object type and a second background object type -; and an operation of determining a second object of a second image associated with a first object based on a second tag for the objects of the second image.
[0363] According to one embodiment, the operation (1650) of determining a second object of a second image associated with a first object corresponding to a first object type among the objects of a first image based on at least a second tag may further include: an operation of identifying a first object corresponding to a main object type as a first object type among the objects of the first image; and an operation of determining a second object of a second image associated with a first object based on user input of an electronic device (101, 201, 301, 510).
[0364] According to one embodiment, a method performed by an electronic device (101, 201, 301, 510) further includes: identifying a fifth object corresponding to a main object type as a first object type among the objects of a first image; and determining a position to place the fifth object in a second image based on user input of the electronic device (101, 201, 301, 510), and the operation (1660) of generating a composite image in which a second object is replaced by a first object in the second image may include the operation of generating a composite image in which the second object is replaced by a first object in the second image and the fifth object is placed at a position in the second image.
[0365] According to one embodiment, when instructions are executed individually or collectively by at least one processor (120), the electronic device (101, 201, 301, 510) is further made to perform an operation to determine the environmental properties of a first image, and the operation (1660) to generate a composite image in which a second object in a second image is replaced with a first object may include an operation to generate a composite image in which a second object in a second image is replaced with a first object and a graphic effect is applied to at least a part of the second image based on the environmental properties.
[0366] A method performed by an electronic device (101, 201, 301, 510) according to one embodiment comprises: an operation (610) of receiving a first image from an external electronic device (102, 104, 520) that has established a communication connection with the electronic device (101, 201, 301, 510); an operation (620) of identifying objects of the first image and objects of a second image acquired using a camera of the electronic device (101, 201, 301, 510); an operation (630) of determining a first tag for objects of the first image and a second tag for objects of the second image; an operation (640) of determining an object type of objects of the first image based on at least the first tag; and an operation (650) of determining a second object of the second image associated with a first object corresponding to a first object type among objects of the first image based on at least the second tag. It may include an operation (660) of generating a composite image in which a second object in a second image is replaced with a first object; and an operation (670) of displaying the composite image through a display (160) of an electronic device (101, 201, 301, 510).
[0367] A method performed by an electronic device (101, 201, 301, 510) according to one embodiment comprises: an operation (1610) of receiving a first image and object region information designated for the first image from an external electronic device (102, 104, 520) that has established a communication connection with the electronic device (101, 201, 301, 510); an operation (1620) of identifying objects of the first image and objects of a second image acquired using a camera of the electronic device (101, 201, 301, 510); an operation (1630) of determining a first tag for objects of the first image and a second tag for objects of the second image; and an operation (1640) of determining the object type of objects of the first image based on the first tag and the object region information of the first image. The method may include: an operation (1650) of determining a second object in a second image associated with a first object corresponding to a first object type among the objects in a first image based on at least a second tag; an operation (1660) of generating a composite image in which the second object in the second image is replaced with a first object; and an operation (1670) of displaying the composite image through a display (160) of an electronic device (101, 201, 301, 510).
[0368] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.
[0369] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0370] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0371] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0372] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0373] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0374] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0375] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and software applications executed on the operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0376] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave in order to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on computer-readable recording media.
[0377] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination, and the program instructions recorded on the medium may be those specifically designed and configured for the embodiment or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0378] The hardware device described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0379] Although the embodiments described above have been explained with reference to limited drawings, those skilled in the art can apply various technical modifications and variations based thereon. For example, appropriate results can be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0380] Therefore, other implementations, one embodiment, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
1. In an electronic device (101, 201, 301, 510), Display (160); At least one processor (120) including processing circuitry; and It includes a memory (130) comprising one or more storage media for storing instructions, and When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101, 201, 301, 510) is made to: An operation (610) of receiving a first image from an external electronic device (102, 104, 520) that has established a communication connection with the above electronic device (101, 201, 301, 510); An operation (620) to identify objects of the first image and objects of the second image acquired using the camera of the electronic device (101, 201, 301, 510); An operation (630) for determining a first tag for objects in the first image and a second tag for objects in the second image; At least an operation (640) for determining the object types of the objects of the first image based on the first tag; At least an operation (650) of determining a second object of the second image associated with a first object corresponding to a first object type among the objects of the first image based on the second tag; The operation (660) of generating a composite image in which the second object in the second image is replaced with the first object; and The operation (670) of displaying the above composite image through the display (160) causing to perform, Electronic device (101, 201, 301, 510).
2. In Paragraph 1, At least the operation (640) of determining the object type of the objects of the first image based on the first tag is, The operation (710) of receiving object region information designated for the first image from the above external electronic device; and Operation (720) of determining the object type of the objects of the first image based on the object region information of the first tag and the first image including, Electronic device (101, 201, 301, 510).
3. In either Paragraph 1 or Paragraph 2, At least the operation (640) of determining the object type of the objects of the first image based on the first tag is, An operation (910) for determining context information regarding the first image above; and An operation (920) to determine the object type of the objects of the first image based on the context information regarding the first tag and the first image. including, Electronic device (101, 201, 301, 510).
4. In any one of paragraphs 1 through 3, Contextual information regarding the first image above is, at least one of scene information for the first image determined using a trained model, subject information for voice acquired during a call connection with the external electronic device, eye-tracking information received from the external electronic device, or location information received from the external electronic device. Electronic device (101, 201, 301, 510).
5. In any one of paragraphs 1 through 4, At least the operation (640) of determining the object type of the objects of the first image based on the first tag is, An operation (1010) for determining a respectful score of the objects of the first image according to the similarity between the first tag of the objects of the first image and the context information of the first image; and An operation (1020) to determine the object type of the objects in the first image according to a defined priority criterion based on the individual scores of the objects in the first image. including, Electronic device (101, 201, 301, 510).
6. In any one of paragraphs 1 through 5, The operation (1010) of determining individual scores of objects in the first image according to the similarity between the first tag of the objects in the first image and the context information of the first image is, An operation to add a score for an arbitrary object based on the determination that a tag for an arbitrary object in the first image includes a predefined keyword. including more, Electronic device (101, 201, 301, 510).
7. In any one of paragraphs 1 through 6, The operation (650) of determining the second object of the second image associated with the first object corresponding to the first object type among the objects of the first image based on at least the second tag is, An operation (1210) for identifying the first object corresponding to the background object type as the first object type among the objects of the first image above - the background object type includes the first background object type and the second background object type -; and The operation (1220) of determining the second object of the second image associated with the first object based on the second tag for the objects of the second image including, Electronic device (101, 201, 301, 510).
8. In any one of paragraphs 1 through 7, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101, 201, 301, 510) is made to: An operation to determine whether a third object exists within a closed area within the region of the second object, based on the determination that the second object is associated with the first object corresponding to the second background object type; Based on the determination that the third object included in the closed region within the region of the second object exists, and based on a comparison between the first tag for at least the objects of the first image and the tag of the third object, determining whether the corresponding comparison result satisfies the first predetermined criterion for the fourth object of the first image; and Based on the determination that the fourth object exists in the first image, the operation of determining that the third object in the second image is replaced by the fourth object, and based on the determination that the fourth object does not exist in the first image, identifying the third object in the second image as part of the second object. to make it perform more, Electronic device (101, 201, 301, 510).
9. In any one of paragraphs 1 through 8, The operation (650) of determining the second object of the second image associated with the first object corresponding to the first object type among the objects of the first image based on at least the second tag is, An operation (1310) of identifying the first object corresponding to the main object type as the first object type among the objects of the first image; and An operation (1320) of determining the second object of the second image whose corresponding comparison result satisfies a second predetermined criterion based on a comparison between the tag of the first object and the second tag for the objects of the second image. including, Electronic device (101, 201, 301, 510).
10. In any one of paragraphs 1 through 9, The operation (650) of determining the second object of the second image associated with the first object corresponding to the first object type among the objects of the first image based on at least the second tag is, An operation (1410) for identifying the first object that does not correspond to a background object type among the objects of the first image; and An operation (1420) of determining the second object of the second image such that the corresponding comparison result satisfies a third predetermined criterion based on a comparison between the tag of the first object and the second tag for the objects of the second image. including, Electronic device (101, 201, 301, 510).
11. In any one of paragraphs 1 through 10, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101, 201, 301, 510) is made to: An operation to determine whether an object corresponding to the user in the second image approaches within a threshold distance of the second object; and An operation to update the composite image so that the first object is replaced by the second object, based on the determination that the object corresponding to the user has approached the second object within the threshold distance, and based on the determination that the first object does not correspond to the main object type. to make it perform more, Electronic device (101, 201, 301, 510).
12. In any one of paragraphs 1 through 11, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101, 201, 301, 510) is made to: Operation to determine the environmental attributes of the first image above To perform further, The operation (660) of generating the composite image in which the second object in the second image is replaced with the first object is, The operation of replacing the second object with the first object in the second image and generating the composite image with graphic effects applied based on the environment attributes to at least a portion of the second image. including, Electronic device (101, 201, 301, 510).
13. In any one of paragraphs 1 through 12, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101, 201, 301, 510) is made to: An operation to determine an object satisfying a fourth defined criterion in the second image as a non-replacement object type based on the second tag. To perform further, The operation (650) of determining the second object of the second image associated with the first object corresponding to the first object type among the objects of the first image based on at least the second tag is, An operation to determine the second object associated with the first object among the objects of the second image that are not identified as the non-replacement object type based on the second tag. including, Electronic device (101, 201, 301, 510).
14. In an electronic device (101, 201, 301, 510), Display (160); At least one processor (120) including processing circuitry; and It includes a memory (130) comprising one or more storage media for storing instructions, and When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101, 201, 301, 510) is made to: An operation (1610) of receiving a first image and object region information designated for the first image from an external electronic device (102, 104, 520) that has established a communication connection with the electronic device (101, 201, 301, 510); An operation (1620) of identifying objects of the first image and objects of the second image acquired using the camera of the electronic device (101, 201, 301, 510); An operation (1630) for determining a first tag for objects in the first image and a second tag for objects in the second image; An operation (1640) to determine the object types of the objects of the first image based on the object region information of the first tag and the first image; At least an operation (1650) of determining a second object of the second image associated with a first object corresponding to a first object type among the objects of the first image based on the second tag; The operation (1660) of generating a composite image in which the second object in the second image is replaced with the first object; and The operation (1670) of displaying the above composite image through the above display causing to perform, Electronic device (101, 201, 301, 510).
15. A method performed by an electronic device (101, 201, 301, 510), An operation (610) of receiving a first image from an external electronic device (102, 104, 520) that has established a communication connection with the above electronic device (101, 201, 301, 510); An operation (620) to identify objects of the first image and objects of the second image acquired using the camera of the electronic device (101, 201, 301, 510); An operation (630) for determining a first tag for objects in the first image and a second tag for objects in the second image; At least an operation (640) for determining the object types of the objects of the first image based on the first tag; At least an operation (650) of determining a second object of the second image associated with a first object corresponding to a first object type among the objects of the first image based on the second tag; The operation (660) of generating a composite image in which the second object in the second image is replaced with the first object; and The operation (670) of displaying the above composite image through the display (160) of the electronic device (101, 201, 301, 510) including, method.