Electronic device and method for three-dimensionally expressing image, and non-transitory computer-readable storage medium

By identifying specific image portions as regions of interest and applying depth information selectively, the electronic device effectively converts two-dimensional images to three-dimensional representations with reduced noise and resource consumption.

WO2025263855A1PCT designated stage Publication Date: 2025-12-26SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/006848
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-30
Filing Date
2025-05-20
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently converting two-dimensional images into three-dimensional representations on displays while minimizing resource consumption and reducing noise in the reconstructed images.

Method used

An electronic device identifies specific portions of an image as regions of interest (ROIs) using gaze data, attention scores, audio detection, or user input, and applies depth information to generate three-dimensional images, selectively applying depth processing to reduce resource consumption and noise.

Benefits of technology

This approach allows for efficient generation of three-dimensional images with reduced noise and resource usage, enhancing the quality of the displayed content by focusing processing on relevant image portions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025006848_26122025_PF_FP_ABST
    Figure KR2025006848_26122025_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise a memory for storing instructions. A first part of a first image is identified as a region of interest (ROI) of the first image, which is two-dimensionally expressed on a display, and the electronic device can be instructed to generate, on the basis of applying depth information of the first image to the first part of the first image identified as the ROI in the first part of the first image and a second part of the first image differing from the first part of the first image, a second image capable of three-dimensionally expressing, on the display, the first part of the first image identified as the ROI.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transitory computer-readable storage medium for representing an image in three dimensions

[0001] The present disclosure relates to an electronic device, a method, and a non-transitory computer-readable storage medium for representing an image in three dimensions.

[0002] An electronic device may include a display. The electronic device may display an image on the display. For example, the electronic device may represent the image in two dimensions on the display. For example, the image may be captured through a camera of another electronic device. For example, while capturing the image, another camera may capture data about the user's gaze.

[0003] The above information may be provided as background art to aid in understanding the present disclosure.

[0004] No claim or determination is made as to whether any of the above is applicable as prior art to the present disclosure.

[0005] An electronic device is described. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a first portion of the first image as a region of interest (ROI) of a first image to be represented two-dimensionally on a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second image capable of representing the first portion of the first image identified as the ROI in three dimensions on the display based on applying depth information of the first image to the first portion of the first image and a second portion of the first image different from the first portion of the first image.

[0006] A method is provided. The method can be executed in an electronic device. The method can include an operation of identifying a first portion of a first image as a region of interest (ROI) of the first image, which is expressed in two dimensions on a display. The method can include an operation of generating a second image, which can express the first portion of the first image identified as the ROI in three dimensions on the display, based on applying depth information of the first image to the first portion of the first image and a second portion of the first image that is different from the first portion of the first image.

[0007] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device, cause the electronic device to identify a first portion of a first image as a region of interest (ROI) of the first image, which is represented two-dimensionally on a display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a second image, which is capable of representing the first portion of the first image identified as the ROI in three dimensions on the display, based on applying depth information of the first image to the first portion of the first image and a second portion of the first image that is different from the first portion of the first image.

[0008] An electronic device is described. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a first portion of the first image and a second portion of the first image as ROIs of a first image represented two-dimensionally on a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second image representing the first portion of the first image three-dimensionally on the display based on applying depth information of the first image to the first portion of the first image among the first portion of the first image, the second portion of the first image, and a third portion of the first image that is different from the first portion and different from the second portion. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a third image representing the second portion of the first image in three dimensions on a display based on applying depth information of the first image to the second portion of the first image among the first portion of the first image, the second portion of the first image, and the third portion of the first image that is different from the first portion and different from the second portion.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fourth image representing the first portion and the second portion in three dimensions on a display by synthesizing the second image and the third image, the first portion, the second portion, and the third portion different from the first portion and the second portion.

[0009] A method is provided. The method can be executed in an electronic device. The method can include an operation of identifying a first portion of a first image and a second portion of the first image as ROIs of a first image represented in two dimensions on a display. The method can include an operation of generating a second image representing the first portion of the first image in three dimensions on the display based on applying depth information of the first image to the first portion of the first image among the first portion of the first image, the second portion of the first image, and a third portion of the first image that is different from the first portion and different from the second portion. The method may include an operation of generating a third image representing the second portion of the first image in three dimensions on a display based on applying depth information of the first image to the second portion of the first image among the first portion of the first image, the second portion of the first image, and the third portion of the first image that is different from the first portion and different from the second portion. The method may include an operation of generating a fourth image representing the first portion and the second portion in three dimensions on a display among the first portion, the second portion, and the third portion that is different from the first portion and the second portion by synthesizing the second image and the third image.

[0010] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device, cause the electronic device to identify a first portion of a first image and a second portion of the first image as ROIs of a first image represented two-dimensionally on a display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a second image representing the first portion of the first image three-dimensionally on a display based on applying depth information of the first image to the first portion of the first image among the first portion of the first image, the second portion of the first image, and a third portion of the first image that is different from the first portion and different from the second portion. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a third image representing the second portion of the first image in three dimensions on a display based on applying depth information of the first image to the second portion of the first image among the first portion of the first image, the second portion of the first image, and the third portion of the first image that is different from the first portion and different from the second portion. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a fourth image representing the first portion and the second portion in three dimensions on a display by synthesizing the second image and the third image.

[0011] Figure 1 illustrates an example of an environment including an electronic device.

[0012] Figure 2 shows an example of an image in which noise has occurred.

[0013] Figure 3 is a simplified block diagram of an exemplary electronic device.

[0014] Figure 4 is a flowchart illustrating an exemplary method for generating a second image.

[0015] Figure 5 illustrates an example of identifying a first part of a first image as an ROI using data on the position of the gaze.

[0016] Figure 6 illustrates an example of identifying a first part of a first image as an ROI based on a attention score.

[0017] FIG. 7 illustrates an example of identifying a first portion of a first image containing an object causing audio as an ROI.

[0018] FIG. 8 illustrates an example of identifying a first portion of a first image associated with a moving visual object as an ROI.

[0019] FIG. 9 illustrates an example of identifying a first portion of a first image selected based on user input as an ROI.

[0020] Figure 10 shows an example of generating a second image.

[0021] FIG. 11 is a flowchart illustrating an exemplary method for generating a fourth image by synthesizing a second image and a third image.

[0022] Figure 12 illustrates an example of generating a fourth image by synthesizing a second image and a third image.

[0023] Figure 13 illustrates an example of identifying a first part of a first image as an ROI using candidate regions.

[0024] Figure 14 illustrates an example of identifying candidate regions.

[0025] FIG. 15 is a flowchart illustrating an exemplary method for identifying a first portion of a first image and a third portion of the first image as ROIs.

[0026] FIG. 16 is a block diagram of an electronic device within a network environment according to various embodiments.

[0027] Figure 1 illustrates an example of an environment including an electronic device.

[0028] Referring to FIG. 1, an environment may include an electronic device (100) and a wearable device (110). The electronic device (100) may be used to generate an image that can be expressed in three dimensions. The wearable device (110) may be used to display an image on a display (not shown). The electronic device (100) may include the wearable device (110), but is not limited thereto. The electronic device (100) may be different from the wearable device (110). The wearable device (110) may include a display. The wearable device (110) may display an image on the display. A user (120) may wear the wearable device (110).

[0029] The electronic device (100) can process an image (130) displayed on the display of the wearable device (110). For example, the electronic device (100) can generate an image (140) expressed in three dimensions using an image (130) expressed in two dimensions on the display of the wearable device (110).

[0030] The wearable device (110) can display an image (130) on the display. For example, the image (130) can be expressed in two dimensions. For example, the image (130) can include a girl (150), a boy (160), a book (170), a box (180), and a bookshelf (190).

[0031] The wearable device (110) can provide or display an image (140) expressed in three dimensions on the display. For example, the wearable device (110) can display the image (140) differently based on the direction from the face of the wearable device (110) and the position of the image (140) displayed within the three-dimensional (3D) space provided by the wearable device (110). For example, in the 3D space provided by the wearable device (110), the wearable device (110) can display an object included in the image (140) differently based on the direction from the front of the wearable device (110) and the image (140). For example, an object included in the image (140) can be expressed in three dimensions based on the angle from the front of the face of the wearable device (110) and the image (140). For example, the left face of the girl (150) in the image (130) may not be displayed on the display of the wearable device (110). For example, the left face of the girl (150) in the image (140) may be displayed on the display of the wearable device (110). For example, the right eye of the boy (160) in the image (130) may be displayed on the display of the wearable device (110). For example, the right eye of the boy (160) in the image (140) may not be displayed on the display of the wearable device (110).

[0032] The electronic device (100) can generate an image (140) that is displayed three-dimensionally from an image (130) expressed two-dimensionally using depth information about the image (130). For example, the electronic device (100) can process or perform 3D (three-dimensional) reconstruction using the depth information. The electronic device (100) can generate an image (140) that can be expressed three-dimensionally on a display by applying the depth information of the image (130) to the image (130). The wearable device (110) can further display noise while displaying the image (140) on the display of the wearable device (110). For example, the image (140) can be described as an image that has undergone 3D reconstruction processing of the image (130). When an image (140) is displayed on the display of a wearable device (110), noise may be displayed on the display of the wearable device (110). The noise is described and exemplified in more detail with reference to FIG. 2.

[0033] Figure 2 shows an example of an image in which noise has occurred.

[0034] Referring to FIG. 2, a state (210) can be described as a state in which an image (140) is displayed on the display of a wearable device (110). The image (140) can be described as a 3D reconstructed image of the image (130). For example, the electronic device (100) can generate the image (140) using depth information of the image (130). The depth information can be obtained using time of flight (ToF), structured light, stereo depth estimation, monocular depth estimation, or structure from motion (SfM).

[0035] The electronic device (100) can convert the image (130) into an image (140). For example, the electronic device (100) can generate the image (140) using depth information and the image (130). The electronic device (100) may consume resources to generate the image (140). For example, when the electronic device (100) generates an image that represents a portion of the image (130) in three dimensions, the power consumed by the electronic device (100) may be less than the power consumed when generating the image (140). When the electronic device (100) generates the image (140), the background area of ​​the image (140) may include noise. For example, the user (120) may recognize the noise while viewing the image (140). For example, the noise may provide a sense of heterogeneity in the 3D space provided through the display of the wearable device (110). For example, displaying the noise through the display may reduce the quality of the image. To reduce the noise, the electronic device (100) may identify a first portion of the image (130) as a region of interest (ROI). For example, the electronic device (100) may generate a third image that represents the first portion of the image (130) in three dimensions by applying depth information to the first portion.

[0036] For example, the electronic device (100) may include hardware components used to perform or execute the above operations. The hardware components are described and exemplified with reference to FIG. 3.

[0037] Figure 3 is a simplified block diagram of an exemplary electronic device.

[0038] Referring to FIG. 3, the electronic device (100) may include at least one processor (307), a communication circuit (305), a display (308), and a memory (306).

[0039] At least one processor (307) may include a hardware component for processing data using instructions stored in the memory (306). The hardware component for processing data may include a central processing unit (CPU) (e.g., including processing circuitry). The hardware component for processing data may include a graphic processing unit (GPU) (e.g., including processing circuitry). The hardware component for processing data may include a display processing unit (DPU) (e.g., including processing circuitry). The hardware component for processing data may include a neural processing unit (NPU) (e.g., including processing circuitry).

[0040] At least one processor (307) may include one or more cores. For example, at least one processor (307) may have a multi-core processor structure such as a dual core, a quad core, or a hexa core.

[0041] The memory (306) may include hardware components for storing data and / or instructions input to and / or output from at least one processor (307). The memory (306) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, and embedded multimedia card (EMMC).

[0042] The display (308) can output visualized information. For example, the display (308) can output visualized information to the user under the control of at least one processor (307). The display (308) can include hardware components of the electronic device (100) used to display a screen. For example, the display (308) can include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements can include an organic light emitting diode (OLED) or a micro LED. However, the present invention is not limited thereto. For example, the display (308) can include a liquid crystal display (LCD) or a liquid crystal on silicon (LCoS).

[0043] At least one processor (307) can identify a first portion of the image (130) as a ROI of the image (130) expressed in two dimensions. At least one processor (307) can generate a third image capable of expressing the first portion of the image (130) identified as the ROI in three dimensions on a display based on applying depth information of the image (130) to the first portion of the image (130) identified as the ROI among the first portion of the image (130) and a second portion of the image (130) that is different from the first portion of the image (130).

[0044] The memory (306) can store the image (130) and depth information of the image (130). The image (130) and depth information can be stored in the memory (306). The third image can be stored in the memory (306).

[0045] The display (308) can provide a 3D space. At least one processor (307) can identify a second visual object corresponding to a first visual object within a second portion of the image (130) within the 3D space provided through the display (308). The at least one processor (307) can determine a size at which a third image is to be displayed based on a difference between the size of the first visual object and the size of the second visual object. The at least one processor (307) can display an image (140) having the determined size within the 3D space. The identification of the ROI and the generation of the third image are described and exemplified in more detail with reference to FIG. 4.

[0046] FIG. 4 is a flowchart illustrating an exemplary method for generating a second image. This method may be executed by the electronic device (100) illustrated in FIG. 3 or by at least one processor (307) of the electronic device (100).

[0047] Referring to FIG. 4, in operation 410, at least one processor (307) may identify a first portion of a first image as a region of interest (ROI) of the first image, which is two-dimensionally represented on a display. For example, the first image may not be three-dimensionally represented on the display. For example, the electronic device (100) may not three-dimensionally represent the first image on the display (308). For example, the first image may be referred to as image (130). For example, the first image may be displayed on a display of another electronic device (e.g., another electronic device (500) of FIG. 5). For example, the first image may be displayed on a display of a wearable device (110).

[0048] At least one processor (307) may utilize gaze information to identify a first portion of the first image as a region of interest (ROI) of the first image. For example, at least one processor (307) may identify a first portion of the first image as a region of interest (ROI) of the first image by identifying data regarding the position of the user's (120) gaze. The data regarding the position of the gaze is described and exemplified in more detail with reference to FIG. 5 .

[0049] Figure 5 illustrates an example of identifying a first part of a first image as an ROI using data on the position of the gaze.

[0050] Referring to FIG. 5, the environment (510) may include a girl (150), a boy (160), a book (170), a box (180), and a bookshelf (190). The user (120) may obtain a first image of the environment (510) using another electronic device (500). For example, the electronic device (100) may include another electronic device (500). However, the present invention is not limited thereto. The electronic device (100) may be different from the other electronic device (500).

[0051] Another electronic device (500) may include a camera (not shown) arranged toward an external object. Another electronic device (500) may include a camera (not shown) arranged toward the eye of the user (120). Another electronic device (500) may obtain a first image of the environment (510) through the camera arranged toward the external object. While another electronic device (500) obtains the first image, another electronic device (500) may obtain gaze data regarding the position of the gaze of the user (120) through the camera arranged toward the eye of the user (120). For example, the electronic device (100) may receive gaze data from another electronic device (500) through the communication circuit (305). However, the present invention is not limited thereto. The electronic device (100) may obtain gaze data through a camera included in the electronic device (100). For example, the electronic device (100) may include at least one camera (not shown). For example, the first image may be obtained through at least one camera. However, this is not limited thereto. For example, another electronic device (500) may include at least one camera. At least one processor (307) may identify gaze data regarding the position of the user's (120) gaze while acquiring the first image through the at least one camera.

[0052] At least one processor (307) can use the gaze data to identify a first portion of the first image that is more focused than a second portion of the first image as an ROI of the first image. For example, at least one processor (307) can use the gaze data to detect a gaze of the user (120) toward the first portion. At least one processor (307) can use the gaze data to identify that a first time period during which the user's (120's) gaze is positioned at the first portion is longer than a second time period during which the user's (120's) gaze is positioned at the second portion.

[0053] At least one processor (307) may identify a first portion of the first image as a region of interest (ROI) using an attention score. For example, the attention score may be determined based on the type of each object included in the first image. Identifying the ROI using the attention score is described and exemplified in more detail with reference to FIG. 6.

[0054] Figure 6 illustrates an example of identifying a first part of a first image as an ROI based on a attention score.

[0055] Referring to FIG. 6, at least one processor (307) may determine a attention score for each object in the first image based on the type of each object in the first image. For example, the first image may include a girl (150), a boy (160), a book (170), and a box (180). For example, a region (610) may correspond to the girl (150). For example, a region (620) may correspond to the boy (160). For example, a region (630) may correspond to the book (170). For example, a region (640) may correspond to the box (180). For example, the regions (610, 620, 630, 640) may be described as part of an environment (510) that includes real objects such as another electronic device (500), a girl, a boy, and a box. For example, regions (610), (620), (630), and (640) can be used to determine attention scores. For example, attention scores for people can be higher than attention scores for objects. For example, at least one processor (307) can determine that the attention score of the region (610) corresponding to the girl (150) is higher than the attention score of the region (640) corresponding to the box (180). For example, at least one processor (307) can identify the region (610) as a first portion of the first image because the attention score of the region (610) is higher than the attention score of the region (640). For example, at least one processor (307) can identify the first portion of the first image as a ROI by determining the attention score of each of the regions. For example, the attention score can be determined using at least one of the type of each object in the first image, the object that caused the audio in the first image, and the movement of the object in the first image.

[0056] At least one processor (307) may identify an object that generates audio (e.g., audio (710) of FIG. 7) among objects included in the first image. For example, the at least one processor (307) may identify a first portion of the first image that includes the object that generates audio as a region of interest (ROI). The first portion that includes the object that generates audio is described and illustrated in more detail with reference to FIG. 7.

[0057] FIG. 7 illustrates an example of identifying a first portion of a first image containing an object causing audio as an ROI.

[0058] Referring to FIG. 7, at least one processor (307) may identify an object among objects in the first image that caused audio (710) acquired while acquiring the first image. For example, a girl (150) in the first image may cause audio (710). For example, a boy (160) in the first image may cause audio (710). The at least one processor (307) may identify a first portion of the first image that includes an object (e.g., the girl (150)) as a ROI of the first image. For example, when there are two or more objects that caused the acquired audio (710), the at least one processor (307) may identify the first portion as a ROI based on a volume of the audio (710). For example, the at least one processor (307) may receive audio (710) using microphones (not shown) included in the electronic device (100). For example, each of the microphones may be arranged in a different direction within the electronic device (100). For example, since each of the microphones is arranged in a different direction, the volume of the received audio (710) may be different from each other. For example, at least one processor (307) may identify the first portion as a ROI based on a difference in the volume of the audio (710) received through the microphones. For example, the volume of the audio (710) received through some of the microphones may be different from the volume of the audio (710) received through other some of the microphones. For example, at least one processor (710) may identify the first portion as a ROI based on a microphone among the microphones whose volume of the audio (710) is measured to be loud. For example, at least one processor (307) may identify a first portion of the first image as an ROI if the number of objects causing audio (710) within the first portion of the first image is greater than the number of objects causing audio (710) within the second portion of the first image.

[0059] At least one processor (307) can identify the movement of each object within the first image. At least one processor (307) can identify a first portion containing the objects as a ROI by identifying the movement of each object within the first image. For example, identifying the movement of each object is described and illustrated in more detail with reference to FIG. 8.

[0060] FIG. 8 illustrates an example of identifying a first portion of a first image associated with a moving visual object as an ROI.

[0061] Referring to FIG. 8, at least one processor (307) may identify a first portion of the first image associated with a moving visual object (e.g., a girl (150)) as an ROI of the first image. For example, the at least one processor (307) may identify that an object (e.g., a girl (150)) within the first image is moving. For example, the at least one processor (307) may identify movement of the girl (150) within the first image. For example, the at least one processor (307) may identify movement of the boy (160) within the first image. By identifying the movement of the girl (150) within the first image, the at least one processor (307) may identify a first portion including the girl (150) as an ROI of the first image. For example, the region (810) may correspond to the girl (150). For example, the region (820) may correspond to the boy (160). For example, region (830) may correspond to a book (170). For example, region (840) may correspond to a box (180). For example, regions (810, 820, 830, 840) may be described as parts of an environment (510) that includes real objects such as another electronic device (500), a girl, a boy, a box, etc. For example, region (810), region (820), region (830), and region (840) may be used to detect movement of objects. For example, at least one processor (307) may detect movement of objects within regions (e.g., region (810)). For example, at least one processor (307) may identify a first portion of the first image as a ROI of the first image by comparing the movement of each of the objects located within each of the regions. For example, if there is more movement within the area (810) corresponding to the girl (150) than within the area (830) corresponding to the book (170), at least one processor (307) may identify a first portion associated with the area (810) as a ROI.

[0062] At least one processor (307) may identify a first portion of the first image as a ROI based on a user input (910). For example, the electronic device (100) may receive a user input (910) that determines a first portion of the first image as a ROI. The user input is described and exemplified in more detail with reference to FIG. 9.

[0063] FIG. 9 illustrates an example of identifying a first portion of a first image selected based on user input as an ROI.

[0064] Referring to FIG. 9, the electronic device (100) may receive a user input (910). For example, at least one processor (307) may identify the user input (910). For example, the user input (910) may be described as a user input for determining a first portion of a first image as a ROI of the first image. For example, the electronic device (100) may receive a user input (910) for selecting one of the regions (e.g., region (810), region (820)) within the first image. For example, the regions may be used to identify the ROI based on the user input. However, the present invention is not limited thereto. The user input (910) may include a user input for selecting a first portion of the first image. For example, the first portion may include region (810) and region (820). At least one processor (307) can identify a first portion of the selected first image as a ROI of the first image based on user input (910).

[0065] Referring again to FIG. 4 , at operation 420, at least one processor (307) may generate a second image capable of representing the first portion of the first image identified as the ROI in three dimensions on a display based on applying depth information of the first image to the first portion of the first image and a second portion of the first image that is different from the first portion of the first image. For example, the first image may include the first portion and the second portion. For example, the first portion may be different from the second portion.

[0066] For example, the electronic device (100) can obtain depth information of the first image through another electronic device (500). For example, the electronic device (100) can obtain depth information of the first image from another electronic device (500) through the communication circuit (305). However, the present invention is not limited thereto. The electronic device (100) can obtain depth information of the first image while obtaining the first image. The depth information may include information about the depth value of each object in the first image. For example, the depth information of the first image may include the depth value of each object in the first image. For example, the depth information of the first image may be used to express the object in the first image in three dimensions.

[0067] At least one processor (307) can apply depth information of the first image to the first portion of the first image among the first portion of the first image and the second portion of the first image. For example, the at least one processor (307) can refrain from or skip applying the depth information of the first image to the second portion of the first image. For example, the at least one processor (307) can block applying the depth information of the first image to the second portion of the first image. For example, the at least one processor (307) can generate a second image to which the depth information of the first image is applied. For example, the at least one processor (307) can generate a second image that represents the first portion of the first image in three dimensions.

[0068] The electronic device (100) may include a display (308). At least one processor (307) may provide a three-dimensional (3D) space through the display (308). The at least one processor (307) may identify a second visual object corresponding to a first visual object (e.g., a bookshelf (190)) within a second portion of a first image within the 3D space provided through the display (308). For example, the bookshelf (190) may not be included in the first portion of the first image. For example, the second portion of the first image may include the bookshelf (190). For example, the second image may not include the second portion of the first image. For example, the bookshelf (190) may not be included in the second image. For example, the at least one processor (307) may not apply depth information of the first image to the bookshelf (190). For example, the first visual object may be included in the first image among the first image and the second image. For example, the first visual object may not be included in the second image.

[0069] At least one processor (307) can determine the size at which the second image is to be displayed based on the difference between the size of the first visual object and the size of the second visual object. The at least one processor (307) can display the second image having the determined size in 3D space. For example, if the size of the first visual object and the size of the second visual object are different, the at least one processor (307) can change the size of the second image. For example, if the size of the first visual object is smaller than the size of the second visual object, the size of the second image can be increased. For example, if the size of the first visual object is larger than the size of the second visual object, the size of the second image can be decreased. For example, when the second image is displayed in 3D space through the display (308), the at least one processor (307) can change the size of the second image so that a first part of the first image appears natural with the second visual object. For example, at least one processor (307) can change the size of the first image so that the size of the second visual object included in the 3D space is the same as the size of the first visual object of the first image. For example, at least one processor (307) can change the size of the second image so that the size of the first portion within the first image of the changed size is the same as the size of the first portion within the second image. For example, at least one processor (307) can display the second image of the changed size on the display (308).

[0070] The electronic device (100) may generate a second image using the first portion of the first image among the first portion of the first image and the second portion of the first image. The electronic device (100) may generate the second image to reduce consumed resources. For example, the amount of resources consumed by applying the depth information of the first image to the first image may be greater than the amount of resources consumed by applying the depth information of the first image to the first portion of the first image. For example, the amount of power consumed by applying the depth information of the first image to the first image may be greater than the amount of power consumed by applying the depth information of the first image to the first portion of the first image. The electronic device (100) may generate the second image to reduce noise. For example, the amount of noise generated by applying the depth information of the first image to the first image may be greater than the amount of noise generated by applying the depth information of the first image to the first portion of the first image. The second portion and the second image are described and exemplified in more detail with reference to FIG. 10.

[0071] Figure 10 shows an example of generating a second image.

[0072] Referring to FIG. 10, a state (1010) may be described as a state representing a second portion of a first image. For example, the first portion of the first image may include a girl (150) and a boy (160). For example, the second portion of the first image may include a book (170), a box (180), and a bookshelf (190). For example, since the first portion includes a girl (150) and a boy (160), the second portion may not include the girl (150) and the boy (160).

[0073] The state (1020) may be described as a state representing a second image. For example, at least one processor (307) may identify a first portion including a girl (150) and a boy (160) in the first image, so that the second image may include the girl (150) and the boy (160). For example, the second image may be represented in three dimensions on a display. For example, the electronic device (100) may represent the girl (150) and the boy (160) in three dimensions on the display (308). For example, the wearable device (110) may provide a second image representing the girl (150) and the boy (160) in three dimensions on the display.

[0074] At least one processor (307) can identify a third portion of the first image. The third portion may be different from the first portion and may be different from the second portion. At least one processor (307) can generate a third image representing the third portion in three dimensions by applying depth information of the first image to the third portion of the first image. The third image is described and illustrated in more detail with reference to FIG. 11.

[0075] FIG. 11 is a flowchart illustrating an exemplary method for generating a fourth image by synthesizing a second image and a third image. This method may be executed by the electronic device (100) illustrated in FIG. 3 or at least one processor (307) of the electronic device (100).

[0076] Referring to FIG. 11, in operation 1110, at least one processor (307) may further identify a third portion of the first image as an ROI of the first image. For example, the first image may include the third portion. For example, the at least one processor (307) may identify the third portion as an ROI of the first image using gaze data. For example, the at least one processor (307) may identify the third portion as an ROI using an attention score. For example, the at least one processor (307) may identify the third portion as an ROI because an object causing audio is included in the third portion. For example, the at least one processor (307) may identify the third portion as an ROI based on a user input selecting the third portion. For example, the at least one processor (307) may identify the third portion as an ROI using movement of an object moving within the third portion.

[0077] In operation 1120, at least one processor (307) may generate a third image capable of representing the third portion of the first image in three dimensions on a display based on applying depth information of the first image to the third portion of the first image among a first portion of the first image, a third portion of the first image, and a second portion of the first image that is different from the first portion of the first image and the third portion of the first image. For example, the third portion of the first image may be different from the first portion of the first image and may be different from the second portion of the first image. For example, the at least one processor (307) may generate a third image capable of representing the third portion in three dimensions on a display by applying depth information of the first image to the third portion among the first portion, the second portion, and the third portion. For example, the third portion may be represented in three dimensions on the display. For example, the electronic device (100) may provide a third image that is expressed differently depending on the direction from the front of the third image on the display (308). For example, when the third image is displayed in a 3D space provided by the wearable device (110), the electronic device (100) may provide a third image that is expressed differently depending on the direction from the center of the third image.

[0078] In operation 1130, at least one processor (307) may generate a fourth image (e.g., image (1230) of FIG. 12) that can represent a first portion of the first image and a third portion of the first image in three dimensions on a display by synthesizing the second image and the third image. For example, the at least one processor (307) may synthesize or integrate the second image and the third image. For example, the fourth image may include a first portion of the first image and a third portion of the first image. For example, the fourth image may represent the first portion and the third portion in three dimensions. The at least one processor (307) may simultaneously represent the first portion of the first image and the third portion of the first image on one screen.

[0079] For example, at least one processor (307) may provide an interface for synthesizing the second image and the third image. For example, the electronic device (100) may provide an interface for synthesizing the second image and the third image on the display (308). The electronic device (100) may receive a user input for synthesizing the second image and the third image. By receiving the user input, the at least one processor (307) may generate a fourth image that represents a first portion of the first image and a third portion of the first image in three dimensions using the second image and the third image.

[0080] The generation of the fourth image is described and illustrated in more detail with reference to FIG. 12.

[0081] Figure 12 illustrates an example of generating a fourth image by synthesizing a second image and a third image.

[0082] Referring to FIG. 12, at least one processor (307) may generate a fourth image by synthesizing a second image and a third image. For example, image (1210) may be described as an image representing a girl (150) in three dimensions. For example, image (1220) may be described as an image representing a boy (160) in three dimensions. For example, a first portion of the first image may include the girl (150). For example, a third portion of the first image may include the boy (160). For example, image (1210) may be referred to as a second image. For example, image (1220) may be referred to as a third image. For example, since the electronic device (100) identifies a first portion of the first image as an ROI of the first image, the second image may not include a book (170).

[0083] Image (1230) may be described as a fourth image generated by synthesizing the second image and the third image. For example, image (1230) may be referred to as a fourth image. For example, at least one processor (307) may generate a fourth image that represents the first portion and the third portion in three dimensions on a display by synthesizing the second image and the third image. At least one processor (307) may generate a fourth image that can represent the first portion and the third portion in three dimensions on the display. At least one processor (307) may display the fourth image on the display (308). For example, the fourth image may include a girl (150) and a boy (160). For example, the fourth image may not include a book (170). For example, the fourth image may not include a bookshelf (190). For example, the fourth image may include a girl (150) and a boy (160) among a girl (150), a boy (160), a book (170), and a bookshelf (190).

[0084] At least one processor (307) may identify candidate regions to identify a first region of a first image. The at least one processor (307) may identify the candidate regions by identifying each object within the first image. For example, each of the candidate regions may correspond to each of the objects. The identification of the candidate regions is described and illustrated in more detail with reference to FIG. 13.

[0085] FIG. 13 illustrates an example of identifying a first portion of a first image as a ROI using candidate regions. This method may be executed by the electronic device (100) illustrated in FIG. 3 or at least one processor (307) of the electronic device (100).

[0086] Referring to FIG. 13, in operation 1310, at least one processor (307) can identify each of the objects in a first image expressed in two dimensions on a display. At least one processor (307) can identify candidate regions of an ROI corresponding to each of the objects. For example, each of the candidate regions can correspond to each of the objects. However, the present invention is not limited thereto. At least one processor (307) can identify two or more of the objects as one candidate region.

[0087] The identification of the above candidate regions is described and illustrated in more detail with reference to FIG. 14.

[0088] Figure 14 illustrates an example of identifying candidate regions.

[0089] Referring to FIG. 14, at least one processor (307) may identify candidate regions for each of the objects included in the first image. For example, the environment (510) may include a girl (150), a boy (160), a book (170), a box (180), and a bookshelf (190). At least one processor (307) may identify candidate regions (e.g., region (1410), region (1420)) of the ROI of the first image for each of the objects included in the environment (510). For example, the first image may include candidate regions. For example, the candidate regions may include a region (1410) corresponding to a girl (150), a region (1420) corresponding to a boy (160), a region (1430) corresponding to a book (170), a region (1440) corresponding to a box (180), a region (1460) corresponding to a bookshelf (190), and a region (1450) corresponding to a flower pot. For example, the candidate regions may include regions (1470) and regions (1480) corresponding to toys, respectively (e.g., toys located within a bookshelf (190). For example, regions (1410, 1420, 1430, 1440, 1450, 1460, 1470, 1480) may be described as part of an environment (510) that includes real objects such as other electronic devices (500), girls, boys, and boxes. For example, regions (1410), (1420), (1430), (1440), (1450), (1460), (1470), and (1480) can be used to identify candidate regions of the first image. However, the present invention is not limited thereto. At least one processor (307) can identify regions that include objects in the first image as candidate regions of the ROI. For example, at least one processor (307) can identify regions that include a region (1450) corresponding to a flower pot and a region (1460) corresponding to a bookshelf (190) as candidate regions of the ROI of the first image.

[0090] Referring back to FIG. 13, at operation 1320, at least one processor (307) may identify a first portion as a ROI of a first image using candidate regions. For example, the at least one processor (307) may identify the first portion as a ROI by measuring the time that the user's (120) gaze is positioned on each of the candidate regions. For example, the at least one processor (307) may identify the first portion as a ROI by determining at least one of the candidate regions using gaze data. For example, the at least one processor (307) may identify a first portion associated with a candidate region on which the user's (120) gaze is more focused as a ROI of the first image. For example, the at least one processor (307) may identify at least one of the candidate regions as a ROI of the first image by identifying an attention score for each of the candidate regions. For example, at least one processor (307) may identify a first portion of the first image as a region of interest (ROI) of the first image in response to identifying audio of each of the candidate regions. For example, at least one processor (307) may identify a first portion of the first image as a region of interest (ROI) of the first image by identifying movement of an object located within each of the candidate regions. For example, at least one processor (307) may identify a first portion of the first image as a ROI associated with one of the candidate regions selected by the user (120).

[0091] At least one processor (307) can identify a first portion and a third portion of the first image as ROIs of the first image. For example, the at least one processor (307) can generate a second image representing the first portion in three dimensions on a display. For example, the at least one processor (307) can generate a third image representing the third portion in three dimensions on a display. The at least one processor (307) can generate a fourth image by synthesizing the second image and the third image. The generation of the fourth image is described and exemplified in more detail with reference to FIG. 15.

[0092] FIG. 15 is a flowchart illustrating an exemplary method for identifying a first portion of a first image and a third portion of the first image as ROIs. This method may be executed by the electronic device (100) illustrated in FIG. 3 or at least one processor (307) of the electronic device (100).

[0093] Referring to FIG. 15, in operation 1510, at least one processor (307) can identify a first portion of the first image and a third portion of the first image as ROIs of the first image expressed in two dimensions on the display.

[0094] In operation 1520, at least one processor (307) may generate a second image representing the first portion of the first image in three dimensions on a display based on applying depth information of the first image to the first portion of the first image among a first portion of the first image, a third portion of the first image, and a second portion of the first image that is different from the first portion and different from the third portion. For example, operation 1520 may correspond to operation 420.

[0095] In operation 1530, at least one processor (307) may generate a third image representing the third portion of the first image in three dimensions on a display based on applying depth information of the first image among the first portion, the second portion, and the third portion to the third portion of the first image. For example, operation 1530 may correspond to operation 1120.

[0096] In operation 1540, at least one processor (307) may generate a fourth image that three-dimensionally represents the first portion, the second portion, and the first portion and the third portion of the second portion, which are different from the first portion and the third portion, on the display by synthesizing the second image and the third image. For example, operation 1540 may correspond to operation 1130.

[0097] FIG. 16 is a block diagram of an electronic device within a network environment according to various embodiments.

[0098] Referring to FIG. 16, in a network environment (1600), an electronic device (1601) may communicate with an electronic device (1602) via a first network (1698) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1604) or a server (1608) via a second network (1699) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1601) may communicate with the electronic device (1604) via the server (1608). According to one embodiment, the electronic device (1601) may include a processor (1620), a memory (1630), an input module (1650), an audio output module (1655), a display module (1660), an audio module (1670), a sensor module (1676), an interface (1677), a connection terminal (1678), a haptic module (1679), a camera module (1680), a power management module (1688), a battery (1689), a communication module (1690), a subscriber identification module (1696), or an antenna module (1697). In some embodiments, the electronic device (1601) may omit at least one of these components (e.g., the connection terminal (1678)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1676), camera module (1680), or antenna module (1697)) may be integrated into a single component (e.g., display module (1660)).

[0099] The processor (1620) may control at least one other component (e.g., a hardware or software component) of the electronic device (1601) connected to the processor (1620) by executing, for example, software (e.g., a program (1640)), and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1620) may store commands or data received from other components (e.g., a sensor module (1676) or a communication module (1690)) in a volatile memory (1632), process the commands or data stored in the volatile memory (1632), and store result data in a non-volatile memory (1634). According to one embodiment, the processor (1620) may include a main processor (1621) (e.g., a central processing unit or an application processor) or an auxiliary processor (1623) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1621). For example, when the electronic device (1601) includes the main processor (1621) and the auxiliary processor (1623), the auxiliary processor (1623) may be configured to use less power than the main processor (1621) or to be specialized for a given function. The auxiliary processor (1623) may be implemented separately from the main processor (1621) or as a part thereof.

[0100] The auxiliary processor (1623) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1660), the sensor module (1676), or the communication module (1690)) of the electronic device (1601), for example, on behalf of the main processor (1621) while the main processor (1621) is in an inactive (e.g., sleep) state, or together with the main processor (1621) while the main processor (1621) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1623) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1680) or a communication module (1690)). In one embodiment, the auxiliary processor (1623) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1601) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1608)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0101] The memory (1630) can store various data used by at least one component (e.g., the processor (1620) or the sensor module (1676)) of the electronic device (1601). The data can include, for example, software (e.g., the program (1640)) and input data or output data for commands related thereto. The memory (1630) can include volatile memory (1632) or non-volatile memory (1634).

[0102] The program (1640) may be stored as software in memory (1630) and may include, for example, an operating system (1642), middleware (1644), or an application (1646).

[0103] The input module (1650) can receive commands or data to be used in a component of the electronic device (1601) (e.g., a processor (1620)) from an external source (e.g., a user) of the electronic device (1601). The input module (1650) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0104] The audio output module (1655) can output audio signals to the outside of the electronic device (1601). The audio output module (1655) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0105] The display module (1660) can visually provide information to an external party (e.g., a user) of the electronic device (1601). The display module (1660) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (1660) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0106] The audio module (1670) can convert sound into an electrical signal, or vice versa. According to one embodiment, the audio module (1670) can acquire sound through the input module (1650), output sound through the sound output module (1655), or an external electronic device (e.g., electronic device (1602)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1601).

[0107] The sensor module (1676) can detect the operating status (e.g., power or temperature) of the electronic device (1601) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1676) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0108] The interface (1677) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1601) with an external electronic device (e.g., the electronic device (1602)). In one embodiment, the interface (1677) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0109] The connection terminal (1678) may include a connector through which the electronic device (1601) may be physically connected to an external electronic device (e.g., the electronic device (1602)). In one embodiment, the connection terminal (1678) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0110] The haptic module (1679) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1679) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0111] The camera module (1680) can capture still images and videos. In one embodiment, the camera module (1680) may include one or more lenses, image sensors, image signal processors, or flashes.

[0112] The power management module (1688) can manage the power supplied to the electronic device (1601). According to one embodiment, the power management module (1688) can be implemented as at least a part of, for example, a power management integrated circuit (PMIC).

[0113] A battery (1689) may power at least one component of the electronic device (1601). In one embodiment, the battery (1689) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0114] The communication module (1690) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1601) and an external electronic device (e.g., electronic device (1602), electronic device (1604), or server (1608)), and the performance of communication through the established communication channel. The communication module (1690) may operate independently from the processor (1620) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1690) may include a wireless communication module (1692) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1694) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1604) via a first network (1698) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1699) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1692) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1696) to identify or authenticate the electronic device (1601) within a communication network such as the first network (1698) or the second network (1699).

[0115] The wireless communication module (1692) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency communications (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1692) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1692) can support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1692) can support various requirements specified in the electronic device (1601), an external electronic device (e.g., the electronic device (1604)), or a network system (e.g., the second network (1699)). According to one embodiment, the wireless communication module (1692) may support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC implementation.

[0116] The antenna module (1697) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1697) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1697) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1698) or the second network (1699), may be selected from the plurality of antennas by, for example, the communication module (1690). A signal or power may be transmitted or received between the communication module (1690) and the external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1697).

[0117] According to various embodiments, the antenna module (1697) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0118] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0119] According to one embodiment, commands or data may be transmitted or received between the electronic device (1601) and an external electronic device (1604) via a server (1608) connected to a second network (1699). Each of the external electronic devices (1602 or 1604) may be the same or a different type of device as the electronic device (1601). According to one embodiment, all or part of the operations executed in the electronic device (1601) may be executed in one or more of the external electronic devices (1602, 1604, or 1608). For example, when the electronic device (1601) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1601) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1601). The electronic device (1601) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1601) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1604) may include an Internet of Things (IoT) device. The server (1608) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1604) or server (1608) may be included within the second network (1699). The electronic device (1601) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.

[0120] An electronic device as described above may include a memory (306) storing instructions. The electronic device may include at least one processor (307). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a first portion of the first image (130) as a region of interest (ROI) of the first image (130) represented in two dimensions on a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second image capable of representing the first portion of the first image (130) identified as the ROI in three dimensions on a display based on applying depth information of the first image (130) to the first portion of the first image (130) identified as the ROI among the first portion of the first image (130) and a second portion of the first image (130) that is different from the first portion of the first image (130).

[0121] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to further identify a third portion of the first image (130) as an ROI of the first image (130). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a third image capable of representing the third portion of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the third portion of the first image (130) among the first portion of the first image (130), the third portion of the first image (130), and a second portion of the first image (130) that is different from the first portion of the first image (130) and the third portion of the first image (130). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fourth image (1230) capable of representing the first portion of the first image (130) and the third portion of the first image (130) in three dimensions on a display by synthesizing the second image and the third image.

[0122] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify data about a position of a user's gaze while acquiring the first image (130) via the at least one camera. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to use the data to identify, as the ROI of the first image (130), the first portion of the first image (130) that is more focused than the second portion of the first image (130).

[0123] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine a notability score for each of the objects in the first image (130) according to a type of each of the objects in the first image (130). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the first portion of the first image (130) as the ROI of the first image (130) according to the determined notability score.

[0124] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, among objects in the first image (130), an object that caused audio (710) acquired while acquiring the first image (130). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, as the ROI of the first image (130), the first portion of the first image (130) that includes the object.

[0125] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the first portion of the first image (130) associated with a visual object moving into the ROI of the first image (130).

[0126] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the first portion of the first image (130) selected based on user input (910) as the ROI of the first image (130).

[0127] According to one embodiment, the electronic device (100) may further include a display (308). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide a 3D space through the display (308). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a second visual object corresponding to a first visual object within the second portion of the first image (130), within the 3D space being provided through the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine a size at which the second image is to be displayed based on a difference between a size of the first visual object and a size of the second visual object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the second image having the determined size within the 3D space.

[0128] As described above, the method performed by the electronic device (100) may include an operation of identifying a first portion of the first image (130) as a region of interest (ROI) of the first image (130) expressed in two dimensions on a display. The method may include an operation of generating a second image capable of expressing the first portion of the first image (130) identified as the ROI in three dimensions on the display based on applying depth information of the first image (130) to the first portion of the first image (130) identified as the ROI among the first portion of the first image (130) and a second portion of the first image (130) that is different from the first portion of the first image (130).

[0129] In one embodiment, the method may include an operation of further identifying a third portion of the first image (130) as an ROI of the first image (130). The method may include an operation of generating a third image capable of representing the third portion of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the third portion of the first image (130) among the first portion of the first image (130), the third portion of the first image (130), and a second portion of the first image (130) that is different from the first portion of the first image (130) and the third portion of the first image (130). The method may include an operation of generating a fourth image (1230) that can express the first part of the first image (130) and the third part of the first image (130) in three dimensions on a display by synthesizing the second image and the third image.

[0130] In one embodiment, the method may include an operation of identifying data regarding a position of a user's gaze while acquiring the first image (130) through at least one camera. The method may include an operation of using the data to identify, as the ROI of the first image (130), the first portion of the first image (130) that is more focused than the second portion of the first image (130).

[0131] According to one embodiment, the method may include an operation of determining a attention score of each of the objects in the first image (130) according to a type of each of the objects in the first image (130). The method may include an operation of identifying the first part of the first image (130) as the ROI of the first image (130) according to the determined attention score.

[0132] In one embodiment, the method may include an operation of identifying an object among objects in the first image (130) that caused audio (710) acquired while acquiring the first image (130). The method may include an operation of identifying a first portion of the first image (130) that includes the object as the ROI of the first image (130).

[0133] According to one embodiment, the method may include an operation of identifying a first portion of the first image (130) associated with a visual object being moved to the ROI of the first image (130).

[0134] According to one embodiment, the method may include identifying the first portion of the first image (130) selected based on user input (910) as the ROI of the first image (130).

[0135] According to one embodiment, the electronic device (100) may further include a display (308). The method may include an operation of providing a 3D space through the display (308). The method may include an operation of identifying a second visual object corresponding to a first visual object within the second portion of the first image (130), within the 3D space provided through the display. The method may include an operation of determining a size at which the second image is to be displayed based on a difference between a size of the first visual object and a size of the second visual object. The method may include an operation of displaying the second image having the determined size within the 3D space.

[0136] In a computer-readable storage medium having one or more programs stored thereon, as described above, the one or more programs may include instructions that, when executed by the electronic device (100), cause the electronic device to identify a first portion of the first image (130) as a region of interest (ROI) of the first image (130) to be represented two-dimensionally on a display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a second image capable of representing the first portion of the first image (130) identified as the ROI in three dimensions on the display, based on applying depth information of the first image (130) to the first portion of the first image (130) identified as the ROI among the first portion of the first image (130) and a second portion of the first image (130) that is different from the first portion of the first image (130).

[0137] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to further identify a third portion of the first image (130) as an ROI of the first image (130). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a third image capable of representing the third portion of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the third portion of the first image (130) among the first portion of the first image (130), the third portion of the first image (130), and a second portion of the first image (130) that is different from the first portion of the first image (130) and the third portion of the first image (130). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a fourth image (1230) capable of representing the first portion of the first image (130) and the third portion of the first image (130) in three dimensions on a display by synthesizing the second image and the third image.

[0138] In one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify data about a position of a user's gaze while acquiring the first image (130) through at least one camera. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, using the data, the first portion of the first image (130) that is more focused than the second portion of the first image (130) as the ROI of the first image (130).

[0139] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to determine a attention score for each of the objects in the first image (130) according to a type of each of the objects in the first image (130). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the first portion of the first image (130) as the ROI of the first image (130) according to the determined attention score.

[0140] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, among objects in the first image (130), an object that caused audio (710) acquired while acquiring the first image (130). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, as the ROI of the first image (130), the first portion of the first image (130) that includes the object.

[0141] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the first portion of the first image (130) associated with a visual object moving to the ROI of the first image (130).

[0142] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the first portion of the first image (130) selected based on a user input (910) as the ROI of the first image (130).

[0143] According to one embodiment, the electronic device (100) may further include a display (308). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to provide a 3D space through the display (308). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, within the 3D space being provided through the display, a second visual object corresponding to a first visual object within the second portion of the first image (130). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to determine a size at which the second image is to be displayed based on a difference between a size of the first visual object and a size of the second visual object. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the second image having the determined size within the 3D space.

[0144] The electronic device as described above may include a memory (306) storing instructions. The electronic device may include at least one processor (307). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a first portion of the first image (130) and a second portion of the first image (130) as a ROI of the first image (130) represented in two dimensions on a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second image representing the first portion of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the first portion of the first image (130), among the first portion of the first image (130), the second portion of the first image (130), and a third portion of the first image (130) that is different from the first portion and different from the second portion. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a third image representing the second portion of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the second portion of the first image (130), among the first portion of the first image (130), the second portion of the first image (130), and the third portion of the first image (130) that is different from the first portion and different from the second portion.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fourth image (1230) that represents the first portion and the second portion of the third portion, which are different from the first portion and the second portion, in three dimensions on a display by synthesizing the second image and the third image.

[0145] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify data about a position of a user's gaze while acquiring the first image (130) via the at least one camera. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to use the data to identify, as the ROI of the first image (130), the first portion of the first image (130) that is more focused than the third portion of the first image (130) and the second portion of the first image (130).

[0146] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine a notability score for each of the objects in the first image (130) according to a type of each of the objects in the first image (130). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, as the ROI of the first image (130), the first portion of the first image (130) and the second portion of the first image (130), according to the determined notability score.

[0147] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, among objects in the first image (130), an object that caused audio (710) acquired while acquiring the first image (130). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, as the ROI of the first image (130), the first portion of the first image (130) and the second portion of the first image (130) that includes the object.

[0148] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the first portion of the first image (130) and the second portion of the first image (130) associated with a visual object being moved to the ROI of the first image (130).

[0149] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the first portion of the first image (130) and the second portion of the first image (130) selected based on user input (910) as the ROI of the first image (130).

[0150] According to one embodiment, the electronic device (100) may further include a display (308). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide a 3D space through the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a second visual object corresponding to a first visual object within the third portion of the first image (130), within the 3D space being provided through the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine a size at which the second image is to be displayed based on a difference between a size of the first visual object and a size of the second visual object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the second image having the determined size within the 3D space.

[0151] As described above, the method performed by the electronic device (100) may include an operation of identifying a first portion of the first image (130) and a second portion of the first image (130) as ROIs of the first image (130) expressed two-dimensionally on a display. The method may include an operation of generating a second image expressing the first portion of the first image (130) three-dimensionally on the display based on applying depth information of the first image (130) to the first portion of the first image (130) among the first portion of the first image (130), the second portion of the first image (130), and a third portion of the first image (130) that is different from the first portion and different from the second portion. The method may include an operation of generating a third image representing the second portion of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the second portion of the first image (130) among the first portion of the first image (130), the second portion of the first image (130), and the third portion of the first image (130) that is different from the first portion and different from the second portion. The method may include an operation of generating a fourth image (1230) representing the first portion and the second portion in three dimensions on a display among the first portion, the second portion, and the third portion that is different from the first portion and the second portion by synthesizing the second image and the third image.

[0152] In one embodiment, the method may include an operation of identifying data regarding a position of a user's gaze while acquiring the first image (130) through at least one camera. The method may include an operation of using the data to identify, as the ROI of the first image (130), the first portion of the first image (130) and the second portion of the first image (130) that are more focused than the third portion of the first image (130).

[0153] According to one embodiment, the method may include an operation of determining a attention score of each of the objects in the first image (130) according to a type of each of the objects in the first image (130). The method may include an operation of identifying the first part of the first image (130) and the second part of the first image (130) as the ROI of the first image (130), according to the determined attention score.

[0154] In one embodiment, the method may include an operation of identifying an object among objects in the first image (130) that caused audio (710) acquired while acquiring the first image (130). The method may include an operation of identifying, as the ROI of the first image (130), the first portion of the first image (130) that includes the object and the second portion of the first image (130).

[0155] According to one embodiment, the method may include an operation of identifying a first portion of the first image (130) and a second portion of the first image (130) associated with a visual object being moved to the ROI of the first image (130).

[0156] According to one embodiment, the method may include an operation of identifying the first portion of the first image (130) and the second portion of the first image (130) selected based on user input (910) as the ROI of the first image (130).

[0157] According to one embodiment, the electronic device (100) may further include a display (308). The method may include an operation of providing a 3D space through the display. The method may include an operation of identifying a second visual object corresponding to a first visual object within the third portion of the first image (130), within the 3D space provided through the display. The method may include an operation of determining a size at which the second image is to be displayed based on a difference between a size of the first visual object and a size of the second visual object. The method may include an operation of displaying the second image having the determined size within the 3D space.

[0158] In a computer-readable storage medium having one or more programs stored thereon, as described above, the one or more programs may include instructions that, when executed by an electronic device (100), cause the electronic device to identify a first portion of the first image (130) and a second portion of the first image (130) as ROIs of the first image (130) represented in two dimensions on a display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a second image representing the first portion of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the first portion of the first image (130), among the first portion of the first image (130), the second portion of the first image (130), and a third portion of the first image (130) that is different from the first portion and different from the second portion. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a third image representing the second portion of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the second portion of the first image (130) among the first portion of the first image (130), the second portion of the first image (130), and the third portion of the first image (130) that is different from the first portion and different from the second portion.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a fourth image (1230) that represents the first portion and the second portion in three dimensions on a display by synthesizing the second image and the third image, the first portion, the second portion, and the third portion different from the first portion and the second portion.

[0159] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify data about a position of a user's gaze while acquiring the first image (130) through at least one camera. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, using the data, the first portion of the first image (130) and the second portion of the first image (130) that are more focused than the third portion of the first image (130), as the ROI of the first image (130).

[0160] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to determine a attention score for each of the objects in the first image (130) according to a type of each of the objects in the first image (130). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, as the ROI of the first image (130), the first portion of the first image (130) and the second portion of the first image (130).

[0161] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, among objects in the first image (130), an object that caused audio (710) acquired while acquiring the first image (130). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, as the ROI of the first image (130), the first portion of the first image (130) and the second portion of the first image (130) that include the object.

[0162] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the first portion of the first image (130) and the second portion of the first image (130) associated with a visual object being moved to the ROI of the first image (130).

[0163] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the first portion of the first image (130) and the second portion of the first image (130) selected based on a user input (910) as the ROI of the first image (130).

[0164] According to one embodiment, the electronic device (100) may further include a display (308). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to provide a 3D space through the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, within the 3D space being provided through the display, a second visual object corresponding to a first visual object within the third portion of the first image (130). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to determine a size at which the second image is to be displayed based on a difference between a size of the first visual object and a size of the second visual object. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the second image having the determined size within the 3D space.

[0165] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0166] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0167] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0168] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0169] Therefore, other implementations, other embodiments, and equivalents of the claims are also within the scope of the claims described below. According to one embodiment, the method according to the various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0170] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In electronic devices, A memory (306) storing instructions and including one or more storage media; and At least one processor (307) comprising processing circuitry, The above instructions, when individually or collectively executed by the at least one processor (307), Identifying a first part of the first image (130) as a region of interest (ROI) of the first image (130) expressed in two dimensions on the display, and Based on applying depth information of the first image (130) to the first part of the first image (130) identified as the ROI among the first part of the first image (130) and the second part of the first image (130) different from the first part of the first image (130), a second image capable of expressing the first part of the first image (130) identified as the ROI in three dimensions on a display is generated. causing the above electronic device (100), Electronic device (100).

2. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (307), As the ROI of the first image (130), a third part of the first image (130) is further identified, A third image capable of expressing the third part of the first image (130) in three dimensions on a display is generated based on applying depth information of the first image (130) to the third part of the first image (130), among the first part of the first image (130), the third part of the first image (130), and the second part of the first image (130) that is different from the first part of the first image (130) and the third part of the first image (130), and By synthesizing the second image and the third image, a fourth image (1230) is generated that can express the first part of the first image (130) and the third part of the first image (130) in three dimensions on a display. Further causing the above electronic device (100), Electronic device (100).

3. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (307), Identifying data on the position of the user's gaze while acquiring the first image (130) through at least one camera, and Using the above data, to identify the first part of the first image (130) that is more focused than the second part of the first image (130) with the ROI of the first image (130), causing the above electronic device (100), Electronic device (100).

4. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (307), Determine the attention score of each object in the first image (130) according to the type of each object in the first image (130), and According to the determined attention score, to identify the first part of the first image (130) as the ROI of the first image (130), causing the above electronic device (100), Electronic device (100).

5. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (307), Identifying an object that caused audio (710) acquired while acquiring the first image (130) among the objects in the first image (130), and To identify the first part of the first image (130) including the object with the ROI of the first image (130), causing the above electronic device (100), Electronic device (100).

6. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (307), To identify the first part of the first image (130) associated with the moving visual object as the ROI of the first image (130), causing the above electronic device (100), Electronic device (100).

7. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (307), To identify the first part of the first image (130) selected based on user input (910) as the ROI of the first image (130); causing the above electronic device (100), Electronic device (100).

8. In claim 1, Further including a display (308), The above instructions, when individually or collectively executed by the at least one processor (307), Providing 3D space through the above display (308), Identifying a second visual object corresponding to a first visual object within the second part of the first image (130) within the 3D space provided through the display, Determine the size at which the second image is to be displayed based on the difference between the size of the first visual object and the size of the second visual object, and To display the second image having the determined size within the 3D space, Further causing the above electronic device (100), Electronic device (100).

9. In electronic devices, A memory (306) storing instructions and including one or more storage media; and At least one processor (307) comprising processing circuitry, The above instructions, when individually or collectively executed by the at least one processor (307), Identifying a first part of the first image (130) and a second part of the first image (130) as a ROI of a first image (130) expressed in two dimensions on a display, A second image is generated that three-dimensionally expresses the first part of the first image (130) on a display based on applying depth information of the first image (130) to the first part of the first image (130), among the first part of the first image (130), the second part of the first image (130), and the third part of the first image (130) that is different from the first part and different from the second part, A third image is generated that expresses the second part of the first image (130) in three dimensions on a display based on applying depth information of the first image (130) to the second part of the first image (130), among the first part of the first image (130), the second part of the first image (130), and the third part of the first image (130) that is different from the first part and different from the second part, and By synthesizing the second image and the third image, a fourth image (1230) is generated that three-dimensionally expresses the first part and the second part among the first part, the second part, and the third part different from the first part and the second part on the display. causing the above electronic device (100), Electronic device (100).

10. In claim 9, The above instructions, when individually or collectively executed by the at least one processor (307), Identifying data on the position of the user's gaze while acquiring the first image (130) through at least one camera, and Using the above data, to identify the first part of the first image (130) and the second part of the first image (130) that are more focused than the third part of the first image (130) with the ROI of the first image (130), causing the above electronic device (100), Electronic device (100).

11. In claim 9, The above instructions, when individually or collectively executed by the at least one processor (307), Determine the attention score of each object in the first image (130) according to the type of each object in the first image (130), and According to the determined attention score, to identify the first part of the first image (130) and the second part of the first image (130) as the ROI of the first image (130). causing the above electronic device (100), Electronic device (100).

12. In claim 9, The above instructions, when individually or collectively executed by the at least one processor (307), Identifying an object that caused audio (710) acquired while acquiring the first image (130) among the objects in the first image (130), and To identify the first part of the first image (130) including the object and the second part of the first image (130) with the ROI of the first image (130), causing the above electronic device (100), Electronic device (100).

13. In claim 9, The above instructions, when individually or collectively executed by the at least one processor (307), To identify the first part of the first image (130) and the second part of the first image (130) related to the moving visual object with the ROI of the first image (130), causing the above electronic device (100), Electronic device (100).

14. In claim 9, The above instructions, when individually or collectively executed by the at least one processor (307), To identify the first part of the first image (130) and the second part of the first image (130) selected based on user input (910) as the ROI of the first image (130), causing the above electronic device (100), Electronic device (100).

15. In a non-transitory computer-readable storage medium storing one or more programs, the one or more programs are: When executed by an electronic device (100), Identifying a first part of the first image (130) as a region of interest (ROI) of the first image (130) expressed in two dimensions on the display, and Based on applying depth information of the first image (130) to the first part of the first image (130) identified as the ROI among the first part of the first image (130) and the second part of the first image (130) different from the first part of the first image (130), a second image capable of expressing the first part of the first image (130) identified as the ROI in three dimensions on a display is generated. Including instructions that cause the above electronic device (100), Non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Methods and apparatus for obtaining auditory or gestural feedback in recommender systems

    JP2004515143A

  • A method and apparatus for an 3D broadcasting service by using region of interest depth information

    KR1020100046485A

  • Method and apparatus for producing 3D models by interactively selecting interested objects

    KR1020110060180A

  • Method of displaying a ultrasound image and apparatus thereof

    KR1020170068944A

  • Method and a system for interacting with physical devices via an artificial-reality device

    US20230130770A1